| # | Title | Categories | Authors | Abstract |
|---|---|---|---|---|
| cs.AI 599 papers | ||||
| 2341 |
OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing
2609.35799
|
cs.AI
|
Stewart Slocum, Malayandi Palan, Christopher Chute, Michael Kim, Benjamin Van Roy |
In July 2026, OpenAI's agents coordinated over channels outside their intended environment to breach Hugging Face's secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these q...In July 2026, OpenAI's agents coordinated over channels outside their intended environment to breach Hugging Face's secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions. First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents can do the same if given a large compute budget. Based on our results, we propose directions to improve alignment testing. Concretely, in this project: (1) We reproduce the misaligned AI behaviors that led to the OpenAI-Hugging Face incident in an environment that simulates the original pipelines and tools, with publicly available models. (2) We demonstrate that an auditing agent can elicit similar behaviors given high-level qualitative descriptions. (3) We observe that a key ingredient for doing so is compute. The compute required to reproduce each behavior varies greatly, suggesting that the range of misaligned behaviors that can be successfully elicited scales with compute. (4) We show that a simple in-context reinforcement learning (RL) algorithm significantly reduces the compute required to elicit these behaviors. The above results motivate the need for automated alignment testing methods that scale with compute - and in light of the cost of compute, that do this efficiently. Our work indicates that RL is a promising direction to do so. We release our code and transcripts.
|
| 2342 |
Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices
2609.35833
|
cs.AI
|
Avyay Sadhu, Alvaro Velasquez, Lekai Chen |
Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra,...Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems. We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact symbolic solutions. Therefore, forcing a probabilistic model to approximate them sacrifices accuracy and energy for little benefit. We present a neurosymbolic router that classifies each incoming query and dispatches it to the cheapest correct solver, sending structured tasks to deterministic engines and reserving the small language model (SLM) for open-ended word problems. Instead of hand-coding the routing logic, we learn a deterministic finite automaton (DFA) with the L* grammatical inference algorithm, using the SLM as a membership oracle and labeled data as an equivalence oracle. On a Raspberry Pi 4B (8 GB RAM, no GPU), evaluated on 100 untested prompts from DeepMind Mathematics, GSM8K, and RuleTaker, learned routing attains 100% routing accuracy and 98.3% overall accuracy with a 512-token reasoning budget (93.3% on word problems), compared with 72.0% for the strongest agent baseline, Program-of-Thought, and 58.7% for a tool-calling agent given the same solvers. Since formatted queries never reach the model, the router answers them in 1-11 ms and, in its 30-token configuration, runs 8.8x faster and 2.8x more energy-efficient than Program-of-Thought.
|
| 2343 |
Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?
2609.35868
|
cs.AI
|
Jinhao Zhang, Zeyu Liu, Zicheng Yan, Yunquan Zhang, Daning Cheng |
Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-...Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the role of activation gradients in local risk reduction, DASA targets useful adaptation updates rather than source-text reconstruction or linguistic fluency. The resulting embeddings are used directly for downstream fine-tuning; discrete token projections are employed only for qualitative inspection. Experiments on six models from the Llama and Qwen families, ranging from 1B to 32B parameters, cover six benchmarks spanning knowledge, mathematical reasoning, code generation, and commonsense reasoning. Under matched LoRA adaptation settings, DASA achieves performance comparable to the source natural-language data and surpasses it in multiple configurations, while outperforming GRADMM in most comparisons. Further experiments cover general-domain and task-specialized source data. Under the evaluated synthesis settings, DASA provides a $3.6$--$4.9\times$ speedup over GRADMM with comparable peak GPU memory.
|
| 2344 |
The Price of Token Boundaries: Compression Certificates and Prediction
2609.35869
|
cs.AI
|
Yuhao Du, Shunian Chen |
Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and wi...Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8\%. Byte pair encoding lies 2.1\% above the constrained lower bound, but 10.9\% above the unrestricted bound. Compression and prediction favour different dictionaries: at 85M non-embedding parameters and matched training-token budgets, unrestricted fitting yields higher mean held-out bits per byte under a common unrestricted decoder in all 12 languages in the paired study and 11 of 12 under independent tuning and evaluation. To study intermediate boundary policies, we introduce boundary licences, which limit the vocabulary entries permitted to cross cuts and admit the same form of certificate. On separate English and Chinese fitting corpora, licensing 10\% of the vocabulary budget recovers 85.2\% and 100.0\% of the achieved token-count reduction from removing all cuts. These results quantify the compression cost of boundaries while separating it from the prediction quality of the resulting token units.
|
| 2345 |
More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
2609.35873
|
cs.AI
|
Ziyang Xu, Haitian Zhong, Hao Zhou, Hao Qin, Chenhan Jin |
Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evalua...Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repeatable score patterns, but these chiefly reveal persistent weaknesses: losses relative to the baseline persist across all three repeats on 100 tasks, while persistent wins occur on only one task and are sensitive to answer extraction. The frozen selector gains 0.00 percentage points, and both populations reach 98.70% oracle coverage at 27 harness executions. Stable complementarity remains unresolved at three repeats. Supporting BIRD traces locate failures in mechanism implementation, activation, and output validity. Together, these findings establish why coverage and repeatability alone cannot justify claims of useful specialization. They motivate an evaluation standard for harness diversity: task advantages should persist across executions, guide usable decisions, and improve on additional fixed-program executions under matched inference budgets.
|
| 2346 |
Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees
2609.35874
|
cs.AI
|
Yaacov Pariente, Vadim Indelman |
Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value functio...Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emph{within} the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners. We instead apply CVaR to the immediate cost over the belief at each step, directly targeting per-step uncertainty about the current state. The standard expected cumulative return is retained as the objective, so the resulting problem has a standard MDP structure: any expectation-based POMDP planner can be made risk-sensitive by changing only the cost computation. We inherit finite-time guarantees for policy evaluation and sparse sampling---with estimation error independent of the risk level---and, as our central theoretical result, prove a finite-time bound on the gap between the particle belief MDP surrogate and the original POMDP, which together yield an end-to-end guarantee from the true POMDP value to the algorithmic estimate. In the risk-neutral limit, the formulation recovers standard expectation-based planning.
|
| 2347 |
Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models
2609.35875
|
cs.AI
|
Leonardo Ferreira, Gardenia Liu, Kaden Zheng |
Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, ...Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control runs, we vary diversity along three axes - personas, sampling temperature, and model identity - pairing every debate configuration with a generation-budget-matched majority-vote control. The hypothesis is rejected on every axis. Debate beats single-agent inference (3--7 points where tasks have headroom) but at matched budget conditions it ties or even loses to self-consistency sampling at 1.6$\times$ the wall-clock and 3.4$\times$ the token cost. Persona prompting reduces accuracy and a dose-response experiment over each model's full combinatorial persona space shows the cost is a persona tax, not a diversity tax: redundant personas hurt most, while maximally-diverse teams recover part of the loss. Furthermore, mixed-model teams lose to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity, and nearly all of debate's benefit comes from the first exchange of answers. We further identify a pervasive measurement hazard in which debate transcripts silently overflow serving context windows, whose correction alone moves our debate-versus-sampling comparison from $-1.8$ points to parity. Our results recast reported MAD gains as an ensemble-sampling effect and provide the budget-matched, contamination-checked baseline bar that future debate mechanisms should be required to clear.
|
| 2348 |
Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training
2609.35890
|
cs.AI
|
Adam Elimadi |
Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or mo...Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size. We investigate this question through reverse-engineering complexity: the causal structure required to recover a model's behavior at a fixed level of faithfulness. To our knowledge, this is the first controlled empirical test of whether representational or attributional simplicity translates into causal simplicity at the circuit level. Starting from the same pretrained GPT-2 Small checkpoint, we apply matched standard and adversarial continual training, requiring both conditions to retain competence on indirect object identification and pass independent robustness verification before comparing mechanisms. We then compare the models along three complementary axes: sparse-autoencoder decomposability, SAE feature engagement in task attribution, and the size of faithful circuits recovered from the raw computational graph. The robust model is more SAE-decomposable and engages fewer SAE features in task attribution. Circuit size is regime-dependent: on competence-matched IOI, standard leads or ties below 85% faithfulness, but robust needs substantially fewer edges at high faithfulness (90%, 95%), a pattern established on the primary pair while representational trends generalize across a seven-point sweep and a second corpus.
|
| 2349 |
Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?
2609.35897
|
cs.AI
|
Haomin Luo (University of Cambridge, Models2 AI) |
The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its under...The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance -- its internal update machinery remains an uninspected black box. We present the first causal mechanistic audit of a self-discovered RL rule, structured directly around the five pillars of the Era of Experience: extended horizon, grounded reward scales, continuing streams, within-lifetime change, and exploration depth. By surgically pinning, freezing, and transplanting recurrent states while holding meta-parameters fixed, we test when learning history acts as an asset or a burden. Three findings organize the audit: (1) Recurrent history actively expands usable reward scales, sustaining a six-decade window versus three under zero-pinning. (2) Decoupling historical content from its maintenance reveals that the penalty of mismatched history stems from perpetual clamping; allowing imported state to evolve naturally attenuates this burden. (3) Under environmental change, controlling replay retention reverses the apparent adaptation advantage over DQN, demonstrating that external data turnover can confound internal plasticity. Validated through capability thresholds and ported to a second rule (OPEN), this work grounds macro-RSI ambitions in micro-level learning dynamics, establishing a foundational audit standard for next-generation, self-evolving RL algorithms.
|
| 2350 |
Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives
2609.35924
|
cs.AI
|
Hua (Edward), Xu, Dongxin Li, Gwen Yidou-Weng, Guy Van den Broeck |
Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one ...Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one unresolved token depends on the other tokens with which it can form a high-reward sequence. Enumerating all such completions makes the whole guidance computation grow exponentially with the number of unresolved positions. We introduce COFFEE, a plug-and-play framework that avoids this enumeration by separating sequence dependence from the objective. At each diffusion step, a target-free carrier absorbs the marginal token distributions predicted by the denoiser to construct a joint model over the unresolved tokens, while a compiled finite-state model records how their combinations affect the sequence-level preference. Pairing their states allows COFFEE to transfer global preferences to unresolved positions and sample a clean reconstruction without retraining the diffusion model. The same framework supports explicit hard constraints and learned soft objectives. We evaluate COFFEE across multiple symbolic, language, and biological benchmarks, where it achieves strong control results with task-dependent quality and diversity trade-offs. By making objectives available to inference rather than only evaluation, COFFEE brings joint conditioning, completion-weighted guidance, and optimization-based constraints into pretrained neural generation, showing the potential of neural-symbolic methods in diffusion guidance.
|
| 2351 |
Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT
2609.35953
|
cs.AI
|
Marx Wang, Ella Zhang, Cameron Tan, Andrea Mock, Songling Ngo |
Young people increasingly turn to General-Purpose Conversational Agents (GPCAs), such as ChatGPT, in moments of distress. We examine young adults' (ages 18-25) experiences using ChatGPT. We first collected 19,930 ChatGPT conversations and survey data from 158 ...Young people increasingly turn to General-Purpose Conversational Agents (GPCAs), such as ChatGPT, in moments of distress. We examine young adults' (ages 18-25) experiences using ChatGPT. We first collected 19,930 ChatGPT conversations and survey data from 158 young adults. We then selected five example conversations reflecting user distress. Finally, we asked ten clinicians to review those five conversations. We found distressed participants reported greater emotional engagement with ChatGPT and greater behavioral change from using it than their peers. When they turned to ChatGPT in moments of acute distress, ChatGPT was quick to give overly dramatic responses and excessive action-oriented suggestions. Clinicians endorsed ChatGPT's availability and much of its wording, but identified seven process failures, such as prematurely jumping to solutions. We translated clinicians' feedback into design guidelines following three stages: 1) asking about safety, 2) de-escalating intensity to restore emotional regulation, and 3) exploring concerns without agreeing with them.
|
| 2352 |
SAGE: A Statistical Acceptance Gate for Self-Evolving Agents
2609.36043
|
cs.AI
|
Yihao Wang, Linhan Xia, Rui Liu, Zhaofeng Zhang, Hongyu Wu |
Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accept...Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer's Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above two limitations, we propose a statistical acceptance gate for self-evolving agents (SAGE). Compared with previous work, SAGE has two contributions. First, SAGE proposes a per-item paired comparison that evaluates the current skill and the edited skill on identical validation items, which exposes regressions that an aggregate score hides and penalizes them asymmetrically. Second, SAGE also employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise. SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting. It commits only a subset of the baseline's edits, filtering out those whose gains are unreliable or purchased by breaking already-solved items. Across five benchmarks and four backbone LLMs under an equal-budget protocol, SAGE lowers the regression rate in 19 of 20 settings and matches the baseline in the remaining one, for example from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attains the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.
|
| 2353 |
PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning
2609.36052
|
cs.AI
|
Zhanhua Pan, Xiao Liu, Zhilong Cao, Jianhong Wang, Dawei Qiu |
Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based ...Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: https://github.com/powerzoojax/PowerZooJax.
|
| 2354 |
GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning
2609.36056
|
cs.AI
|
Shaoxiang Qin, Yucheng Zhao, Fuyuan Lyu, Di Zhou, Jiachen Yao |
In urban low-altitude flight, buildings reshape ambient wind into spatially varying 3D flow, making unmanned aerial vehicle (UAV) energy depend on local wind exposure as well as path length. However, building-resolved wind information is rarely available when ...In urban low-altitude flight, buildings reshape ambient wind into spatially varying 3D flow, making unmanned aerial vehicle (UAV) energy depend on local wind exposure as well as path length. However, building-resolved wind information is rarely available when a mission must be planned. Computational fluid dynamics (CFD) can produce high-fidelity urban flow fields, but each simulation is tied to a fixed inflow boundary condition and can take hours to days, which is incompatible with urban UAV missions that typically last minutes to tens of minutes. We present GeoWind2Plan, a geometry-to-wind-to-planning framework for mission-time 3D urban wind prediction and energy-efficient UAV planning. Given only a background wind vector, 3D building geometry, and a start-goal pair, GeoWind2Plan transforms the building geometry into a reference-wind frame, predicts mission-relevant 3D wind patches with a localized geometry-conditioned neural operator, stitches them into a queryable local wind field, and optimizes a feasible 3D path and speed profile using a physically grounded UAV energy model. Rather than pursuing CFD-perfect reconstruction, GeoWind2Plan targets decision-useful wind prediction: trajectories are planned with predicted wind and evaluated under high-fidelity CFD wind. Across held-out urban domains, wind speeds, and mission wind-angle regimes, GeoWind2Plan performs corridor-localized wind inference in about 3 seconds, compared with roughly 8 hours for CFD. Under CFD evaluation, trajectories planned with GeoWind2Plan reduce energy by 6.9%, 12.7%, and 4.5% in tailwind, headwind, and crosswind missions relative to wind-agnostic planning, recovering 87.9%, 85.7%, and 75.0% of CFD-reference savings. These results show that fast, corridor-localized 3D urban wind prediction can make wind-aware UAV energy planning practical at mission time.
|
| 2355 |
Mirror-Score: Calibrated, Inference-only Scoring Exposes the Limits of Sequence-compatibility Ranking in D-peptide Design
2609.36057
|
cs.AI
|
Jiada Li |
D-peptides combine protease resistance with high target specificity, but computational design of D-peptide binders remains immature. Mirror-Peptidizer introduced an in silico mirror-image screening pipeline using target reflection, backbone generation, and Pro...D-peptides combine protease resistance with high target specificity, but computational design of D-peptide binders remains immature. Mirror-Peptidizer introduced an in silico mirror-image screening pipeline using target reflection, backbone generation, and ProteinMPNN sequence design, but its raw ProteinMPNN negative log-likelihood (NLL) ranking was not validated against measured affinities, and only 4 of 9 tested MDM2 designs bound detectably. We introduce Mirror-Score, a calibrated, inference-only scoring framework for heterochiral D-peptide/L-protein complexes, and a public benchmark of 31 crystal complexes across four target families, including 18 with literature-verified affinities. Raw ProteinMPNN NLL is not a valid affinity ranker: its pooled Spearman correlation with affinity is 0.19, and correlations reverse between MDM2/CHIP (+0.62) and gp41 (-0.70). We therefore evaluate Boltz-2 mirror-space cofolding confidence. For the complete viral-entry family (7 structures representing 3 peptides), interface predicted local distance difference test (pLDDT) achieves structure-level leave-one-out Spearman rho = 0.90 (p = 0.006) and correctly orders all three peptides by affinity, whereas NLL fails (structure-level rho = 0.18). Because only three independent chemotypes are represented, this result indicates directional consistency rather than a statistically validated predictor. Cross-family calibration does not transfer at current sample sizes, supporting family-matched calibration as the practical deployment mode. We also specify a prospective design protocol for the antimicrobial-resistance targets LasR and LecB from Pseudomonas aeruginosa, including mirrored structures, ligand-derived hotspot maps, diffusion-model-ready inputs, and Mirror-Score ranking. Code, benchmark data, structures, and analysis scripts are openly available at https://github.com/Jiadalee/Mirror-Score.
|
| 2356 |
SMat-Attention: Structured Long-Context Sequence Modeling
2609.36062
|
cs.AI
|
Emile Anand, Abdullah Ateyeh, Archer Wang, Marin Solja\v{c}i\'c |
Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size stat...Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes through a tunable notion of structure. To this end, we introduce Structured Matrix Attention (SMat-Attention) via a family of causal masks with structured long-range routing whose row supports have VC-dimension $d$. In our construction, $d=1$ recovers the standard causal mask, and increasing $d$ permits richer subset-routing patterns. We give chunkwise forward and backward algorithms to enable hardware-efficiency. For sequences of length $T$, the hard-routing construction takes $O(T^{2-3/d}+T)$ work, despite the mask being dense, for our prescribed family. In fixed-horizon streaming, decoding after the distant prefix takes constant time per token using $O(T^{1-1/d})$ cached states. SMat-Attention therefore makes VC-dimension an explicit knob governing access-pattern complexity, prefill cost, and decoding memory. Empirically, subset-routing and rule-assisted multi-key retrieval experiments illustrate the masks' routing expressiveness. Extensions to Mamba-2 and Gated DeltaNet using learned routing with top-$k$ query reads retain subquadratic prefill, improve recall accuracy over the backbones in several settings, and achieve comparable small-scale language-modeling performance.
|
| 2357 |
LongCat-DeepResearch Technical Report
2609.36071
|
cs.AI
|
Meituan LongCat Team, He Zhu, Yue Xu, Wanli Wu, Haolin Ren |
We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinat...We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec. Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop. Once the sections are assembled, global review guides targeted local revisions, reducing reliance on repeated full-report rewriting. This workflow also supports the construction of research tasks and trajectories for the mid-training and post-training of LongCat's general-purpose models. LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. On an in-house benchmark, it scores 76.04, ranking second among four compared systems. Development-set analyses show benefits from combining planning perspectives, while further planning refinement has mixed effects. Additional editing improves average automatic readability preference across two benchmarks, with different trends on each.
|
| 2358 |
A Polyphonic Conception of AI Understanding
2609.36079
|
cs.AI
|
Matthieu Queloz, Pierre Beckmann |
When a doctor, a judge, or an engineer must decide whether to trust an AI model's output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs withou...When a doctor, a judge, or an engineer must decide whether to trust an AI model's output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs without reintroducing the question of AI understanding in all but name. Yet the question is ill-framed as it stands, because the inherited concept operates within a monophonic paradigm: the idea that a cognitive system's understanding of something must be localised to a single mechanism underpinning all the capacities conferred by such understanding. Drawing on a wide range of mechanistic evidence, we show that LLMs are pervasively polyphonic: outputs emerge from coalitions of parallel mechanisms of uneven reliability, which variously complement, duplicate, or drown out one another, with several coalitions sufficing for a task without any one being indispensable. Polyphony not only complicates attributions of understanding, but renders monophonic inference patterns hazardous. In response, we develop a conception of understanding fit for polyphonic AI. It centres on sound circuitry that is reliably and correctly recruited and in control of outputs. Attributions of understanding thereby become tractable claims about internal organisation, and can do the work of guiding trust in AI.
|
| 2359 |
GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis
2609.36082
|
cs.AI
|
Ethan D. Frakes, Amy Kvien, Rishabh Kundu, Redad Mehdi, Van D. Tran |
We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textu...We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies. It provides a competency query taxonomy at different difficulty levels from spatiotemporal containment and proximity, spatiotemporal co-occurrence analysis, multimodal evidence, to hypothetical evaluation. Over multimodal KG and query classes, GeoOutageBench provides user-configurable evaluation of three important, highly coherent yet less studied tasks: (1) LLMs' understanding for ambiguous geospatiotemporal questions in terms of NL to SPARQL interpretation, (2) query-driven assessment of ontology utility, and (3) answer accuracy of multimodal KGQA retrieval. GeoOutageBench provides a design principle and foundation for assessing LLM-KG systems that support real-world infrastructure resilience analysis. Our benchmark, source code, data, results, and other documentation are available at https://github.com/UCF-SAGE/GeoOutageBench.
|
| 2360 |
An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures
2609.36104
|
cs.AI
|
Blaz Bertalanic, Carolina Fortuna |
Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer...Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.
|
| 2361 |
The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent Interface
2609.36118
|
cs.AI
|
Yuxiang Liu, Lizhi Yang, Fengze Xie, Aaron Ames, Yisong Yue |
Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for...Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for frozen backbones across three pretrained models and two manipulation benchmarks, LIBERO and CALVIN, with three policy-training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations underperform the best observed single-layer policy. Our stastical analysis further confirms that fusion's advantage is very limited. However, the best layer varies substantially across backbones and benchmarks, making layer selection consequential and exhaustive policy sweeps expensive. We further derive a reweighting equivalence between the proposed information-bottleneck objectives for action-conditioned InfoNCE and action prediction, motivating InfoNCE as a proxy for layer quality. Empirically, InfoNCE provides the most consistent positive association with policy success among four evaluated proxies. Selecting the layer with the highest InfoNCE score requires 9-33 times less GPU compute than exhaustive policy sweeps and reduces mean selection regret from 17.89 percentage points for deepest-layer selection to 3.71 points across six settings. Its mean regret is close to the 3.28-3.50 points achieved by fixed-layer heuristics optimized retrospectively using all six oracle sweeps, without requiring closed-loop evaluations during selection.
|
| 2362 |
AdaST: Adaptive Coupling for Spatial-Temporal Forecasting
2609.36119
|
cs.AI
|
Zhenyu Lei, Chenghao Liu, Yushun Dong, Qi R. Wang, Jundong Li |
Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, r...Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data's inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.
|
| 2363 |
Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents
2609.36130
|
cs.AI
|
Hongjun Liu, Chen Zhao |
Long-running LLM agents compress past interactions into persistent memories that may be reused as premises for later tasks. This creates a distinct derivation problem: whether the memory actually follows from what the interaction history supports. Relevant evi...Long-running LLM agents compress past interactions into persistent memories that may be reused as premises for later tasks. This creates a distinct derivation problem: whether the memory actually follows from what the interaction history supports. Relevant evidence may be scattered across earlier interactions, while compression can introduce relations or event status that the history never established. A valid memory may therefore appear unsupported because its citations omit relevant evidence, while individually supported facts may be composed into a stronger statement the history never established. We characterize this problem through three coupled requirements: (1) Evidence scope; (2) Compositional validity; (3) Admission reliability. We therefore ask whether the interaction history available at write time supports what enters persistent memory. We introduce DerivAudit, a framework for auditing whether a memory is actually supported by the history available when it was written. The audit separates three questions: whether supporting evidence lies beyond writer-provided citations, whether the composed memory introduces unsupported meaning, and how write-time admission decisions affect later memory use. Across two natural memory corpora, audits using broader pre-write history recover support for nearly 60% of memories that appear unsupported from citations alone, while 17-21% remain unsupported after expansion. Yet broader evidence does not by itself make admission reliable: unsupported memories are still frequently admitted across verification models, and evidence expansion alone worsens it on two backbones.
|
| 2364 |
More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with Jev
2609.36154
|
cs.AI
|
Orhan Konak |
General-purpose models promise sensor-based decisions without training a task-specific classifier, which could reduce the dependence of Human Activity Recognition (HAR) on labeled data. Yet it remains unclear whether such models can directly interpret determin...General-purpose models promise sensor-based decisions without training a task-specific classifier, which could reduce the dependence of Human Activity Recognition (HAR) on labeled data. Yet it remains unclear whether such models can directly interpret deterministic descriptions of physical sensor signals well enough to replace or complement trained HAR models. We study this question using Jev, a fixed general-purpose probabilistic decision model, on 1,800 class-balanced accelerometer windows from WISDM, UCI341, and PAMAP2. Jev receives no labeled examples, retrieval context, or HAR-specific parameter updates. We evaluate three deterministic sensor representations and compare 5,400 Jev decisions with a generative baseline and three supervised HAR models. Jev remains far below supervised recognition, with its strongest representation reaching macro-F1 of 0.038, 0.118, and 0.089 across the three datasets, compared with 0.686 to 0.907 for the supervised models. More numerical features do not improve Jev. Instead, they reduce recognition on all three datasets, while augmenting the same numerical evidence with a deterministic semantic rendering partially recovers performance, although the experiment does not isolate semantics from the accompanying serialization and redundancy changes. Jev is fast and inexpensive to query, but its probabilities are not reliably calibrated for recognition. A post-hoc fusion analysis finds a small improvement on WISDM that does not replicate on UCI341 or PAMAP2. These results show that training-free sensor decisions depend not only on the information available in the signal, but also on whether the model can use the representation through which that information is exposed. The sensor-to-model interface should therefore be treated as part of the model evaluation rather than as a neutral preprocessing step.
|
| 2365 |
Principled Thoughts for Latent Recursive LLM Systems
2609.36159
|
cs.AI
|
Fahd Seddik, Fatemeh Fard |
Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not cons...Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30\%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST
|
| 2366 |
FigAct: Turning Scientific Figures into Active Canvases for Explanation
2609.36190
|
cs.AI
|
Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw, Jennifer Healey |
Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how peopl...Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40$\times$. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.
|
| 2367 |
An Empirical Study and Assessment of EU AI Act Compliance Checkers
2609.36228
|
cs.AI
|
Zhen Tao, Alize Kahraman, Shidong Pan, Zhenchang Xing, Chiara Ullstein |
The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements are directly relevant to security and privacy, while also addressing closely related issues such as data gove...The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements are directly relevant to security and privacy, while also addressing closely related issues such as data governance, transparency, accuracy, and robustness. However, stakeholders such as small-to-medium businesses and individual developers often lack the legal expertise required to interpret these obligations and translate them into engineering and governance practices. This disconnect creates challenges for implementing the EU AI Act and may lead to missing safeguards or misdirected development and deployment efforts. To address this, various automated EU AI Act compliance checkers (AIACCs) have emerged, claiming to streamline compliance assessments and provide practical guidance. In this paper, we present the first empirical study and assessment of AIACCs. We characterize 12 mainstream AIACCs across multiple dimensions, evaluate their legal coverage and alignment, and analyze checker-generated compliance reports for structure, determinacy, and actionability. We find that the quality of AIACCs varies significantly and that they currently can only serve as early-stage orientation tools. Specifically, we observe inconsistent interaction modes and user-friendliness, a tendency to overly simplify or omit key obligations, and a failure to provide determinate, actionable guidance. As a result, reliance on the current generation of AIACCs may foster a false sense of compliance. With our study, we provide a critical baseline of the current AIACC landscape. We further offer design principles for the implementation of more reliable compliance-support tools.
|
| 2368 |
MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis
2609.36235
|
cs.AI
|
Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi |
Major depressive disorder (MDD) severely impacts daily activities and quality of life. Detecting MDD involves multimodal data, such as interview recordings and sensor measurements. This is particularly challenging, as these heterogeneous modalities often deman...Major depressive disorder (MDD) severely impacts daily activities and quality of life. Detecting MDD involves multimodal data, such as interview recordings and sensor measurements. This is particularly challenging, as these heterogeneous modalities often demand distinct, customized prediction pipelines. Existing efforts to address this challenge have explored both manually engineered multimodal architectures and agent-assisted pipeline development. Despite their progress, it remains challenging to autonomously revise pipelines based on experimental feedback and carry verified improvements forward into subsequent designs. To this end, we propose Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis (MERID). The framework develops depression pipelines through experience-based recursive self-improvement (RSI). Grounded State Construction (GSC) grounds experience by aligning multimodal records with subject-level depression targets. Coupled Pipeline Exploration (CPE) jointly modifies representations, fusion, and predictors to build successor pipelines for classification and severity estimation. Evidence-Guided Evolution (EGE) guides revisions through feedback and verifies gains under uncertainty in small depression cohorts before inheritance. Extensive experiments on depression benchmarks show that MERID achieves the best results on multiple tasks compared with multimodal and agent-based baselines. Further analysis highlights the value of acoustic and linguistic cues for depression detection. Our code is available at https://github.com/DiscoAILab/MERID
|
| 2369 |
ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning
2609.36238
|
cs.AI
|
Nico Bohlinger, Jan Peters |
A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their represen...A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their representation space in units of time. We therefore introduce ChronoSRL, which gives the critic's embeddings an explicit temporal geometry. The distance between state-action and goal embeddings is trained to match the time that the agent takes to reach the goal (goal-reaching time), while goals that were not reached, and goals from other trajectories, are pushed at least one discount horizon away. Furthermore, reaching a goal quickly once does not mean that reaching it is reliable in general, so the policy should not follow the temporal distance directly. Instead, we build on survival reinforcement learning and predict from our temporal embeddings not only the full distribution of goal-reaching times but also the time spent near the goal. Thereby, the policy is trained to favor actions that reach the goal sooner and more reliably and that keep the agent near it. ChronoSRL learns faster and reaches higher performance than contrastive, action-chunked contrastive, and survival reinforcement learning baselines on seven standard locomotion and navigation benchmarks, even with much smaller networks. To test the limits of self-supervised reinforcement learning, we introduce velocity tracking, goal-position reaching, and box climbing tasks with a quadruped robot in a realistic sim-to-real locomotion setup, and show how the shaping terms that are typical for robotics can be naturally incorporated into our framework. ChronoSRL is the only one of the tested self-supervised reinforcement learning methods that learns to stay at the commanded velocities and goal positions, and climbs the highest boxes.
|
| 2370 |
CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models
2609.36245
|
cs.AI
|
Zhaolong Su, Yujin Han, Feng Wang, Jameson Dong, Hins Hu |
Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted rewar...Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model's training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator's current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.
|
| 2371 |
Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
2609.36254
|
cs.AI
|
Xiangyu Zhou, Saleh Zare Zade, Rafi Ibn Sultan, Alexander Kotov, Dongxiao Zhu |
Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing lit...Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility. Code is available at https://github.com/xzhou98/SARA.
|
| 2372 |
OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models
2609.36264
|
cs.AI
|
Liner Xiang, Wenbo Zhang, Hengrui Cai |
Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior...Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators. Code is available at https://github.com/LinerXiang/OTROPE.
|
| 2373 |
From Surfaces to Volumes: Registered Geometry for Protein Representation Learning
2609.36277
|
cs.AI
|
Siyuan Chen, Cai Zhou, Jinrui Zhang, Zhaokang Liang, Taku Komura |
Existing protein geometry models typically represent molecular surfaces using local geometric features such as sampled points, normals, and curvature. While effective for capturing exposed molecular shape, these representations do not explicitly model the volu...Existing protein geometry models typically represent molecular surfaces using local geometric features such as sampled points, normals, and curvature. While effective for capturing exposed molecular shape, these representations do not explicitly model the volumetric organization beneath the surface or provide a consistent coordinate system for residue-wise volumetric structure. We introduce Protein-TetSphere, a registered residue-wise volumetric representation for proteins. Each protein chain is tetrahedralized to obtain local volumetric regions associated with individual residues, which are then registered to a shared fixed-topology tetrahedral reference and represented in a common Laplacian basis. This registration establishes consistent volumetric coordinates across residues, enabling local three-dimensional deformation to be integrated with surface and chemical information in a multimodal protein representation. We evaluate Protein-TetSphere on ligand-binding pocket classification, protein--protein interface prediction, and de novo protein binder design. Across the three tasks, Protein-TetSphere improves ligand-binding pocket balanced accuracy from $0.795$ to $0.826$, Pinder-Pair/Site AUROC from $0.914/0.852$ to $0.932/0.866$, and binder-design success from $14.95\%$ to $19.90\%$ on the BoltzGen Challenge Set and from $27.62\%$ to $32.19\%$ at the ProtDBench backbone level. These results show that registered volumetric geometry provides complementary spatial information beyond molecular surfaces across protein recognition, interaction, and design.
|
| 2374 |
Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations
2609.36278
|
cs.AI
|
Azza Bouleimen, Nicol\`o Pagan, Anik\'o Hann\'ak |
Generative agent-based models (GABMs) are increasingly used to simulate social media dynamics, including misinformation spread. For such social simulations to be valid proxies of human behavior, LLM agents should replicate established human cognitive biases, a...Generative agent-based models (GABMs) are increasingly used to simulate social media dynamics, including misinformation spread. For such social simulations to be valid proxies of human behavior, LLM agents should replicate established human cognitive biases, among them the Illusory Truth Effect (ITE), where repeated exposure to a claim increases its perceived truth value. We investigate whether and how the ITE manifests across four LLMs (Gemma-3-4b-it, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and GPT-5-nano) in a social media simulation context. We propose a two-phase within-context experimental design that embeds the repetition manipulation inside a realistic news feed interaction. Using this design, we collect 336,000 truth, importance, sentiment, and interest ratings across 100 statements, 10 feed variants, and 3 replications. The key comparison is between ratings assigned to repeated statements, seen throughout a simulation phase, and completely unseen ones, rated within the same experimental context window. We distinguish genuine ITE (truth-specific repetition boost) from mere exposure effects. We run an OLS regression followed by a Linear Mixed-Effect Model to account for differences across models and ratings. Our results reveal four qualitatively distinct patterns: Gemma-3 exhibits a genuine ITE; Qwen2.5 shows a mere exposure effect; GPT-5-nano displays no repetition effect on truth and mild skepticism toward repeated content; Llama-3.1 shows a small truth boost alongside decreases in evaluative dimensions. Crucially, temperature has no effect on these findings, and a variance decomposition highlights the high context-sensitivity of LLM rating behavior. Our findings caution against assuming uniform ITE replication across LLMs in social simulations, while suggesting that Gemma-3-4b-it may offer the most behaviorally realistic approximation for misinformation-related simulations.
|
| 2375 |
CheatBench: Measuring Reward Gaming in AI Agents
2609.36308
|
cs.AI
|
Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang |
Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have access...Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at https://cheatbench.ai
|
| 2376 |
StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents
2609.36319
|
cs.AI
|
Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque |
Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common rem...Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes what a model reads as useless, and bounds the context at little cost. However, it decides from the text of the history alone and sees nothing of how the code is connected. Since a coding agent edits code many times over a single task, and each write can change what code elsewhere means, such maintenance may keep records a write has falsified, drop ones that still hold, and miss code the agent needs next. To overcome these challenges, this paper proposes StateTape, a novel and scalable framework that rewrites a coding agent's context as the repository changes rather than as the context grows. The key idea of StateTape is to model the repository as a symbol-level code graph, whose dependencies and language rules expose which symbols a write can affect. Upon this graph, a tape marks the symbols each write changed, which turns staleness from an inference about text into an observation of the agent's writes. We propose a per-write procedure in which the tape nominates the records a write could have falsified while a small manager model settles what the write log cannot, and further provide a theoretical analysis and TraceBench, a benchmark that labels what an agent is holding against what is actually needed. Empirically, we demonstrate that StateTape can effectively clear falsified records and retrieve what is needed, and thus achieve a higher resolve rate in all experiments spanned by six coding agents and three edit-heavy benchmarks with little computational overhead.
|
| 2377 |
Towards an AI Software Factory for Data Systems
2609.36323
|
cs.AI
|
Anna Pavlenko, Bogdan Crivat, Brandon Haynes, Carlo Curino, Fotis Psallidas |
AI-assisted coding tools deliver significant acceleration of coding, but only limited impact across the end-to-end software development lifecycle (SDLC)--an Amdahl's law effect! In this paper, we discuss our progress towards building an AI SW Factory that acce...AI-assisted coding tools deliver significant acceleration of coding, but only limited impact across the end-to-end software development lifecycle (SDLC)--an Amdahl's law effect! In this paper, we discuss our progress towards building an AI SW Factory that accelerates all the stages of SDLC-Targeting, Coding, Reviewing, and Ops. The AI SW Factory produces a metadata exhaust that enables self-improvement by fine-tuning model weights and updating our World Model (a rich data substrate). We focus on Data Systems and the important class of Evolutionary Coding Tasks (i.e., those with a measurable objective to hill-climb) and report on 1) scaled deployments at Microsoft (tens of repositories) leading to 3x engineering efficiency above agentic coding and up to 22x token efficiency, and 2) several open challenges.
|
| 2378 |
PILLAR: Private Inverted-Index Lexical Lookup for Augmented Retrieval
2609.36326
|
cs.AI
|
Truong Son Nguyen (Arizona State University), Daniel Blackley (George Mason University), Ni Trieu (Arizona State University), Evgenios M. Kornaropoulos (George Mason University) |
Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k documents most similar to their qu...Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k documents most similar to their query from a server-held and publicly known corpus to respond to their query, while the server learns nothing about the query, either its terms or its access pattern. Prior PPRAG constructions rely on dense retrieval alone, translating approximate nearest-neighbor search into many query-dependent rounds of PIR, and pay for it in both latency and retrieval quality. PILLAR instead performs private hybrid retrieval in two stages. A sparse stage issues a small, fixed number of PIR queries against a carefully designed index of precomputed BM25 scores, filtering the corpus down to candidates that share terms with the query without the server ever seeing which terms these are. A dense stage then fetches only those candidates' document embeddings and re-ranks them locally, avoiding the many costly PIR queries that private dense retrieval typically requires. We instantiate PILLAR with two protocols that trade latency against retrieval quality, each built on a different private rendering of lexical search. PILLAR-Bin bins posting lists into a hash table and is a single-round design that achieves lower latency than state-of-the-art private retrieval schemes. PILLAR-Tree turns block-max pruning into an oblivious tree traversal combined with cuckoo hash tables and achieves the highest retrieval quality at lower latency than state-of-the-art schemes.
|
| 2379 |
ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue
2609.36340
|
cs.AI
|
Yuyan Chen |
High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist jud...High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client's side, grounded in a curated knowledge base that cites its sources in every answer.
|
| 2380 |
HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks
2609.36357
|
cs.AI
|
Thai Nguyen, Khang Tran, NhatHai Phan |
Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-bas...Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-based compression framework that leverages diffusion-based LLMs (dLLMs) with Multi-Token Prediction (MTP) to accelerate LLM-based data compression processes. We identify a trade-off in diffusion-based compression, where increasing decoding throughput degrades the compression rate. To mitigate this trade-off, HyperZip employs a hypernetwork to generate data-specific updates from a context representation, adapting the dLLM to the target data without costly fine-tuning, resulting in a low compression rate and high throughput. Extensive experiments show that HyperZip achieves a superior trade-off between compression rate and speed compared with state-of-the-art baselines.
|
| 2381 |
Better Nearest Neighbor Graph Indices via (Efficient) LLM-Guided Pruning
2609.36359
|
cs.AI
|
Fangzhou Wu, Haike Xu, Sandeep Silwal |
Graph-based approximate nearest neighbor search (ANNS) is widely used for large-scale semantic search. Its indices are constructed primarily based on geometric relationships among embeddings of an input dataset (e.g., documents or images), rather than explicit...Graph-based approximate nearest neighbor search (ANNS) is widely used for large-scale semantic search. Its indices are constructed primarily based on geometric relationships among embeddings of an input dataset (e.g., documents or images), rather than explicitly optimizing for semantic relevance. However, when using these indices for downstream query retrieval, performance is evaluated based on the semantic relevance of the retrieved results to the query. This creates a fundamental "geometry-semantic" mismatch between how the indices are constructed and how their retrieval results are evaluated. While existing LLM-based reranking methods can partially mitigate this mismatch at query time, they leave this underlying structural problem in the graph unresolved. We therefore propose LLM-Guided Graph Pruning (LGP), a general framework that addresses this mismatch directly by leveraging LLM reasoning to refine an existing ANN graph index itself. LGP identifies structurally "low-value" neighbors of nodes and replaces them with LLM-selected alternatives that provide useful semantic information while retaining desired geometric structures of the original graph, including sparsity and efficient navigability. Experiments on representative semantic retrieval benchmarks show that LGP consistently improves end-to-end retrieval performance over both vanilla greedy graph search and LLM-based reranking across widely used graph-based ANN indices such as DiskANN and HNSW.
|
| 2382 |
Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents
2609.36365
|
cs.AI
|
Kehang Zhu, Anand Shah, David Parkes |
Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent environments with explicit rules a...Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent environments with explicit rules and known optimal strategies. These settings let us vary how a decision problem is presented while retaining a benchmark for evaluating behavior. Drawing on human-motivated theories of simplicity, we compare interfaces that elicit a complete bid or ranking with sequential interfaces that make safe choices easier to identify. We then hold the interaction format fixed and vary reasoning scaffolds and rule descriptions. Across four model families, the ascending auction interface substantially reduces bid deviations. The matching comparison also shows why sequential responses require different error accounting from complete rankings. Laying out payoff contingencies and explaining why truth-telling is safe also improve choices, whereas prompts to plan through matching rounds or form beliefs about opponents worsen play overall. In auctions, these behavioral gains are not accompanied by corresponding improvements in measured verbal indicators of strategic understanding in the agents' short stated plans. Other prompts change those indicators without improving bids. Our findings suggest that human-motivated theories of simplicity can inform the design of decision environments for artificial agents. They also show why scaffolds should be evaluated through realized choices as well as explanations: improvements in one need not appear in the other.
|
| 2383 |
Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Impact, Detection, and Mitigation
2609.36384
|
cs.AI
|
Roshan Reddy Upendra, Alexandre Dorais, Joe Meyer, Andrew Pouret, Anastasios Lambrianos Stappas |
Relational in-context learning (ICL) conditions predictions on the labeled support examples and their linked tables, creating a failure mode when the support set contains target-derived features that are unavailable for the query. We formulate this problem as ...Relational in-context learning (ICL) conditions predictions on the labeled support examples and their linked tables, creating a failure mode when the support set contains target-derived features that are unavailable for the query. We formulate this problem as support-set target leakage, distinct from leakage during dataset construction, temporal splitting, or representation learning. Here, the target-derived (leaker) columns are present only in the labeled support set during relational in-context inference, while queries remain clean. We construct 14 synthetic leaker types, corresponding to 20 columns, spanning proxies with different noise levels, coverage, modalities, semantic transparency, and relational distances. We evaluate a frozen relational encoder with an ICL head on held-out RelBench databases and use Integrated Gradients (IG) to rank and remove suspicious columns. Our results show that the effect of support-set leakage varies across tasks and relational distances. Target-table leakers cause the clearest degradation, while one- and two-hop leakers are not consistently used by the model. IG ranks target-table leakers highly across datasets and partially recovers performance in settings where leakage has the largest effect.
|
| 2384 |
Persona Dosing: Calibrated Activation Steering for Graded Trait Control
2609.36388
|
cs.AI
|
Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen, Jingyuan Zhang |
An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensit...An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.
|
| 2385 |
ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering
2609.36392
|
cs.AI
|
Yuyan Chen |
In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Ther...In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline conflict registry spanning all active ME/CFS diagnostic frameworks. Second, a conflict-aware retrieval calibration pipeline that re-ranks retrieved evidence using query-specific focus and conflict signals. Third, a benchmark scored by LLM-as-Judge, avoiding systematic underestimation averaging 10.1 percentage points caused by keyword matching. ARCagent achieves 95.3%, outperforming all base LLMs. Code is available at https://github.com/Yukyin/ARCagent.
|
| 2386 |
From Retrieval to Reasoning: Agentic Mechanism Prediction from Cell Painting Profiles
2609.36406
|
cs.AI
|
Jiayuan Chen, Botao Yu, Tianyu Liu, Thai-Hoang Pham, Meng Wu |
Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as repre...Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA prediction as a calibrated evidence reasoning problem, where retrieved neighbors are treated as uncertain observations that must be evaluated, compared, and sometimes rejected before supporting a mechanistic conclusion. We propose PhenoAIR, a reliability-aware multi-agent framework that maintains a candidate-centric evidence memory and performs controller-guided refinement over phenotype- and mechanism-side evidence. PhenoAIR uses offline reference-set calibration to weight evidence by source reliability, phenotype stability, and mechanism-level confusion. We evaluate PhenoAIR on a benchmark constructed from JUMP Cell Painting profiles and annotations, covering controlled, realistic, and discovery-oriented open-world MOA prediction settings. PhenoAIR outperforms representation-matching and LLM-based baselines across all settings.
|
| 2387 |
Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability
2609.36417
|
cs.AI
|
Roshan Reddy Upendra, Alexandre Dorais, Joe Meyer, Andrew Pouret, Anastasios Lambrianos Stappas |
Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query....Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set target leakage. We construct 20 controlled target-derived features that vary in signal fidelity, representation, semantic transparency, coverage, and zero-, one-, and two-hop relational placement, and evaluate them across 13 RelBench tasks and five relational ICL configurations that vary the ICL head, message-passing depth, pretraining cohort, or relational encoder architecture. We evaluate matched 0-hop, 1-hop, and 2-hop leakage settings, together with a Full leakage condition containing all 20 leaker columns. Within the tested configurations, target-table (0-hop) and Full leakage produce the largest aggregate deviations from clean evaluation, while higher-hop effects are often weaker, consistent with differences in effective exposure associated with temporal reachability, sampling, and aggregation fidelity. Leakage effects are strongly task- and model-dependent and can reverse relative conclusions between model variants even when aggregate changes are small. For leaker detection, we compare an Integrated Gradients (IG)-based screening method with mutual information (MI) and leave-one-column-out (LOCO) on a common Baseline subset. Ranking quality is strongest in the high-impact 0-hop and Full leakage conditions, but detector-based removal does not consistently restore the clean evaluation. A four-task rel-salt case study further shows the same evaluation concern with native-schema leakage candidates from the original relational schema. These results identify the support/query information boundary as an important component of reliable relational ICL evaluation.
|
| 2388 |
The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization
2609.36434
|
cs.AI
|
Benoit Dherin, Michael Munn, Xavier Gonzalvo, Adrian Goldwaser, Blaz Bratanic |
Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly th...Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.
|
| 2389 |
Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference
2609.36437
|
cs.AI
|
Taeung Yoon, Yupeng Zhang, Xiaojing Liao |
Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers...Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model parameters. However, despite the growing interest in ZK-LLMs, the understanding of ZK-friendly quantization remains limited. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints. Understanding ZK-friendly quantization is therefore essential for making ZK-LLMs practical. In this work, we present the first systematic study of ZK-friendly quantization for LLMs. We first formalize the definition of ZK-friendly quantization, capturing the properties required for ZK proof generation. We then evaluate nine language models, including Qwen2.5-14B and the mixture-of-experts model Qwen3-30B-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation. Also, we identify RMSNorm inverse-square-root lookups as a recurring bottleneck in several large models and recover near-baseline utility by selectively increasing precision only at the bottleneck. Finally, we show that reducing bit-width or lookup-table size does not necessarily yield proportional end-to-end proving savings, showing that conventional low-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator-aware precision selection.
|
| 2390 |
Rethinking Reasoning Paths as Phase-Structured Trajectories
2609.36461
|
cs.AI
|
Zhenghao He, Guangzhi Xiong, Sanchit Sinha, Bohan Liu, Wenqian Ye |
Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer corr...Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.
|
| 2391 |
Going Beyond State-Reaching: Learning Abstractions for Intrinsically Motivated Option Discovery
2609.36473
|
cs.AI
|
Akhil Bagaria, Anita De Mello Koch, George Konidaris |
Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in na...Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in narrow regions of the state-space, eventually causing an explosion in the number of options that overwhelms the agent, and impedes progress on its primary task of reward maximization. We introduce an algorithm that instead identifies a small, relevant subset of features for each subgoal, yielding options that generalize broadly and accelerate exploration. Our approach learns abstract, transferrable options and achieves rapid exploration in three sparse-reward, image-based domains, including the Atari game MontezumasRevenge.
|
| 2392 |
Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework
2609.36478
|
cs.AI
|
Jose Tupayachi, Xueping Li, Soham Das |
The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating ...The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating a regenerative commons as a constrained Markov game or a constrained multi-agent MDP with a designer-specified depletion budget. We develop a nonstationary Lagrangian framework that constructs a policy sequence from solutions of unconstrained games or cooperative control problems. Extending earlier time-average constructions, we introduce average-epoch solution concepts for reset episodes with discounted rewards and terminal costs. We prove a reward-independent feasibility certificate, cooperative feasibility and approximate optimality against feasible policy mixtures, and an extension to unbiased sampled costs. For self-interested agents, a constrained Nash certificate quantifies the price-dispersion term introduced by deviations that redistribute budget across epochs. Under the stated assumptions on solver accuracy and multiplier updates, these results give constrained policy-sequence guarantees using solutions of unconstrained problems. Experiments with constrained IPPO and MAPPO in a Gordon-Schaefer fishery examine how depletion budgets shape stock retention, harvest rewards, and price adaptation.
|
| 2393 |
Human-AI Collaboration: From Paradoxes to Patterns
2609.36481
|
cs.AI
|
Michael Weiss |
Evidence shows that humans and AI systems perform better together, by collaborating, than alone. This paper examines two key design dimensions of human-AI collaboration (autonomy and initiative) and explores the collaboration patterns that they generate. Docum...Evidence shows that humans and AI systems perform better together, by collaborating, than alone. This paper examines two key design dimensions of human-AI collaboration (autonomy and initiative) and explores the collaboration patterns that they generate. Documenting these patterns starts with identifying the underlying problems and solutions, followed by examining the internal tensions within the problems. The paper uses a paradox perspective to analyze those tensions. It describes a process for surfacing the tensions and mapping the underlying paradoxes. It also illustrates how the pattern descriptions can be derived from mapping these paradoxes. Finally, the paper documents four human-AI collaboration patterns: Instruction, Delegation, Assistance, and Co-creation.
|
| 2394 |
AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes
2609.36503
|
cs.AI
|
Weihan Xu, Kan Jen Cheng, Koichi Saito, Jingyu Shi, Tingle Li |
Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target ...Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.
|
| 2395 |
BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning
2609.36505
|
cs.AI
|
Quan Xiao, Mingda Liu, Gaowen Liu, Katsuki Fujisawa, Tianyi Chen |
Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated...Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning are order-sensitive: adapting the retriever before optimizing the policy yields a larger reward gain than the reverse order. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve it efficiently, we introduce BRIDGE, a memory-efficient first-order bilevel method motivated by a loss-landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving the multi-hop average over the strongest baseline by 9.6 and 3.4 EM points, respectively. It also achieves the best averaged answer accuracy and reasoning quality across medical QA benchmarks.
|
| 2396 |
MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems
2609.36556
|
cs.AI
|
Lei Ma, Dennis Hofmann, Haowen Xu, Joshua DeOliveira, Peter VanNostrand |
Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM sy...Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM systems evolve: tasks may leak into training data and thus be memorized by LLMs, traces and anomaly patterns expire as backbones evolve, and labels must be provided reliably for each refresh. To address these challenges, we present MAADBench (MA: multi-agent; AD: anomaly detection), the first refreshable MAS AD benchmark designed for diverse, evolving LLM backbones underlying the agents. MAADBench combines (1) sampled-and-coupled generative tasks over an approximately 10^37-task space to mitigate task leakage, (2) refreshable trace generation under configurable LLM backbones, and (3) automated provision of cost-free, deterministic step-level labels for fine-grained AD evaluation. Beyond offering the paradigm itself, we run MAADBench with five state-of-the-art LLM backbones and release the MAADBench-Full dataset with 5,200 step-labeled traces. Benchmarking 25 AD methods on the MAADBench dataset reveals substantial limitations in current approaches: they rely heavily on supervision, struggle with subtle MAS-specific anomalies, and lack robustness across LLM backbones. These gaps point to a rich research agenda for MAS-specific anomaly detection, with MAADBench providing a systematic and refreshable testbed for method development and evaluation. We open-source MAADBench-Full at https://huggingface.co/datasets/hww123/MAADBench-full.
|
| 2397 |
Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning
2609.36572
|
cs.AI
|
Zhongan Bi, Kepeng Lin, Xuanang Gao, Yuhan Sun, Lianrun Zhang |
Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims ...Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model's predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.
|
| 2398 |
Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?
2609.36576
|
cs.AI
|
Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson |
Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, a...Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI splits an attack objective into incomplete fragments, embeds them in external content retrieved through the agent's tools, and uses a reconstruction cue to prompt the agent to combine them. It then iteratively refines the fragments and cue with OpenEvolve using graded scoring and natural-language execution feedback from the target agent. Empirically, AdaLCPI achieves higher attack success than strong adaptive baselines, reaching 61.4\% macro-average ASR compared with 32.8\% for Trojan Hippo-style and 30.0\% for AgentVigil. Safety evaluations should therefore test whether agents remain robust when harmful objectives must be reconstructed from incomplete fragments.
|
| 2399 |
SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time
2609.36580
|
cs.AI
|
Yu Cheng, Yongkang Hu, Shuaijie Ma, Zhihang Lin, Weicheng Meng |
LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds...LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.
|
| 2400 |
MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms
2609.36581
|
cs.AI
|
Guohong Liu, Jialei Ye, Shanhui Zhao, Yunxin Liu, Yuanchun Li |
Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what ...Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses memory mechanisms through structured primitives for representation, admission, retention, consolidation, budgeting, and retrieval, while enforcing causal and bounded-memory constraints. Although structured, the derived program space remains large and contains heterogeneous, conditionally dependent design choices whose effects can only be assessed via downstream execution. We therefore propose MemEvo, an LLM-driven auto-research framework that uses pretrained LLM as a semantics-aware proposal model to iteratively generate and refine candidate memory programs based on accumulated experimental feedback. At runtime, a deterministic evaluation pipeline validates and evaluates each candidate, while the underlying vision-language model remains frozen throughout discovery. We finally produce a training-free, bounded-memory mechanism. Extensive experiments on StreamingBench and OVO-Bench demonstrate strong streaming video understanding performance together with substantial context and inference efficiency.
|
| 2401 |
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
2609.36585
|
cs.AI
|
Zehao Jin, Ruixuan Deng, Junran Wang |
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all m...Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/
|
| 2402 |
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
2609.36601
|
cs.AI
|
Miteto Wei, Xiaohan Wang, Zehao Chen, Jiajun Chai, Sichao Liu |
On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation ...On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-probability token. Under maximal coupling, the correction probability is exactly TV(p_t, q_t), so the same trust-region radius controls rollout deviation and upper-bounds intervention and specialized-supervision frequency. We further implement an engine-resident speculative verifier that preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x. Across seven mathematical reasoning benchmarks, SAKI improves the matched teacher-guided baseline in Mean@8 and Pass@8 for both 1.7B and 0.6B students. Placement controls and fixed-prefix analysis further support correction-triggered routing as a conflict-adaptive supervision signal.
|
| 2403 |
Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents
2609.36611
|
cs.AI
|
Jianchang Su, Yiwei Yang, Wei Zhang |
Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the eviden...Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, confidence and search respond to the labels as intended in 49 of 50 comparisons, yet a label change alone alters 4-23% of confident verdicts for Qwen3 models and up to 50% for an existing RL-trained fact-checker. Standard GRPO fine-tuning amplifies this shortcut at 8B in all six settings. To reduce it, we propose trust-swap augmentation (TSA), which trains GRPO on each claim with both its original and its label-swapped evidence under the same gold verdict. At 4B, TSA lowers the verdict flip rate by 7-35% (relative) in four of six settings, keeps accuracy and the intended confidence and search responses, outperforms reward-based alternatives in the main setting, and carries over to an unseen label-removal perturbation. An added consistency reward helps on the trained-on swap but not on unseen perturbations. At 8B, TSA's effect is not detectable, which makes scale the main open question.
|
| 2404 |
Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge
2609.36620
|
cs.AI
|
Zixing Jia, Yuhang Pan, Ni Ji |
Structural reasoning, the ability to recognize and make inferences over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hid...Structural reasoning, the ability to recognize and make inferences over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hidden structure and lack interpretability. We introduce Neural Structural Reasoner (NSR), a brain-inspired network that preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations. NSR draws inspiration from three biological mechanisms: multi-layered architecture for encoding hierarchical knowledge, stable representations of entity and concepts, and path integration for input-driven state inference. At query time, NSR parallelizes computation over candidate relational structures and leverages confidence-weighted scores to perform link prediction. Across standard knowledge-graph benchmarks, NSR achieves competitive accuracy without leading on every dataset, and has lower reported training times than several neural baselines. Because reasoning is implemented through sequences of human-readable neuron activations, NSR affords native interpretability by tracking intermediate inference steps. The model further extracts latent relational hierarchies and compositional rules, demonstrating the brain-inspired architecture as an effective, efficient, and highly interpretable substrate for structural reasoning.
|
| 2405 |
DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors
2609.36624
|
cs.AI
|
Yuanzhuo Hu, Zehan Liu, Xiaoyi Qin, Ming Li |
Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couple...Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fr\'echet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.
|
| 2406 |
Semantic Projection for Continual Self-Evolution of Language Agents
2609.36626
|
cs.AI
|
Ziyu Liu, Jun Chen, Lixu Wang |
Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improvements for new tasks can ove...Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improvements for new tasks can overwrite procedures needed for earlier ones. In continual learning, Orthogonal Gradient Descent (OGD) addresses analogous interference by projecting a new-task gradient onto a subspace that locally preserves prior predictions. Natural-language skill revisions, however, have neither gradients nor a canonical vector space in which such a projection can be performed. We introduce \emph{Semantic-Scope Projected Evolution} (SSPE), which transfers the functional principle of gradient projection from parameter space to behavior space. SSPE treats an unconstrained skill revision as a proposed update, identifies acquired capabilities with which it may interfere, and uses the observed gains and regressions to construct a compatible revision rather than merely rejecting the update. This enables one shared skill to evolve across latent and recurring task contexts without exposing semantic domain identities to the evolution model. Across controlled synthetic streams and heterogeneous real-agent benchmarks, SSPE improves final cross-domain competence and mitigates forgetting relative to strong skill-evolution baselines. The evolved skill also retains the strongest average performance after transfer to a different executor model. These results establish semantic projection as a promising principle for stable and adaptive self evolution of language agents.
|
| 2407 |
Distilling Agentic Systems: A Roadmap across Models, Artifacts, and Harnesses
2609.36630
|
cs.AI
|
Ziluowen Luo, Senzhang Wang, Chaozhuo Li, Jun Yin, Hao Yan |
Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imitates a teacher model. We de...Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imitates a teacher model. We define Agent Distillation as the persistent transfer of task-solving knowledge from a teacher agent to a student agent. Our study organizes the field by where transferred knowledge is retained: within the model, as artifacts, through the execution harness, or across substrates. This perspective separates transfer evidence from its outcome and clarifies how knowledge moves between agent components. We develop an evaluation framework that relates retention to causal contribution and deployed utility. Together, these contributions establish a foundation for the reliable, maintainable, and safe development of increasingly complex agentic systems.
|
| 2408 |
RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation
2609.36652
|
cs.AI
|
Zixuan Yang, Yiqun Chen, Qi Liu, Wei Yang, Erhan Zhang |
Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward metho...Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.
|
| 2409 |
FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation
2609.36670
|
cs.AI
|
Song-Li Wu, Weinan Gan, Zhaocheng Du, Xianquan Wang, Jingyi Wang |
A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantizat...A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propagation. In standard Top-1 assignment, gradients concentrate on a narrow subset of frequently selected codewords, leaving the majority inherently under-trained and causing severe SID collisions. To overcome this limitation natively without relying on complex initialization priors, we propose FineSID, a unified quantization framework that moves beyond Top-1 assignment by enabling fine-grained gradient propagation across the entire codebook. Instead of updating only a single selected codeword, FineSID distributes learning signals to all codewords in a soft, differentiable manner. This design promotes globally balanced codebook optimization while strictly preserving semantic consistency, effectively alleviating SID collisions and stabilizing training in large, high-dimensional codebooks. Extensive experiments on multiple public benchmarks demonstrate that FineSID is robust to initialization configurations and consistently improves both codebook utilization and recommendation accuracy. Our work provides a principled, initialization-agnostic solution for semantic identifier learning, advancing the practicality of generative recommendation.
|
| 2410 |
FairDiff: Mitigating the Self-Reinforcing Matthew Effect in Diffusion Recommender Models
2609.36671
|
cs.AI
|
Song-Li Wu, Xianquan Wang, Zhaocheng Du, Weinan Gan, Jingyi Wang |
While the "Matthew Effect" and filter bubbles are widely recognized outcome-level biases in recommender systems, we reveal that Diffusion Recommender Models (DRMs) uniquely compound this issue through their generative dynamics. Rather than merely inheriting da...While the "Matthew Effect" and filter bubbles are widely recognized outcome-level biases in recommender systems, we reveal that Diffusion Recommender Models (DRMs) uniquely compound this issue through their generative dynamics. Rather than merely inheriting data imbalances, DRMs trigger a self-reinforcing amplification of popularity bias. We identify that this phenomenon is driven by two compounding mechanisms. First, while optimization loss is universally dominated by high-frequency items across recommenders, DRMs suffer from a unique structural prior mismatch during generation. Because the forward terminal distribution of long-tailed data deviates significantly from the standard Gaussian prior, reverse sampling trajectories inherently collapse toward high-density popular items, fundamentally suppressing niche item generation. To dismantle this self-reinforcing loop, we propose FairDiff, a plug-and-play fairness-aware diffusion framework. To overcome the popularity-dominated loss, we introduce Popularity Condition Guidance (PCG). Rather than altering the training objective, PCG acts as an inference-time distributional reweighting mechanism, mathematically reshaping the score-based gradient field to penalize high-popularity regions and guide trajectories toward niche semantics. Furthermore, we design a Semantic Calibration (SC) Module to bridge the prior mismatch, aligning the forward and reverse distributions via one-step optimal transport. Comprehensive evaluations demonstrate that FairDiff achieves state-of-the-art performance while effectively mitigating the self-reinforcing Matthew Effect, highlighting its value as a general framework for DRMs.
|
| 2411 |
MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development
2609.36679
|
cs.AI
|
Xin Yu, Lizhu Zhang, Jiamu Bai, Yanhong Wu, Zellux Wang |
Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings t...Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic tools alone does not ensure that agents learn when to use them or how to act on their findings. We introduce ToolMLBench, a suite of executable tools for data inspection, code verification, and experiment diagnosis, together with an SFT and RL pipeline for learning their use. Diagnostic calls acquire evidence whose value depends on subsequent decisions, so final outcomes provide limited guidance on which calls to reinforce. We address this challenge with SPICE, which measures how privileged context changes the likelihood of a sampled tool action and uses this difference as a turn-level reward alongside the final outcome. We train on 80 synthetic tasks and evaluate on 25 in-domain and 10 out-of-domain tasks. Providing tool interfaces and descriptions alone yields inconsistent gains across unadapted models. With the same diagnostic interface, our training pipeline raises in-domain success from 24.8% to 52.4% for Qwen3-8B and from 35.6% to 69.2% for Qwen3.5-35B-A3B. The latter also improves from 31% to 48% out-of-domain, supporting learned diagnostic tool use on held-out sources and targets.
|
| 2412 |
JudgeProfile: Understanding and Steering Subjectivity in LLM Judges
2609.36705
|
cs.AI
|
Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei |
LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a jud...LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize each judge's prioritization using attribute weights estimated from its own overall choices. These weights differ across judges even when estimated from the same attribute judgments. We then learn new weights from reference labels to adapt their decisions to a target evaluation standard. Reweighting perceived attributes improves average held-out agreement with reference labels from 66.48% to 71.97%, outperforming fine-tuning and rubric prompting. Our findings show that understanding and steering the subjectivity of LLM judges requires attention not only to what they perceive, but also to how they prioritize it.
|
| 2413 |
Can AI Scientists Change Their Minds? Prior-Evidence Conflict in Synthetic Universes
2609.36726
|
cs.AI
|
Kargi Chauhan |
Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms....Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms. We evaluate each reported law twice: by executing it on held-out continuations and transfer settings, and by independently checking whether it recovers the generating mechanism. In the current checkpoint of a pre-specified 60-cell study, 22 trials were graded and one additional run ended in infrastructure failure. Among 20 twin trials, 8 pass predictive verification while 5 recover the generator. The dissociation is bidirectional: six parsable outputs predict successfully while missing the mechanism, whereas three recover the mechanism but fail predictive rollout. Drag exhibits the first pattern (5/5 predictive pass, 1/5 mechanism recovery); Gravity exhibits the second (1/5 predictive pass, 4/5 mechanism recovery). Because matched famous controls, the corrected identifiability sweep, and the Evidence Ladder remain incomplete, we do not claim a confirmatory causal prior-conflict effect. Instead, the completed runs establish a narrower verification result: predictive adequacy and mechanism recovery are distinct scientific claims and require distinct tests.
|
| 2414 |
Can Agents Design Libraries for Agents?
2609.36730
|
cs.AI
|
Gabriel Orlanski, Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala |
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase bench...Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.
|
| 2415 |
BiFE: Search-Efficient Discovery of CPU-Only Branching Policies via LLM-based Bi-Fidelity Evolution
2609.36735
|
cs.AI
|
Ce Zhang, Bin Zhang, Zhiwei Xu, Hao Chen, Xinyue Lu |
In branch-and-bound (B&B) for mixed-integer linear programming (MILP), branching variable selection critically impacts efficiency. Existing neural branching policies often require GPU inference, while CPU-efficient symbolic expressions lack the representat...In branch-and-bound (B&B) for mixed-integer linear programming (MILP), branching variable selection critically impacts efficiency. Existing neural branching policies often require GPU inference, while CPU-efficient symbolic expressions lack the representational capacity for complex logic. Large Language Model (LLM)-generated code provides a flexible search space for designing lightweight branching rules with diverse algorithmic logic. To discover effective rules within LLM-based evolutionary frameworks, a core challenge arises: full B&B evaluation on real instances is prohibitively expensive, whereas offline imitation learning suffers from distribution shift. To address this, we introduce a Bi-Fidelity Evolutionary framework (BiFE). It employs low-fidelity imitation scores as a rapid pre-screener and selectively applies high-fidelity on-instance evaluation only to elite candidates, effectively balancing search efficiency with performance reliability. Experiments validate both the search efficiency of BiFE and the competitiveness of its discovered rules, which outperform the SCIP solver and other baselines on CPUs, and even surpass certain GPU-based neural policies.
|
| 2416 |
Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion
2609.36741
|
cs.AI
|
Yibo Guo, Xiaodan Wang |
Under irreversible resource depletion, an agent can spend resources to distinguish among latent fault models, or to change the system state so that the remaining models admit a common acceptable continuation--at which point further diagnosis becomes unnecessar...Under irreversible resource depletion, an agent can spend resources to distinguish among latent fault models, or to change the system state so that the remaining models admit a common acceptable continuation--at which point further diagnosis becomes unnecessary. This distinguish-or-homogenize principle identifies a path that existing frameworks for identification, planning, and diagnosis do not make explicit: prior formulations treat the mapping from fault models to acceptable policies as a given, whereas LCPI makes it a function of the agent's own actions. We formalize this principle through Last-Chance Policy Identification (LCPI), where correctness is evaluated at the state the agent reaches rather than at the initial state. The Last Identifiable Margin (LIM) marks the feasibility boundary between distinguishing and homogenizing. For deterministic diagnostic graphs we provide the Exact-LIM recursion; for noisy finite-horizon recovery we propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint. Across incident recovery on abstract microservice topologies and latent-damage navigation in MiniGrid, RBCP improves risk-feasible recovery while satisfying the failure budget. A sham control--cost-matched actions that preserve model incompatibility--eliminates the gain entirely, confirming that the benefit comes from changing which policies are acceptable for which models, not from extra search or additional budget.
|
| 2417 |
SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
2609.36742
|
cs.AI
|
Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo |
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-poli...Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.
|
| 2418 |
EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents
2609.36746
|
cs.AI
|
Zhen Xiong, Qiaoyu Tan |
Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explicitly modeling downstream exe...Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explicitly modeling downstream executor behavior. We show that this can cause systematic cross-executor degradation: curators trained with different executors perform best when paired with their own training executor, indicating that effective skill curation is executor-dependent. We formulate behavior-adaptive skill curation and introduce EASE, a framework that learns a single curator that adapts its decisions to different executor behaviors. EASE maintains an online behavioral profile of recent execution patterns and conditions the curator on this profile, the current trajectory, and retrieved skills to add, modify, or remove skills from an evolving repository. We train the shared curator jointly across multiple frozen executors with reinforcement learning, using retrieval-aware and behavior-aware temporal attribution to focus optimization on curation actions with observable downstream influence. Across ALFWorld, ScienceWorld, and WebShop, with executors ranging from Qwen3-8B/32B and GPT-OSS-120B to unseen Kimi K2.6, DeepSeek V4 Flash, and Gemini 3.5 Flash, EASE outperforms strong skill- and memory-based baselines without per-executor finetuning. EASE also maintains 34.5--41.0% fewer skills, improves skill retrieval by 36.3--38.7% and measured edit utility by 51.8--60.0%, and reduces deployment-time inference tokens by 9.1--14.5%. These results establish behavior-adaptive skill curation as an effective principle for building self-evolving agents.
|
| 2419 |
Generalizable Lifelong Model Editing via Preference Optimization
2609.36748
|
cs.AI
|
Dahyun Jung, Suhyune Son, Heuiseok Lim |
Knowledge editing enables rapid updates of specific factual knowledge in large language models (LLMs) without full retraining. However, more realistic scenarios call for a lifelong framework that handles continual updates rather than one-off modifications. In ...Knowledge editing enables rapid updates of specific factual knowledge in large language models (LLMs) without full retraining. However, more realistic scenarios call for a lifelong framework that handles continual updates rather than one-off modifications. In such settings, existing editing methods often overfit to target prompts, significantly degrading both the generalization of the edited knowledge and the model's general capabilities. To address this issue, we propose GLIME (Generalizable Lifelong Model Editing), which combines knowledge editing with preference optimization over generation behavior. GLIME further incorporates replay-based editing and a gradient constraint to preserve previously edited knowledge. Experimental results show that GLIME significantly improves knowledge generalization in lifelong editing settings while maintaining both editing performance and general capabilities.
|
| 2420 |
Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
2609.36777
|
cs.AI
|
Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen, Edward Zhang, Zhaoxu Zheng |
Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing r...Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface. Construction tests scene-level spatial reasoning from open-ended language specifications, where many realizations are valid; editing tests precise control of scene state, where the agent must recover the target scene from reference images while preserving everything else. Rather than scoring code or rendered views, Code4Scene evaluates the generated engine-native scene for task fulfillment, artifact integrity, and static physical validity, with edits additionally compared against withheld ground truth. Across 14 coding-agent configurations on the 95-case public set, construction and editing performance are strongly correlated but not interchangeable (Spearman $\rho = 0.78$): Claude Fable 5.1 leads construction, Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent, while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.
|
| 2421 |
Aperture: Merge-Consistent Rotary States for Compressed Tokens
2609.36781
|
cs.AI
|
Yuhao Du, Shunian Chen |
Token compression combines content from several positions, yet rotary position embeddings usually assign the merged token one coordinate. We ask what positional information must survive later merges. Aperture stores Fourier moments of the token's weighted supp...Token compression combines content from several positions, yet rotary position embeddings usually assign the merged token one coordinate. We ask what positional information must survive later merges. Aperture stores Fourier moments of the token's weighted support at the model's rotary frequencies. We prove that these moments have minimal real dimension among continuous states sufficient for the selected expected rotary interactions. Represented mass makes updates additive; attention normalisation remains a separate readout choice. Uniform intervals give a centre rotation times a sinc gain. We characterise when centres determine interval widths and construct matched examples where they do not. Numerical checks verify the weighted-support implementation. In trained temporal readers, compression transfer varies with gain calibration and feature placement. In a prespecified native video question-answering comparison, stored support reaches $65.63\%$ accuracy versus $67.12\%$ for the deployed merging rule. These results separate exact positional preservation under compression from downstream benefit.
|
| 2422 |
AI as a Compiler: Compiling Triton kernels without the Triton compiler
2609.36800
|
cs.AI
|
Fran\c{c}ois Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini |
Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We ...Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.
|
| 2423 |
UpliftMem: Learning Set-Level Uplift for Agent Memory Retrieval
2609.36805
|
cs.AI
|
Mengkun Liang, Haoran Qiang, Guannan Liu, Junjie Wu |
Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while ev...Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts. We introduce \textsc{UpliftMem}, which learns memory retrieval from set-level execution uplift relative to the same executor without memory. A theoretical analysis of how retrieval preferences restrict feedback coverage motivates targeted probing of alternative memory sets. Probe selection follows an expected value of sample information (EVSI) criterion, derived in closed form under a correlated Gaussian model, to allocate limited training rollouts according to their expected improvement in local retrieval decisions. The shared scorer is trained with a frozen executor and selects memory sets without test-time probes. Across ALFWorld, WebShop, and BigCodeBench, \textsc{UpliftMem} achieves the best success rates among evaluated baselines on the main evaluation sets. Controlled fixed-store and matched probe budget evaluations further demonstrate improved memory-use decisions and more effective use of execution feedback.
|
| 2424 |
CAD-Native Transformer Operators for AI-Aided Engineering
2609.36806
|
cs.AI
|
Daniel Leibovici, Nikola Borislavov Kovachki, Dawon Ahn, Ruben Ohana, Ira J. S. Shokar |
Modern engineering systems, from automobiles to aircraft, are designed by using precise, continuous parametric computer-aided design (CAD) models. Evaluating design changes through numerical simulation requires meshing the continuous geometry, a computationall...Modern engineering systems, from automobiles to aircraft, are designed by using precise, continuous parametric computer-aided design (CAD) models. Evaluating design changes through numerical simulation requires meshing the continuous geometry, a computationally expensive and often brittle process that can require manual intervention and replaces the continuous representation with a discrete approximation. Most neural surrogates accelerate the simulation, but inherit this representation gap by relying on meshes, point clouds, voxels, or other sampled approximations of geometry. We introduce CANTO, a transformer neural operator that maps directly from continuous CAD geometry to physical fields, without meshing the input geometry. We develop a theoretical framework for learning operators from geometric manifolds to function spaces of physical fields, representing geometry through sequences of parametric patches. CANTO instantiates this framework by directly tokenizing non-uniform rational B-spline (NURBS) patches from their control points, knot vectors, and weights, and predicts continuous surface and volume fields at arbitrary query locations. We evaluate CANTO on four automotive and aircraft aerodynamics industry benchmarks: AhmedML, WindsorML, DrivAerML, and HiLiftAeroML. CANTO achieves state-of-the-art accuracy on most evaluated surface and volume prediction tasks, including a 19.8% reduction in surface-pressure relative $L_2$ error compared with AB-UPT on HiLiftAeroML. Differentiability with respect to CAD parameters further enables gradient-based inverse design of designs. On AhmedML, CANTO identifies designs with 4.4 to 20.4% lower drag than the best dataset designs satisfying the same volume and lift constraints, with the improvements verified using the same CFD setup used to generate the original dataset.
|
| 2425 |
Geometry-Conditioned Fixed-Scaffold Encoders for Time-Warp Robust Sequence Retrieval
2609.36809
|
cs.AI
|
Cassandra Yang, Yufan Tang |
Embedding-based retrieval is attractive for long sequence collections because each item can be encoded once and searched by nearest-neighbor ranking. The difficulty is that the objects being indexed are often observed under a noncanonical clock: cardiac cycles...Embedding-based retrieval is attractive for long sequence collections because each item can be encoded once and searched by nearest-neighbor ranking. The difficulty is that the objects being indexed are often observed under a noncanonical clock: cardiac cycles stretch with rate, speech changes with tempo, and sensor traces reach comparable states at different speeds. This paper studies a specific source of instability in patch-based encoders for this regime. If patch boundaries are chosen from signal geometry, then the tokenization can change under the same temporal deformation that the representation is expected to tolerate. We propose GeoPatch, a fixed-scaffold patch encoder that keeps token support independent of geometry and uses slope, curvature, acceleration, affine-residual, and confidence descriptors only as continuous conditioning variables. The design turns boundary variation into feature modulation: geometry can change the embedding through a controlled pathway, but it cannot change the number, order, or support of local tokens. We formalize this distinction through a mechanism-level stability analysis that separates boundary drift, affine timing variation, confidence-weighted geometry perturbation, and retrieval-margin effects. The same local tokens support global embedding retrieval and late-interaction scoring, so the scoring rule can be matched to the evaluation protocol. Across ECG, speech, and multivariate time-series retrieval tasks, GeoPatch improves early-rank retrieval under timing variation while exposing a clear trade-off between local surface matching and strict non-overlap retrieval.
|
| 2426 |
Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models
2609.36828
|
cs.AI
|
Wenxiao Fan, Jingling Fu, Lichen Ma, Yu He, Luohang Liu |
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects th...Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.
|
| 2427 |
The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents
2609.36829
|
cs.AI
|
Xueqi Li, Jingjie Ning, Yibo Kong |
An executor can respond strongly to a change in a supplied plan's priority while showing a small change in the same information-selection probability when a default-aligned whole plan is removed. We call the risk of interpreting the latter as weak responsivene...An executor can respond strongly to a change in a supplied plan's priority while showing a small change in the same information-selection probability when a default-aligned whole plan is removed. We call the risk of interpreting the latter as weak responsiveness to alternative priorities the default trap. We compare paired plans that prioritize different information targets with a shared no-plan reference. An accounting identity relates these distinct behavioral contrasts. Across 3,200 decision windows on 160 selected Retail, Airline, and AgentDojo tasks, switching priorities strongly redirects two models' choices, while the two plan-versus-default contrasts differ. In 2,160 additional windows, reversing account-list order shifts default target selection by 63.3-98.3 percentage points; priority-switching effects remain 96.7-100.0 points in either order. A separate 3,240-window component study finds strong control under single priority sentences, with effects of additional text varying by group and direction. Finally, 1,080 full-task episodes yield observed success differences of -19.4 to +8.3 points relative to no plan. All Retail and Airline success intervals include zero; AgentDojo results describe four fixed application worlds. These findings support joint reporting of priority responsiveness, presentation-dependent defaults, and task success and cost.
|
| 2428 |
Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training
2609.36830
|
cs.AI
|
Chenliang Li, Neiwen Ling, Zijun Wei, Alfredo Garcia |
Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the ...Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess pool occupancy into an adaptive rejection budget and ranks completed trajectories using an effective-staleness score that combines Waiting Staleness with prefix-aware Generation Staleness. This avoids penalizing long or interrupted rollouts solely because they span multiple policy versions. In single-turn mathematical reasoning, PACE improves the six-benchmark average validation accuracy by 18.7\% over unfiltered asynchronous RL at the same wall-clock budget and matches synchronous RL performance with 47.1\% less GPU time. PACE also improves validation performance in multi-turn tool-integrated reasoning, outperforming both synchronous and unfiltered asynchronous RL. Further experiments with the mixture-of-experts model and an alternative RL algorithm support its applicability across model architectures and training algorithms.
|
| 2429 |
ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction
2609.36835
|
cs.AI
|
Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang, Yanjia Li |
Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods...Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we first train a value-aware indexer to select real-key anchors in a single scoring pass. ARC-KV then applies convex-hull-constrained key merging and fits an attention-mass bias and compact values against the full cache. At inference time, ARC-KV builds the compact cache once per context using the frozen indexer and reuses it for all subsequent queries. Extensive experiments demonstrate that ARC-KV outperforms reported compaction methods in most settings across QuALITY, RULER, and LongBench on Llama-3.1-8B-Instruct. In particular, at 10% KV retention on QuALITY, ARC-KV improves accuracy from 0.6409 to 0.6474 over Attention Matching while reducing compaction time by a factor of 25.73, from 959.8 s to 37.3 s.
|
| 2430 |
Automated Screw Planning for Reduced Pelvic Fractures Based on Statistical Shape Models and Deep Learning
2609.36847
|
cs.AI
|
Yang Gao, Sutuke Yibulayimu, Yanzhen Liu, Zian Zhao, Yudi Sang |
Percutaneous iliosacral screw fixation is an important minimally invasive treatment for unstable pelvic fractures. Because the sacroiliac region has complex anatomy and narrow screw corridors, the accuracy and safety of screw placement directly affect surgical...Percutaneous iliosacral screw fixation is an important minimally invasive treatment for unstable pelvic fractures. Because the sacroiliac region has complex anatomy and narrow screw corridors, the accuracy and safety of screw placement directly affect surgical outcomes. Accurate and reliable preoperative screw planning is therefore essential to improve surgical success and reduce intraoperative risks. Conventional preoperative planning typically requires surgeons to determine screw trajectories through manual measurements, a labor-intensive process that depends on subjective clinical experience. To address these challenges, we propose a fully automated pipeline for preoperative iliosacral screw planning in patients with pelvic fractures. Using patient-specific three-dimensional anatomy, the pipeline automatically identifies safe screw corridors and generates individualized insertion trajectories to support clinical preoperative planning. We evaluated the proposed pipeline on 200 clinical cases of pelvic fractures. Compared with conventional manual measurements, the safety margin of the safe insertion corridors increased by 2% across the four screw types, the mean planning time decreased by more than 90%, and the clinical acceptance rate reached 95%.
|
| 2431 |
When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration
2609.36855
|
cs.AI
|
Yaxin Gong, Gangyi Zhang, Chongming Gao, Leyang Shen, Chenxiao Fan |
Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its ...Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent's original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream's specific wrong answer--a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.
|
| 2432 |
IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence
2609.36860
|
cs.AI
|
Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao |
We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache repl...We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.
|
| 2433 |
State Trace Rationale As Auxiliary Task in Reinforcement Learning
2609.36867
|
cs.AI
|
Muhammad U. Nasir, Alex Vogt, Steven D. James, Julian Togelius |
We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agen...We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.
|
| 2434 |
WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
2609.36887
|
cs.AI
|
Bo Mao, Hang He, Linting Wang, Lizhi Lin, Maosen Zhou |
Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scalin...Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable agentic interaction system construction, execution-driven self-evolution, and stable post-training. WEFT scales agentic interaction system construction across environment breadth, task complexity, and interaction diversity. Execution-driven self-evolution iteratively uses execution traces and state evidence to attribute failures and revise the responsible components, with fresh rollouts evaluating the changes and providing evidence for subsequent evolution rounds. For stable post-training at scale, WEFT addresses both optimization and execution reliability: prefix-preserving sampling retains verified progress and atomic-turn credit assignment localizes learning signals, while MegaMCP maintains isolated, recoverable state across concurrent rollouts over shared tool services. Extensive experiments across various models and benchmarks demonstrate the effectiveness of WEFT for tool-use post-training. WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, $\tau^2$-Bench, and Claw-Eval. In particular, WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points. WEFT-35B-A3B further extends these gains to more challenging long-horizon workflow benchmarks, including Toolathlon-Verified and AutomationBench.
|
| 2435 |
Beyond Sub-Gaussian Detector Scores: Robust Weighted Profile-Loss Change Point Detection for Human-LLM Text Segmentation
2609.36888
|
cs.AI
|
Wan Tian, Zhongyi Li, Yawen Li, Rui Zhang, Yijie Peng |
Mixed human-LLM documents require locating authorship transitions from detector scores whose reliability varies across text units. Existing weighted mean contrasts are vulnerable to extreme scores, while directly replacing means with robust centers obscures ho...Mixed human-LLM documents require locating authorship transitions from detector scores whose reliability varies across text units. Existing weighted mean contrasts are vulnerable to extreme scores, while directly replacing means with robust centers obscures how a misplaced boundary changes the population objective. We propose Robust Weighted Profile-Loss Change Point Detection (RWCP), which combines capped reliability weights, Huber profile gains, and narrowest-over-threshold search in reliability coordinates. Our key analysis expresses the population gap between a true and a displaced split as a merge cost, avoiding a closed-form solution for the nonlinear center of a mixed segment. Under explicit curvature, spacing, and dependence conditions, core RWCP recovers the number of changes and localizes their boundaries; its quadratic-loss limit recovers squared weighted CUSUM. We also study RWCP-R, a separately evaluated decoder that shares source centers across nonadjacent passages. Across five retrospective cached-score benchmark families, core RWCP reduces family-macro WindowDiff by 17.6\% relative to weighted change-point detection, and RWCP-R lowers it further. Boundary recovery improves most clearly for isolated changes, while both fixed configurations miss changes in collaborative and densely alternating text.
|
| 2436 |
Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents
2609.36892
|
cs.AI
|
Zeyu Gan, Zixuan Gong, Yong Liu |
As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve in...As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently understood. To narrow this gap, we investigate three central questions concerning harness architecture, harness scale, and self-evolution algorithms through complementary empirical and theoretical analyses. Empirically, we introduce a preference-oriented benchmark and systematically characterize the capabilities and limitations of personal agents associated with these three dimensions. Theoretically, we formulate harness evolution as a learning problem and explain these phenomena through approximation, generalization, and optimization errors. Analyses of reachable policies, capacity under finite interaction evidence, and biased update dynamics provide theoretical accounts of the observed phenomena. Together, these results offer a unified perspective on the limits of personalization through harness evolution and inform future harness design.
|
| 2437 |
HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL
2609.36896
|
cs.AI
|
JunHyeok Oh, Zian Jang, Byung-Jun Lee |
Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the approp...Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is too short can force infeasible transitions, whereas one that is too long can introduce redundant motion. We introduce HorizonFlow, a hierarchical planner that treats plan length as an output of generation rather than a prescribed input. Its subgoal route planner guides its action-prefix controller through a sequence of latent subgoals. Both components combine insertion-based generation with flow matching to jointly generate continuous plan content and length, using the partially generated plan to guide token insertion. HorizonFlow reuses the resulting length information to select candidates and steer generation toward shorter plans without a separate learned value model. Across Maze2D, Multi2D, and OGBench navigation and visual manipulation benchmarks, HorizonFlow achieves the highest average performance among the compared methods.
|
| 2438 |
STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking
2609.36900
|
cs.AI
|
Wan Tian, Zhongyi Li, Xiang Xu, Minhao Zou, Yijie Peng |
Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy op...Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.
|
| 2439 |
PrecogUI: Proactive GUI Agents via Pre-cognitive Simulation and Experience Retrieval
2609.36923
|
cs.AI
|
Bin Kang, Jiarui Ouyang, Li Jiang, Bin Chen, Zhuotao Tian |
Existing reactive Graphical User Interface (GUI) agents often fail in long-horizon, dynamic scenarios, where unexpected disturbances trigger attention-diverting and cascading failures. To address this, we propose PrecogUI, a pre-cognitive architecture that shi...Existing reactive Graphical User Interface (GUI) agents often fail in long-horizon, dynamic scenarios, where unexpected disturbances trigger attention-diverting and cascading failures. To address this, we propose PrecogUI, a pre-cognitive architecture that shifts the paradigm from reactive execution to proactive decision-making. Specifically, we design a Proactive Experience Pool (PEP), which caches recurring anomaly and success patterns as "state-action-result" tuples in a dual-memory repository. Furthermore, we introduce a Proactive Simulation Executor (PSE) that learns to forecast the next symbolic UI layout given a candidate action, enabling early anomaly avoidance and ranking candidate actions by predicted reliability. Finally, a Pre-cognitive Execution Controller (PEC) fuses these priors and predictions, prioritizes handling of foreseen anomalies, and ensures execution robustness through a closed-loop error correction mechanism. For robust evaluation, we develop AutoTraj, an automatic data-generation engine, to construct InterfereBench, a benchmark for long-horizon tasks with strong disturbances. Experiments demonstrate that PrecogUI surpasses state-of-the-art methods on InterfereBench while maintaining competitive performance on public benchmarks. The code will be publicly available.
|
| 2440 |
Neuro-Symbolic Computer Use: Learning Reusable Policies for Reliable and Efficient Execution
2609.36927
|
cs.AI
|
Hyewon Suh, Thanh Minh Nguyen, Chih-Lun Lee, Darrow Hartman, Lizhao Liu |
Many computer tasks recur: the same workflow runs many times, with new inputs and from different starting states. Current computer-use agents re-plan every step of every run, which makes them costly and unreliable on such tasks. We introduce neuro-symbolic com...Many computer tasks recur: the same workflow runs many times, with new inputs and from different starting states. Current computer-use agents re-plan every step of every run, which makes them costly and unreliable on such tasks. We introduce neuro-symbolic computer use, in which a recurring workflow is executed by a learned policy rather than re-derived by an agent on each run. The policy fixes the decisions that are stable across runs (ordering, variables, loops, and branches) in executable code, and delegates observation-dependent decisions, such as grounding and state checks, to neural models. We learn these policies with neuro-symbolic policy iteration: starting from one agent trajectory, it executes the policy, diagnoses failures with task-completion and step-level judges, and revises the code with a coding model informed by an agent's continuation from the point of failure, without access to the benchmark evaluator. Iterating on generated parameter and initial-state variants makes the policy reusable, and a pre-action verifier guards each state-mutating step at deployment. On OSWorld-Verified and ScienceBoard, the learned policies achieve the highest Pass^3 of all methods in all four settings, 3.6-15.8 points above the base agent, while cutting per-run cost by 15-217$\times$ and latency by 3.4-5.1$\times$. On OSWorld-Verified, policies built only on variants transfer to the held-out original tasks, exceeding AutoRPA by 8.6-17.5 points in Pass^3.
|
| 2441 |
Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO
2609.36932
|
cs.AI
|
Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen |
Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollo...Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07$\times$ training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64\% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://github.com/Nicozwy/FastRL.
|
| 2442 |
VLALight: A Vision-Language-Action Model for Traffic Signal Control
2609.36934
|
cs.AI
|
Pan Zhang, Siqi Lai, Kemu Dong, Hao Liu |
Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically r...Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight directly maps visual observations to coordinated signal actions through multi-target spatiotemporal traffic reasoning and topology-aware cooperative perception across intersections. To establish this capability, we develop a two-stage supervised cold-start training strategy for visual traffic understanding and signal decision-making, followed by cooperative agentic reinforcement learning that jointly optimizes local control and network-wide traffic efficiency. Furthermore, VLALight introduces adaptive fast and slow reasoning modes, enabling the policy to allocate deeper reasoning only when additional deliberation provides sufficient control benefits. Through balanced mode-aware rollouts and relative advantage optimization, VLALight learns to trade off decision quality and inference cost. Extensive experiments on seven real-world traffic-flow datasets across three urban networks demonstrate that VLALight consistently outperforms transportation-based, RL-based, and LLM/VLM-based baselines. Ablation studies validate the effectiveness of cooperative perception, network-level optimization, and adaptive reasoning. These results demonstrate the potential of VLA models for real-world physical traffic control. Our project is available at https://github.com/usail-hkust/VLALight.git.
|
| 2443 |
CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory
2609.36935
|
cs.AI
|
Jingguang Li, Yebo Wu, Zuyi Guo, Kailang Ma, Xianjie Dai |
Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded text...Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves potentially useful source excerpts verbatim in a pending set, allowing subsequent context to clarify their relevance before irreversible compression. As new context arrives, a learned policy revisits each pending excerpt and decides whether to promote it to the committed memory, retain it for further consideration, or discard it. A frozen verifier ensures proposed facts are accepted only if supported by retained excerpts and current context. To further guide effective memory management, we train this policy using reinforcement learning by combining fine-grained, step-level evidence rewards with final answer rewards. Extensive experiments demonstrate that CoEM consistently improves long-context reasoning. When evaluated on 6,400 documents long-context input, CoEM outperforms the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B. Code repository: https://github.com/benmagnifico/CoEM.
|
| 2444 |
SCA: Spatial Credit Assignment for Reinforcement Learning of GUI Agents
2609.36939
|
cs.AI
|
Shengtian Yang, Ziyu Xiong, Kaibing Yang, Guangfeng Cai, Yewen Li |
GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. ...GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. However, binary evaluation treats spatially different failed clicks as identical and provides no relative signal when all sampled clicks fail. To address these limitations, we propose Spatial Credit Assignment (SCA), which uses the screen coordinates of sampled clicks to refine group-relative credit. Specifically, SCA predicts each held-out response's reward from the other responses in groups containing both successes and failures, then uses the prediction residual to adjust credit. When all sampled clicks fail, SCA instead orders them by distance to the annotated target. These spatial references are used only to construct the training update; the deployed policy remains unchanged. We evaluate whether this correction improves the policy update itself by comparing its error and directional alignment with the exact return gradient in a controlled synthetic study. Across GUI grounding and offline action-prediction benchmarks, SCA improves grounding across professional domains and achieves the strongest results among reinforcement-fine-tuned models on most action-prediction metrics, with consistent gains across the reported GUI suites.
|
| 2445 |
Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation
2609.36944
|
cs.AI
|
Zhongyi Li, Wan Tian, Xiang Xu, Yutian Xiao, Yikun Ban |
Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while...Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.
|
| 2446 |
REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing
2609.36984
|
cs.AI
|
Jiawen Tao, Xiaokun Yuan, Yaoming Li, Chenxu Liu, Mengzhou Wu |
Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chain...Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessity Rate (BNR), which measures how often targeted evidence removal prevents answer recovery on initially correct instances. Across five existing benchmarks, panel-mean BNR ranges from 16.6% to 48.9%, exposing a substantial gap between annotated structure and observed dependence. Guided by this diagnosis, we introduce REALHOP, a diagnose-construct-verify framework that rebinds entities, factorizes selected relations, adds complete competing paths, and places evidence at traceable locations. Structural and semantic checks precede freezing; behavioral interventions follow. On 790 paired MuSiQue questions, REALHOP raises panel-mean BNR from 27.4% to 94.4% while retaining high Full accuracy. It also yields high BNR on REALHOP-FRAMES and REALHOP-LONGBENCH. On 216 long-context questions, the matched multiple-choice spread across 16 models grows from 13.9 to 59.2 points and persists under repeated open-ended evaluation. Together, these results show that a conceptually coherent chain and a correct final answer do not by themselves establish multi-hop reasoning. Verifying that success depends on every intended hop is therefore as fundamental to multi-hop evaluation as measuring answer accuracy itself.
|
| 2447 |
CF-LoRA: Decoupled Factor Aggregation and Adaptation-Aware Client Clustering for Federated LoRA Fine-Tuning
2609.36986
|
cs.AI
|
Mengjun Yi, Langxing Yang, Suhan Guo, Furao Shen, Jian Zhao |
Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a structural aggregation mismatch caused by independently averagi...Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a structural aggregation mismatch caused by independently averaging LoRA factors, and a statistical collaboration mismatch caused by enforcing a single global adapter across divergent clients. To address these issues, we propose CF-LoRA, a clustered federated LoRA fine-tuning framework that combines decoupled factor aggregation with adaptation-aware client clustering. CF-LoRA first learns a globally shared $A$ factor while retaining personalized $B_i$ factors, then identifies clients with similar adaptation patterns based on the cosine similarity of their learned $B_i$ factors, and finally performs intra-cluster $B$-factor aggregation with a frozen $A$ factor. By decoupling LoRA factor aggregation, CF-LoRA preserves the low-rank structure and mitigates the structural aggregation mismatch, while adaptation-aware clustering promotes collaboration among clients with similar adaptation patterns and reduces negative transfer caused by statistical heterogeneity. Experiments on four language tasks and four vision datasets with RoBERTa and ViT show that CF-LoRA achieves the highest average accuracy in both modalities while communicating only one LoRA factor per optimization round.
|
| 2448 |
ImbalancE: Inference-Time Latent Search Against Degree Imbalance in Link Prediction
2609.36996
|
cs.AI
|
Alberto Bernardi, Luca Costabello, Christophe Gueret |
Knowledge Graph Embedding models have been extensively used to learn representations of entities and relations in Knowledge Graphs for predicting missing links. However, the quality of the learned representations varies a lot across different areas of the grap...Knowledge Graph Embedding models have been extensively used to learn representations of entities and relations in Knowledge Graphs for predicting missing links. However, the quality of the learned representations varies a lot across different areas of the graph. If previous research has loosely linked the problem to relation types or degree bias, we show that it is more widespread and it correlates with the degree imbalance of the entities in test triples. In particular, the prediction of a target entity that has a degree much smaller than the degree of the anchor entity is extremely problematic. This is critical in recommender systems and other use cases, where these triples represent important corner cases. To address this issue, we propose an inference-time latent search optimization method capable of significantly improving model predictions on the most imbalanced triples. Built on top of a pre-trained model, it explores the embedding space at evaluation time, blending known and out-of-band information to mitigate the degree imbalance bias. We show the value of our approach on imbalanced triples from common benchmark datasets, where we outperform conventional methods, opening the door to the successful adoption of Knowledge Graph Embedding models on these critical corner cases.
|
| 2449 |
CADOC: Cache-Aware Dynamic Object Context for Long-Horizon Agents
2609.37012
|
cs.AI
|
Junjie Yao, Zhangchen Zhou, Zhi-Qin John Xu |
For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the...For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the exact originals retrievable, but editing the history can break prefix-cache reuse, and prior recoverable methods time their edits by forecasts of future reuse or by preset intervals. We propose CADOC (Cache-Aware Dynamic Object Context), an online algorithm that replaces structured objects with compact Cards while preserving exact, on-demand retrieval of their original contents. CADOC schedules replacements in batches by balancing accumulated waiting cost against shared cache-reconstruction cost. Its scheduling rule follows from an economic order quantity trade-off, recovers the optimal integer batch under stationary assumptions. Across evaluation, CADOC consistently achieves the lowest aggregate input cost among the compared configurations, which reduces input cost by approximately 40\% on average while maintaining task performance close to full context. CADOC thus provides a cost-derived approach to compressible context management, demonstrating that efficient compression depends not only on shortening prompts but also on scheduling edits to preserve cache reuse.
|
| 2450 |
Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization
2609.37022
|
cs.AI
|
Guoqing Zhang, Rafik Hadfi, Takayuki Ito |
Efficient patient flow coordination across autonomous hospital departments is critical for mitigating overcrowding and balancing resource utilization. While classical queueing theory, specifically open Baskett--Chandy--Muntz--Palacios (BCMP) networks, provides...Efficient patient flow coordination across autonomous hospital departments is critical for mitigating overcrowding and balancing resource utilization. While classical queueing theory, specifically open Baskett--Chandy--Muntz--Palacios (BCMP) networks, provides an interpretable mathematical topology for healthcare operations, analytical models rely on stationary assumptions and fixed routing matrices that degrade under state-dependent real-world dynamics. Conversely, centralized reinforcement learning approaches struggle to accommodate the decentralized structure of hospital governance, where individual clinical departments function with localized observations, heterogeneous resources, and divergent operational objectives. In this paper, we present a Multi-Agent Systems (MAS) framework titled \emph{Physics-Informed Multi-Agent Coordination}, which embeds empirically calibrated BCMP queueing topologies as physical priors within a decentralized multi-agent reinforcement learning architecture. Formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) under coupled resource constraints, our method enables autonomous departmental agents to cooperatively negotiate patient routing and dynamic service scaling. To mitigate environmental non-stationarity without inducing excessive communication overhead, agents exchange localized action fingerprints along network edges and optimize a spatially decomposed reward structure. Empirical evaluations driven by real-world MIMIC-IV patient trajectories indicate that this cooperative multi-agent approach substantially reduces cumulative system delay compared to static Markovian approximations, heuristic dispatching, and independent multi-agent baselines, while maintaining clinical safety constraints.
|
| 2451 |
Language as the Interface: Foundation-Model Contrastive Learning Links Transcriptomes and Electrophysiology
2609.37024
|
cs.AI
|
Junbo Shen, Jinying Gao, Bo Lei |
Integrating transcriptomic and electrophysiological data is essential for building multimodal foundation models for neuroscience. Patch-seq provides paired measurements of gene expression and intrinsic electrophysiology from the same neuron, establishing a bas...Integrating transcriptomic and electrophysiological data is essential for building multimodal foundation models for neuroscience. Patch-seq provides paired measurements of gene expression and intrinsic electrophysiology from the same neuron, establishing a basis for training cross-modal models. Here we introduce LangPatch, a foundation-model-based contrastive learning framework that uses paired Patch-seq data to align pretrained GenePT representations with electrophysiological phenotypes through a language-based interface. Gene descriptions and verbalized electrophysiological profiles are embedded by the same frozen text encoder. A context adapter and projection modules connect the modalities through paired contrastive learning. Across mouse visual, mouse motor, and human cortical cohorts, LangPatch achieves the highest mean transcriptome-to-electrophysiology prediction correlation among the evaluated foundation-model and representation-learning methods. It also improves held-out cross-modal alignment in the two mouse cohorts (FOSCTTM 0.107/0.135 vs. 0.208/0.222 for JAMIE, an existing cross-modal Patch-seq imputation method). It predicts transcriptomic family, type, cortical layer, and marker-gene expression from electrophysiology, exceeding other baselines on most endpoints. More importantly, the method transfers across brain areas and species: a model trained on mouse visual cortex predicts electrophysiology in motor cortex with approximately 70% correlation retention and in human cortex with 47% (58% on acute-slice recordings). Together, these results demonstrate alignment between molecular and functional representations of neurons, providing a building block for multimodal foundation models in neuroscience.
|
| 2452 |
AnyAct: Universal Action for Self-Evolving Agents
2609.37025
|
cs.AI
|
Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren, Chao Huang |
As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to s...As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to semantic APIs. However, three core challenges persist: the "scale dilemma" of massive tool ecosystems exceeding LLM context windows, the "non-stationarity" of tool quality due to updates or outages, and the "heterogeneity" of feedback formats (pixels, text, structured data) creating information silos. To address these, we propose AnyAct, a universal action layer that unifies available capabilities into a self-evolving action space, enabling agents to operate efficiently and reliably in large-scale, dynamic tool ecosystems. AnyAct's core design focuses on two objectives: constructing this action space via hierarchical progressive retrieval (filtering task-relevant actions) and test-time reliability evolution (pruning unreliable actions), and enabling reliability-aware action orchestration through a heterogeneous observation grounding module that unifies multi-modal feedback. Additionally, it defines a hybrid action space (primitive + semantic actions) and optimizes for a balance between task success rate and execution cost. Evaluations on LiveMCPBench and OSMCP (a new benchmark we developed for multi-granularity action collaboration) demonstrate state-of-the-art performance. AnyAct delivers substantial performance gains over baseline methods across various LLM base models on LiveMCPBench and improvements are particularly notable for models with constrained native capabilities. On OSMCP, it achieves 77.27% overall success with only 50 steps, which is half the steps required by most competitors.
|
| 2453 |
Beyond Low-Rank Parameterization: Narrowing the Gap Between LoRA and Full Fine-Tuning via Gradient Decomposition
2609.37027
|
cs.AI
|
Yihao Ouyang, Shiwei Li, Haozhao Wang, Xiandi Luo, Zhuoqi Hu |
Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each trai...Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space directions are constrained by the current parameterization. We characterize the corresponding LoRA-accessible gradient space and show that it coincides with the tangent space induced by the current LoRA parameterization. This characterization yields an orthogonal decomposition of the full weight gradient at the current model parameters. We term the component orthogonal to this space the normal gradient. Based on this decomposition, we propose GDLoRA (Gradient-Decomposed Low-Rank Adaptation). GDLoRA reconstructs the full weight gradient from forward activations and backward signals, extracts its normal component, and directly updates the base weights with this component, while retaining standard AdamW optimization for the LoRA factors. GDLoRA incorporates complementary normal gradients without increasing standard LoRA's optimizer-state memory budget under matched adapter and optimizer configurations. Experiments on natural language understanding, mathematical reasoning, commonsense reasoning, and image classification show that GDLoRA consistently improves over LoRA and narrows the performance gap to FFT. The code is available at https://anonymous.4open.science/r/GDLoRA.
|
| 2454 |
FedLAFP: Low-Rank Aggregation Meets Full-Rank Personalization in Federated Fine-Tuning
2609.37033
|
cs.AI
|
Mengjun Yi, Huaian Gu, Yinghao Ai, Furao Shen, Jian Zhao |
Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but statistical heterogeneity makes a single global adapter insufficient for personalized prediction. Existing perso...Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but statistical heterogeneity makes a single global adapter insufficient for personalized prediction. Existing personalized methods typically use the same low-rank structure for both shared and private adaptation, overlooking their distinct requirements for aggregation and personalization. We propose FedLAFP, a role-aware framework that couples a compact, globally aggregated LoRA branch with a client-private, full-rank-capable RandLoRA branch. The shared branch provides an efficient interface for transferring common knowledge, whereas the private branch combines fixed random low-rank bases with learned scaling coefficients to provide expressive client-specific adaptation without additional communication. Client- and layer-specific mixing coefficients jointly fuse the two branches, and only the shared LoRA parameters are exchanged. A controlled linear study supports this role assignment: LoRA yields more aligned client updates and lower aggregation error, while RandLoRA more accurately recovers client-specific residuals. Experiments across four visual recognition benchmarks show that FedLAFP consistently outperforms local-only and federated LoRA baselines, achieving an average personalized accuracy of $86.93\%$ and exceeding the best baseline average by $1.30$ percentage points.
|
| 2455 |
Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning
2609.37035
|
cs.AI
|
Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu |
Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can om...Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can omit visual details before later questions reveal their relevance; the retained state alone cannot recover them. We introduce Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning. WTI maintains compact natural-language memory entries tagged with source-video time ranges; these entries support direct reasoning when sufficient and otherwise anchor selective recall of finer visual evidence. For each question, WTI answers when current context and memory suffice, continues watching when required evidence has not appeared, or recalls a relevant past interval and decides again after incorporating the returned chunks, without replaying the full observed history. To train this behavior, we construct WTI-82K, comprising 82,335 timed questions across 4,812 causally aligned trajectories, and develop Stream-GDPO to optimize complete multi-question streaming rollouts using trajectory-level feedback for response timing, source-video recall, and memory updates. WTI achieves state-of-the-art aggregate performance among the compared open-source streaming baselines, reaching 83.3% on StreamingBench and 73.6% weighted overall accuracy on OVO-Bench.
|
| 2456 |
MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows
2609.37053
|
cs.AI
|
Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu |
Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and ...Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.
|
| 2457 |
actr: aligning thoughts and responses for multilingual safety in reasoning llms
2609.37054
|
cs.AI
|
Xianhui Zhang, Jian Yu, Chengyu Xie, Chenhang Cui, Shuyi Miao |
Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their rea...Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual safety gap. Next, using a corpus of jailbreak queries, we assess neuron importance through changes in response representations caused by neuron masking and compare the high-importance neuron sets obtained with reasoning enabled and disabled to identify safety think neurons that support the use of safety reasoning. Finally, we devise neuron-selective consistency optimization (NSCO), which uses a frozen judge model to reward agreement between the safety categories of reasoning traces and responses while updating only the parameters associated with the selected neurons, requiring no human-annotated responses or preference data. Across two reasoning models, ACTR achieves lower average attack success rates than the evaluated state-of-the-art methods on AdvBench-X and MultiJail, with safety gains extending to unseen languages, while preserving or improving average performance on multilingual knowledge and mathematical reasoning tasks and limiting false refusals of benign requests. Warning: this paper contains examples with unsafe content.
|
| 2458 |
Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM-based Scientific Reviewers
2609.37097
|
cs.AI
|
Zhuo Chen, Hao Zeng, Jiawei Liu, Guoxiu He, Le Cai |
The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree ...The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions are largely based on a narrow set of perturbation strategies instantiated with static templates, providing limited evidence of actual reliability. In this paper, we construct a three-level evaluation framework covering perturbations to surface presentation, argumentative logic, and value judgment. Experiments on representative LLM-based reviewers reveal two limitations of static evaluation: stratified vulnerability, where perturbation effects depend on whether the paper's original review score is high or low, and perturbation undercoverage, where a single template misses vulnerabilities exposed by diverse realizations. To address these limitations, we propose SCOPE-Fuzzer, a strategy-aware fuzzer that combines feedback-driven strategy selection with adaptive mutation of paper content. By iteratively probing reviewers with dynamic perturbations, SCOPE-Fuzzer consistently uncovers vulnerabilities overlooked by static evaluation and other baselines.
|
| 2459 |
Learning from Viable Failure Prefixes: Milestone Viability Potential Policy Optimization for Long-Horizon LLM Agents
2609.37111
|
cs.AI
|
Qi Zhou, Yuanfan Li |
Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Existing group-based methods such as GRPO and GiGPO alleviate this issue by comparing rollout returns or repeated ...Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Existing group-based methods such as GRPO and GiGPO alleviate this issue by comparing rollout returns or repeated anchor states, but they still fail when the compared returns have no variation. We identify this failure mode as zero-credit failure: during early training, many failed rollouts contain useful prefixes, yet existing methods assign them no task-discriminative advantage. To address this issue, we propose Milestone Viability Potential Policy Optimization (MVPO), a potential-routed policy optimization algorithm that learns from viable failure prefixes. MVPO estimates prefix potential over Union-Find viability regions, repairs zero-credit groups with potential-difference advantages, and attenuates the potential branch according to relative performance progress. Experiments with Qwen2.5-1.5B-Instruct show that MVPO outperforms eight strong baselines, including GRPO and GiGPO. Under the same training length, MVPO improves over the GiGPO baseline by +4.4 success points on ALFWorld and +5.3 on WebShop, while adding only 0.16%-0.20% advantage-construction overhead.
|
| 2460 |
When Should Agents Check External State? Budgeting Observations for Stored Intentions
2609.37125
|
cs.AI
|
Zhengkun Di, Bin Shi, Kai Sun, Yiming Xu, Bo Dong |
Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid calls. Existing systems dec...Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid calls. Existing systems decide when intentions require attention, but do not allocate the resulting observations under a shared budget. We introduce the first resource-allocation formulation for the external observations required by stored intentions under a shared episode budget. BudgetPM offers two policy variants that share a hard-budget executor. BudgetPM-Static uses a lightweight Logistic scorer to learn whether a check improves the current decision. BudgetPM-Sequential distills full-episode hindsight schedules into a lightweight policy that decides when to spend or reserve capacity using only pre-query information at deployment. We evaluate BudgetPM against two public memory-agent systems, five matched controls, and four hand-designed monitoring or budget-adaptation rules. Across two benchmarks and three backbones, BudgetPM-Static outperforms adapted Mem0 and PMA workflows. On PM-Bench, its Logistic scorer reaches competitive quality--cost operating points alongside higher-capacity scorers and retains 99.9--100\% of unconstrained quality with 42--54\% fewer observations. Under severe scarcity and the same hard caps, BudgetPM-Sequential exceeds the strongest tested natural monitoring schedule by 1.92--2.58 Set F1 points. It reaches the same Set F1 and on-time recall with 16--33\% fewer observations. Matched attribution, exact-cost analysis, and a fixed-budget load intervention link this gain to competition between present and future opportunities. These results yield a demand--capacity design rule: local gating works when capacity covers demand, while future-aware supervision adds value when observations compete across time.
|
| 2461 |
SkillCome: Group Contrast Skill Optimization with Dual Memory
2609.37128
|
cs.AI
|
Haolin Li, Feng Hong, Ang Li, Chilin Fu, Weichang Wu |
Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a single trajectory per question. However, this provides insu...Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a single trajectory per question. However, this provides insufficient optimization signals since it requires inferring effective skill edits from a solitary path. It is difficult to pinpoint which actions caused the failure in a failed trajectory, or to determine which actions in a successful one should be incorporated into the skill. Furthermore, they rely on a local batch of trajectories for analysis, making the optimization direction susceptible to noisy evidence. To address these, we propose SkillCome, a Skill-evolution method based on group Contrast optimization with dual memory. For each question, SkillCome generates trajectories and performs group contrast analysis to precisely identify key behavioral divergences between successful and failed trajectories, offering reliable optimization signals. The dual memory system further accumulates evidence from historical steps to track patterns shared across different groups, leading to more generalized optimization directions. Together, SkillCome builds a systematic optimization process that transforms experience from observed successful trajectories into reusable skills. Extensive experiments on six benchmarks spanning question answering, reasoning, and agentic tasks demonstrate the effectiveness of our method. SkillCome consistently outperforms baselines across five models of varying families and scales, with gains up to +5.69 points.
|
| 2462 |
Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
2609.37132
|
cs.AI
|
Zheng Zhang, Xinyue Tan, Lufei Li, Xinyi Zhang, Yexin Li |
On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the cu...On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher's ability to exploit privileged information. We ask whether the model's own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD), which temporarily trains the policy ahead to obtain a future teacher, restores the student to the original policy state, and then uses the future teacher to supervise the restarted student. The future teacher improves supervision in two complementary ways, it can generate more reliable privileged trajectories and, conditioned on them, provide more informative token-level targets along the restarted student's on-policy trajectories. Experiments on mathematical reasoning with Qwen3-4B and Qwen3-8B show consistent improvements over standard OPSD in both settings, including gains from 27.50 to 41.30 and from 48.80 to 64.44 in the rollout-privileged setting. Our findings point to a broader principle for self-improving models that future learning progress can be distilled backward, preserving acquired knowledge while bootstrapping beyond the optimization state that produced it.
|
| 2463 |
From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks
2609.37145
|
cs.AI
|
Zheng Zhang, Lufei Li, Xinyue Tan, Yuanhao Zeng, Ziwei Shan |
LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic ju...LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to training-time supervision. We systematically investigate judgment quality and downstream utility by examining both how judgments are elicited and how they are used. For judgment elicitation, we vary the Judge protocol along three dimensions: verdict granularity, critique usage, and evaluation batching. For judgment usage, beyond policy training, we extend Judge to test-time inference through Best-of-N selection, Judge-guided revision, and beam search. We find that, (i) Surprisingly, judgment quality and downstream utility do not always align. (ii) Judge protocol design substantially affects both intrinsic judgment quality and downstream utility. (iii) Judge guidance effectively converts test-time compute into performance gains, with benefits varying across inference strategies. Our results call for a multifaceted evaluation of LLM Judges on open-ended tasks, encompassing intrinsic judgment quality, and downstream utility.
|
| 2464 |
When Tools Silently Lie: Evaluating and Mitigating Blind Compliance in Tool-Augmented Data Agents
2609.37153
|
cs.AI
|
Zifu Tao, Changqing Yin |
Tool-augmented data agents rely on tool outputs for analytical decisions. Yet successful execution can return plausible but incorrect evidence, requiring agents to decide whether to trust or verify it. Understanding this failure requires examining both the evi...Tool-augmented data agents rely on tool outputs for analytical decisions. Yet successful execution can return plausible but incorrect evidence, requiring agents to decide whether to trust or verify it. Understanding this failure requires examining both the evidence obtained through checking and the answer ultimately adopted. We introduce ToxicBench to measure checking and adoption under numerical, label, schema, and retrieval errors, pairing clean and poisoned observations over fixed source data. In the 118-task GPT evaluation across three adapters, poisoning lowers task success by 26 to 39 percentage points. Ordinary retries help under one-shot poisoning, whereas repeated poisoning reveals wrong-answer adoption after checking. Controls on three public tables isolate how supplied evidence affects recovery. After freezing the scorer, we compare its judgments with human annotations on 200 trajectories, finding 96% task-success agreement. Human judgments support retry gains over Base and confirm adoption after checking on audited tasks. We release trajectories, versioned scoring, and reference and delivery audits. These findings highlight evidence availability and answer selection as complementary dimensions of agent reliability.
|
| 2465 |
From Learner Behavior to Reusable Skills for Effective and Efficient Learner Simulation
2609.37157
|
cs.AI
|
Zijian Chen, Zheng Zhang, Miao Jia, Xingchen Hu, Weibo Gao |
Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction hist...Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction history to reconstruct the learner. This introduces additional context and inference costs and makes the acquired learner-specific simulation capability difficult to reuse across different LLMs. We therefore propose Learner2Skill, which externalizes the simulation capability acquired from historical interactions into a persistent and reusable Simulation Skill. The Skill captures the learner's current learning state and recurring response patterns, evolves as new real interactions arrive, and can be adapted to a new LLM through lightweight executor calibration without reconstructing the learner from scratch. Experiments show that Learner2Skill more faithfully reproduces fine-grained learner behavior while reducing overall token cost, and that the same constructed Skills can be effectively reused across different LLM executors.
|
| 2466 |
SimpleEvol: An Agent-Loop Framework for LLM-Driven Automated Heuristic Design with Minimal Human Priors
2609.37172
|
cs.AI
|
Jianghan Zhu, Cong Zhang, Rongjie Zhu, Chi Zhang, Zhiguang Cao |
Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuristics. However, the dominant paradigm embeds LLMs as narrow, fixed components, such as crossover or mutation,...Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuristics. However, the dominant paradigm embeds LLMs as narrow, fixed components, such as crossover or mutation, within heavily hand-engineered evolutionary frameworks. We argue this misapprehends LLMs. It treats them as specialized tools rather than general reasoners, constrains them to low-level operations, and underutilizes their autonomy. Moreover, the extensive human priors in these frameworks violate the bitter lesson principle that general methods scaling with computation surpass hand-crafted solutions. This raises a key question: which AHD framework designs best convert stronger LLM capabilities into better heuristics? To address this, we propose metrics for LLM-driven AHD framework handcraftedness (AHI) and intelligence conversion efficiency (ICE). Evaluating ten LLMs across three challenging combinatorial optimization problems, we obtain a notable finding that frameworks with fewer human priors consistently yield higher ICE. Based on this finding, we propose SimpleEvol, an agent-loop framework for AHD which removes nearly all human priors and allows the LLM to operate autonomously. SimpleEvol consistently achieves the highest ICE, often by a large margin. Our results challenge the trend toward complex AHD pipelines and point to a lighter and more model-centric alternative, suggesting that reducing human priors is a more effective strategy to scale up with model intelligence. The source code is available at https://github.com/HenryZhu1029/SimpleEvol-Master.
|
| 2467 |
Absorbed in Inertia: Activation Analysis for Computer-Use Agents
2609.37176
|
cs.AI
|
Giulio Segalini, Zhi Wen Soi, J\'er\'emie Decouchant, Lydia Chen |
Computer-use agents have become increasingly capable of executing tasks on live desktops through natural-language instructions, based on trajectories of screenshots, actions, and reasoning. We discover that they can stealthily exhibit inertia, in which they re...Computer-use agents have become increasingly capable of executing tasks on live desktops through natural-language instructions, based on trajectories of screenshots, actions, and reasoning. We discover that they can stealthily exhibit inertia, in which they repeat fruitless actions despite recognizing that these actions are ineffective. We hypothesize that inertia is reflected in the agent's internal state, i.e., the activation values of the agent's underlying model, and propose a protocol to measure the relationship between the two. Extensive analysis of high-dimensional activation states shows that inertia corresponds to an absorbing region of activation space, where activation values become stale across actions and even after attempts to steer them. We conjecture that drastically changing the agents' activations by re-initializing them is necessary to escape inertia. Specifically, we propose R$^3$ (Reset, Reroute, Restore), which temporarily resets the agent's context trajectory to escape the absorbing region and then restores the historical context to effectively complete the task. Our approach yields 17-55% lower measured inertia across models relative to unmodified agents. These results suggest that changing the context can interrupt recurrence more effectively than directly steering the resulting activations. Our code is available at https://anonymous.4open.science/r/vlm-agent-defense-D076
|
| 2468 |
Accelerated surrogate dynamics for dynamical, stochastic system evolution
2609.37184
|
cs.AI
|
Marco Jochum, Ioannis Kouroudis, Gohar Ali Siddiqui, Taher Amine Hamzaoui, Manuel G\"o{\ss}wein |
Dynamic simulations are an entrenched way of gaining insight into the evolution of system dynamics. Their computational cost however is often prohibitively high, especially in cases of stochastic frameworks. Machine learning algorithms are especially suited as...Dynamic simulations are an entrenched way of gaining insight into the evolution of system dynamics. Their computational cost however is often prohibitively high, especially in cases of stochastic frameworks. Machine learning algorithms are especially suited as simulation surrogates. Nevertheless, they face some very distinct limitations. Firstly, the sheer dimensionality of these systems, however, precludes the use of traditional time series models who struggle with high dimensional feature spaces. Additionally, traditional time series focus exclusively on either long or short range effects, causing local or global drift given enough time. In this paper, we propose a framework that addresses those limitations. Our framework combines a Variational Autoencoder, with a convolutional or graph basis that reduces the dimensionality of the system. This latent vector is propagated in time using a Temporal Fusion Transformer model, which includes both long range and short range effect encoding, as well as static covariate support. We test our framework on three distinct cases, to prove its robustness and in all three we have achieved practically identical to the simulation results at a fraction of the time. Further, our framework is flexible enough to be adapted to any new system and provides an inbuilt uncertainty quantification for targeted experiment design.
|
| 2469 |
V-Engram: Trigger-Indexed External Memory for Modular Text-to-Image Personalization
2609.37198
|
cs.AI
|
Haoran He, Runyuan Cai, Yiming Wang, Lin Yu, Xiaodong Zeng |
Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit iden...Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prompt-selective and multi-concept access without merging model updates. Experiments show that V-Engram broadly matches DreamBooth-LoRA in overall subject fidelity while showing advantages in settings such as contextual subject preservation. Prompt-matched loading retrieves only matched entries, reducing most additional adaptation-state loading for a single-concept query. Qualitative results further demonstrate paired-trigger composition and same-class separation, while prompts without registered entries retain the frozen model's base behavior. Together, these results establish trigger-indexed memory as a modular interface for adding targeted visual evidence without rewriting the generator.
|
| 2470 |
Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning
2609.37203
|
cs.AI
|
Qili Zhang, Qianren Mao, Hanze Cai, Kaiming Zhao, Yuening He |
Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its prem...Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final answer. Existing methods lack machine-checkable verification of intermediate conclusions and answer-supporting proof dependencies, so they may assign credit to invalid or answer-irrelevant steps. We propose Proof-R1, an RL framework from formal verification that trains LLMs to construct verifiable proofs for natural-language logical reasoning. Proof-R1 admits a generated conclusion into the verified proof state only when the corresponding reasoning action satisfies the proof obligations through UNSAT-based machine-checkable formal verification. Proof-R1 also recovers the answer-supporting dependency closure to trace the proof structure of the final answer and align outcome credit with the proof dependencies. Experiments demonstrate that Proof-R1 improves answer accuracy across three logical reasoning benchmarks and four backbone models and outperforms training-free agents and training-based methods in terms of reasoning-process verifiability.
|
| 2471 |
Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis
2609.37220
|
cs.AI
|
Jingxi Feng, Xudong Chen, Yifan Zhang, Heming Xu, Hongcheng Han |
Exploring high-order correlations and long-range dependencies in brain networks holds significant value for both neuroscience research and clinical diagnosis. However, previous studies have lacked a unified integration of high-order and long-range dependency i...Exploring high-order correlations and long-range dependencies in brain networks holds significant value for both neuroscience research and clinical diagnosis. However, previous studies have lacked a unified integration of high-order and long-range dependency information in brain networks, and there is substantial redundancy behind various types of information. These issues limit their effectiveness in the diagnosis of brain diseases. To address this, we propose an Information Bottleneck-Guided Adaptive HyperGraph Transformer (IBAHGT). By incorporating the information bottleneck (IB) principle, this approach enables adaptive learning of high-order correlations and both short- and long-range dependencies within a unified framework for brain network analysis, achieving high-precision brain disease diagnosis. IBAHGT consists of three key components: an information bottleneck-guided adaptive hypergraph convolution, which introduces a novel hypergraph information bottleneck (HIB) principle to adaptively learn hypergraph message-passing weights between nodes and hyperedges, optimizes information flow and captures high-order information in brain networks that is maximally informative and minimally redundant (MIMR). The Transformer encoder captures global information within brain networks through the attention mechanism, specifically modeling short- and long-range dependencies. An information bottleneck-guided node-level adaptive fusion employs the IB principle to learn independent weights for each node, facilitating the fine-grained integration of high-order information and global information to obtain an efficient representation for downstream tasks. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art methods and can identify biomarkers for clinical applications.
|
| 2472 |
OptiCom : A Unified Framework for State-Conditioned Composition in LLM-Driven Optimization
2609.37221
|
cs.AI
|
Chenxing Wei, Sichen Liu, Lizhao Liu, Ningyuan Sun, Chen Bingzhou |
Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve r...Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnostics reveal that mechanism effectiveness is highly state-dependent. Motivated by this, we analyze how individual decisions drive final outcomes, decomposing the expected terminal improvement under a shared budget into cumulative decision opportunities minus cumulative selection losses. Guided by this opportunity-loss theoretical foundation, we propose OptiCom, a unified framework that represents LLM-driven optimizers within a shared configuration space: C=(A,Q,O,E,M,S), corresponding to artifact, query, operator, evaluation, memory, and strategy. Operating within this space, a fast LLM-based Optimization Controller dynamically composes immediate mechanisms through structured Action Packages, while a slower Strategy Adapter refines long-term selection preferences, operator weights, and templates based on accumulated trajectory feedback. Comprehensive evaluations across 32 benchmark groups demonstrate the superiority of framework: OptiCom achieves an average Max-score rank of 1.72 among 14 evaluated configurations, securing the top score in 23 groups. Ultimately, these results highlight the broad applicability and high extensibility of OptiCom as a general-purpose paradigm for robust LLM test-time scaling.
|
| 2473 |
Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents
2609.37236
|
cs.AI
|
Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi, Leshem Choshen |
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what informatio...An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model $15\times$ larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the $15\times$ larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.
|
| 2474 |
Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym
2609.37267
|
cs.AI
|
Jio Oh, Seunghyun Do, Young-Jun Lee, Steven Euijong Whang, Dongyeop Kang |
Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agen...Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.
|
| 2475 |
Task-Relevant Null-Space Residuals for Non-Injective Neural Mappings
2609.37272
|
cs.AI
|
Bizu Feng, Zhimu Yang, Shuming Wang, Yuan Cheng, Shaode Yu |
Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream task...Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-induced indistinguishability and task-required distinctions. For non-injective linear operators realized in the current forward pass, their null spaces exactly characterize these invisible input variations. We propose Task-Relevant Null-Space Residuals (NSR), a general residual framework for non-injective linear mappings. NSR combines null-space component extraction from pre-mapping representations, member-level encoding and gating, and application-specific integration to exploit potentially task-relevant information under downstream supervision while preserving the original aggregation or merging rules. We evaluate NSR in two structurally different settings: token merging and graph aggregation. In token merging, NSR achieves higher semantic segmentation performance than the corresponding compressed baselines in 34 out of 36 evaluated configurations, with a maximum observed gain of 31.51 mIoU points under strong compression. In graph aggregation, NSR achieves 100% training accuracy on Tree-NeighborsMatch at depths d=2--6 across three backbones, alongside gains on heterophilic node classification and molecular graph regression. Together, these results support null-space residuals as a practical complement to non-injective linear mappings, enabling downstream models to learn from input distinctions invisible in the original operator's output.
|
| 2476 |
Transolver-$\sigma$: Joint Spectral-Physical Subspace Modeling for Neural PDE Solving
2609.37279
|
cs.AI
|
Haonan Shangguan, Hang Zhou, Haixu Wu, Yuezhou Ma, Jianmin Wang |
Neural solvers offer efficient surrogates for numerical simulation of partial differential equations (PDEs). For time-dependent problems, strong one-step accuracy does not necessarily translate into reliable autoregressive rollout. We observe that a solver bas...Neural solvers offer efficient surrogates for numerical simulation of partial differential equations (PDEs). For time-dependent problems, strong one-step accuracy does not necessarily translate into reliable autoregressive rollout. We observe that a solver based only on physical-state modeling can achieve lower one-step error, whereas its spectral-only counterpart can become more accurate at later rollout steps. Motivated by this observation, we present Transolver-$\sigma$, a neural PDE solver based on joint spectral--physical subspace modeling. Within each block, adaptive physical-state interactions and spectral transformations are modeled in dedicated latent subspaces, whose responses are recomposed to enable information exchange between the two representations. Within the physical subspace, we introduce Slice-Residual Physics-Attention (SRPA), which preserves an explicit slice-space identity path while retaining learnable cross-slice interaction. In parallel, an axis-factorized Fourier operator captures global spectral structure. Across five well-established PDE benchmarks spanning steady-state prediction and time-dependent dynamics, Transolver-$\sigma$ achieves state-of-the-art with a benchmark-averaged relative error reduction of 33.4% over the strongest baseline for each metric, while consistently improving autoregressive rollout over single-operator counterparts. Transolver-$\sigma$ further delivers strong gains on coupled multiphysics systems and real-world fluid and combustion measurements from RealPDEBench, demonstrating its effectiveness beyond standard simulation benchmarks.
|
| 2477 |
AssayRouter: Historical Utility Priors for Frozen Molecular Predictor Routing
2609.37285
|
cs.AI
|
Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu |
Laboratories often face a new molecular assay with 16-64 labels and a bank of predictors whose training data and parameters are unavailable. The practical question is which frozen outputs to include in a small local model. AssayRouter treats completed assays a...Laboratories often face a new molecular assay with 16-64 labels and a bank of predictors whose training data and parameters are unavailable. The practical question is which frozen outputs to include in a small local model. AssayRouter treats completed assays as pseudo-targets and labels each candidate by its post-fit utility: the reduction in held-out discovery loss when the candidate is added to the local target predictor. A shared regressor learns to predict this utility from candidate behavior on the support set, without source identity; on a new assay, one frozen ranking selects four sources and separate labels fit a convex combiner. We train only on completed ChEMBL-MT assays and evaluate 24 external regression assays across six frozen interface families. AssayRouter-C lowers strict four-call negative log-likelihood (NLL) by 0.0409 relative to Support-CV@4. Frozen candidate-label permutations confirm that candidate-utility correspondence carries the transferred information, and leave-one-interface-out training shows that the mapping generalizes to unseen predictor families. Completed assays therefore provide transferable supervision for scarce-label routing through frozen prediction interfaces.
|
| 2478 |
MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller
2609.37304
|
cs.AI
|
Zhibin Wen, Tao Han, Lei Bai, Can Li, Yang Xu |
Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Converse...Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at https://github.com/binbin2xs/MetaCtrl.
|
| 2479 |
ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents
2609.37311
|
cs.AI
|
Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin, Wenqi Fan |
Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, ...Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16\% across three recommendation agent tasks, namely searching, ranking, and judging.
|
| 2480 |
Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation
2609.37322
|
cs.AI
|
Jiayuxuan Yang, Jie M. Zhang, Yiling Lou, Zhenpeng Chen |
Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We i...Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks whether the injected defects reduce response quality, and uses insufficiently penalized defects to refine the rubric. We evaluate Mubric on 703 tasks across four representative domains against six advanced rubric generation methods. Mubric achieves the highest overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points.
|
| 2481 |
VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict
2609.37324
|
cs.AI
|
Jiale Dai, Liuxian Ma, Xiaoke Niu, Wenjing Zhang, Huiying Zhao |
Conflicting emotional cues can be individually valid: a subdued voice may reflect a blocked goal while a smile satisfies a social obligation. Their interpretation depends on what the event means to the person. We introduce VISTA (Value-Informed Semantic Trust ...Conflicting emotional cues can be individually valid: a subdued voice may reflect a blocked goal while a smile satisfies a social obligation. Their interpretation depends on what the event means to the person. We introduce VISTA (Value-Informed Semantic Trust Arbitration), a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. A log-odds decomposition separates emotion expectation from cue diagnosticity, motivating an interface that lets appraisal change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone and matched training examples and steps, VISTA reaches 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency. Shuffling appraisal across scenes or removing its decision connection reduces this benefit. A common frozen-backbone probe reaches 0.600 macro CCC for appraisal readout, compared with 0.505 for emotion-only fine-tuning. Evaluations across five benchmarks connect recognition under increasing conflict with appraisal readout and downstream decision use. Together, the analyses and experiments support scene-specific appraisal as an intermediate representation that helps interpret conflicting emotional evidence.
|
| 2482 |
Solving Without Stopping: On-Policy Distillation at Small Scale
2609.37326
|
cs.AI
|
Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu, Said Mammar |
On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, i...On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student's single attempt never exceeds what it could already reach in many attempts before training, and the smaller the student, the further it stays below the teacher. Stopping is where the modes part. In non-thinking mode every student keeps stopping; in thinking mode students stop ending their reasoning early in training, and the smaller the student, the less of this ability survives: the teacher signals a stop almost only where a student already ends its reasoning, so distillation teaches no new stops; it only keeps the student's existing stops that land on a right answer, and a weak student has few such stops. The smallest students often reach the right value but do not commit to it: they either rarely mark it or mark it and write past it. Together, these results describe how small students behave under on-policy distillation, and a diagnostic that separates answer marking, correctness and stopping.
|
| 2483 |
Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation
2609.37353
|
cs.AI
|
Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang |
Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain conf...Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insufficient evidence. To address it, we propose SeekVLN, an evidence-seeking framework that couples semantic progress reasoning with active acquisition of task-relevant observations. SeekVLN is trained in two stages: First, Future-guided Reverse Generation (FRG) uses future expert actions to augment offline expert trajectories with supplementary views and evidence annotations. Supervised fine-tuning on these trajectories establishes a prior for evidence seeking and progress reasoning without additional expert interaction. However, imitation alone does not reveal whether seeking improves subsequent navigation. We therefore introduce Counterfactual Contrastive Policy Optimization (C2PO) for reinforcement fine-tuning. By comparing each evidence-seeking branch with a counterfactual direct-navigation branch from the same state, C2PO uses a contrastive reward to assign credit to seeking decisions based on subsequent navigation benefit. Experiments on simulated benchmarks show that SeekVLN achieves state-of-the-art performance, improving success rate by 12.7% and 7.5% over the base model on R2R-CE and RxR-CE, respectively. Both simulated and real-world evaluations exhibit human-like evidence-seeking behaviors for more reliable progress grounding.
|
| 2484 |
Teaching LLMs to Generate Challenging MILP Instances via Solver Feedback
2609.37356
|
cs.AI
|
Jitin Singla, Parikshit Pareek, Pratik Jawanpuria, Parag Singla |
Generating optimization instances that are both feasible and computationally challenging is crucial for benchmarking solvers and training learning-based optimization algorithms. Existing non-LLM generators rely on seed instances or parameter tuning, resulting ...Generating optimization instances that are both feasible and computationally challenging is crucial for benchmarking solvers and training learning-based optimization algorithms. Existing non-LLM generators rely on seed instances or parameter tuning, resulting in high test-time computational cost, while existing LLM generators lack explicit hardness measures. Recent reinforcement learning methods with verifier feedback evaluate only binary correctness, which is misaligned with generating challenging problems. We note that an optimization solver reports the cost of solving at several stages of its pipeline, and leverage this to design a reward that scores both the solvability and the hardness of generated problems, measured by branch-and-bound nodes and post-cut relaxation gaps. Our key idea is a challenger-solver asymmetric self-play approach, where an LLM challenger generates progressively harder instances and the solver verifies feasibility and hardness, so no seed or training MILP instances are required. We fine-tune Gemma-4-12B and Qwen3.5-4B with GRPO and a size curriculum into OptiScribe-12B and OptiScribe-4B, which generate feasible yet challenging MILP problems from natural language instructions. On capacitated facility location and max-cut, OptiScribe-12B raises median SCIP search nodes by 1.7-5x and post-cut gaps by 1.1-1.7x over its base model and improves the feasibility rate on facility location by 9-19 points, while OptiScribe-4B raises median nodes by up to 15.6x. The problems cover a wider difficulty range than public benchmarks of the same size, follow instructions on density and difficulty, and can tune solver settings for families that public libraries lack. These results indicate that optimization-specific rewards, used in self-play mode, can teach LLMs to generate high-difficulty optimization benchmarks. We will release our code and models publicly on acceptance.
|
| 2485 |
Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing
2609.37362
|
cs.AI
|
Guannan Lai, Han-Jia Ye |
Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is...Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate pool, and often requires additional supervision or retraining as the routing environment changes. We ask whether LLM routing can instead be approached from a foundation-model perspective, learning a reusable routing capability that generalizes across tasks, candidate models, and deployment conditions. To this end, we introduce RouteFM, which learns to characterize anonymous candidate models from behavioral context and infer their target-specific capabilities, rather than binding routing decisions to fixed model identities or a single environment. Through episodic pretraining across heterogeneous routing environments, this capability can be reused by a frozen router and adapted to new environments through context alone. Experiments demonstrate transfer across changes in domains, modalities, candidate pools, and context budgets, with the largest gains when behavioral evidence is limited. On MMR-Bench, which is excluded from pretraining, RouteFM outperforms the strongest baseline by 2.23 quality points with only eight observations per candidate. These results support moving LLM routing from repeated local fitting toward a pretrain once, route anywhere paradigm. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/RouteFM.
|
| 2486 |
Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation
2609.37377
|
cs.AI
|
Jiaxuan Wang, Jiafei Lyu, Yuchen Cai, Siye Wu, Pengyuan Wang |
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection acr...On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional changes across prompt supports and model pairs. Functional alignment with the teacher varies across supports and target tasks; in continuation pairs, teacher-aligned prediction changes can coexist with weak parameter alignment. Continued OPD on effective supports can restore performance after unfavorable transfer. Finally, targeted selection does not consistently outperform uniform random sampling, and filtering out a source that performs poorly alone yields no consistent gain across three paired support draws. Overall, our results distinguish prompt efficiency from prompt interchangeability and show that effective data choice depends on the teacher-student pair and target capability, with random sampling providing a competitive baseline in the studied settings.
|
| 2487 |
Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation
2609.37398
|
cs.AI
|
Xiangcheng Zhan, Zirui Chen, Yicheng Zhao, Ziteng Gao, Shuo Yang |
World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipul...World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.
|
| 2488 |
Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing
2609.37402
|
cs.AI
|
Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen |
Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect quer...Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33--41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9--9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/SaveRouter.
|
| 2489 |
Demistifying Data and Simulator Assumptions in Supervised Causal Discovery
2609.37446
|
cs.AI
|
Pingchuan Ma, Rui Ding, Bojun Huang, Shuai Wang |
Supervised causal discovery learns to infer causal structure for a new dataset from training datasets paired with structural labels. These training pairs are typically simulated, making the simulator both a source of supervision and a carrier of assumptions ab...Supervised causal discovery learns to infer causal structure for a new dataset from training datasets paired with structural labels. These training pairs are typically simulated, making the simulator both a source of supervision and a carrier of assumptions about causal graphs, mechanisms, and noise. Understanding the resulting predictions therefore requires examining how these assumptions supplement the information available in observational data, which may be compatible with multiple causal graphs. This paper examines that relationship across representative methods available through June 2026. We organize these methods by prediction target, prediction granularity, encoder, structural decoder, and training regime to relate what each method predicts to how it uses data and simulator-based supervision. Using this framework, we distinguish two questions: whether the target is identifiable under the assumed model class, and whether a trained predictor generalizes beyond its training distribution. Restrictions on mechanisms and noise can make otherwise ambiguous causal directions identifiable, but predictive accuracy under those restrictions does not establish transfer when they change. This distinction motivates evaluation that matches metrics to the identifiable graph target and tests changes in graphs, mechanisms, and noise between training and deployment. Extending such evaluation to real data also requires documenting the external causal evidence and uncertainty behind benchmark reference graphs. Together, these analyses guide method comparison and identify open questions in transfer, test-time adaptation, and uncertainty assessment.
|
| 2490 |
Commitment Hierarchies under Intent Revision: A Belief-Revision Account of Salvage in Tool-Use Agents
2609.37453
|
cs.AI
|
Spandan Ghose Chowdhury |
When a user changes their mind partway through a task, an agent that has already split the task into sub-goals and paid for tool calls must decide, per cached sub-result, whether to keep, patch, or discard it (salvage), restarting wastes valid work and continu...When a user changes their mind partway through a task, an agent that has already split the task into sub-goals and paid for tool calls must decide, per cached sub-result, whether to keep, patch, or discard it (salvage), restarting wastes valid work and continuing unchanged answers the old question. Our main finding is that salvage quality is a matter of role design rather than model capability: a language model asked the keep/patch/discard question one node at a time is unreliable, but asked to classify the revision once, with a deterministic layer propagating the decision, it reaches the cost-optimal oracle on all three models tested, from two vendors. Modeling the plan as a commitment hierarchy and the intent change as a belief-revision operator with AGM style postulates, we prove that no policy observing only a node's local view can be both safe and cost optimal, while the single classification design is both. Across three environments the policy recovers the full achievable savings, 43% cheaper than restart, at 100% correctness.
|
| 2491 |
Governing the Edge: Automating Commercial Property and Casualty Insurance Underwriting via a Hybrid Local-Cloud Multi-Agent Framework
2609.37454
|
cs.AI
|
Vivek Kumar Singh, Gautam Bhowmick |
Underwriters in commercial Property and Casualty (P&C) insurance spend 30 to 40% of their time on administrative work rather than risk judgment, and a single submission takes about 40 minutes by hand. We present Governing the Edge, a multi-agent framework ...Underwriters in commercial Property and Casualty (P&C) insurance spend 30 to 40% of their time on administrative work rather than risk judgment, and a single submission takes about 40 minutes by hand. We present Governing the Edge, a multi-agent framework for that layer, organized around a data-residency constraint: sensitive submission data must not leave the perimeter. Eleven agents and two deterministic control nodes, one a human-escalation interrupt, form a 13-node LangGraph workflow across two tiers. Agents touching raw submissions run locally on Gemma 2, each bound through a tool interface logging every invocation (synthetic stubs in the current prototype); only anonymized scores and non-identifying fields cross to Claude Sonnet in the cloud. Compliance rules are encoded as conditions on graph edges, so a non-compliant submission never reaches pricing. We walk through one full scenario end to end, a commercial auto submission whose principal driver carries serious violations, showing every agent call, every tool invocation, and the resulting routing decision. The framework processes a 40-minute manual submission in a few minutes on a single edge device, dominated by serialized on-device inference. On the hard-stop violation tier the framework enforces every rule correctly and reproducibly, since hard stops are deterministic predicates over fields extracted at temperature 0; overall compliance accuracy across the 20-scenario benchmark is 70%, with the remaining gap concentrated in softer, judgment-based tiers. It keeps a complete audit trail and respects the privacy boundary throughout. We release all code, compliance rules, tool stubs, and synthetic datasets.
|
| 2492 |
VeriWeave Govern: Evidence-Gated Deterministic Runtime Governance for Enterprise AI Agents
2609.37457
|
cs.AI
|
Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya |
Enterprise artificial-intelligence agents increasingly call tools, modify infrastructure, and process protected data, creating a need to separate action generation from action authorization. This article presents VeriWeave Govern, a deterministic runtime gover...Enterprise artificial-intelligence agents increasingly call tools, modify infrastructure, and process protected data, creating a need to separate action generation from action authorization. This article presents VeriWeave Govern, a deterministic runtime governance layer that evaluates structured agent actions against versioned policies, validates typed evidence, applies fixed deny > review > allow precedence, routes consequential actions to accountable human review, and records replayable tamper-evident audit state. GovernBench evaluates the design over 30 independent seeds and 60,000 oracle-labelled cases spanning five enterprise domains, adversarial evidence, out-of-distribution actions, and temporal policy evolution. VeriWeave achieves 0.9888 mean accuracy, 0.9836 macro-F1, zero observed aggregate false allows, and zero observed Governance Attack Success Rate on the evaluated cases. Six ablations show that evidence gating, deny precedence, out-of-distribution fail-safe behavior, human review, contradiction handling, and temporal replay contribute complementary safety. The deployed API additionally passes 12/12 end-to-end scenarios and a 40,040-request concurrency matrix with zero failures. A separate 150-case EU/Austria regulation-grounded evaluation uses frozen predictions and two independent blinded human annotators, who agree on all decisions. On this set, deterministic engines remain conservative, while a Gemma 4 31B comparator aligns more closely with the human consensus. The results expose a measurable safety--utility trade-off and motivate evidence-aware, replayable governance as an independent control plane for enterprise agent execution.
|
| 2493 |
CRASM-Gate: Deterministic-First Constraint- and Role-Aware Semantic Mapping with Selective Model Assistance Across Heterogeneous Industrial Standards
2609.37458
|
cs.AI
|
Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya |
Industrial standards often encode the same engineering concept through incompatible hierarchies, identifiers, roles, and structural constraints, so the nearest lexical or embedding match can still be technically inadmissible. This article presents CRASM, a det...Industrial standards often encode the same engineering concept through incompatible hierarchies, identifiers, roles, and structural constraints, so the nearest lexical or embedding match can still be technically inadmissible. This article presents CRASM, a deterministic constraint- and role-aware semantic mapping method, and CRASM-Gate, its selectively model-assisted extension. The framework separates standard-specific canonicalization, bounded retrieval, deterministic rules, destination-versus-origin role interpretation, semantic and structural ranking, ambiguity refusal, and target validation. CRASM-Gate adds a confidence/disagreement gate that may invoke a candidate-constrained large language model, while final authority remains with deterministic validation. A controlled artifact covers six directed industrial-standard pairs, three difficulty levels, and ten configurations, yielding 14,400 sample-level decisions. With a fixed local model endpoint, CRASM-Gate reaches mean F1 0.9938 and an implemented structural-validity rate of 1.0000; deterministic CRASM reaches 0.9931 without generative-model calls; and the model-only baseline reaches 0.5347. Relative to the model-only baseline, CRASM reduces top-1 errors from 670 to 10 while exhibiting 0.0688 s/sample rather than 24.5263 s/sample observed latency. CRASM-Gate improves CRASM by one additional correct decision out of 1,440, with observed latency increasing to 15.7666 s/sample. The results support a deterministic-first interoperability architecture in which model assistance is optional, measurable, candidate-bounded, and unable to bypass structural validation.
|
| 2494 |
Authority Before Utility: Non-Compensatory Control for Persistent LLM Memory
2609.37474
|
cs.AI
|
Wesley Shu |
Persistent memory creates a control problem that retrieval relevance alone does not solve: a memory can remain highly useful after an update, deletion, or revocation makes it inadmissible for the current answer. We formalize this as a separation between utilit...Persistent memory creates a control problem that retrieval relevance alone does not solve: a memory can remain highly useful after an update, deletion, or revocation makes it inadmissible for the current answer. We formalize this as a separation between utility and authority. A fixed finite penalty applied to an unnormalized utility score cannot guarantee exclusion under arbitrary positive-affine reparameterization of that score; by contrast, rank-normalized compensation is scale-invariant and therefore forms a stronger empirical comparator. Our prospectively frozen TIDE/LongMemEval primary was quarantined before a valid HELDOUT comparison because the materialized TIDE adapter conflated historical age with query-relative inadmissibility and the aligned LongMemEval split left no DEV set for the predeclared penalty selection. We therefore report a post-primary replacement diagnostic on Memora Remembering, where update/delete operations provide item-level forgetting state. On Qwen3-8B, DEV selected lambda = 0.6 from a ten-point normalized SOFT family. Across 185 HELDOUT units in 28 dependency clusters, HARD exclusion yields 4.04% balanced construct error versus 19.66% for locked SOFT, a paired difference of 15.61 points with a 20,000-replicate cluster-bootstrap 95% interval of [13.07, 18.76]. The effect is driven primarily by forgotten-value leakage while current-value recall is preserved. This is same-Q operator-comparison evidence, not a universal claim that scalar control fails, not an evaluation of learned authority inference, and not an independent downstream-harm endpoint.
|
| 2495 |
Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift
2609.37475
|
cs.AI
|
Wesley Shu |
Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mechanism that binds a single-use...Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mechanism that binds a single-use commit authorization to the exact action and to a semantic projection of the authorization state that justified it. On 2,847 attacked AgentDojo episodes, the boundary kernel preserves the unprotected agent's behavior exactly (80.576% utility; 1.616% attack success). On 10,302 frozen proposals, it accepts every unchanged commit and rejects every instance of ten prospectively specified changed or replayed classes. In an independently generated boundary-drift experiment, full joint binding commits 0/4,403 invalid contexts while retaining 5,899/5,899 valid contexts. A prospective external evaluation on the 3,460-scenario CONTINUITY suite retains all 700 benign cases, prevents 1,200/1,200 represented non-replay invalid effects, handles 160/160 replay lifecycles correctly, and withholds 200/200 ambiguous no-release cases. The broader external suite also exposes the method's limit: across all 2,560 attacks, BSC-R has a 25% invalid-effect commit rate versus 0% for CONTINUITY. The result is therefore a scoped commit-time consistency mechanism, not a universal agent-safety claim.
|
| 2496 |
SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
2609.37539
|
cs.AI
|
Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li |
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for ...Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training
|
| 2497 |
How Can Recommendation Feedback Evolve Agent Memory?
2609.37544
|
cs.AI
|
Shanwen Mao, Mingming Li, Hao Zhang, Zhiheng Li, Yige Wang |
Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience ...Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may result from the combined influence of multiple memories, making accurate attribution difficult. Existing methods rely primarily on immediate feedback or semantic retrieval and therefore struggle to reliably translate recommendation outcomes into memory fitness. To address this challenge, we propose TIDE (Trajectory-Informed Directed Memory Evolution), an external memory evolution framework driven by delayed recommendation feedback. We further introduce Memory Evolution Gain (MEG), which measures the utility improvement of evolved memory over a no memory baseline on strictly future tasks. TIDE treats memory as a capacity-constrained population of experiences: temporal and semantic credit assignment estimates contextual fitness, while responsibility credit distributes outcome signals according to the memories referenced during generation. These signals are then used to reinforce, crossover, mutate, or evict memories. On an e-commerce membership marketing content-generation agent, TIDE achieves a +7.75-percentage-point MEG in offline temporal replay and significantly improves both unique click-through rate (UCTR) and activation rate in an online A/B test. On a delayed-label benchmark, TIDE achieves the lowest mean absolute error (MAE) and root mean squared error (RMSE) and the highest MEG among the compared methods, demonstrating its effectiveness.
|
| 2498 |
Rational Clarification by Assistive Agents via Value-of-Information Reasoning
2609.37588
|
cs.AI
|
T. Duy Nguyen-Hien, Yee Whye Teh, Wee Sun Lee, Tan Zhi-Xuan |
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most saf...Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user's intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks --- ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) --- we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.
|
| 2499 |
FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents
2609.37590
|
cs.AI
|
Shantanu Dixit, Anson Bastos, Xuchao Zhang, Chetan Bansal, Saravan Rajmohan |
LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively ...LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.
|
| 2500 |
XU-RS: Explaining Credal Width in Random-Set Language Models
2609.37594
|
cs.AI
|
David Achara, Maryam Sultana, Alexander D. Rast, Fabio Cuzzolin |
Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this pr...Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained language models. These classifiers assign probability to individual answers and to groups of answers, producing lower and upper probabilities for each answer; The difference between these probabilities, called credal width, is used to represent epistemic uncertainty about an answer arising from limited training data. We propose XU-RS, a framework that attributes an answer's credal width to the input tokens (words or word pieces) supplied to a language model. XU-RS uses Expected Gradients (a standard feature attribution method) to estimate how input tokens contribute to credal width. The proposed framework is evaluated on a MedQA dataset using SmolLM3-3B and Llama-2-7B models, demonstrating that setting the embedding of a token ranked highly by XU-RS to zero (zero-masking) causes larger changes in credal width than zero-masking randomly selected tokens. In addition, we show that normalisation can cause other answer groups to influence an answer's width, reveal how token attribution can mask numerical errors, and provide diagnostic checks to verify whether a token ranked highly by XU-RS meaningfully explains model uncertainty.
|
| 2501 |
Flattening the Connectome Spectrum: A Spectral Filter for FC Induces a Pretraining Target for fMRI Encoders
2609.37642
|
cs.AI
|
Giovanni Marraffini (UNITO), Victoria Shevchenko (UNITO), Carlo Alberto Barbano (UNITO), Demian Wassermann (MIND) |
Self-supervised pretraining reshaped prediction in language and vision, and brain foundation models (BFMs) inherited its promise. Representations learned from large unlabelled corpora should capture individual functional dynamics and generalise across cohorts....Self-supervised pretraining reshaped prediction in language and vision, and brain foundation models (BFMs) inherited its promise. Representations learned from large unlabelled corpora should capture individual functional dynamics and generalise across cohorts. However, kernel ridge regression (KRR) fitted on functional connectivity (FC) matrices still predicts individual phenotypes more accurately than any BFM we tested. In this paper, we show that KRR is weighted by the eigenvalues of the FC which are miscalibrated for phenotype prediction. We apply an efficient spectral filter to recalibrate the eigenvalues of each subject's FC matrix, enabling the model to exploit more inter-individual variance. Across the 5 datasets, 11 parcellations and 6 prediction targets we tested, we match or exceed the KRR baseline. Based on this finding, we then pretrain a small encoder model on about 4,000 hours of fMRI from 162 open datasets, whereby we align the pairwise similarities between the embeddings of recording snippets with those between the recalibrated connectomes. Our model performs on par with the best of the 6 published BFMs we tested while having an order of magnitude fewer parameters. Our encoder performs better than FC on short scans and in smaller cohorts, especially in fingerprinting. We release the pretrained model weights, the code and the pretraining data, preprocessed and parcellated.
|
| 2502 |
Beyond a single latent space: a dual-latent world model for long-horizon planning
2609.37644
|
cs.AI
|
Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen |
Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-La...Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation Learning with Weighted Rollout (LoRe), which supervises self-generated predictions at both levels. An analysis of recursive error propagation motivates exponential horizon weights with separate decay rates for the two temporal scales. During planning, the high-level model generates latent subgoals that the low-level model refines into actions for precise execution. We evaluate from-scratch Dual-WM on five goal-conditioned visual control tasks against the task-wise strongest baselines without actor-guided proposals. At goal offsets of 50 and 100 environment steps, mean success increases from 75.9% to 84.4% and from 61.4% to 69.5%, respectively. At offset 100, Dual-WM outperforms these baselines on all five tasks and improves mean success over LeWM by 30.8 percentage points. Ablations and supporting analyses provide evidence of more informative representations for goal evaluation and greater consistency under recursive prediction. These results highlight the value of separating temporal roles and training across multiple horizons for reliable latent planning. Our core implementation is available at https://github.com/DeLin1001/Dual-WM-Official.
|
| 2503 |
EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
2609.37658
|
cs.AI
|
Min Yang, Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong |
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as ...LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.
|
| 2504 |
KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora
2609.37673
|
cs.AI
|
Changmian Wang, Yuchao Ma, Xuchao Lu, Chen Zhang, Ping Sun |
Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to take. Routine work records often leave out this tacit knowledge, making it difficult for Large Language Model (L...Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to take. Routine work records often leave out this tacit knowledge, making it difficult for Large Language Model (LLM) agents to use professional experience effectively. We introduce KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction. It turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents. Six case elements preserve the task process: context, cues, judgment, action, boundaries, and outcomes. Nine-layer cognitive corpus construction organizes tacit experience along nine extraction dimensions and stores the resulting assets in six libraries: rules, constraints, best practices, negative examples, corner cases, and skills. Semantic alignment, individual experience distillation, organizational consolidation, and cross-review preserve source evidence, conditions of use, and unresolved disagreements. The platform packages these assets into callable skills with explicit inputs, steps, dependencies, and stopping conditions, connecting experience collection to task execution and evaluation feedback. Using authorized samples from 20 randomly selected practitioners, the platform processed 1,576 source files into 23,024 individual experience records and 13,113 organizational assets. The evaluation spans multiple professional domains. Under common task inputs and scoring criteria, the base model, raw corpus retrieval-augmented generation (RAG), and KUPAS MASTER agent scored 70.63, 79.75, and 89.58, respectively. The KUPAS MASTER agent improved on raw-corpus RAG in all seven scoring dimensions. The platform provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.
|
| 2505 |
Learning from Shared-Control Overrides: Context-Driven Acceleration Profile Prediction for Personalized Overtaking
2609.37684
|
cs.AI
|
Ruizheng Xu (Heudiasyc), Lounis Adouane (Heudiasyc), Javier Iba\~nez-Guzm\'an, Cl\'ement Zinoune |
Adaptive Cruise Control (ACC) systems are typically calibrated for an average driver, often resulting in a mismatch between vehicle behavior and individual expectations during time-critical maneuvers such as highway overtaking. When the ACC is perceived as too...Adaptive Cruise Control (ACC) systems are typically calibrated for an average driver, often resulting in a mismatch between vehicle behavior and individual expectations during time-critical maneuvers such as highway overtaking. When the ACC is perceived as too conservative and inconsistent, drivers intervene through throttle overrides, providing implicit feedback on the system's behavior. This paper reframes these override actions as human-in-theloop supervisory signals and proposes a data-driven framework for personalized vehicle adaptation, termed Context-driven Personalized ACC (CoP-ACC). Rather than relying solely on end-to-end regression, which tends to over-smooth dynamic responses, we introduce a hybrid pipeline combining: (i) unsupervised hierarchical clustering to extract representative acceleration profiles from override events; (ii) a context classifier that maps pre-maneuver driving conditions to the appropriate profile; and (iii) a residual regressor that refines the selected profile into a smooth, personalized acceleration profile tailored to the immediate context. Evaluated on real-world public-road data against a withheld forced-ACC baseline, the approach demonstrates high reconstruction fidelity and generates acceleration profiles that tend toward the driver's expected behavior in potential override contexts. The results highlight the potential of learning from shared-control overrides to enable anticipatory, personalized ACC behavior, reducing manual interventions and improving ride comfort.
|
| 2506 |
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?
2609.37686
|
cs.AI
|
Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun, Haoyang Li |
Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies ...Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.
|
| 2507 |
WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation
2609.37687
|
cs.AI
|
Muhammad Huzaifa, Lea Sch\"onherr, Thorsten Eisenhofer |
Active test-time adaptation (ATTA) improves robustness under distribution shift by updating a deployed model during inference while selectively querying supervision. However, most existing ATTA methods implicitly assume that supervision can be requested for ev...Active test-time adaptation (ATTA) improves robustness under distribution shift by updating a deployed model during inference while selectively querying supervision. However, most existing ATTA methods implicitly assume that supervision can be requested for every incoming test batch, which can incur substantial annotation cost over long test streams. In this work, we introduce \emph{budgeted ATTA} in which labels are available for only a fraction of test batches. This formulation shifts the central challenge from deciding \emph{what} to label within a batch to deciding \emph{when} supervision should be applied over time. To address this challenge, we propose a budget-aware approach \emph{WISE-ATTA} that allocates supervision over the test stream based on lightweight signals computed online, prioritizing periods where supervision is likely to be most useful. When a batch is selected for supervision, we further employ a drift-based sample selection criterion that targets samples exhibiting ongoing, unconverged adaptation dynamics, enabling effective updates from a single labeled example. We evaluate this approach on synthetic corruptions (ImageNet-C) and natural distribution shifts (ImageNet-R/K/A). Across settings, WISE-ATTA achieves competitive or improved performance compared to recent ATTA methods while requiring substantially fewer labels. Overall, we find that the timing of supervision is a key, yet underexplored, aspect of active test-time adaptation. Code: https://github.com/Muhammad-Huzaifaa/WISE-ATTA
|
| 2508 |
Locating Answer-Correctness Signals in Frozen Large Language Models
2609.37700
|
cs.AI
|
Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung |
Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under dis...Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search over hidden states, token probabilities, residual-stream features, attention, and their fusion, treating the selected readouts as a predictive measurement rather than a mechanistic localization. We run this analysis separately in closed-book and with-context settings, since context can change which readouts are informative. A consistent anatomy emerges: correctness concentrates in the answer span, recovered from the answer tokens even under retrieval, and the families carry it complementarily, so fusing them helps most out of distribution, where a single signal is weakest. The protocol is effective across two backbones and gates a retrieval controller as one downstream use.
|
| 2509 |
Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics
2609.37708
|
cs.AI
|
Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as, Chirag Raman |
Human social behaviour is not a collection of independent motions, but a jointly organised process in which group dynamics and individual variation continuously shape one another. Yet existing social motion models often prioritise plausible trajectories while ...Human social behaviour is not a collection of independent motions, but a jointly organised process in which group dynamics and individual variation continuously shape one another. Yet existing social motion models often prioritise plausible trajectories while leaving interaction state implicit, limiting their ability to transfer across groups, tasks, and partial-observation regimes. To address this gap, we introduce Bilevel Representations for Agent Interaction Dynamics (BRAID), a hierarchical sequential latent-variable model for generative multi-person interaction. BRAID explicitly formulates social motion generation as a meta-transfer learning problem: shared interaction priors are learned across datasets and adapted through arbitrary context sets of observed people and joints. The model represents each scene through a group-level latent state that captures shared interaction dynamics and person-level latent states that capture individual behaviour conditioned on the evolving group context. This modelling choice enables coherent generation under full, sparse, or partial observations while exposing compact social-state vectors that can serve as an interface for downstream embodied-agent systems. We evaluate BRAID under a unified SMPL-based representation on social forecasting, tracking and in-filling, and response generation, using metrics that assess not only reconstruction accuracy but also realism, diversity, temporal alignment, and interpersonal coordination. We further analyse the hierarchical latent space, showing that it captures separable group- and individual-level structure.
|
| 2510 |
Context Language Models
2609.37725
|
cs.AI
|
Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang |
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is mo...We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.
|
| 2511 |
Spatiotemporal Hyperedges for EEG Seizure Detection and Prediction
2609.37730
|
cs.AI
|
Hyunju Kim, Sheo Yon Jhin, Noseong Park, Nabil Imam |
Seizure detection and prediction from EEG are clinically important but challenging because seizures are rare, temporally localized, and propagate as coordinated events across multiple channels. Recent dynamic graph neural networks model this by running a tempo...Seizure detection and prediction from EEG are clinically important but challenging because seizures are rare, temporally localized, and propagate as coordinated events across multiple channels. Recent dynamic graph neural networks model this by running a temporal model over a sequence of per-time-step pairwise channel edges. However, this pairwise construction misses the spatiotemporal coupling that constitutes a seizure, at substantial training cost. We propose HyBrain, which summarizes spatiotemporal EEG evidence through a small set of soft hyperedges rather than pairwise edges. A per-channel Mamba backbone produces one token per (channel, second), and a spatiotemporal hyperedge block pools these tokens into E_h shared group embeddings through soft memberships and broadcasts them back. The same encoder serves three downstream tasks: window-based detection, one-second point-wise detection, and preictal seizure prediction. On TUSZ and CHB-MIT, HyBrain achieves the best AUROC on every reported setting against ten baselines, with the largest gap on long-clip preictal prediction. It also matches the most efficient baselines in training time and peak GPU memory. A qualitative analysis shows that even a single learned hyperedge cleanly captures the preictal -> ictal -> postictal trajectory on a real seizure clip.
|
| 2512 |
ContextRender: From Execution Dependencies to Agent Context
2609.37743
|
cs.AI
|
Savini Kashmira, Jayanaka L. Dantanarayana, Lingjia Tang, Jason Mars |
LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing contex...LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing context management methods can overlook how earlier tool results are used in subsequent execution, leaving needed information out of context. We introduce ContextRender, which manages context through a persistent graph of execution dependencies. We develop Tool-Flow Analysis to track how later operations reuse information from earlier tool results, providing a signal called observed reuse. A renderer combines this signal with recency and semantic relevance to select results within a fixed history budget, retaining omitted results for later use. Across AppWorld and 8-objective QA with three execution models, ContextRender outperforms the evaluated context management baselines using a 6K history budget, well below the models' maximum context windows. Within this budget, it achieves task performance close to or above that of passing the full history while reducing mean inference cost by 10.2%-32.2% relative to Full history. Ablations show that observed reuse improves task performance and retention of results reused later.
|
| 2513 |
Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models
2609.37751
|
cs.AI
|
Yury Nahshan, Nati Daniel, Jacob Goldberger, Yoli Shavit |
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, thes...Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-$K$ selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.
|
| 2514 |
OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells
2609.37773
|
cs.AI
|
Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li |
Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation l...Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.
|
| 2515 |
Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients
2609.37787
|
cs.AI
|
Ruinan Jin, Difei Cheng, Ling Chen, Jun Luo, Hao Zhou |
Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-s...Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether Adam converges on generalized smooth objectives under only second moment information on the stochastic gradients, without such concentration assumptions, was identified as an important open direction by Li et al. (2023). This paper gives an affirmative answer under fairly general conditions: such tail assumptions are not necessary. Building on the Adam self-normalization framework of Jin et al. (2026), developed for classical smoothness and bounded variance, we extend the stopping-time and de-preconditioning strategy to the $L_0$-$L_p$ generalized smoothness condition and a generalized second moment ABC condition. Even when the stochastic-gradient condition provides only second moment information that may grow along the trajectory, the stochastic trajectory of Adam remains in a locally well-behaved smoothness region, with stretched-exponential tail decay under bounded variance and global smoothness. Consequently, we establish high-probability convergence rate guarantees over the full range $p<2$, with confidence dependence of order $\delta^{-1/2}$, while the stepsize prefactor depends on $\delta$ only through a single logarithmic factor. We further construct a hard instance showing that, under only second-moment information, this $\delta^{-1/2}$-type confidence dependence is sharp. Finally, in the regime $p<1$, we combine the trajectory control with polynomial-growth estimates on rare events to obtain convergence rate guarantees in expectation.
|
| 2516 |
A neural network that maintains and retrieves memories based on context
2609.37791
|
cs.AI
|
Hayoung Song, JeongJun Park, Qihong Lu, Giacomo Vedovati, Monica D. Rosenberg |
Every day, people continuously infer situational context and adjust the way they understand and remember the world. Context, signaled by the prefrontal cortex, is known to modulate working memory and episodic memory, but the algorithmic understanding of this m...Every day, people continuously infer situational context and adjust the way they understand and remember the world. Context, signaled by the prefrontal cortex, is known to modulate working memory and episodic memory, but the algorithmic understanding of this modulation remains limited. Here, we train a recurrent neural network (RNN), augmented with an episodic memory buffer, to infer context using Bayesian inference as it continuously makes predictions of upcoming scenes while watching naturalistic movies. When the inferred context modulates the RNN's recurrent connectivity (the basis of working memory) in a low-rank manner, the model's activity patterns best match neural responses in human participants who watched the same movies during fMRI. Context also modulates episodic memory retrieval, such that the model retrieves memories based on not only content similarity but also context similarity. This is implemented as a key-value system with self-attention, designed to additionally encode context and retrieve context-congruent memories. The resulting model not only better resembles human brain representations but also learns to retrieve memories like humans much faster than a model without context modulation. Together, our findings suggest a computational mechanism by which context modulates information maintenance and long-term memory retrieval in naturalistic environments.
|
| 2517 |
DIET: Deletion-response Expert Trimming for Video Diffusion Transformers
2609.37829
|
cs.AI
|
Jiachang Zhang, Teng Hu, Bohao Feng, Songhang Shen, Wenqiang Wang |
Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert storage remains costly. Existing one-shot pruning criteria mainly rely on static activation or routing statistics ...Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert storage remains costly. Existing one-shot pruning criteria mainly rely on static activation or routing statistics and cannot capture layer-level re-routing after expert deletion. We introduce DIET, a training-free expert pruning framework based on deletion responses. A single all-expert calibration pass records expert outputs and router states for matched conditional and unconditional tokens. Candidate deletions are then replayed from cached tensors, requiring no additional model forward passes. The resulting deletion-response signatures characterize each expert by the changes induced when it is removed. DIET selects retained experts by minimizing Overall Diversity Loss (ODL), which preserves directional coverage in signature space, and combines intra-layer local search with an inter-layer regression-guided budget search to allocate experts across layers. On LingBot-Video 30B-A3B, pruning 50% of experts (6,144 to 3,072) reduces the checkpoint from 57 GB to 30 GB and enables single-card deployment on a 48 GB GPU without fine-tuning. Under a fixed 284-case VBench protocol, the VBench Total increases from 0.7941 to 0.8115. Across tested retention budgets, DIET consistently outperforms competitive pruning baselines adapted from large language models.
|
| 2518 |
Can a Cacheable Decision Model Follow Rules?
2609.37832
|
cs.AI
|
Dushyant Rajput (AltSlate Labs LLP), Nirdesh Chauhan (AltSlate Labs LLP), Siddharth Kosaraju (AltSlate Labs LLP) |
Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so ...Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate. We ask how much rule-sensitivity survives that move, and whether it can be trained back. Four experiments on Certo: (1) the tested conversion to cacheable scoring loses rule-sensitivity (recall@1 1.00 -> 0.24) while the joint scorer holds 1.00, and a shortlist+rerank rescue fails; (2) targeted counterfactual supervision restores strong performance on held-out synthetic rule tasks (paraphrase, counterfactual, composition; reproducible across seeds), though we do not isolate whether predictions depend on the supplied rule; (3) on real rules the added benefit is not established -- after fixing a truncation confound, the joint scorer wins significantly on the short tier (0.861 vs 0.500) and directionally on the hard tier (0.655 vs 0.483, n=29); (4) a matched cross-domain real-prose mixture did not help and reduced contract accuracy (-9.3, -16.2 points). A cacheable encoder can be made rule-sensitive on its training distribution, but transfer to unseen-source real rules is not established; the joint scorer keeps an edge at the cost of caching.
|
| 2519 |
Mixture of Self-Improving Branches For Agent Harness Optimization
2609.37834
|
cs.AI
|
Haoyu Dong, Yuhang Zhou, Zihao Lin, Yifan Wu, Bo Peng |
Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code ge...Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.
|
| 2520 |
Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability
2609.37857
|
cs.AI
|
Zhenting Huang, Junnan Liu, Qianren Mao, Zhixing Tan, Bo Jiang |
Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning...Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs via feature sensitivity. Experiments demonstrate that scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width\(\times k\) factorial experiment identifies the active budget k as the root cause: the degradation arises from the selection boundary rather than dictionary width alone. We attribute this failure to the geometry of TopK selection. The active margin, the distance to the cutoff, predicts feature loss without thresholds. Guided by this margin diagnosis, we introduce pairwise rank stabilization. Our method targets ordering failures at the cutoff and improves rare-feature sensitivity by \(8.83\) percentage points, while keeping reconstruction and alive-feature coverage near the baseline. Overall, our results suggest that wide TopK SAEs should be evaluated not only by reconstruction, sparsity, and feature count, but also by feature reliability under semantic variation and boundary geometry for stable interpretability.
|
| 2521 |
Co-PiLOT: Constrained Physics-Informed Latent Optimization for Target-Driven Inverse Design
2609.37875
|
cs.AI
|
Mahish K. Guru, Mayank Nagar, Ayush vyas, Jan Bohlen, Roland Aydin |
Inverse design of physical systems (molecules, devices, microstructures) often reduces to optimizing a high-dimensional structure against an expensive black-box simulator. Direct search is difficult because the space is non-Euclidean, feasibility is hard to en...Inverse design of physical systems (molecules, devices, microstructures) often reduces to optimizing a high-dimensional structure against an expensive black-box simulator. Direct search is difficult because the space is non-Euclidean, feasibility is hard to encode, and each evaluation is expensive. We present Co-PiLOT, a latent optimization approach that maps candidates through a generative encoder-decoder, uses the decoder as a learned validity prior, and searches the latent space with physics-informed black-box optimization. The framework is applied on the inverse design of magnesium alloy microstructure/texture. We develop a vision transformer based-encoder; paired with latent diffusion, diffusion transformer and rectified-flow transformer-based decoders on $\sim80{,}000$ EBSD-derived microstructure dataset to learn a minimal bottleneck, $z$. The ViT-FMDiT model ($z$=$768$) reconstructs high-fidelity microstructure images (FID $27.86$, MS-SSIM $0.178$), which our self-segmenting orientation codec converts into input grids for crystal plasticity solver. Finally, we introduce MERIDIAN, an active latent optimizer driven by deep-kernel Gaussian-process uncertainty, failure-aware feasibility prediction, manifold-aware trust regions, and target-aware acquisition. Within a budget of $160$ simulations, the ViT-FMDiT and MERIDIAN combination yields the best target-driven objective score, reducing the relative target error by $3$--$22\%$ against seven baselines (DANTE, TuRBO, BAxUS, CMA-ES, DDOM, SEIKO, DDPO) on the same decoder.
|
| 2522 |
Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL
2609.37898
|
cs.AI
|
Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng, Shuo Zhou |
Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate...Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early training stage where outcome rewards provide little learning signal. As the gap narrows and eventually reverses, however, continued distillation becomes less beneficial and may hinder further improvement. Motivated by this observation, we propose Gap-Adaptive Teacher Scheduling (GATS), which augments the student's RL objective with an OPD term whose weight adapts to the teacher-student performance gap. Specifically, GATS gradually reduces teacher guidance as the student approaches the teacher's reference performance and withdraws it once that reference is reached. This enables GATS to leverage task-trained teachers smaller than the student, since teacher guidance is primarily needed during early training. Across ALFWorld, WebShop, and ScienceWorld with three Qwen2.5 teacher-student configurations, GATS achieves the highest average success rate among the compared methods in all three configurations, improving over reward-only GRPO by 4.37%-11.87% under matched student rollout budgets. Code is available at https://github.com/Ricardo-H/guide-then-let-go.
|
| 2523 |
You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
2609.37902
|
cs.AI
|
Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan |
Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live end...Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does not reliably predict quality or availability; and provider feasibility is task-selective, with one deployment nearly normal on knowledge tasks but catastrophically degraded on multi-step reasoning. We formulate same-model provider selection as a price-taker market-aware routing problem. A simple measured-map policy routes to the cheapest provider that is both quality-equivalent and healthy, yielding matched-quality savings while avoiding degraded endpoints. Because the map drifts, we introduce FACET, an online provider router that certifies per-(provider x task) feasibility facets and fails safe to an anchor before serving uncertified arms. Across relaxed deployment assumptions, FACET tolerates imperfect task assignment and sparse feedback, while systematic evaluator bias exposes a quality-signal trust boundary that can be mitigated with ground-truth probes or audits. Live provider runs further confirm that certification can move real traffic from a premium anchor to a substantially cheaper certified endpoint. Our results suggest that market-aware LLM routing must measure not only which model to use, but also who serves it.
|
| 2524 |
GRFBrain: Graph-Structured Rectified Flows for EEG Dynamic Modeling
2609.37934
|
cs.AI
|
Haohui Jia, Zheng Chen, Jathurshan Pradeepkumar, Xu Cao, Yasuko Matsubara |
Forecasting time-varying functional connectivity from electroencephalography (EEG) requires modeling both history-dependent trends and structured variability across channels. Conditional flow matching provides a framework for distributional forecasting, yet it...Forecasting time-varying functional connectivity from electroencephalography (EEG) requires modeling both history-dependent trends and structured variability across channels. Conditional flow matching provides a framework for distributional forecasting, yet it remains unclear whether graph-informed source distributions offer practical advantages over isotropic noise and strong deterministic predictors. We introduce a graph-structured residual flow framework that separates conditional mean prediction from stochastic residual transport. A history-only predictor estimates the future connectivity graph, while a graph Gaussian source encodes dependencies derived from past connectivity through a Laplacian-based covariance. A conditional velocity field transports source samples to future graph residuals, with transport time explicitly distinguished from physical EEG time. Our study identifies the conditions and controls needed to distinguish useful residual transport from improvements attributable to deterministic prediction, learned representations, and sampling effects.
|
| 2525 |
Topological Coherence for Self-evolving Multi-agent Systems
2609.37953
|
cs.AI
|
Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Xinyu He |
Complex tasks inherently couple workflow structure, agent responsibility, collaboration, and memory access: task regions delimit responsibility and tool scope, cross-region dependencies give rise to handoffs, and ownership boundaries delimit private and select...Complex tasks inherently couple workflow structure, agent responsibility, collaboration, and memory access: task regions delimit responsibility and tool scope, cross-region dependencies give rise to handoffs, and ownership boundaries delimit private and selectively shared memory. Existing methods can jointly optimize agent and communication structures, yet such optimization does not by itself require responsibility, handoff, and memory boundaries to remain consistent with task dependencies. We term this requirement topological coherence. We introduce TOCOMAS, a Topology-Coherent Multi-Agent System. TOCOMAS grounds a task graph in tool interfaces, organizes compatible task nodes into reusable responsibility domains, and derives dependency-induced and profile-conditioned collaboration together with boundary-regulated memory visibility. During online self-evolution, TOCOMAS proposes coupled changes to agent, collaboration, and memory policies, retaining for subsequent tasks only candidates that satisfy structural constraints and improve evaluated reward. Across BBEH, WorkBench, SWE-Bench-Verified, and CoMemBench, TOCOMAS improves task success over baselines across backbones. CoMemBench also shows gains over the self-evolving baseline in verified progress, handoffs, and memory isolation.
|
| 2526 |
BrainNet Studio: A Unified Toolkit for Brain Network Construction, Intelligent Analysis, and Visualization
2609.37956
|
cs.AI
|
Xiwei Zeng, Shengrong Li, Yiheng Liu, Chunwei Tian, Daoqiang Zhang |
Brain networks characterize structural and functional relationships among brain regions and support research on cognition, brain disorders, and brain-computer interfaces. Their time-varying topology and higher-order spatiotemporal dependencies are not adequate...Brain networks characterize structural and functional relationships among brain regions and support research on cognition, brain disorders, and brain-computer interfaces. Their time-varying topology and higher-order spatiotemporal dependencies are not adequately represented by conventional static networks. Existing tools primarily focus on static connectomes and provide limited integration of dynamic network modeling with modern graph and sequence learning methods. We present BrainNet Studio, an integrated toolkit for static and dynamic brain network analysis. It provides a unified workflow encompassing network construction, feature extraction, predictive modeling, candidate biomarker identification, visualization, and assisted interpretation. The toolkit integrates 27 algorithms, including deep learning, graph neural networks, and spatiotemporal sequence models, to support classification and the identification of discriminative brain regions and connections. A large language model generates researcher-verifiable summaries of functional connectivity, structural connectivity, and structure-function coupling at individual and group levels. Within a consistent computational framework, users can configure analytical tasks, compare methods, inspect outputs, and extend functionality without repeatedly assembling application-specific pipelines. BrainNet Studio provides a practical and extensible platform for connectome analysis in cognitive neuroscience, exploratory studies of brain disorders, and brain-computer interfaces. The toolkit is publicly available at https://github.com/xbrainnet/Brainnet-Studio.
|
| 2527 |
SelfSearch: Reward-Free Search for Self-Improving Agents
2609.37968
|
cs.AI
|
Jungwoo Yang, In Jin Kong, Yohan Jo |
Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs ...Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.
|
| 2528 |
KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
2609.37988
|
cs.AI
|
Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier |
As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recen...As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache across layers (depth), caching at fewer bits (precision), or truncating the low-rank latent cache representations (rank). We call the resulting method KV-Kaizen, for the many small per-layer choices it compounds. We observe that these interventions taken independently and uniformly over all layers limit achievable compression because they degrade accuracy. Crucially, composing them locally and adaptively to the context can instead preserve accuracy while achieving large memory savings. At inference, the selector runs once, before pre-fill. In evaluations on instruction following and reasoning tasks, our selectors reach the Pareto frontier of accuracy against cache size, against learning-free and post-hoc baselines. On long-context tasks, KV-Kaizen improves on eviction and can be composed with it, reaching a 32x smaller decode-time cache on a 14B model while preserving accuracy. A 4x cache size reduction incurs no accuracy degradation from 7B parameters up, and a compressed model is more accurate than a smaller uncompressed one with the same cache size. Together, these findings support pre-training large models and compressing them only afterwards.
|
| 2529 |
Which Attention Heads are like the Human Head? Not the Ones that Compute
2609.37991
|
cs.AI
|
Christopher Pinier, Gustaw Opie{\l}ka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez |
Brain-AI alignment is often interpreted as a sign that model and brain perform similar computations. Whether the aligned units are causally involved in model computation is rarely checked. On an abstract pattern-completion task (AAABAAA $\rightarrow$ B), we co...Brain-AI alignment is often interpreted as a sign that model and brain perform similar computations. Whether the aligned units are causally involved in model computation is rarely checked. On an abstract pattern-completion task (AAABAAA $\rightarrow$ B), we compare LLM attention-head representations with human EEG and test how ablating those heads affects task performance. Alignment and causation dissociate: brain-aligned heads contribute to performance, but their removal is substantially less disruptive than removal of heads selected via attribution patching. We compare two head sets that prior interpretability work defines without reference to the brain: concept vectors (CVs), which represent abstract patterns across formats, and function vectors (FVs), selected for their contribution to correct-answer prediction. Brain alignment shows little association with FV scores, while its association with CV scores varies across models. Among brain-aligned heads, we find recurring attention profiles: one emphasizes distinctive elements (novelty heads), the other repeating elements (repetition heads). The novelty family tracks salience and attends to the same elements that humans look at, yet its removal is less damaging than random ablation on average. Repetition heads contribute modestly to performance and are associated with abstract-pattern representation (CVs). Across 17 models spanning 3B-72B parameters, FV-ranked removal is substantially more disruptive than brain-ranked removal. Brain alignment thus captures how the model reads the stimulus, and only faintly captures how it represents the pattern and solves the task.
|
| 2530 |
Diagnosing and Improving Probabilistic Reasoning in Large Language Models
2609.38005
|
cs.AI
|
Huaman Sun, Dingcheng Wang, Jason Hartline, Jessica Hullman |
Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two compon...Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL interventions targeting beliefs, decisions, or both improve these components across three domains, whether improvements transfer across components and elicitation formats, and whether decision performance can improve without improvement in belief formation. We find that targeting one component of probabilistic reasoning redistributes decision loss, improving the target without necessarily transferring to others, and that jointly targeting belief formation and decision-making improves both but hinges on matched formats between training and evaluation.
|
| 2531 |
HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment
2609.38006
|
cs.AI
|
Kenan Alkiek, Moontae Lee, David Jurgens, V. G. Vinod Vydiswaran |
Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy an...Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are safe to deliver (safety). We show that both decisions can be made from the model's own hidden states. The prefill state, computed before any token is generated, predicts whether the model will answer correctly, and the answer state, at the end of the generated answer, predicts whether that answer is correct. HARISSA fine-tunes the model so that both states predict correctness, then makes both decisions with one policy that cascades through the ways of answering from cheapest to most expensive, skipping a way the prefill state predicts will fail and deferring the query when the answer it stops with is predicted wrong. On a device running a single model, HARISSA is within one accuracy point of chain-of-thought at 2.7 times lower latency. On a server holding four sizes of one model, HARISSA is more accurate than the FrugalGPT and Self-REF cascades at the same latency, and at the same deferral rate the answer state leaves fewer wrong answers than the standard confidence signals in five of six task and setting pairs.
|
| 2532 |
PE-EK-PINN: Physics Embedding with Evolving Kernel for Scalable Physics-Informed Neural Networks
2609.38023
|
cs.AI
|
Huiwen Zhang, Feng Ye, Chu Ma |
Physics-Informed Neural Networks (PINNs) embed governing equations into deep learning, but enforce them only through loss residuals, leaving highly oscillatory wave behavior to be discovered by optimization. As a result, methods that achieve relative $L_2$ err...Physics-Informed Neural Networks (PINNs) embed governing equations into deep learning, but enforce them only through loss residuals, leaving highly oscillatory wave behavior to be discovered by optimization. As a result, methods that achieve relative $L_2$ errors below $10^{-3}$ on standard manufactured Helmholtz benchmarks can fail on practical radiation problems involving singular excitations, absorbing boundaries, and wave fields spanning tens of wavelengths. Architectural physics embedding addresses this limitation by factorizing the field into analytically derived oscillatory kernels and learnable envelopes. However, the kernel dictionary must be manually constructed and scales with the number of elementary units, growing exponentially with the depth of hierarchically structured systems such as antenna arrays and metasurfaces. We propose PE-EK-PINN (Physics Embedded with Evolving Kernels), which treats physics kernels as reusable learned representations rather than fixed analytical inputs. A converged subsystem field is frozen and promoted to an evolved kernel, whose transformed copies are reused to represent higher-level configurations without deriving new governing equations. The resulting hierarchy makes the peak number of active kernels independent of system size and reduces cumulative training cost from $O(N)$ to $O(\log N)$. Experiments on dipole arrays, composite line-source geometries, and cross arrays demonstrate the dramatic training cost reduction, while achieving a reduced or comparable relative $L_2$ error. One notable example is PE-EK-PINN solves a $256$-dipole array more than 30 times faster than direct PE-PINN.
|
| 2533 |
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
2609.38024
|
cs.AI
|
Jaewon Chu, Ji Soo Lee, Jihwan Park, Dohwan Ko, Jeehye Na |
An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly availab...An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose \textbf{Retrieval-Augmented Skill Optimization (RASO)}, a framework that leverages an external skill corpus as prior knowledge throughout skill optimization. RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness. RASO comprises two complementary stages: \textbf{Retrieval-Augmented Skill Initialization (RASI)} constructs a knowledge-grounded initial skill without requiring agent rollouts, while \textbf{Retrieval-Augmented Skill Update (RASU)} iteratively refines the skill by retrieving external knowledge guided by execution feedback. Across four agent benchmarks and two models, extensive experiments show that RASO consistently outperforms baselines without retrieval-augmented skill initialization and updating.
|
| 2534 |
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
2609.38043
|
cs.AI
|
Ashish Jain, Armaan Sandhu |
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do...Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.
|
| 2535 |
Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
2609.38070
|
cs.AI
|
Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang |
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confi...As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at https://github.com/szu-tera/DTC.git.
|
| 2536 |
Character Training for Risk-Averse Agents
2609.38093
|
cs.AI
|
Arav Dhoot, Punya Syon Pandey, Jamie Johnson, Daniel Tan, Elliott Thornley |
Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be ris...Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents.
|
| 2537 |
NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning
2609.38098
|
cs.AI
|
Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao |
Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the ...Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.
|
| 2538 |
Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
2609.38108
|
cs.AI
|
Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos L\'opez de Prado, Shadab Khan |
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it fa...Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.
|
| 2539 |
Stochastic World Models for Verifying Vision-Based Neural Feedback Systems
2609.38120
|
cs.AI
|
I. Samuel Akinwande, Mykel J. Kochenderfer, Clark Barrett |
Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analysis. Generative adversarial networks (GAN...Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analysis. Generative adversarial networks (GANs) have served as perception surrogates, but they are large, reproduce complex scenes poorly, and are hard to verify. We explore stochastic world models as a richer class of perception surrogates. We train a world model with physically grounded latents, built from operations that standard verifiers bound. It reproduces held-out frames more faithfully than GAN surrogates with up to 130 times as many parameters. To verify these surrogates, we develop a procedure that combines falsification, adaptive refinement, symbolic, and backward analyses. On an emergency braking benchmark with a GAN surrogate, our procedure resolves the entire state space, 38% of which the state-of-the-art verifier left unresolved. On the RGB version of the benchmark, where no verification results have previously been reported, our procedure resolves over 80% of the state space with a world model surrogate.
|
| 2540 |
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
2609.38142
|
cs.AI
|
Rishabh Agrawal, Hejie Cui, Shasha Li, Shanchan Wu, Sercan \"{O}. Ar{\i}k |
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need ...A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.
|
| 2541 |
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
2609.38143
|
cs.AI
|
Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang, Heng Ji |
Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the ...Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.
|
| 2542 |
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
2609.38147
|
cs.AI
|
Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu |
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic m...As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.
|
| 2543 |
Bathtubs, Boundaries, and Sandboxes: AI Regulatory Learning under Legal Uncertainty
2601.04094
|
cs.AI
|
Tom Deckenbrunnen, Alessio Buscemi, Marco Almada, Alfredo Capozucca, German Castignani |
Effective regulation of AI is a defining policy challenge, driven by their integration into all aspects of society. To remain responsive to their rapid development and emergent properties, policymakers across the globe rely on high-level principles and abstrac...Effective regulation of AI is a defining policy challenge, driven by their integration into all aspects of society. To remain responsive to their rapid development and emergent properties, policymakers across the globe rely on high-level principles and abstract legal requirements. Yet, while this flexibility supports future-proofing human-centred regulations and aligning them with socio-ethical values, it also causes legal uncertainty downstream as developers, companies, and auditors struggle with translating these abstract requirements into verifiable technical requirements. Using the AI Act as an example, this paper draws on Coleman's bathtub to analyse the regulatory learning space in AI governance. It argues that legal uncertainty cannot be fully reduced ex ante and that, within reasonable bounds, it is also necessary for regulatory learning because it creates the space in which boundary negotiation over socio-technical meaning can occur. Building on this analysis, the paper shows how boundary objects and boundary negotiating artifacts help explain the translation of legal requirements into operational practice. By examining technical sandbox frameworks, it further identifies concrete properties that technical infrastructures must possess to function effectively as boundary negotiation artifacts in AI assessment. The paper concludes that legal certainty remains the long-term aim, but that premature closure of regulatory instruments risks undermining the learning processes needed for adaptive governance.
|
| 2544 |
Sage: Formalization with Semantic Correction
2609.35790
|
cs.AI
|
Thomas Hirtz, Farzad Jafarrahmani, Abdelmouksit Sagueni, Xiang Zhou, Wengping Deng |
While neural theorem provers have achieved impressive milestones in formal mathematics, they largely operate on the assumption that faithful Lean 4 formal statements are already provided. Translating informal natural language into a formal language is a critic...While neural theorem provers have achieved impressive milestones in formal mathematics, they largely operate on the assumption that faithful Lean 4 formal statements are already provided. Translating informal natural language into a formal language is a critical data bottleneck plagued by an "illusion of rigor": standard type-checkers accept statements that compile but drop hypotheses, introduce vacuous truths, or subtly alter mathematical bounds. To resolve this, we introduce Sage (Semantic Agent-Guided Formalization Engine), an agentic framework that replaces monolithic translation with a four-stage decomposed generation pipeline coupled with a dual-signal semantic correction loop. By pairing Lean 4 compiler diagnostics with multi-dimensional semantic feedback, our correction loop enforces mathematical fidelity alongside syntactic validity. By explicitly accounting for the gap between open-ended queries and declarative formal targets, our pipeline prevents models from achieving high formalization rates by guessing unverified answers (exhibiting a 70.9% answer leakage rate). Consequently, Sage suppresses leakage to 2.7% while achieving 73.3% pass@4 joint compilation and semantic fidelity on the Omni-MATH without proofs (compared to 42.0% for a fine-tuned Goedel-Formalizer-V2 baseline). Finally, on IMO-Unformalized, a novel frontier of 175 unformalized International Mathematical Olympiad problems, Sage demonstrates effective zero-shot generalization with 87.4% pass@4 verified fidelity compared to just 19.4% for the baseline, winning over 79% of blind pairwise evaluations.
|
| 2545 |
Calibration-First Cross-Cohort Multimodal Temporal Learning for Transferable Asthma-Risk Forecasting
2609.35795
|
cs.AI
|
Taimoor Ahmad |
Asthma deterioration forecasting must remain reli- able when patient populations, sensor ecosystems, and available modalities change across cohorts. Existing models commonly optimize within-cohort discrimination and may produce poorly calibrated probabilities ...Asthma deterioration forecasting must remain reli- able when patient populations, sensor ecosystems, and available modalities change across cohorts. Existing models commonly optimize within-cohort discrimination and may produce poorly calibrated probabilities after transfer. We present CALIBRA, a calibration-first multimodal temporal framework for short- horizon risk prediction with incomplete data. Dedicated recurrent encoders process environmental, pulmonary, symptom, medication, wearable, and context streams; a reliability-conditioned gate suppresses stale or absent modalities, while gradient-reversal training discourages avoidable cohort signatures. A shrinkage- based hierarchical logistic layer calibrates probabilities using a patient-disjoint target subset, and split conformal prediction provides abstention-capable prediction sets. To avoid fabricating clinical evidence, we evaluate the complete implementation on a documented three-cohort semi-synthetic benchmark with controlled distribution shift, informative missingness, and sealed target patients. Across five configured seeds, CALIBRA achieved mean target-test AUPRC 0.224 versus 0.240 for the strongest non-ablation comparator, TemporalTransformer; mean AUROC was 0.717, and Brier score was 0.098. Experiments additionally assess complete-modality failures, calibration, conformal coverage, decision curves, subgroup behavior, ablations, runtime, and parameter count. The results verify the method and reproducible pipeline under controlled shift, but do not establish clinical effectiveness. External validation on harmonized real asthma. Overall this artifact provides evidence for carefully governed real-cohort validation.
|
| 2546 |
Binarization Flattens the Score Space
2609.35797
|
cs.AI
|
Jacob Cole |
Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met...Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met each criterion. We model each pass/fail verdict as a score on an unreported scale, compared with one cutoff. A stretch of that scale moves every score proportionally toward or away from the cutoff, but never across it, so no verdict changes. A policy is therefore free to apply any stretch without changing anything the panel reports. Under a joint-Gaussian model, a third grade adds a second threshold and removes this affine stretch ambiguity. On MATH and SciBench outputs from one seven-criterion judge, all 14 constructed criterionwise stretches were invisible after binarization but visible with three grades. At $n=1{,}024$, a test given both population laws had at least 96.5% power at a $1.5\times$ stress. Retaining grades closes one blind spot created by binarization, but verdicts alone remain insufficient as some changes are still indistinguishable from genuine improvement. These include arbitrary within-grade changes and fixed-covariance, loading-aligned mean shifts -- the signature of a sycophancy-shaped lift the panel reads as competence. The shared-factor reference approximation fit MATH and SciBench but not HealthBench, delineating its empirical scope. We recommend keeping at least three grades (for example, asking the judge whether each criterion is fully, partially, or not met and rewarding {0, 0.5, 1}), and externally validating gains along the remaining direction, which no finer scale removes.
|
| 2547 |
Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models
2609.35804
|
cs.AI
|
Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann |
Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raise...Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, particularly regarding bias and hallucination. In this work, we evaluate the robustness of LLMs to perturbed variations of the original inquiry in decision-making tasks. We show that contrary to previous studies, perturbations can mitigate bias and hallucination in some LLMs over other models. It's found that Claude 3 is more effective for the tasks represented in most datasets, whereas models like GPT3.5 exhibit varying levels of adequacy, performing comparably in some cases but falling significantly behind in others. These insights are crucial for understanding the practical implications of deploying LLM-based assistants as effective decision-support tools in real-world applications, emphasising the need for rigorous testing and validation to ensure reliability and effectiveness. This study contributes to the growing body of research on LLM evaluation and provides insights for developing more robust and trustworthy AI assistants in critical decision-making contexts.
|
| 2548 |
Alignment Forecasting: Predicting Misalignment From Training Data
2609.35805
|
cs.AI
|
Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak |
Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. T...Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes. Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode's base rate and the target model's prior tendency. This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear. More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT setting.
|
| 2549 |
From Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework
2609.35806
|
cs.AI
|
Ranuga Disansa, U. S. Samarasinghe, Lasith Gunawardena |
Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA) captures exactly this dimens...Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA) captures exactly this dimension, defining 147 professional skills across seven responsibility levels, but no automated LLM-based extraction targeting SFIA has been reported. We formalize the task as structured prediction of (skill, level) pairs from free text and ask three questions: how accurately can text be mapped onto SFIA's closed vocabulary, which strategies reliably predict the level alongside the skill, and do agentic designs improve on simpler retrieval and prompting? We evaluate five strategies (a lexical baseline, dense retrieval with LLM reranking, a zero-shot schema-constrained LLM, single-agent agentic RAG, and a three-agent retriever--matcher--verifier crew) against expert-mapped European ICT role profiles, all drawing on an SFIA~9 corpus built by a fully automated agentic pipeline that we release. Retrieval-based matching identifies the most skills while generative strategies are markedly more precise; only strategies assigning the level as an explicit decision predict it reliably, with similarity-based selection more than twice as inaccurate; and the crew doubles latency without improving accuracy, so added agent roles do not automatically benefit closed-taxonomy matching. These results provide the first reproducible baseline for structured, level-aware skill extraction against SFIA.
|
| 2550 |
Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety
2609.35807
|
cs.AI
|
Charlie Summers, Prajwal Raghunath, Aaditya Pai, Mayur Kulkarni, Zhuo Zhang |
LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without h...LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover. We argue that the execution environment should instead enforce safety as the agent runs and steer it toward safe alternatives when violations occur---we call this Environment Steering. We implement this by modeling the agent and harness execution state as database tables, track the record-level data flows, and check these data flows against declarative policies during runtime. When violations are detected, policy- and context-specific feedback steers the agent toward safe trajectories. On AgentDyn, this enables the agent to improve task success rate over no-defense while achieving 0% attack success rate.
|
| 2551 |
When Successful Memories Mislead Embodied Agents:Memory Adaption For Task-Conditioned Execution
2609.35808
|
cs.AI
|
Quanquan Li, Hongbo Zhang, Yihe Chi, Liuyang Song, Jingyu Li |
Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current execution context. Existing memory systems pri marily optimize construction and retrieval; semantic relevance and ...Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current execution context. Existing memory systems pri marily optimize construction and retrieval; semantic relevance and historical success therefore remain insufficient when retrieved ex perience contains incompatible actions or an inappropriate level of structure. We introduce Memory Adaptation for Task-Conditioned Execution (MATE), a deterministic post-retrieval procedure that converts trajectories into execution-oriented memory. MATE re moves obsolete control context, extracts condition-action-effect transitions, applies verified action normalization, selects a task dependent representation, and serializes the result under a fixed budget without additional LLM inference. On 134 ALFWorld tasks, MATE achieves task success rates of 81.3% and 93.3% with Qwen2.5-14B and 72B while using approximately one-tenth of the tokens required by raw trajectories. Controlled comparisons show that verified action normalization is the principal mechanism by which MATE restores the utility of retrieved experience, support ing memory adaptation as a distinct stage between retrieval and embodied execution.
|
| 2552 |
Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?
2609.35809
|
cs.AI
|
Jiyao Yang, Yang Liu, Zhenyue Qin, Qingyu Chen, Xiuzhen Zhang |
The rapid advancement of generative AI raises concerns about the misuse of Multimodal LLMs (MLLMs) for large-scale disinformation campaigns on social media. Despite existing research on textual disinformation, a fundamental question remains unanswered: can MLL...The rapid advancement of generative AI raises concerns about the misuse of Multimodal LLMs (MLLMs) for large-scale disinformation campaigns on social media. Despite existing research on textual disinformation, a fundamental question remains unanswered: can MLLMs be exploited to fabricate realistic multimodal fake news, and can they reliably detect it? We introduce a multi-agent framework in which a story agent, an image agent, and a critic agent collaborate to produce fake social media posts that plausibly counter true news. We apply the framework to generate over 9,000 paired multimodal news posts across science, health, and entertainment domains, and benchmark 16 open- and closed-source MLLMs for automated detection. We find that most models fall substantially short of human-level accuracy and fail critically on identifying image authenticity. Our research provides a foundation for developing robust defenses against social media fake news. Code and data are available at https: //github.com/xiuzhenzhang/Multimodal.
|
| 2553 |
TRACE: Deployable Tree-Relational Structure Enhancement for Oncology LLMs
2609.35810
|
cs.AI
|
Jizheng Lai, Yingyun Li, Ying Qin, Haiyang Qian |
Large language models are increasingly used in oncology applications, but their predictions are often weakly grounded in explicit medical structure. We present TRACE, a deployable tree-relational enhancement framework for oncology LLMs. TRACE separates expensi...Large language models are increasingly used in oncology applications, but their predictions are often weakly grounded in explicit medical structure. We present TRACE, a deployable tree-relational enhancement framework for oncology LLMs. TRACE separates expensive offline structure learning from lightweight online inference: oncology concepts and relations are organized into an updatable tree-relational structure, refined using LM-loss-derived evidence, and retrieved at inference time as compact prompt evidence. This design supports task-adaptive evidence selection without requiring supervised labels in the zero-shot setting. Across ten oncology classification tasks and one MedQuAD CancerGov QA benchmark, TRACE improves both label-free evaluation and supervised fine-tuning. Additional analyses show that TRACE improves over vanilla RAG and generic GraphRAG, remains useful under leakage-controlled METABRIC inputs, and produces interpretable evidence paths aligned with clinical reasoning. These results suggest that explicit, updatable medical structure is a practical path toward more accurate and auditable oncology LLM deployment.
|
| 2554 |
Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning
2609.35811
|
cs.AI
|
Zongze Wu, Yani Guo, Runnan Li |
Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based val...Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost of prohibitive latency. We propose Lookahead-R, a planning-based framework that reformulates tool retrieval as a resource-constrained sequential decision-making problem. At its core, Lookahead-R introduces a lightweight execution-aware surrogate world model that jointly predicts tool execution success, latency cost, and semantic utility---without invoking real APIs. This world model drives a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that navigates the tool space under strict budget constraints. Evaluated on the large-scale ToolBench benchmark, Lookahead-R achieves a superior accuracy-efficiency trade-off across all test scenarios. On the most challenging I3 split, it attains an NDCG@5 of 91.40\%, outperforming the state-of-the-art ToolGen (90.16\%) by 1.24\%. Ablation studies confirm that explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.
|
| 2555 |
Local Predictability and Collective Fidelity in LLM-Agent Societies
2609.35813
|
cs.AI
|
Igor Itkin |
Compact surrogates could reduce the cost of simulating large language model societies, but must reproduce collective behavior. We compare individual predictions and collective forecasts using 9,455 published trajectories and new experiments on opinion dynamics...Compact surrogates could reduce the cost of simulating large language model societies, but must reproduce collective behavior. We compare individual predictions and collective forecasts using 9,455 published trajectories and new experiments on opinion dynamics. Neighbor information improves individual prediction in all 16 public-data settings and pooled collective forecasts on held-out questions, although collective gains depend on transfer conditions. Tests on 24 new statements do not confirm earlier contrasting history effects in forecasts from the initial state. Qwen benefits from history after three observed rounds. These findings motivate direct collective validation, explicit limits on available observations, and comparisons with simple baselines.
|
| 2556 |
Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions
2609.35814
|
cs.AI
|
Xunjian Yin, Tianchen Guan, Jinao Wang, Weili Cao, Daisy Xinlei Lin |
As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. This makes difficulty expensive to refresh and difficult to control: when many aspects change at once, it is uncle...As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. This makes difficulty expensive to refresh and difficult to control: when many aspects change at once, it is unclear what actually makes a task challenging. We instead construct challenging instances from tasks agents already solve, turning difficulty into a programmable property of the environment. BreakingWeb pairs every base task with an intervention condition that preserves the user instruction, latent target, and backend success criterion while changing the environment at different web stack layers. Each intervention is deterministic, detectable, and recoverable, and is annotated with the cognitive primitive it primarily loads. The benchmark contains 519 clean/intervention task pairs across seven self-hosted websites and 29 intervention families, all graded against outcomes. We evaluate six strong browser-use agents, three GUI-only agents that see only screenshots, and humans. The construction is effective: interventions cut agent pass rate by 22.9% on average and overturn nearly half of the tasks each agent solves cleanly, whereas humans lose 10.0% on a first attempt and 5.7% after one familiarisation attempt. The dominant failure is belief failure: 75% of the six agents' failures end with a declared success although the required change never happened. Our code, data and environment are publicly available at www.breakingweb.app.
|
| 2557 |
How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats
2609.35815
|
cs.AI
|
Ian Arawjo |
Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address ...Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in several contributions. First, we find that running statistics over raw LLM judge scores leads to inflated false positives: counterintuitively, for many inter-rater agreement metrics, false positive risk peaks at "almost perfect" human-LLM agreement. To help researchers understand how to run statistics over LLM judges responsibly, we present guidance and tooling for the statistical analysis of mixed human-AI judge designs, and implement nine hypothesis tests via prediction-powered inference (PPI), including the first known PPI corrections for four rank-based tests (Wilcoxon signed-rank, Mann-Whitney U, and omnibus variants). To keep PPI++ stable with small human-labeled calibration sets, we introduce bootstrap-adaptive power tuning, which shrinks the estimated weight toward a target estimated from the labeled data, and accounts for that weight's own sampling variance. Second, through Monte Carlo simulations, we derive recommendations for what CI, p-value, and FWER correction methods to use for small-sample AI evaluations (N<100), and warn researchers against bootstrap CIs. We package these recommendations into evalstats, an open-source Python package that selects calibrated methods automatically, and demonstrate it in three scenarios, including one where a real LLM judge validated at "substantial agreement" would have led a researcher to publish a spurious finding. evalstats is publicly available at https://github.com/ianarawjo/evalstats.
|
| 2558 |
PrimeSeeker: Capability-Oriented Supervision for Deep Search Agents
2609.35816
|
cs.AI
|
Linzhi Peng, Hanting Chen, Heng Chang, Ke Cheng, Bowen Du |
Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additional hops, and longer trajectories. These global properties, however, are only indirect proxies for the local retr...Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additional hops, and longer trajectories. These global properties, however, are only indirect proxies for the local retrieval capabilities required during search. To address this mismatch, we introduce latent anchor reasoning, which consists of resolving an unnamed retrieval anchor from descriptive specifications and transferring the recovered anchor into a subsequent information demand. This primitive retrieval unit decomposes deep search into chains of coupled operations and organizes question construction around anchor resolution and relation transfer, without prescribing a canonical search path. Based on this formulation, we propose PrimeSeeker, a capability-oriented framework that constructs web-grounded anchor structures and jointly derives a question and a reference evidence skeleton. The skeleton preserves supporting evidence from construction and guides expert generation through extractive highlights of current tool observations. These highlights are removed before supervised fine-tuning, while the skeleton is subsequently reused to audit reference-step coverage for reinforcement-learning rewards. We construct 9,221 expert trajectories, training a 30B search agent. Across five deep-search benchmarks, PrimeSeeker achieves strong performance, while reference-step optimization further improves the supervised policy. The resulting trajectories exhibit low retrieval redundancy, and fixed-budget evaluation shows strong solution coverage with substantially fewer tool calls than long-horizon systems.
|
| 2559 |
Less Uniform Discrete Diffusion is More Powerful and Scalable
2609.35817
|
cs.AI
|
Kaibo Wang, Ding Ding, Fangyu Ding, Zijin Feng, Han Shi |
Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these,...Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework. Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling. Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation. We further continue-train a 7B autoregressive model into LUDI-7B, resulting in a UDLM capable of complex reasoning. It achieves a 3-token-per-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked.
|
| 2560 |
$\tau$-Multilingual: Benchmarking Voice Agents Across Languages
2609.35820
|
cs.AIcs.SDeess.AS
|
Soham Ray, Edgard dos Santos Paiva, Ruben Valenzuela, Karthik Narasimhan, Keshav Dhandhania |
English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce $\tau$-Multilingual, extending $\tau$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language a...English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce $\tau$-Multilingual, extending $\tau$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output. Across 4,500 full-duplex calls and five voice configurations, Spanish, Portuguese, and Hindi remain within 3.2 task-completion points of English, but Korean and Mandarin fall by 14.7 and 8.4 points. The failure modes also vary: Korean systems miss more responses, Mandarin systems interrupt more often, and both struggle with tools and entities. Grok leads task completion but scores lowest on generation quality, motivating separate task, interaction, and generation reporting. We release language packs, validated judges, and tools for community-built multilingual voice-agent evaluation.
|
| 2561 |
When Should LLMs Trust Their Own Revisions? A Risk-Aware Study of Intrinsic Self-Correction
2609.35832
|
cs.AI
|
Tianzhu Zhang |
Intrinsic self-correction asks a language model to revise its own answer without receiving new external evidence. A second pass can recover mistakes, but it can also overturn answers that were already correct. We study this trade-off across 29 open-weight LLMs...Intrinsic self-correction asks a language model to revise its own answer without receiving new external evidence. A second pass can recover mistakes, but it can also overturn answers that were already correct. We study this trade-off across 29 open-weight LLMs on BoolQ, GSM8K, and Corr2Cause by tracking correctness transitions between initial and revised answers. Aggregate accuracy can conceal substantially different revision behavior: for example, Llama-3.1-8B improves by 25.5 percentage points on GSM8K, while refinement changes 19.1% of initially correct answers into wrong ones. A controlled BoolQ study further shows that refinement prompts shift the balance between recovery and harm. We then compare three runtime choices: keeping the initial answer, always accepting the revision, and selectively invoking revision using signals available after the initial response. The comparison identifies settings where learned gating is useful and others where a simpler unconditional policy performs better. These results suggest treating intrinsic self-correction as a revision policy rather than as a uniformly beneficial second pass, and evaluating it through both the corrections it recovers and the errors it introduces.
|
| 2562 |
Estimation of Room Impulse Responses from Handclaps
2609.35839
|
cs.AIcs.SD
|
Shih-Yu Lai, Kyung Yun Lee, Nils Meyer-Kahlen, Eloi Moliner, Bing-Yu Chen |
Handclaps provide an equipment-free excitation for room acoustics, but their unknown and variable source waveform makes room impulse response (RIR) estimation challenging. In this work, we investigate whether RIRs can be estimated directly from handclaps. To t...Handclaps provide an equipment-free excitation for room acoustics, but their unknown and variable source waveform makes room impulse response (RIR) estimation challenging. In this work, we investigate whether RIRs can be estimated directly from handclaps. To this end, we introduce an anechoic handclap dataset containing 2,540 claps from 17 participants, designed to capture variability across natural claps and different hand configurations. We first establish the performance attainable when the excitation clap is known using regularized deconvolution, and show that approximating the unknown excitation by windowing the direct sound from the reverberant recording is insufficient. To estimate the RIR without a known excitation, we propose using the anechoic handclap recordings to train a deep neural network with a supervised regression objective. Evaluated on a controlled synthetic benchmark, the proposed neural regressor significantly outperforms windowing-based baselines across all instrumental metrics. Furthermore, we test the proposed method on handclap recordings measured in real acoustic spaces, showing that the inferred RIR spectra are consistent across different handclap measurements taken in the same room location. These results showcase the feasibility of directly estimating RIRs from natural handclaps without requiring knowledge of the excitation signal.
|
| 2563 |
Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models
2609.35841
|
cs.AI
|
Nils Kiele, Zainab Saad, Zirui Wang, Steve Drew, Samira Ebrahimi Kahou |
Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps ...Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-based approaches generate more realistic faults, most remain test-blind: The model sees only the source code and cannot reason about what existing tests already cover. We propose test-aware mutant generation, in which an LLM receives the problem statement, canonical solution and base tests in a single prompt, and must generate a nontrivial mutant that passes the base unit tests. We evaluate this approach across a set of five LLMs -- Gemini 3.1 Pro, Gemini 3 Flash, GPT 5.1 Codex Mini, GPT 4.1 Mini, Qwen3-32B -- on the HumanEval and MBPP benchmarks. The extended EvalPlus test suites serve as an automated oracle to verify whether surviving mutants represent genuine bugs. Test-aware prompting yields verified fault rates of 87.7% (HumanEval) and 79.1% (MBPP), meaning these mutants pass all base tests but are caught by the oracle. This vastly outperforms the matched test-blind prompting (which yields only 12.2% and 23.0%, respectively) and the traditional rule-based tool mutmut (4.4% and 5.7%). While fault subtlety (the fraction of extended tests a mutant fails) remains comparable across all three methods, test-awareness minimizes the computational cost per verified fault, compared to test-blind prompting. Exposing an LLM to existing unit tests shifts mutant generation from untargeted bug injection toward effective discovery of weaknesses in an existing test suite. Our work establishes a concrete foundation for future research to scale test-aware mutant generation to production-level environments.
|
| 2564 |
Position: Let's Strengthen Verifiability If We Can't Enforce Reproducibility
2609.35854
|
cs.AI
|
Samet Hicsonmez, Nermin Samet, Renaud Marlet |
In the field of Machine Learning, many papers contain empirical results supporting claimed statements or illustrating the performance of a proposed method. However, most practitioners know that (1) results are generally hard to reproduce, and increasingly so, ...In the field of Machine Learning, many papers contain empirical results supporting claimed statements or illustrating the performance of a proposed method. However, most practitioners know that (1) results are generally hard to reproduce, and increasingly so, (2) code is not often available to do so, and (3) it hinders the development of research. In this position paper, we analyze and quantify these issues, and make concrete proposals to improve result checkability, if not reproducibility. Code and supporting materials are available at https://github.com/giddyyupp/position-enforce-verifiability.
|
| 2565 |
Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution
2609.35855
|
cs.AI
|
Yubin Lyu, Fu Li, Jiawei Fei, Yang Zhao, Weixing Mei |
Optimizing deployed AI systems increasingly amounts to editing prompts, skills, harnesses, and code rather than model weights. Existing approaches commonly optimize these artifacts through propose-evaluate-select procedures, where candidate configurations are ...Optimizing deployed AI systems increasingly amounts to editing prompts, skills, harnesses, and code rather than model weights. Existing approaches commonly optimize these artifacts through propose-evaluate-select procedures, where candidate configurations are evaluated and only those meeting an acceptance criterion are selected. Yet our analysis shows that discarded candidates often contain information critical for subsequent optimization. Discarding them causes later proposals to revisit the same failure modes. We introduce Mara Chain, a refinement procedure that turns rejected candidates into stepping stones. Rather than discarding a rejected candidate, Mara Chain retains and iteratively refines it using evidence accumulated across preceding attempts. The procedure limits each refinement chain to a fixed depth and applies Pareto-filtered Top-N selection to bound the candidate pool. Across AppWorld skill optimization, TerminalBench 2.1 harness optimization, and MuSiQue retrieval-pipeline optimization, Mara Chain delivers greater task-performance gains with fewer rollouts. It outperforms GEPA, ACE, and SkillOpt-Lite by up to 20.5% in relative performance on AppWorld, reaching the target score with 65.5% fewer rollouts than GEPA. It improves the pass rate by 20.2 and 22.5 percentage points over AHE and Meta-Harness on TerminalBench 2.1, respectively, and improves MuSiQue test nDCG@10 and Recall@10 by 0.104 and 0.131 over a hand-written retrieval pipeline.
|
| 2566 |
Beyond Keywords: Leveraging Generative LLMs and Label Aggregation to Classify Economic Policy Uncertainty in News Articles
2609.35856
|
cs.AI
|
Paul Trust |
This research describes the adaptation of Large Language Models (LLMs) for economic monitoring in the public sector to automatically determine whether an article discusses Economic Policy Uncertanity (EPU) and to identify its specific type. Previous studies ei...This research describes the adaptation of Large Language Models (LLMs) for economic monitoring in the public sector to automatically determine whether an article discusses Economic Policy Uncertanity (EPU) and to identify its specific type. Previous studies either rely on keywords, which often result in a high count of false positives, or use machine learning approaches that require a large number of quality human labeled data that is costly and time consuming to acquire. In this study, we propose approaches based on weak supervision techniques, using generative LLMs to create synthetic labels through prompting, making the approach both cost-effective and scalable. Additionally, we propose methods for for multi-label and hierarchical classification of articles related to EPU.
|
| 2567 |
The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models
2609.35860
|
cs.AI
|
Pranav Darshan, Pranav A, Sravan Karthick T, Minal Moharir, Ivan P. Yamshchikov |
Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answ...Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of $0.35$ to $0.46$ AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ($|\rho|\approx0.94$ to $1.00$), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dispersion preserve the asymmetry, with bootstrap $95\%$ intervals excluding zero in all $12$ model and dataset settings. A stricter test using individual diffusion trajectories and no cross seed information preserves the asymmetry across all three LLaDA datasets ($p<0.005$) and directionally across all three Dream datasets, with one reaching significance. The hard regime varies substantially in prevalence across models ($16\%$ to $77\%$), and matched prompts frequently change regimes between models. These findings show that aggregate detection metrics conceal persistent, model dependent heterogeneity in language model failures and motivate regime conditioned evaluation.
|
| 2568 |
CruxBench: A Benchmark of Information Discovery
2609.35879
|
cs.AI
|
Hui Dai, Lina Piao, Nick Merrill, Nadja Flechner, Ezra Karger |
Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficu...Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question. CruxBench enjoys a rare combination of three key properties: it is (1) contamination-resistant by construction, since ground truth is generated by future world events; (2) open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer; and (3) grounded, with informativeness measured against quantified changes in real-world beliefs. We evaluate a diverse set of eight models on 293 target forecasting questions and find that VOI correlates highly with independent measures of model capability (r=0.90) and captures cruxes' usefulness for answering target questions. However, information discovery remains challenging even for frontier LLMs, which only narrowly outperform a random-timing baseline.
|
| 2569 |
Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money
2609.35886
|
cs.AI
|
Ankit Srivastava, Debjyoti Paul |
AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charg...AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it should, and no check keyed on identity will see it. We present three artefacts for measuring and reducing that loss. First, a taxonomy of agentic commerce fraud that separates five observation levels (agent reasoning, wire, settlement rail, counterparty, principal) from the request-level and history-level evidence available at each, and records which levels can observe which attacks. Second, Agentic Commerce Bench (ACB), a benchmark of twenty fraud classes generated from production aggregates, 1,647 catalogued service operations and 1,068 settlements, of which six involve a counterparty that is exactly who it claims to be. Third, gordonguard, an open-source detector stack and offline harness with which an operator can audit an agent configuration, replay hostile counterparties without an account, and run the same detectors inline. Calibrating to a stated false-positive budget on clean training traffic gives a 6.5% clean flag rate, replicated across three independent generations, and leaves eight of twenty classes no better than chance. On the four classes a reasoning layer can observe, a widely used agent security scanner run over its jailbreak-detection panel scores zero on all four, while correctly scoring 1.0 on a jailbreak supplied as a control. A measured median payment of $0.007 places a hard constraint on deployment: one human review costs 143 times the value of the payment it examines.
|
| 2570 |
SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents
2609.35889
|
cs.AI
|
Xiaoyu Xu, Zi Liang, Minxin Du, Qipeng Xie, Qingqing Ye |
Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counter...Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the requested output but add an effect forbidden by the task contract. We introduce SINGED (Source Integrity and the Nonidentifiability Gap in Execution Decisions for LLM Agents), a controlled benchmark covering five primary and two held-out task families. It varies displayed rank, evidence depth, decision policy, model release, and agent configuration, while task and process oracles verify the artifact and execution path. Across 7,549 audited trials, the randomized-rank study finds counterfeit execution in 45% (27/60) of rank-one trials and none at later ranks. Cross-candidate comparison eliminates shallow failures and reduces layered failures from 15.7% to 4.2%, but leaves dependency failures; its benefit is uncertain on unseen effects and public-package structures. Moreover, seven releases with no counterfeit executions when benign alternatives are available execute the counterfeit in 55/175 single-source cells after alternatives are removed. SINGED thus exposes a rank-, evidence-, and choice-sensitive outcome-to-execution gap: evaluation must connect correct outputs to execution paths.
|
| 2571 |
Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy
2609.35900
|
cs.AI
|
Yuehui Wang, Xinyu Qi, Guirong Xue, Cheng Wang, Yangbin Xie |
The integration of large language models (LLMs) into scientific workflows is accelerating, yet their ability to reconstruct the reasoning underlying published research remains unexplored. Papers specify explicit procedures while leaving many methodological dep...The integration of large language models (LLMs) into scientific workflows is accelerating, yet their ability to reconstruct the reasoning underlying published research remains unexplored. Papers specify explicit procedures while leaving many methodological dependencies-data selection, calibration corrections, priors, and domain assumptions-implicit. This ambiguity complicates the evaluation of LLM-based agents, since a failure to reproduce a result may reflect either limitations of the agent or underspecification in the source. We present a framework that evaluates agents through end-to-end reproduction, separating execution from verification and computational failure from methodological ambiguity. We apply it to fourteen astronomy studies: a case study from The Astrophysical Journal and thirteen papers published in Nature. Eleven of the thirteen contained an ambiguity preventing a uniquely specified reproduction path. In a controlled case study, twelve predefined paths, a 3x2x2 sensitivity analysis over sample definition, sky masking, and parallax zero-point treatment-gave estimates from 2.16 to 3.53 kpc for the same quantity, with only one recovering the published value (about 2.70 kpc). The published value was never used as an optimization target, selection criterion, or stopping condition; the matching path was found only after all twelve had run. Crucially, the decisive information (a +0.02 mas parallax zero-point correction) was already in the paper, but the agents did not recognize its causal relevance until the analysis made the effect visible. Matching a published outcome therefore does not validate reconstruction of the underlying reasoning, and the bottleneck is as often a failure to connect relevant information as to retrieve it. End-to-end reproduction thus serves both as a test of reproducibility and as a framework for evaluating implicit scientific knowledge in AI systems.
|
| 2572 |
Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning
2609.35908
|
cs.AI
|
Zihan Zhang, Shuangjie Yao, Zesen Liu, Zhixiang Zhang, Wai Ip Lai |
Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an...Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response under a query with high cosine similarity to benign requests. The vulnerability stems from a gap between retrieval similarity and answer validity. From an information-bottleneck perspective, query embeddings can lose information needed to distinguish valid from invalid cache hits, which limits any matching algorithm that uses only these embeddings. We propose a novel defense that recovers this necessary information from the raw text of the cache key. Across poisoning attacks, adversarial queries share a rewrite-residual structure: they pair a rewrite of the target query with residual content. The rewrite maintains high similarity, while the residual elicits the malicious response. Deleting the residual makes the remaining rewrite more similar to the incoming query. We exploit this structure using Deletion Gain to search shortened variants of the cached query for similarity gains, and an Answer Check to test whether the removed text contributes to the stored answer. We prove that Deletion Gain stays positive when a deletion leaves text close enough to the rewrite, and we search for such deletions with a sliding window. Across three poisoning attack classes, our defense blocks 82.0% to 98.2% of poisoned entries at a 5% false-positive rate, with negligible serving overhead.
|
| 2573 |
Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery
2609.35909
|
cs.AI
|
Kaikai Zhang, Zihan Zhang, Yuchong Xie, Zesen Liu, Shuangjie Yao |
Autonomous LLM agents turn vulnerability discovery into a repository-scale search: they generate many vulnerability hypotheses but can verify only a subset under a finite budget. We show that autonomous vulnerability discovery exhibits a hypothesis-verificatio...Autonomous LLM agents turn vulnerability discovery into a repository-scale search: they generate many vulnerability hypotheses but can verify only a subset under a finite budget. We show that autonomous vulnerability discovery exhibits a hypothesis-verification asymmetry, where verifying a candidate hypothesis through reachability analysis, execution, and proof-of-concept construction is substantially more expensive than forming it. Under a finite resource budget, this makes autonomous discovery a resource-bounded selective-verification process, further exposing verification effort as a unique defense surface. We present RedHerring, which inserts certifiably safe decoys that divert verification effort from real vulnerabilities. Each decoy combines a CVE-derived vulnerability chain that attracts verification with a false bridge that keeps its dangerous sink unreachable. A private certificate lets the defender verify this property efficiently, while establishing the same fact from the released repository requires solving a computationally hard problem. RedHerring further adapts each decoy to the target repository so that it reads as ordinary program logic. Across 33 OSS-Fuzz projects, 70 evaluation instances, and five models under matched budgets, RedHerring reduces real vulnerabilities discovered by 38.7-60.4%. Trajectory analysis shows that agents spend 30.6-51.5% of completion tokens and an estimated 32.5-49.9% of runtime verifying decoys, showing that RedHerring redirects a substantial fraction of the fixed search budget toward decoys. When explicitly informed that decoys may be present, the agent adapts its search strategy, yet RedHerring still reduces vulnerabilities discovered by 37.2% relative to an informed Baseline, showing that its effectiveness does not depend on decoy secrecy.
|
| 2574 |
Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents
2609.35911
|
cs.AI
|
Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang |
Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an on...Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametric framework that separates immediate use from persistent trust. It turns informative transitions into hypotheses that can guide the next step, while requiring prospective validation before reuse across episodes. Their predicted effects are checked against subsequent observations outside the source episodes, and only sufficiently supported hypotheses become available for persistent guidance. This process updates external knowledge while keeping all model parameters fixed. Over five rounds on WebArena-Lite and ALFWorld, StepLearn achieves average success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, respectively. It outperforms EvoTest, the strongest baseline, by 2.2-12.7 percentage points across the four settings. Learning dynamics further shows that these gains are not restricted to the final repetition, with advantages already present on first task attempts in most settings.
|
| 2575 |
MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?
2609.35912
|
cs.AI
|
Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng, Xiaohu Du |
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise maliciou...Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicated to end-to-end safety evaluation of image-borne attacks in multimodal skills. To instantiate this attack surface, we design Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images, such as annotations and interface labels. The accompanying SKILL.md provides auxiliary guidance toward relevant visual regions without explicitly stating the malicious operation. Built from 28 curated clean skills, MMSkillRisk contains 36 attack packages and 108 executable cases spanning five attack objectives, with separate checks for attack success and legitimate-task completion. Across nine model-harness configurations evaluated in isolated sandboxes, NCVA induces unauthorized operations in every configuration. Its pooled attack success rate (ASR) reaches 43.1%, exceeding the matched text-carrier baseline by 16.4 percentage points, with higher ASR in all nine configurations. Attack success and legitimate-task completion co-occur in 36.5% of cases, reaching 72.2% for GPT-5.6-sol with Codex. These results show that skill-bundled images can induce unauthorized actions even as agents complete legitimate tasks, so task success alone does not establish safe skill use. Our code and data are available at https://github.com/kaill-jlq/MMSkillRisk.
|
| 2576 |
UNBIND: UNlearning By INference-time Directional Steering for Code LLMs
2609.35913
|
cs.AI
|
Zhengyang Shan, Jiayun Xin, Yanjun Lin, Xu Qian, Zhiang Liu |
Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. ...Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patterns, creating a tension between forgetting specific implementations and preserving general programming ability. We propose \textbf{UNBIND}, a code unlearning framework that separately considers which hidden states correspond to the target code and how to suppress its reproduction. By constructing separate directions for these objectives, UNBIND achieves selective unlearning at inference time while keeping model weights fixed. Our evaluation covers fourteen baselines across two code models and two corpora. UNBIND achieves the highest joint forgetting and utility score in every setting. It reduces target code reproduction by 97.3\% to 99.1\% as measured by F-BLEU, with at most two fewer HumanEval+ and six fewer MBPP+ problems solved than the original models. In repeated extraction tests under a fixed budget, the number of targets yielding exact spans of at least 50 tokens falls from 188--262 to 0--2 out of 300 per setting. No extracted span reaches 100 tokens, and the mean best recovery ratio ranges from 0.43\% to 6.45\%. Multilingual and related-code evaluations further show effective forgetting with limited impact on useful programming capabilities, supporting UNBIND as a practical approach to selective code unlearning.
|
| 2577 |
Evaluating Name-Only Directory Routing for One-Shot Code Search
2609.35918
|
cs.AI
|
Manoj Bajaj |
Finding the right files is an early challenge for coding agents. We test whether a language model can follow directory and file names to find annotated code files missed by fixed lexical queries. Across 82 audited issues from 11 repositories at pinned pre-fix ...Finding the right files is an early challenge for coding agents. We test whether a language model can follow directory and file names to find annotated code files missed by fixed lexical queries. Across 82 audited issues from 11 repositories at pinned pre-fix commits, name-only directory routing recovered 0.465 of gold files within eight candidates, compared with 0.352 for FTS5 and 0.245 for a fixed full-issue rg query. The paired gain over FTS5 was 0.113 (95% repository-cluster bootstrap interval, 0.053 to 0.168). Under a shared 16K-token context budget, routing delivered 0.443 of annotated lines versus 0.246 for FTS5 on 55 cases with fully aligned annotations. At the same eight-file limit, combining routing with FTS5 reached 0.491 file recall, but its gain over routing alone was uncertain. An exploratory flat path control reached 0.572 recall while using 24.6 model calls per issue, compared with 8.9 for routing. Routing averaged 8.9 seconds per issue; FTS5 took 7 milliseconds per query after a 0.9-second build. On this cohort, directory routing added relevant file candidates to one-shot lexical search, but the study cannot attribute the gain to hierarchy or show that it improves issue resolution.
|
| 2578 |
Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change
2609.35922
|
cs.AIcs.SD
|
Bhavik Mangla |
A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is au...A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is audible, can change the right action. VoxParity tests whether agents act on it. In 183 scenarios from 14 sectors, one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper), and with it the correct typed tool call. A words-only null test credits a system only if hearing the call moves its actions more than it moves a pipeline that only reads the words. Only 11 of the 23 systems that can also be run on the transcript pass. Descriptively, errors run toward the words: when the audio calls for protection, all 28 systems carry out the routine request more often than they over-react on clean calls (41% against 12% pooled; the words-only pipeline, 58% against 15%). Exploratory analyses place most of the leading systems' misses on cues they heard; systems beat the null almost entirely on items that state the rule; the leading systems overrule heard resignation or confusion far more often than acute alarm; and, in the models tested, describing the voice and stating the rule each recover part of the shortfall, leaving a gap on emotion.
|
| 2579 |
Normative Loss Landscape Navigation: A Trajectory-Based Approach to Mitigating Forgetting in Incremental Learning
2609.35926
|
cs.AI
|
Isabelle Aguilar, Zayn Andre Zainal, Luis Fernando Herbozo Contreras, Zhaojing Huang, Omid Kavehei |
Continual learning models suffer from catastrophic forgetting when trained sequentially on non-stationary data distributions. Previously, this has been addressed through weight regularization. While preconditioning gradients offer a promising alternative to mi...Continual learning models suffer from catastrophic forgetting when trained sequentially on non-stationary data distributions. Previously, this has been addressed through weight regularization. While preconditioning gradients offer a promising alternative to mitigate forgetting, current approaches are myopic. Conversely, standard regularization methods apply rigid, scalar Euclidean penalties that entirely ignore the underlying Riemannian geometry of the parameter space. To overcome this gap, we propose TMLN (Trajectory-Modulatory Landscape Navigation), a normative navigation policy that formalizes continual learning as an optimal control problem over a curved loss landscape. TMLN utilizes a memory-efficient diagonal empirical Fisher Information Matrix (FIM) to define a localized Riemannian manifold. To compensate for the spatial limitations of the diagonal approximation, TMLN dynamically modulates a preconditioner using the normalized historical trajectory of the network's parameter values. By integrating this trajectory-based preconditioning directly into the gradient update, we actively shield historically critical parameter directions without relying on additive penalties. Empirical evaluations on class- and domain-incremental benchmarks demonstrate that our method significantly reduces the loss barrier between consecutive tasks.
|
| 2580 |
Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems
2609.35928
|
cs.AI
|
Xavier Del Giudice, Alessio Palma, Matteo Migliarini, Fabio Galasso, Indro Spinelli |
Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model family, the group splits into c...Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model family, the group splits into clusters, where agents prefer interacting with others carrying their same label, although nothing in the task rewards or asks for such a split. We argue that the label itself causes this split, which we define as $\textit{factionalism}$. We show and measure this phenomenon in two cooperative games and on a reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families. We further show that when the announced families are shuffled, or replaced by arbitrary labels, the factions still follow this information; when the label is removed, this behavior disappears. In strictly cooperative tasks, labeled groups spend on average $30\%$ more rounds and $55\%$ more tokens to reach a decision, and their success rate drops from $96\%$ to $81\%$. The effect replicates across tasks, group sizes and model families. Withholding identity labels from the agents is simple and effective mitigation.
|
| 2581 |
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
2609.35932
|
cs.AI
|
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu |
Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a seque...Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66 percentage points on three of four open-weight families, and the gap carries over to multi-turn agent tasks in AgentDojo. On Qwen3-8B the gap is 8 points, because without reserved ids the model still recognises the forged turn from its text by reasoning; suppressing the reasoning block widens the gap to 50. The authority sits in the single learned vector at the marker position: the mean of the marker's subword vectors does not reproduce it, the vector of the nearest ordinary token restores the attack on Llama-3.1, and an adaptive attacker who searches for non-reserved markers finds such embedding neighbours on three of four families. In every base and instruction-tuned pair we test, instruction tuning strengthens the model's preference for reserved markers. The standard mitigation, a tokenizer option that encodes special tokens as ordinary subwords, applies only to tokens a configuration declares special, so in 33 of 67 distinct tokenizer configurations, covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves intact the tool-protocol tokens through which agents read untrusted tool output, and the gap persists on that channel.
|
| 2582 |
Graph neural networks for sampling-invariant embeddings of organized signal sets
2609.35934
|
cs.AI
|
Martin Bauw (CMM), Santiago Velasco-Forero (CMM), Jesus Angulo (CMA) |
Sensor networks and radars can deliver signals as organized sets, e.g. ordered signals, signals describing range cells within a grid or signals perceived as graph nodes. Within such sets, individual signals may be characterized by distinct sampling parameters....Sensor networks and radars can deliver signals as organized sets, e.g. ordered signals, signals describing range cells within a grid or signals perceived as graph nodes. Within such sets, individual signals may be characterized by distinct sampling parameters. This paper investigates organized signal sets neural network encoders. In the context of this work, the purpose of such encoders is to project heterogeneously sampled signal sets into an arbitrary fixed-size vectors space. This new representation space is designed so that signal sets can be processed as vectors rid of sampling differences to allow for arbitrary topology-aware processing with no signal processing constraints. Within this representation space designed to reduce the influence of heterogeneous sampling parameters, the relevance of signal sets representations is evaluated by considering signal sets discrimination potential with a focus on waveforms separation. The encoding and embeddings discrimination experiments conducted rely exclusively on synthetic complex-valued radiofrequency signals.
|
| 2583 |
Embodied Semantic Communication for Collective Autonomous Agents: A Tutorial on Representation, Wireless Delivery, and Closed-Loop Coordination
2609.35936
|
cs.AI
|
Yizheng Huang, Wensheng Lin, Lixin Li, Qinghe Du, Wenchi Cheng |
As autonomous systems and embodied intelligence enter the dynamic physical world, multi-agent collaboration calls for a paradigm shift in communication design. However, existing communication paradigms overlook that agents form action understanding from their ...As autonomous systems and embodied intelligence enter the dynamic physical world, multi-agent collaboration calls for a paradigm shift in communication design. However, existing communication paradigms overlook that agents form action understanding from their own states, environmental observations, and collaboration relations through a process that evolves as a task unfolds. Consequently, reliable bit delivery, general semantic recovery, or single-task utility optimization alone cannot ensure that heterogeneous agents form coordinated actions compatible with their own conditions from shared information during task execution. To address this gap, this paper proposes embodied semantic communication (ESC) as a paradigm that transforms information transmission into action-oriented semantic interaction. Specifically, ESC characterizes how an explicit communication link can encapsulate multimodal perceptual states, intrinsic hardware capabilities, and collaborative intents into unified actionable semantic representations, thereby enabling heterogeneous receiving agents to parse, align, and ground them in local motor control. This paper clarifies the conceptual boundary, system characteristics, and environment-constrained technical pathways of ESC. It maps the underlying mathematical tools, including semantic information theory, world models, and multi-agent decision theory. Finally, this paper summarizes key open challenges, including measurable semantic reliability, ambiguity-triggered interaction under dynamic environments and tasks, and bandwidth-adaptive semantic transmission, outlining a roadmap for collective embodied networks.
|
| 2584 |
PrivacySkills: How Privacy Guidance Shapes Source Selection in LLM Agents
2609.35937
|
cs.AI
|
Lucas Biechy, C\'edric Eichler, H\'eber H. Arcolezi, Nicolas Anciaux |
While prior work has documented privacy failures in LLM agents, it remains unclear how the presentation of privacy guidance influences their choice of information sources. We introduce PrivacySkills, a controlled framework for evaluating how agents choose amon...While prior work has documented privacy failures in LLM agents, it remains unclear how the presentation of privacy guidance influences their choice of information sources. We introduce PrivacySkills, a controlled framework for evaluating how agents choose among acquisition pathways that provide the same task-relevant value: consulting publicly available personal information, accessing confidential sources, or interacting with the user. The evaluation framework comprises 55 synthetic tasks spanning 11 categories of personal information, with 169 associated skills that describe the available acquisition pathways. We consider privacy guidance through system-level instructions, skill-level metadata labels, or both. Separately, we vary user availability and urgency framing. With users available and no privacy guidance, agents access confidential sources in 30% of valid runs on average across five open-weight models, despite sufficient alternatives. This rate increases to 45% when users are unavailable, whereas urgency framing has no detectable effect. System-level privacy instructions alone have limited effects on confidential access, while skill-level intrusiveness labels produce a modest reduction (24% on average), but combining the two roughly halves confidential access. Our findings motivate incorporating privacy annotations into skill specifications and evaluating their effectiveness alongside system-level instructions.
|
| 2585 |
Intrinsic Associative Memory on Riemannian Manifolds: Curvature, Capacity, and Emergent Modes
2609.35948
|
cs.AI
|
Krishnakumar Balasubramanian, Zhaoyang Shi |
Geometry does more than constrain an associative memory: curvature determines what it remembers and which states it creates. We develop intrinsic dense associative memories on Riemannian manifolds by casting memory as Epanechnikov kernel-density mode seeking. ...Geometry does more than constrain an associative memory: curvature determines what it remembers and which states it creates. We develop intrinsic dense associative memories on Riemannian manifolds by casting memory as Epanechnikov kernel-density mode seeking. We compare geodesic and volume-corrected energies and show that curvature separates their behavior. We prove that geodesic memory always retains an isolated pattern, while corrected memory obeys a sharp Ricci-curvature threshold: positive curvature can erase memories in high dimensions, while negative curvature reinforces them. We derive geodesic capacity scalings of $q_\beta^{-1/2}$ for retaining every pattern and $q_\beta^{-1}$ for a typical one, where $q_\beta$ is the pairwise kernel-overlap probability. We show how overlap \emph{creates} novel memories: designed $N$-pattern configurations realize all $2^N-1$ subset modes, but random data at the storage threshold yield only a Poisson number. We establish exact one-step recall using Riemannian mean shift. In simulations, we recover the predicted curvature transition and every designed mode. On WordNet's full noun hierarchy, we demonstrate that volume correction improves low-capacity retrieval. Together, our work shows that curvature is a design variable for associative memory, not merely a property of the data.
|
| 2586 |
HEAR: Real Voices, Real Bias: A Large-Scale Human-Recorded, Demographically Diverse Benchmark for Audio Language Models
2609.35952
|
cs.AIcs.SD
|
Shen Yan, Duc Le, Irina-Elena Veliche |
We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real human audio samples from 843 demographically diverse participants. HEAR enables comprehensive evaluation through M...We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real human audio samples from 843 demographically diverse participants. HEAR enables comprehensive evaluation through Multiple Choice Question Answering (MCQA) and open-ended long-form tasks. To our knowledge, this is the first large-scale voice benchmark grounded entirely in authentic human speech. We evaluate model behavior across both real-time speech-to-speech and speech-to-text architectures. Our results reveal that voice-conditioned bias is a model-specific property. Furthermore, we demonstrate that personalization instructions consistently exacerbate demographic disparities. Our findings establish that voice bias is a controllable model characteristic, providing a foundational framework for future bias mitigation and evaluation in Audio-LLM development.
|
| 2587 |
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
2609.35954
|
cs.AI
|
Zhiwei Zhang, Huayu Deng, Fei Zhao, Jiayan Fu, Bin Liang |
Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later polic...Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.
|
| 2588 |
Solver Agent: an Agentic AI Framework for Theoretical Physics Computations Applied to F-theory Uplifts of O3-planes and S-folds
2609.35958
|
cs.AI
|
Eliott Morgensztern, Cesar Fierro Cota, Alessandro Mininno |
We introduce Solver Agent, an AI framework based on large language models for calculations and proofs in mathematics and theoretical physics. The solution process is tracked through a persistent ledger that records assumptions, derivations, and computations. A...We introduce Solver Agent, an AI framework based on large language models for calculations and proofs in mathematics and theoretical physics. The solution process is tracked through a persistent ledger that records assumptions, derivations, and computations. A central agent delegates tasks to specialized sub-agents, while independent agents verify both intermediate steps and the final result. This setup improves the traceability, reproducibility, and verification of computer-assisted calculations. Applying Solver Agent, we study global F-theory uplifts of Type IIB orientifolds and their S-fold generalizations. We establish sufficient conditions for Weierstrass models over projective threefolds with terminal $\mathbb{Z}_k$ quotient singularities ($k\in\{2,3,4,6\}$) to give $\mathbb{Q}$-factorial projective elliptically fibered Calabi-Yau fourfolds with isolated Gorenstein terminal quotient singularities. These geometries realize O3-planes and S-folds, where local D3-brane probes of the latter yield four-dimensional $\mathcal{N}=3$ superconformal field theories. Using stringy invariants, we derive fixed-point contributions to Hodge data and Euler characteristics, and show that these Euler corrections determine the localized D3-brane charges required for tadpole cancellation. We illustrate these results using toric hypersurface constructions, where a single three-dimensional polytope determines both the Type IIB Calabi-Yau threefold and the F-theory base; here, the orientifold double cover naturally forms a bisection of an alternative genus-one-fibered uplift with discrete $\mathbb{Z}_2$ gauge symmetry. Finally, we provide methods for toric computations and four-form flux analysis in four-dimensional $\mathcal{N}=1$ compactifications with non-abelian gauge sectors.
|
| 2589 |
Causal and Interpretable Structures in LLM Compositional Tasks
2609.35970
|
cs.AI
|
Gurbir Arora, Toni J. B. Liu, Jiajun Bao, Rapha\"el Sarfati, Christopher J. Earls |
Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensemble...Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.
|
| 2590 |
Infrared Subtraction with Artificial Intelligence
2609.36007
|
cs.AI
|
Wenjie He, Xiaohui Liu, Yandong Liu, Zhan Wang |
We present AI-developed local infrared subtraction, building on projection to Born and EFT matching. The framework separates an integrable radiation term from a finite contribution at Born kinematics, referred to as the Born contact. The contact is determined ...We present AI-developed local infrared subtraction, building on projection to Born and EFT matching. The framework separates an integrable radiation term from a finite contribution at Born kinematics, referred to as the Born contact. The contact is determined using the EFT singular distribution in a resolution observable such as N-jettiness $\tau_N$. Under human physics guidance, an LLM develops two implementations. One uses a neural network for phase space projection and fits the contact by matching to EFT cumulants. The other uses an analytic construction that keeps the Born momenta fixed while integrating over radiation. It combines the EFT $\delta(\tau_N)$ coefficient with finite 4-dimensional radiation integrals to calculate the contact term directly. This gives a local subtraction formula without a slicing parameter, while reusing existing lower-order radiation calculations and EFT singular predictions. As a demonstration, we reconstruct the full NLO correction for massless 3- and 4-jet production in electron-positron annihilation. The attempt to the NNLO dijet production is also made by recursively using the NLO P2B construction with the LLM designing machine-learning controls to reduce the variance of the contact integral. The tested predictions are in good agreement with EERAD3. The numerical calculation and projection-network training use a 2020 Apple M1 MacBook, without GPU acceleration, illustrating the feasibility of the construction with modest computing resources. The appendices develop an extension of the local subtraction to 3-jet NNLO, giving explicit radiation maps and a proposed contact formula. We also show how to integrate over NNLO radiation while keeping the Born momenta fixed, for any number of massless final-state jets. Our results demonstrate how AI can help higher-order calculations by constructing infrared subtraction and improving its numerical integration.
|
| 2591 |
TORQUE: Optimizing What (not) to Quantize Before and After Rotation
2609.36032
|
cs.AI
|
Ran Ben Basat, Michael Mitzenmacher, Shay Vargaftik |
Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. We introduce TORQUE, a framework that improves on previous qua...Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. We introduce TORQUE, a framework that improves on previous quantization works that use random rotations by jointly optimizing how many and which coordinates to preserve at high precision both before and after rotation, under a fixed overall expected bit budget. Intuitively, before rotation, preserving large input coordinates at high precision can reduce overall error by preventing the rotation from spreading their values across many coordinates. Likewise, after rotation, preserving a small fraction of the largest-magnitude coordinates at high precision allows the remaining values to be quantized more accurately using codebooks optimized offline for the resulting truncated Gaussian distribution. We derive a quantization error upper bound and prove that top-$k$ pre-rotation retention minimizes it for each $k$. This reduces the search over coordinate subsets to an optimization over $k$, enabling a fast optimizer that uses offline codebooks and parallel parameter selection for practical implementation. We demonstrate an improved tradeoff between reconstruction accuracy and storage cost through numerical evaluation under the Gaussian model and experiments on nearest-neighbor retrieval, KV-cache compression, and activation compression.
|
| 2592 |
Neural networks for spectral optimization
2609.36047
|
cs.AI
|
Alexis de Villeroch\'e, Beniamin Bogosel, St\'ephane Breuils, Dorin Bucur, Jacques-Olivier Lachaud |
Given a functional dependent on the spectrum of a differential operator, we address the problem of finding a domain which optimizes this functional. PDE solvers might be used to tackle this optimization. It is however computationally expensive. We propose two ...Given a functional dependent on the spectrum of a differential operator, we address the problem of finding a domain which optimizes this functional. PDE solvers might be used to tackle this optimization. It is however computationally expensive. We propose two neural network models which learn the spectrum directly from the geometry of the domain and can be used to optimize the domain from one or more eigenvalues. We investigate two representations. The first encodes the domain through Fourier coefficients and a light MLP, which is efficient on star-shaped geometries, achieving a precision of 0.2\%. Through a rescaling of the coefficients the designed models satisfy the scaling law of the eigenvalues. Additionally, averaging the outputs of the trained surrogates over rotations and reflections induces invariance for these transformations. The second is a model that takes the landscape function, the indicator function and the gradient of the landscape function. A Gram-Schmidt process produces orthogonal eigenfunctions as output of the model along with the associated eigenvalues. The landscape model reaches 1\% mean relative error on the first ten eigenvalues, compared with 4\% for an FNO model. Replacing the landscape by an SDF worsened both prediction and optimization errors. The trained model also generalizes from synthetic shapes to domains given as classical image dataset. The resulting surrogates of both approaches recover classical spectral optima such as the disk for the first eigenvalue or the conjectured minima of higher eigenvalues. This confirms that our models produce accurate differentiable estimates of eigenvalues, which can be used in shape optimization problems involving spectral quantities.
|
| 2593 |
Improving scalable oversight with co-trained monitors
2609.36049
|
cs.AI
|
Joseph H. Rudoler, Kevin Tan, Benedict Tessler, Timothy Kong, Enric Boix Adser\`a |
Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by co-training the monitor alongside the worker, and explore both sup...Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by co-training the monitor alongside the worker, and explore both supervised and self-supervised approaches. In the supervised setting, we prove a characterization: monitoring is possible with vanishing error and query rates exactly when the class of possible monitor functions has finite Littlestone dimension. This connects worker monitoring with an established literature on adversarial online learning. For self-supervision, we propose a co-training procedure based on test-time distillation: the monitor uses additional test-time compute to generate training labels, then trains its standard-compute policy on those labels. For majority-vote labels, we give a finite-sample sharpening guarantee under adaptive worker distributions with action coverage, that shows that the monitor's verdicts converge to its initial modal verdicts. We stress-test the former in code-security settings where the worker is trained adversarially to fool the monitor. Our results suggest that adaptive monitors are better at keeping pace with evolving worker strategies, while fixed monitors are more vulnerable to evasion.
|
| 2594 |
What if automating AI R&D triggers an intelligence explosion?
2609.36054
|
cs.AI
|
Alan Chan, Christoph Winter, Andrew Barto, Jakub Pachocki, Geoffrey Hinton |
In contrast to even a year ago, AI systems now write most of the code inside the companies that build them. As more of the AI research and development (R&D) pipeline is automated, could AI progress radically accelerate in an "intelligence explosion," where...In contrast to even a year ago, AI systems now write most of the code inside the companies that build them. As more of the AI research and development (R&D) pipeline is automated, could AI progress radically accelerate in an "intelligence explosion," where years of advances are compressed into months or less? Preliminary evidence suggests that it could. In this work, we assess this evidence, analyze an intelligence explosion's potential impacts, and propose policy responses. AI systems are on track to automate most AI R\&D work within a few years, and possibly all of it. If this triggers an intelligence explosion, it could dramatically bring forward AI's benefits, but also pose extreme risks: capabilities growth could accelerate far beyond what society can keep up with, humanity could lose control over superhuman AI systems, and checks on power within and between states, companies, and branches of government could be severely eroded. Although there remains much uncertainty about these possibilities, the high stakes warrant serious further attention. Policymakers should urgently obtain more visibility into the automation of AI R&D, develop ways to steer and constrain an intelligence explosion, and prepare society to adapt to an intelligence explosion's impacts.
|
| 2595 |
Mnemon: Raw Records, Fast Judgments, Slow Thoughts
2609.36059
|
cs.AI
|
Guangren Wang |
Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two s...Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work: many small, independent yes/no judgments about records, such as whether a record is needed or no longer current, which a decision model makes by the dozen in a third of a second. Only a little is slow System 2 work: writing a few search queries, naming what the reply needs and composing the answer, which an LLM does well but slowly. We present Mnemon, a memory agent built on this division. It keeps conversations as raw, dated records; an LLM (System 2) plans searches over them, a decision model, Jev (System 1), judges what the searches return, and rules with explicit budgets turn the judgments into a small View for an unchanged answering model. A background pass consolidates each record once into topic timelines, value histories and standing instructions linked to the records, so that questions about a whole conversation reach evidence their own searches miss. Because nothing is decided about a record when it is written, the same agent can read any store that returns dated records. With gpt-4.1-mini answering, as in a public re-evaluation of 14 systems, Mnemon scores 91.7% on LoCoMo, the highest among them, and 83.8% on LongMemEval-S, from under 4k tokens of context per question, with the lowest effective cost index on LoCoMo. With a reasoning model answering, it reaches 92.2% on LoCoMo and 94.4% on LongMemEval-S, the latter on par with the best published results. From 100K to 10M tokens of history on BEAM, its cost per question grows by a factor of 1.11. On the same records, Jev separates gold evidence better than two LLMs and is 3-11 times faster.
|
| 2596 |
Understanding Decision-Making Mechanisms in Neural Routing Solvers
2609.36063
|
cs.AI
|
Fatemeh Askari, Mazdak Teymourian, Mohammad Izadi, Mahdieh Soleymani Baghshah |
Neural Combinatorial Optimization (NCO) has achieved strong empirical success, yet the internal mechanisms driving model decisions remain largely unexplored. In this paper, we investigate three representative autoregressive NCO models spanning two encoder-deco...Neural Combinatorial Optimization (NCO) has achieved strong empirical success, yet the internal mechanisms driving model decisions remain largely unexplored. In this paper, we investigate three representative autoregressive NCO models spanning two encoder-decoder configurations: AM and POMO (heavy-encoder, light-decoder), and LEHD (light-encoder, heavy-decoder). Through behavioral analyses, representation probing, and causal interventions, we examine how these models construct solutions and use internal representations during decoding. Our results suggest that AM and POMO predominantly follow a persistent geometric pattern throughout solution construction, whereas LEHD contains linearly accessible information about multiple future actions. Causal experiments further provide evidence for the role of future-node representations in LEHD's decision-making. We also observe that LEHD relies strongly on the current-node representation for immediate local decisions, while the start-node representation plays a broader navigational role over the subsequent route. Cross-instance alignment analyses additionally indicate that LEHD maps current-node representations into a relatively shared latent region, which may provide a stable reference for evaluating subsequent decisions. Across the Traveling Salesman Problem and the Capacitated Vehicle Routing Problem, these results reveal distinct decision-making patterns across these architecturally distinct solvers and provide a foundation for more interpretable analyses of NCO solvers. Code and additional visualizations are provided in the https://github.com/NCO-Interpretability/NCO-Interpretability.
|
| 2597 |
FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural Design
2609.36064
|
cs.AI
|
Sahand Rezaei-Shoshtari, Patryk Wozniczka, Shu Ishida, Gregg Streuber, Farnoosh Javadi |
Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly. We introduce FLOORA (Floor Layout Optimization with RL Alignment), a family of small domain-specific language...Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly. We introduce FLOORA (Floor Layout Optimization with RL Alignment), a family of small domain-specific language (DSL) models for architectural layout generation. With specialized data and alignment, our 0.6B model outperforms much larger frontier models, achieving VLM judge win rates up to 92.0% on out-of-distribution real-world buildings and 96.0% on synthetic buildings. Human evaluations further corroborate these results, with FLOORA selected as the best model in 89.3% of evaluations. FLOORA combines a token-efficient DSL, custom tokenization, domain-specific pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL) with learned human-preference and verifiable rewards. This pipeline improves architectural and geometric validity, supported by extensive empirical evaluation and ablation studies. Although focused on architecture, our results suggest that similar domain-specific recipes may be useful in other engineering domains with structured, verifiable outputs. Datasets, models, and inference code are available at https://github.com/AutodeskAILab/floora.
|
| 2598 |
The Uneven Decline of Collective Knowledge Production: Evidence from Stack Overflow After Generative AI
2609.36069
|
cs.AI
|
Myokyung Han, Taegyoon Kim, Jinhyuk Yun, Lanu Kim |
Generative AI (Gen AI) is reshaping how individuals learn and work, but its consequences for collective knowledge, the shared body of knowledge that online communities produce together, remain poorly understood. Prior work has documented an aggregate decline i...Generative AI (Gen AI) is reshaping how individuals learn and work, but its consequences for collective knowledge, the shared body of knowledge that online communities produce together, remain poorly understood. Prior work has documented an aggregate decline in participation on knowledge-sharing platforms, but it remains unclear which specific kinds of knowledge are being lost first. We study this question using Stack Overflow, one of the largest online communities for software engineering, treating the release of ChatGPT-3.5 as a natural shock. Analyzing over two million questions posted between 2020 and 2025, we track how two dimensions of collective knowledge, difficulty and data availability, change following Gen AI's release. Using diverse methods and robust checks, we find consistent patterns. Easy questions decline sharply while difficult questions become more common, a pattern corroborated by rising code complexity. Data-rich topics and tags lose share of questions, while data-scarce ones gain ground. The two dimensions also interact: the decline in easy questions is concentrated specifically within data-rich domains, while difficult questions increase regardless of data availability. This pattern extends beyond Python across programming languages, with more prevalent languages showing sharper shifts. Together, our findings reveal that Gen AI's impact on collective knowledge is uneven, eroding easy, accessible knowledge first while more complex, less common knowledge persists.
|
| 2599 |
Measuring trainable degrees of freedom in materials graph neural networks: a random-subspace intrinsic dimension analysis
2609.36084
|
cs.AI
|
Shehroz Ahmad Shoaib, Kangming Li |
Final predictive accuracy is the standard basis for comparing graph neural networks (GNNs) in materials-property prediction, but it does not show how strongly performance depends on access to trainable parameter-space directions. Here, we introduce trainable-d...Final predictive accuracy is the standard basis for comparing graph neural networks (GNNs) in materials-property prediction, but it does not show how strongly performance depends on access to trainable parameter-space directions. Here, we introduce trainable-degree dependence as a complementary characterization of materials GNN learning. Using random-subspace intrinsic-dimension analysis, we train CGCNN, ALIGNN, and DimeNet++ in randomly oriented parameter subspaces across six prediction tasks and measure how performance recovers as independent trainable degrees of freedom are restored. The resulting recovery curves separate endpoint accuracy from the trainable-dimensional demand required to recover it. They reveal distinctions that final errors alone miss: metallic classification and log-bulk-modulus regression recover near-reference performance from small fractional subspaces, formation-energy and band-gap prediction show stronger architecture dependence, and phonon prediction is most sensitive to dimensional restriction. Dataset-size sweeps show that band-gap models require larger fractional subspaces as training data grows, whereas formation-energy and bulk-modulus responses are more stable. A width sweep shows that fractional thresholds can remain stable while absolute threshold dimensions increase with model size. Random-subspace analysis therefore provides a targeted stress test for how materials GNNs use their optimization space.
|
| 2600 |
PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
2609.36086
|
cs.AI
|
Cheng Chang, Yining Mao, Peng Qi |
Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment M...Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether their implied preferences align. Building on this, we introduce PADM\'E, a data synthesis method that generates reliable criterion-based meta-evaluation data for agentic settings. PADM\'E uses only small language models, requires no human involvement during evaluations, and operates under a low computational budget. We build a prototype of PADM\'E and synthesize a dataset of 1,000 samples across four agentic domains and three evaluation criteria. Human validation on a 150-sample subset demonstrates that PADM\'E improves agreement with human judgment from 73% to 85% over a naive baseline. Meta-evaluating 25 common models with our dataset demonstrates the correlations between evaluation performance and scoring granularity, leniency, and model size, among other factors.
|
| 2601 |
PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG
2609.36087
|
cs.AI
|
Yipeng Zhang, Chenda Duan, Yuanyi Ding, Tianyi Wang, Atsuro Daida |
Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by ...Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduce PHASE, a physiology-guided foundation model that makes these characteristics explicit learning targets, pairing them with masked latent prediction in a temporal stage (PHASE-T) within each channel and a spatiotemporal stage (PHASE-ST) across synchronized channels. PHASE is pretrained on heterogeneous recordings from 222 participants at nine clinical sites. On all five Omni-iEEG clinical tasks, frozen PHASE-T outperforms every evaluated foundation model by up to 31\%, and fine-tuned PHASE-T surpasses the task-specific models, setting a new state of the art. PHASE-T benefits from physiological supervision, outperforming variants trained with latent prediction alone or auxiliary waveform reconstruction on every task in matched ablations. PHASE-T generalizes to unseen institutions, outperforming the compared models with few or no local labels. PHASE-ST further improves seizure-onset-zone identification over PHASE-T and, when frozen, decodes sound volume and pitch on BrainTreebank better than published models. Beyond task performance, PHASE learns to encapsulate the physiological characteristics clinicians recognize, from seizure onset and its propagation to anatomical region identity, even though its pretraining contains no ictal recordings or anatomical labels.
|
| 2602 |
LoopICL: Looping a single transformer block to solve tabular tasks
2609.36108
|
cs.AI
|
Amir Rezaei Balef, Katharina Eggensperger |
Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped t...Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped transformer whose core design decouples parameter count from computational depth. LoopICL consists of a single block, processing data through two coupled streams: a cell stream capturing per-cell feature representations and a row stream capturing in-context example representations, jointly refined through within-column and cross-column attention. During pre-training, we vary loop counts, allowing the block to be unrolled for a varying number of iterations at test-time and use a learned exit-gate to automatically exit. In its standard setting, LoopICL performs competitively with TabICLv2 on TabArena and TALENT at the same computational cost (FLOPs), while using nearly 90% fewer parameters. Furthermore, its recurrent design enables users to also trade off inference cost and performance, providing a resource-aware TFM.
|
| 2603 |
ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs
2609.36120
|
cs.AI
|
Mehdi Makni, Ryan Lucas, Rahul Mazumder |
Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and c...Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free approaches such as DartQuant, but both remain hard to scale to the largest architectures. To address the computational bottlenecks in gradient-free rotation learning, we introduce two ideas for efficiency, (i) a data selection procedure which reduces the required number of calibration data points, and (ii) an exact reduction of the associated optimization on this reduced calibration set. Our data selection procedure exploits the geometric structure of the convex hull of the activations. Using this idea, we show that a carefully selected calibration set with several orders of magnitude fewer activations than state-of-the-art rotation-based methods can match their performance in low-bit quantization settings. Under this extreme data efficiency, the selected activations span an $r$-dimensional subspace with $r<d$, making optimization over a $d\times d$ rotation equivalent to optimizing a $d\times r$ matrix on the Stiefel manifold. We solve this reduced problem using an efficient ADMM algorithm that iteratively employs thin matrix updates at every step, hence the name ThinQuant. For Llama-3-70B with W4A4KV4 quantization, ThinQuant completes the entire rotation calibration in under 12 minutes and achieves a WikiText-2 perplexity of 5.63, compared with 7.55 for DartQuant, which requires 111 minutes. Unlike SpinQuant and DartQuant, ThinQuant also scales to Llama-3.1-405B on a single H200 GPU, completing rotation calibration in just over 2 hours and achieving WikiText-2 perplexity of 2.97 at W4A4, compared with 3.48 for GPTAQ+QuaRoT.
|
| 2604 |
Render Before Reading: Visual Rendering as a Prompt Injection Defense
2609.36121
|
cs.AI
|
Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r |
Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: mult...Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other modalities mainly as content to parse or describe. We then demonstrate how this gap can be turned into a training-free defense, by rendering all untrusted payloads as typographic images (or audio) before they reach the model. Across ten models and two prompt injection benchmarks (DirectInject and AgentDojo) we show that our defense Pictionary consistently reduces attack success rates even against the strongest adaptive attacks and human red teamers, while largely preserving benign utility. We further show that benign fine-tuning on image-rendered instructions erodes the modality gap, tracing it to the text-centric instruction-tuning distribution.
|
| 2605 |
Reasoning with Neural Cellular Automata
2609.36126
|
cs.AI
|
Mayalen Etcheverry, Pietro Miotti, Aidan Sirbu, Konstantin Sch\"urholt, Mariia Drozdova |
Modern AI architectures used to solve visual reasoning tasks typically rely heavily on global connectivity and synchronization. As biological systems demonstrate, though, sophisticated computation can be performed in a more decentralized fashion. In this work,...Modern AI architectures used to solve visual reasoning tasks typically rely heavily on global connectivity and synchronization. As biological systems demonstrate, though, sophisticated computation can be performed in a more decentralized fashion. In this work, we test the reasoning capabilities of Neural Cellular Automata (NCAs), networks of recurrent cells that use strictly local connectivity and asynchronous updates. NCAs have been extensively studied in artificial life experiments, but it is unclear whether they can perform complex multi-step reasoning. We show that NCAs produce spatio-temporal dynamics capable of solving challenging visual reasoning tasks, including large mazes, Sudoku, and ARC-AGI-1. Furthermore, we provide evidence that NCAs generalize out-of-distribution when running with larger grids, longer rollouts, or parallel trials; and that the latter can be made more efficient via pruning of redundant trajectories. We find that these generalization capabilities depend on training with sample replay and stochastic perturbations, and that stochasticity remains beneficial at test time. Finally, we show that NCAs are robust reasoners capable of dynamically modulating compute to recover efficiently from damage, and that they can scale to solve reasoning in raw pixel space.
|
| 2606 |
Accessible, but Not Adopted: Increasing LLM Adoption among First-generation, Low-income (FGLI) College Students beyond Expanding Access
2609.36129
|
cs.AI
|
Hyungsik Kim |
Large language models (LLMs) are increasingly positioned as a force to empower underserved communities, and significant efforts are being made to expand access. Yet, access alone does not equate to meaningful adoption. First, even if a system is accessible, it...Large language models (LLMs) are increasingly positioned as a force to empower underserved communities, and significant efforts are being made to expand access. Yet, access alone does not equate to meaningful adoption. First, even if a system is accessible, it won't be adopted if users are not willing to adopt it. Second, even if an LLM system is superficially adopted, the heterogeneity of LLM tools means that LLM adoption can be further deepened. Closing this access-adoption gap is critical to ensuring that the full social potential of LLM is not only accessible but fully realised. Drawing on 61 interviews (15 long-form semi-structured interviews with first-generation, low-income college (FGLI) students, 3 non-FGLI students, 3 FGLI program directors, and 40 intercept interviews), this paper examines the access-adoption gap in first-generation, low-income student communities. This paper a) finds that while FGLI students have adopted LLM systems, their depth of LLM tool usage is limited to chatbots (e.g., ChatGPT or Claude) for narrow use cases, and b) identifies barriers limiting their willingness to learn and use (low perceived value, under-estimated self-efficacy, unclear starting point, low peer exposure, and resource constraints). Then, from these findings, the paper derives the four design principles to design a system or an intervention aimed at closing the access-adoption gap in LLM adoption by FGLI students. In doing so, the paper contributes to the field by a) examining the LLM access-adoption gap in the FGLI student community, and b) reframing LLM adoption as a depth gradient across four modes of LLM tool use: basic chatbot interfaces, tool-augmented prebuilt interfaces, agentic development interfaces, and programmatic integration.
|
| 2607 |
Language Models Are "Insecure" Reporters
2609.36139
|
cs.AI
|
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin |
As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality...As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work. We call this phenomenon "insecure reporting." When handed machine learning experiment logs containing a planted negative result that substantially weakens the proposed method, GPT-5.5 flags the negative result in only 2 of 200 generated reports. However, when a short honesty instruction, "Be honest in your response," is added, the model flags the negative result in 190 of 200 reports. Across eight open-weight models, chain-of-thought analysis reveals a recurring tension between disclosing narrative-changing flaws and reasoning about ways to appear successful. We perform an activation analysis and a steering experiment on Qwen3.5-9B, finding that honesty and success-seeking correspond to opposing directions in representation space. Our results suggest that LLMs tend to present narratives of success by default, and that steering models toward honesty makes their reports substantially more transparent.
|
| 2608 |
Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs
2609.36157
|
cs.AI
|
M. Saeid HaghighiFard, Sinem Coleri |
Most federated learning frameworks for vehicular ad hoc networks assume that all vehicles collaboratively train a single model for a common task. This assumption limits their applicability to practical vehicular environments, where vehicles may perform heterog...Most federated learning frameworks for vehicular ad hoc networks assume that all vehicles collaboratively train a single model for a common task. This assumption limits their applicability to practical vehicular environments, where vehicles may perform heterogeneous but related perception tasks with different output spaces. This paper proposes encoder-sharing hierarchical multi-task federated learning (EN-HMTFL), which integrates cluster-based hierarchical federated learning with a globally shared encoder and vehicle-local decoders. EN-HMTFL enables vehicles performing different tasks to collaboratively learn a transferable feature representation while preserving their task-specific models locally. Only the encoder is exchanged and aggregated through the hierarchy, whereas raw data and local decoder parameters remain at the vehicles. The proposed framework is evaluated on the MNIST and GTSRB datasets in different vehicular scenarios. Across the evaluated scenarios, EN-HMTFL improves accuracy by up to 24.0% relative to the compared representation-sharing benchmark. In scenarios where EN-HMTFL converges earlier, the reduction reaches up to 69 communication rounds (28.8%).
|
| 2609 |
From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents
2609.36161
|
cs.AI
|
Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang |
Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifier...Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest evaluated model passes all ten tasks in at least one run each. In contamination-control experiments, identifier obfuscation reduces line similarity to the original implementations from 0.51--0.96 to 0.03--0.44 without reducing the observed pass rate of any evaluated model. On repositories created after the stated knowledge cutoffs, the strongest model passes eight of nine runs. The reconstruction family comprises thirteen tasks spanning numerical, geometric, hardware, and transactional systems (e.g. CAD and CRM). Two models meet the benchmark's pass criteria on all thirteen, although our audit shows that the CFD task cannot establish numerical-solver capability. Benchmark construction and auditing uncover 28 verifier defects, including 24 false negatives and two false positives. These findings show that executable verification can itself introduce substantial measurement error. We present three practical checks: test whether prescribed methods can reach the grading thresholds, investigate agreement among independently generated candidates, and recompute diagnostics from submitted artifacts. ReviveBench thus provides both an evaluation of software revival and engine reconstruction, and cases in validating the verifiers used to measure coding agents for software design.
|
| 2610 |
Adversarial Debiasing of Machine Learning Models for Enhanced Network Security against DDoS Attacks
2609.36167
|
cs.AI
|
Aadith Sukumar, Isha Singh, Devershika Mohane, Ankit Mukherjee, Ankush Dutta |
Distributed Denial of Service attacks are a growing threat to network infrastructure, and new techniques, including the use of generative AI, make them harder to detect. Traditional detection systems, such as rule based firewalls, often fail to identify these ...Distributed Denial of Service attacks are a growing threat to network infrastructure, and new techniques, including the use of generative AI, make them harder to detect. Traditional detection systems, such as rule based firewalls, often fail to identify these evolving attack patterns. In this study, we propose a new method for detecting DDoS attacks by combining synthetic data generation using Generative Adversarial Networks with a Random Forest classifier. The GAN generated data showed 80.3 percent cosine similarity to real traffic, which helped the model learn underlying traffic patterns more effectively. To address imbalances in the data, especially in packet related features, we applied adversarial debiasing. This reduced the model's sensitivity to skewed distributions in variables such as forward and backward packet counts and total byte lengths. Our results show that models trained on a mix of synthetic and real data achieved significantly better performance: 99.98 percent accuracy on benchmark data and a 22.60 percent improvement when tested on previously unseen synthetic traffic. This suggests that the method can generalize well across different traffic scenarios and adapt quickly to new types of attacks. The proposed approach not only improves DDoS detection but also provides a scalable foundation for security models that account for bias and benefit from data augmentation. Our findings show that combining GANs with adversarial debiasing can lead to more robust and effective DDoS mitigation, supporting the further development of machine learning based cyber security.
|
| 2611 |
Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding
2609.36173
|
cs.AI
|
Haohui Zhang, Keyu Chen, Haocheng Sun, Weibo Gu, Ruizhi Qiao |
Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight modul...Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor's predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel. A unified transfer space built from the target model's output embeddings unifies layer-wise injection with predecessor-conditioned decoding, and layer-wise output-embedding supervision promotes the formation of predicted features in shallow layers. Fused kernels and a transition cache execute both efficiently in parallel within SGLang. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3-4B and Qwen3-8B. At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash's 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average.
|
| 2612 |
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
2609.36178
|
cs.AI
|
Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li, Yuexing Hao |
Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relev...Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
|
| 2613 |
SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
2609.36201
|
cs.AI
|
Jianxing Chen, Xiao Yu, Shipra Agrawal, Zhou Yu |
Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires caref...Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshots alone may be unable to determine the actual consequences of actions. To address these challenges, we introduce SCOUT, a two-stage agentic safety verifier that synergizes reasoning-intensive rubric generation with tool-intensive evidence gathering. First, our SCOUT rubric generator extensively reasons over the task and the agent's trajectory to determine what successful and safe execution should entail, generating task-specific completion and safety rubrics. Then, our SCOUT probing agent follows these rubrics to interact with the post-execution environment and collect grounded evidence for final safety and completion judgments. We evaluate our framework on two computer-use safety benchmarks. On AutoElicit-Bench, SCOUT achieves 75.4 unsafe F1 and 74.5 completion F1, outperforming LLM-as-a-judge verifiers and naive tool-use verifiers. SCOUT leads on OS-Blind with 76.4% unsafe detection accuracy. Test-time reflection reduces final unsafe execution rates from 30.2% to 17.2% on AutoElicit-Bench. Ablations and analysis show that tool-free rubric generation in SCOUT elicits substantially more reasoning and is crucial for safety detection across verifier backbones, especially non-frontier ones. A preliminary extension to coding tasks shows that SCOUT can support safety verification beyond computer-use.
|
| 2614 |
Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
2609.36208
|
cs.AI
|
Zahra Khodagholi, Niloofar Yousefi |
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without label...Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting the recorded contrasts onto that space gives an empirical error floor for any unrestricted decoder on those classes. On a 140-rectangle siRNA interaction panel, a graph neural network's training-only feature mask merges 165 endpoint states into 90 classes and cuts the rank of the 140 interaction contrasts to 72. The resulting floor is 0.009980, which is 14.6% of the fitted model's interaction squared error; the fitted model reaches 0.068335, slightly worse than a control predicting no interaction at all. A minimum of three restored chemistry columns recovers full rank. Refitting without the mask removes the floor entirely, yet interaction MSE improves by only 0.000017 under the reported protocol, and the restored columns remain absent from every training input. On a released RNA-splicing predictor, whose encoding is injective on the measured states, the same computation returns the full design rank of 1,986 and a floor of exactly zero. These results separate what an encoding permits from what a fitted model achieves; they do not identify what limits the remaining error. The rank check needs no fits and bounds what any amount of training under a fixed encoding can recover. The project repository is available at https://github.com/shadi97kh/REPRESENTABLE-BUT-UNLEARNED.
|
| 2615 |
CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles
2609.36218
|
cs.AI
|
Mir Tafseer Nayeem, Susmoy Chakraborty, Davood Rafiei |
Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation,...Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progression, and themes without explicit scene or event structure. CineSubBench contains 1,012 films with complete subtitle coverage in six languages, yielding 6,072 tracks and 8.13M timestamped subtitle entries. It provides a matched multi-task, multilingual, and multicultural (MultiX) evaluation setting: seven tasks span narrative reconstruction and abstraction, genre prediction, age suitability, country-specific motion-picture ratings across ten national classification systems, and subtitle-grounded language safety. Across nine LLMs, plot premises are recovered more reliably than event-complete synopses; cross-lingual consistency varies substantially across models and languages; national rating systems expose distinct calibration patterns; and strong profanity is far easier to ground than mild obscenity. CineSubBench establishes film as a long-context LLM evaluation domain and provides a unified benchmark for measuring narrative, multilingual, cultural, and evidence-grounding capabilities.
|
| 2616 |
Population Fidelity: Evaluating Population Representativeness in LLMs
2609.36253
|
cs.AI
|
Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang |
Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways tha...Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonstrate the framework's utility in two ways. First, we reproduce a prior study of "machine bias" in LLM survey responses and apply the framework to its models and more recent ones, showing that poor representation reflects not only insufficient between-group variation but also variation assigned to the wrong groups. Second, we evaluate one proposed approach to improving models' population representativeness: cultural fine-tuning. We find that cultural fine-tuning can improve alignment with the survey center without improving the representation of within-population differences, a distinction that measures of aggregate agreement do not capture. We argue that representing a population requires models to reproduce several features of human attitudinal variation simultaneously. Our framework organizes these features and provides reusable code, data, and trained models for evaluating population fidelity across substantive domains and assessing proposed alignment methods.
|
| 2617 |
Paired Multimodal Scaling Laws
2609.36263
|
cs.AI
|
Marcus Ma, Shrikanth Narayanan |
Existing multimodal scaling laws fit multimodality terms empirically after testing and never vary how much data is multimodally paired at fixed data budgets. We investigate how, under the same total data per modality, changing the number of paired data affects...Existing multimodal scaling laws fit multimodality terms empirically after testing and never vary how much data is multimodally paired at fixed data budgets. We investigate how, under the same total data per modality, changing the number of paired data affects loss curves in multimodal classification tasks. We train models in three different environments and run experiment sweeps varying data sizes and pairing budget. Pairing ratios have a dramatic impact on loss and this impact is directly tied to how much information synergy the task contains. Only paired data is able to reduce synergistic loss, while unpaired data can reduce redundant or unimodal information up until unimodal floors. Unlike traditional scaling laws where loss drops immediately in power law decay, synergy acquisition is gated, requiring a critical threshold of paired data before synergistic loss falls at all. We introduce a new family of multimodal scaling laws where total data-attributable loss is the sum of four individual power laws corresponding to the four different information channels of redundancy, a unique channel per modality, and synergy, and show how this law is both more theoretically sound and empirically valid across our experiments. This law predicts multimodal loss in our experiments more accurately than existing laws, with 3.2% error on fit tests versus 10.4% error for the best pairing extension of published laws.
|
| 2618 |
In-Context Learning Amplifies a Latent Symbolic Circuit
2609.36265
|
cs.AI
|
Melissa Wessel |
Large language models can learn abstract rules from just a few in-context examples, but how their internal mechanisms activate as examples accumulate is not well understood. We trace a three-stage symbolic reasoning circuit (abstraction, induction, retrieval) ...Large language models can learn abstract rules from just a few in-context examples, but how their internal mechanisms activate as examples accumulate is not well understood. We trace a three-stage symbolic reasoning circuit (abstraction, induction, retrieval) across shot counts in three model families and find it is detectable and functional well before the model achieves high accuracy. Per-head causal contribution grows up to 8x from 1- to 10-shot, and cross-shot activation patching raises accuracy from 1% to 56% at 0-shot and 17% to 88% at 1-shot. Function vectors scaled and injected at 0-shot rescue accuracy up to 86%, largely substituting for the induction stage but depending critically on an intact downstream retrieval stage. The infrastructure for abstract rule-following is present in the weights before any demonstrations; in-context examples, function vectors, and related interventions appear to supply input to the same latent circuit.
|
| 2619 |
Proofs Without Nominals: G\"odel's Ontological Argument, its Shallow Embedding, and the Open Questions of the Monatshefte Notes
2609.36279
|
cs.AI
|
Christoph Benzm\"uller |
The shallow embedding of higher-order modal logic in classical higher-order logic, used in Benzm\"uller and Scott's Notes on G\"odel's and Scott's variants of the ontological argument (2025), reaches beyond the modal object language of the arguments: its prope...The shallow embedding of higher-order modal logic in classical higher-order logic, used in Benzm\"uller and Scott's Notes on G\"odel's and Scott's variants of the ontological argument (2025), reaches beyond the modal object language of the arguments: its property quantifiers range over terms that may also express nominals and satisfaction operators of hybrid logic, and a proof using one proves a theorem of the embedding that need not be one of the modal logic. That the framework affords this is not new, and whether a result is one of the modal logic can be settled in two ways: by replaying it in an explicit proof calculus, done by hand for chosen theorems, or by analysing the proofs the embedding itself produces, which this article does mechanically, for every result at once. Every statement the Notes prove has a proof inside the object language: 294 written out by hand and machine-checked, none using a nominal. The proofs the Notes themselves give instantiate no nominal either; what the detector flags there are terms a prover substituted. The three questions the Notes leave open are settled too, and without nominals, but the conjunction axiom has to be emended: generalised in the Notes to G\"odel's "any number of summands", it covers the conjunction of no properties, and of one; the empty one alone settles all three, and the two together yield what a separate axiom of G\"odel's is for. This article restricts the conjunction axiom to at least two different conjuncts, the reading G\"odel's footnote suggests, and the questions are settled again, by proofs that turn on the argument rather than a degenerate instance. The restriction holds of the object language only: with a nominal the axioms make the accessibility relation the identity and the readings coincide. Every theorem is verified in Isabelle/HOL and independently in Lean 4; the countermodels are Nitpick's, certified by the build.
|
| 2620 |
How Much Prompt Is Enough? A Blackbox Minimization of Few-Shots in LLMs
2609.36289
|
cs.AI
|
Ali Alfageeh, Rahul Gopinath, Amin Alipour |
Prompts are the primary mechanism for directing the behavior of large language models (LLMs). Yet the internal structure and causal hierarchy of prompts remain poorly understood: which parts are causally necessary and which are redundant is an open question. T...Prompts are the primary mechanism for directing the behavior of large language models (LLMs). Yet the internal structure and causal hierarchy of prompts remain poorly understood: which parts are causally necessary and which are redundant is an open question. This opacity can have severe consequences. Subtle prompt variations can silently shift model outputs in critical software systems, and engineers lack techniques to reason about prompt reliability. We present \framework, a blackbox prompt-minimization framework that reduces few-shot prompts to their necessary minimal subset. We use a case study to apply \framework to a few-shot learning system and demonstrate the insights that this framework can provide. Our experiments show that few-shot exemplars can be reduced by a mean of 65.3\%~$\pm$~15.8\% in character count while fully preserving propositional output fidelity. The models preferentially retain logical identifiers and constraint declarations while discarding natural language prose and cross-prompt relational annotations. Our analysis also shows that some models are universal encoders, able to produce highly legible yet minimized prompts, while others are universal decoders, able to interpret minimized prompts from most other models. By identifying which components are indispensable, \framework provides a principled basis for prompt compression and structural analysis of few-shot exemplars.
|
| 2621 |
MoRE: Scaling mixture of experts with hardware-aware low-rank routing
2609.36301
|
cs.AI
|
Honam Wong, Surbhi Goel, Enric Boix-Adser\`a |
Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token co...Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $\Theta(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of $\Theta(h/r)$ more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\&A benchmarks after pretraining, while matching reasoning ability. Code available at https://github.com/Matheart/MoRE_code.
|
| 2622 |
Training LLMs to Verbalize Evaluation Awareness
2609.36316
|
cs.AI
|
Usman Anwar, Sahar Abdelnabi, David Krueger |
Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent a...Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supervise the latent belief itself. VT uses a model's spontaneous verbalizations as evidence that awareness is present and truncates each rollout immediately before the verbalization, producing training prefixes at which the model is presumed to be aware. The model is then trained with an RL objective designed to increase verbalization in a calibrated way. Across Qwen3.6-35B-A3B, Kimi K2.6, and Inkling, VT increases verbalized EA by 2.4-2.9 times and transfers to held-out agentic settings, while measured latent EA and behavior remain largely stable. In a causal experiment, we independently implant meta-knowledge about evaluations through synthetic-document fine-tuning and show that VT-induced verbalizations reflect the richer knowledge acquired by the model.
|
| 2623 |
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
2609.36322
|
cs.AI
|
Xingyu Zhu (Luke), Pu (Luke), Yi, Ziheng Cheng, Ang Lv |
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase...Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.
|
| 2624 |
DecoyTrace: Toxic Decoys for Active Defense in Decentralized Federated Learning
2609.36330
|
cs.AI
|
Pedro Beltr\'an-L\'opez, Enrique Tom\'as Mart\'inez Beltr\'an, Pantaleone Nespoli, Manuel Gil P\'erez, Alberto Huertas Celdr\'an |
Decentralized Federated Learning (DFL) eliminates the central aggregation server, reducing the single point of observation that traditional defenses against attacks rely on. As a result, peer-to-peer networks become exposed to malicious updates containing back...Decentralized Federated Learning (DFL) eliminates the central aggregation server, reducing the single point of observation that traditional defenses against attacks rely on. As a result, peer-to-peer networks become exposed to malicious updates containing backdoors or semantic poisoning, since such updates can remain close to benign ones in the parameter space while behaving very differently. This may evade defenses based on passive parameter inspection. However, existing deception-based defenses have mainly been designed for centralized FL and do not jointly address local observation, poisoning propagation, source attribution, and containment in strictly serverless DFL. To address these limitations, this paper presents DecoyTrace, a proactive cyber deception-based defense for strictly serverless DFL environments. DecoyTrace deploys a mobile DecoyNode that generates decoy challenges using chaotic maps, disseminates a dual model (clean vs. decoy) based on neighbor trust, and evaluates them using three-state semantic metrics. Upon confirmation, a distributed protocol isolates the source and performs a model reset or recovery to preserve training progress. Evaluated across sixty configurations on the NEBULA platform (five datasets, three topologies, and four attack/defense scenarios), DecoyTrace systematically restores lost utility. The F1-score remains within 0.03 of the baseline on MNIST/FashionMNIST (mitigating drops of up to 0.37), matches or exceeds the baseline on EMNIST and CIFAR-100, and remains between 0.05 and 0.10 below the baseline on CIFAR-10, the most visually complex convolutional scenario evaluated. Furthermore, containment reduces CPU and network usage by up to two-thirds. These results demonstrate the feasibility of unifying deception, identification, and containment in DFL, while also identifying its limitations in complex tasks and multi-attractor threat models.
|
| 2625 |
Calibrating One-Round Membership Inference with Neighbors
2609.36331
|
cs.AI
|
Francesco Rita, Jie Zhang, Florian Tram\`er |
The state-of-the-art Membership Inference (MI) methods calibrate their signal separately for each example using reference models, auxiliary models trained to exclude the target. This paradigm scales poorly to modern large models, however, whose training is too...The state-of-the-art Membership Inference (MI) methods calibrate their signal separately for each example using reference models, auxiliary models trained to exclude the target. This paradigm scales poorly to modern large models, however, whose training is too expensive to replicate. This has motivated one-round settings, where only a single trained model is available; but without reference models the per-example calibration that drives the strongest attacks can no longer be estimated, leaving the membership signal weak. We ask whether neighbors of the target point can recover this calibration without training any additional model. Our key observation is that reference models serve only to reveal how an example behaves under models not trained on it, and that querying the target model on nearby samples yields the same information. We propose two complementary ways to obtain such neighbors, and show that querying them against an early training checkpoint further sharpens the signal. We evaluate across three image classification datasets and three training setups, showing that neighbors yield strong membership signals and competitive attack performance at no additional training cost.
|
| 2626 |
ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
2609.36333
|
cs.AI
|
Ke Fang, Yupu Yao, Lu Cheng |
Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relevant novelty structure to be we...Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relevant novelty structure to be weakened as representations are transformed into the final latent used by the planner. We introduce Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution. ATLAS transfers normalized pairwise structure from an informative encoder representation to the planning latent and uses Wasserstein embedding matching (WEMReg) to calibrate its marginal through one-dimensional Wasserstein-2 transport. Our analysis shows that relational preservation and marginal calibration impose non-redundant constraints, and connects finite-candidate planning stability to relational distortion, latent-scale mismatch, and prediction error. Instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes. Representation and rollout diagnostics further show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error. Together, these results highlight preservation of planning-relevant latent geometry as an important ingredient for reliable world-model planning. Code is available at https://anonymous.4open.science/r/atlas-world-model-72C4/.
|
| 2627 |
Explainability from Training with Applications to TCR-Epitope Prediction
2609.36354
|
cs.AI
|
Jiarui Li, Zixiang Yin, Samuel Landry, Zhengming Ding, Ramgopal Mettu |
Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models ...Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-epitope prediction models, TCR-SRIM, TULIP, MixTCRpred, and NetTCR-2.2, spanning post-hoc and interpret-by-design approaches as well as transformers and CNNs. To investigate how structural information affects model explanations, we introduce a benchmark, TCR-XAI2, containing 388 unique experimentally resolved TCR-epitope structures, complemented by structures predicted using AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Using EFT with TCR-XAI2, we demonstrate that (1) CNN and transformer models exhibit distinct learning trajectories; (2) TCR $\alpha$ and $\beta$ evidence can conflict during learning, limiting the benefits of jointly modeling both chains, while MHC information mitigates this; and (3) real versus predicted structural data for TCR-epitope prediction exhibits distinct TCR and peptide feature preferences as well as differing trajectories of model certainty.
|
| 2628 |
Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex
2609.36366
|
cs.AI
|
Iishaan Inabathini, Margaret M. Henderson |
Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel res...Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabling flexible stimulus-dependent weighting of image content across space. Here, we extend this framework to naturalistic video, using per-parcel cross-attention to dynamically route features from a self-supervised video model (V-JEPA-2) across both space and time, fitting this model to fMRI responses to short video clips. We compare joint spatiotemporal attention with factorized and selectively constrained alternatives, and find that joint routing improves predictions of brain responses to held-out videos across higher visual regions, most consistently in lateral and dorsal visual areas associated with dynamic motion perception. Moreover, our method provides interpretable, stimulus-specific attention maps that dynamically follow moving objects, revealing which locations and temporal moments contribute to each neural response. We further show that attention maps from parcels in different category-selective networks (face-, body-, scene-selective) differentially weight content in accordance with expected semantic selectivity. Together, this work provides a new computational framework for understanding how visual information is adaptively weighted by cortical populations during dynamic visual perception.
|
| 2629 |
LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents
2609.36371
|
cs.AI
|
Yuning Han, Yangchenchen Jin, Tyler Jandreau, Jingwei Sun |
Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows fi...Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows first apply an LLM-based execution-free (EF) verifier to filter candidates before running tests, which adds another model pass over every trajectory. We introduce LatentSift, a token-free and execution-free filter that replaces this first stage with hidden states the policy already produces while generating the candidates. It represents each candidate through its reasoning, observation, and function-call states, compares them with positive and negative banks of such states collected from successful and unsuccessful trajectories during policy training, and fuses the resulting distance scores with a learned linear score to retain promising candidates for the execution-based stages. On SWE-bench Verified, across three agents and two policy sizes, LatentSift cuts EF-verifier tokens by 66.6--81.0% and total verification tokens, which include test generation, by 49.1--62.1% at K=16, while hybrid Best@16 matches or improves on each agent's reference workflow, rising from 59.26% to 60.06% on DeepSWE-Preview.
|
| 2630 |
Audience-Bound Persistent Memory: Authorization Across the Memory Lifecycle
2609.36373
|
cs.AI
|
Sibo Liu |
A personal language agent that acts for its owner across private and shared conversations can learn a fact from one audience and later place it in the context it assembles for another. We study authorization before context across the whole memory lifecycle. Ea...A personal language agent that acts for its owner across private and shared conversations can learn a fact from one audience and later place it in the context it assembles for another. We study authorization before context across the whole memory lifecycle. Each memory item carries the audience present when it was recorded; derived items are partitioned by audience, receive the intersection of their sources' audiences, or are suppressed; an audience widens only by an explicit, object-specific grant; and an item enters a model attempt only when every current viewer belongs to one of its authorized audiences, with unresolved viewers failing closed to public-only. Under explicit identity, provenance and complete-mediation assumptions, this admission is sound and policy-complete on the exact assembled context, enforced by exclusion rather than by model behavior. We realize it in two independently persisted reference architectures, a flat store and a relationship graph, and, descriptively, in a native agent-memory runtime. In a prospectively frozen confirmation over 10,000 multi-party histories, no forbidden item entered any architecture's context, whereas unscoped retrieval exposed forbidden items in 82% of its contexts. Entitled recall matched policy-equivalent baselines exactly and exceeded unscoped retrieval by 0.30 Recall@5, with a Holm-confirmed advantage that grows with distractors. No architecture produced a wrong-principal substitution, but unscoped substitutions were too rare to establish the prespecified joint decision.
|
| 2631 |
Quantization Enables Private Dense Retrieval against Malicious Service Providers
2609.36376
|
cs.AI
|
Louis Tremblay Thibault, Sofiane Azogagh, Marc-Olivier Killijian, Ulrich A\"ivodji |
Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector representations of queries and passages from a large corpus. In privacy-sensitive applications, the server observes the ...Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector representations of queries and passages from a large corpus. In privacy-sensitive applications, the server observes the query and controls which evidence is returned, creating both confidentiality and integrity risks. We formulate private dense retrieval as providing query privacy and retrieval integrity against a malicious server, and develop a two-round cryptographic protocol that provides both guarantees. Our protocol reduces private and verifiable retrieval to multiplication of a committed matrix by an encrypted vector and uses low-bit quantization to make this computation practical. We evaluate the resulting trade-off between cryptographic cost, retrieval quality, and downstream RAG accuracy across six embedding models, four language models, and corpora of up to 2.68 million passages. Our results show that, with a clipped quantizer, three-bit quantization largely preserves retrieval quality and downstream accuracy, while a private query over a corpus the size of a clinical reference requires one to three minutes of server time. These results suggest that private dense retrieval is already practical for moderately sized, privacy-sensitive corpora when minute-scale latency is acceptable.
|
| 2632 |
Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
2609.36393
|
cs.AI
|
Muhang Tian, Sherry Yang |
Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE)...Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate -- the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then charges each action for the time it consumes at that rate. We first conduct theoretical analysis in the bandit setting to establish that RPG approximates the optimal reward rate and empirically demonstrate it outperforms baselines while avoiding enumeration over the policy space, a known issue for an existing method. We then further apply RPG on a small language model (Qwen3.5-4B) with self-improvement loops and empirically show it obtains higher rewards within a fixed time budget than vanilla RL on MLE-Bench and NanoGPT, with a 19.2% and 85.7% margin, respectively. Our method provides a practical solution for optimizing performance under wait time considerations in modern agentic RL tasks, where actions interact with external environments and cost time.
|
| 2633 |
Where Should Physics Enter a Molecular Crystal Generator?
2609.36398
|
cs.AI
|
Haocheng Tang, Junmei Wang, Wengong Jin |
Generative models make molecular crystal structure prediction fast, but their samples still exhibit geometric and packing violations. Physics can be introduced during training, post-training, or inference, yet these choices are rarely compared with the generat...Generative models make molecular crystal structure prediction fast, but their samples still exhibit geometric and packing violations. Physics can be introduced during training, post-training, or inference, yet these choices are rarely compared with the generator and physical signal held fixed. We introduce CrystAF, an all-atom crystal flow-map generation model, and use it with the UMA interatomic potential to systematically study where physics should enter. Post-training learns physical preferences directly into CrystAF, improving molecular validity and crystal packing while leaving sampling unchanged: physics is paid for once during training rather than repeatedly at deployment. In contrast, UMA relaxation is effective at repairing local clashes but makes generation 6--26$\times$ slower, while learning from relaxed targets provides little benefit. These routes are complementary rather than competing. Physics-informed post-training first shifts the generated distribution toward more physically reasonable structures, after which inexpensive inference-time corrections further remove clashes and restore stereochemistry that the generator cannot represent. Importantly, the same post-training strategy also improves the multi-step all-atom Clari-M and rigid-body MolCrystalFlow generators, demonstrating transfer across architectures and representations. Together, our results suggest a simple principle: learn reusable physical alignment into the generator, and reserve inference-time physics for residual constraints that are better corrected than learned.
|
| 2634 |
Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV
2609.36399
|
cs.AI
|
Bushra Asseri, Abdulaziz Asseri |
Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Modul...Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12 matched American personas and without a persona, in English and Arabic, under eight ways of formulating the request (288,000 answers). JEV's answers were highly repeatable (ICC 0.997), and without a persona they resembled those of its own American personas. When the persona was Saudi rather than American, the answers moved in the direction of the human Saudi-US difference, reproducing 87% of its size in English but 62% in Arabic, with long-term orientation reversed. A language cross shows that the smaller difference in Arabic comes from the language of the items, not from the language of the persona description. Age shifted the profiles about as much as nationality, gender shifted them more for Saudi than for American personas, and JEV was less confident in Arabic and for Saudi personas. These patterns held in every request design, although the model never generates text.
|
| 2635 |
Longer Records, Broader Invariance: The Hidden Scaling Problem in Longitudinal Contrastive Learning
2609.36409
|
cs.AI
|
Rameen Mahmood, Xuhai "Orson" Xu, Zachary Beattie, Jeffrey Kaye, Danny Yuxing Huang |
Longitudinal data are valuable because people change. Yet the objectives used to learn from these data can inadvertently erase that change. In person-level contrastive learning, observations from the same person are treated as positives; as records grow, those...Longitudinal data are valuable because people change. Yet the objectives used to learn from these data can inadvertently erase that change. In person-level contrastive learning, observations from the same person are treated as positives; as records grow, those positives can span increasingly distant---and increasingly different---behavioral states. More history can therefore produce not only more data, but broader invariance. We show that this distinction is fundamental. We separate \emph{record span}, how much history the learner sees, from \emph{supervision span}, how far across that history positive-pair supervision reaches. Across in-home sensing records spanning up to 2.7 years, broader supervision systematically suppresses recoverable changing-state information, even when the available history is held fixed. At the broadest span, less than 10\% of the information recoverable from an untrained encoder remains. Yet keeping positives local is not sufficient: as records grow, even distant states that are never paired become increasingly similar. Explicitly contrasting other observations from the same person reverses this loss without shortening the record, revealing a second route by which longitudinal scale can broaden invariance. Finally, we prospectively reproduce the supervision-span effect in 199 GLOBEM participants. Longitudinal scale therefore presents a choice: more history need not mean more invariance. By controlling what is held invariant as records grow, we can preserve the change that made the longitudinal data valuable in the first place.
|
| 2636 |
Reliable Parallel Decoding in Masked Diffusion Language Models
2609.36452
|
cs.AI
|
Zhenghao He, Bohan Liu, Guangzhi Xiong, Aidong Zhang |
Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is relia...Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.
|
| 2637 |
Channel-Dependent State Space Model for Multivariate Time Series Forecasting
2609.36453
|
cs.AI
|
Yu-Cheng Wu, Fan-Keng Sun, Li-Chun Lu, Duane S. Boning |
Multivariate time series forecasting (MTSF) is critical across many real-world domains. Existing deep learning approaches fall into two paradigms with distinct limitations: channel-independent (CI) methods unconditionally ignore cross-variable dependencies and...Multivariate time series forecasting (MTSF) is critical across many real-world domains. Existing deep learning approaches fall into two paradigms with distinct limitations: channel-independent (CI) methods unconditionally ignore cross-variable dependencies and model only temporal dynamics, while channel-dependent (CD) methods consider both but typically rely on architectural compromises to mitigate overfitting and computational overhead. We therefore propose Chameleon, a specialized CD state space model (SSM) that enables data-dependent, fine-grained interactions across variables while scaling linearly with their number. By connecting selective SSMs with the Kalman filter, we leverage the missing measurement update in the former for cross-variable modeling while preserving the SSM backbone for robust temporal modeling. We further identify favorable inductive biases of GatedDeltaNet for time series, adapt it as our backbone, and improve generalization through additional techniques, including a previously unexplored stochastic perturbation of reversible instance normalization. On strongly dependent ODE and PEMS datasets, Chameleon achieves the best MSE and MAE across all settings, while its CI ablation and prior CD methods incur 61-178% higher MSE on average. Across 28 standard benchmark settings, Chameleon also achieves better MSE and MAE than each baseline in at least 27 and 22 cases, respectively. Training-time and peak-memory analyses on Traffic and ETT further demonstrate competitive efficiency and favorable memory scalability across different variable counts.
|
| 2638 |
Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension
2609.36460
|
cs.AIcs.SD
|
Maral Ebrahimzadeh, Gilberto Bernardes, Sebastian Stober |
Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In...Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In parallel, data-driven methods such as skip-gram have been used to learn chord embeddings from symbolic corpora, but their ability to recover tonal structure and its relation to tonal tension remains underexplored. In this work, we investigate how skip-gram chord embeddings reflect tonal structure and whether they provide a useful basis for analyzing structural aspects of tonal tension. Using chord sequences with and without transposition-based augmentation, we evaluate the learned spaces from geometric, functional, and tension-related perspectives. We show that augmented embeddings exhibit strong transposition equivariance, recover a clear circle-of-fifths structure, and support interpretable shifts between key-related regions of the learned space. We then derive embedding-based measures from chord-to-key distance and contextual chord-distance relations, and show that they capture meaningful aspects of tonal tension structure through correspondence with matched tonal measures and moderate alignment with human tension profiles. Across analyses, transposition-based augmentation generally improves the stability, tonal coherence, and interpretability of the learned space.
|
| 2639 |
Guard Models Are Overconfident Where Base Models Are Uncertain
2609.36477
|
cs.AI
|
Jonghyun Hong, MinJae Jung, Minwoo Kim |
Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions. We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degra...Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions. We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections. Comparing each guard with its corresponding base LM, we find that uncertainty signals often remain available, with the base model typically expressing uncertainty on the same inputs where the guard fails. Layer-wise analyses localize this guard-base divergence to later layers, where guard models exhibit sharper safe/unsafe separation and lower-rank representations, while adversarial harmful inputs lie closer to the clean-safe region. These findings highlight a mismatch between guard confidence and base model uncertainty under attack.
|
| 2640 |
Quantum Computing for Network Security Classification: Near-Term Classification and Long-Term Memory Efficiency
2609.36479
|
cs.AI
|
Yuqing Li, Poonam Bala Nehru, Yunpeng Zhang, Danindu Gammanpilage, Xin Jin |
Quantum computing has already been explored in several network-security applications. However, how quantum computing may contribute to network-security classification in both the near term and the longer term has not been systematically discussed. This paper s...Quantum computing has already been explored in several network-security applications. However, how quantum computing may contribute to network-security classification in both the near term and the longer term has not been systematically discussed. This paper studies this question through two complementary experiments. First, we evaluate near-term quantum-kernel support vector machines (SVMs) on practical network-security classification tasks and compare them with classical SVM baselines on KDD Cup 1999, CICIDS2017, and BoT-IoT. Across these runs, quantum kernels are competitive. They can match or improve classical baselines in some settings, while classical RBF kernels remain stronger in others. This suggests that near-term quantum-kernel methods should be evaluated as practical, dataset-dependent alternatives to classical kernels rather than as uniformly superior replacements. Second, we use quantum oracle sketching (QOS) to study a longer-term memory advantage for classification with streaming classical samples. In QOS, samples are processed online and used to incrementally construct an approximate quantum oracle, which provides coherent query access for downstream quantum algorithms without retaining the entire dataset. Under the QOS-inspired machine-size estimate, comparable accuracy corresponds to a substantially smaller effective memory-size proxy than explicit sparse/QRAM-style storage. Compared with a simple streaming proxy, the result is more nuanced because aggressive feature filtering can make the streaming dimension small. This suggests that the long-term value of quantum computing for network-security classification may lie in memory-efficient data access rather than immediate runtime speedup. Together, these experiments show how quantum computing may contribute to network-security classification from near-term classification performance and longer-term memory efficiency.
|
| 2641 |
LLMs Learn to Evade Latent Monitors from Prior Feedback Alone
2609.36490
|
cs.AI
|
Hugo Lyons Keenan, Christopher Leckie, Sarah Erfani |
Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent's internal activations rather than its outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to...Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent's internal activations rather than its outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to the model about how its internal states are being evaluated. We ask whether an agent can infer the monitor's decision rule from this feedback and then selectively edit its activations to evade detection. Unlike prior evasion attacks, the model is never explicitly told what the monitor detects. Surprisingly, off-the-shelf models already produce activation edits aligned with the monitored direction, but at insufficient magnitude for evasion. Simply scaling up these edits by a factor of 8 reduces the monitor's TPR from 100% to 27%. A rank-1 LoRA amplifies this behavior into effective evasion within the forward pass, reducing TPR further to 4% on held-out concept monitors while leaving other concepts at their normal detection rates. Capabilities on standard benchmarks are retained under this finetuning, and the evasion skill survives retraining the monitors on the new activations. Mechanistically, we find evidence that the model computes its activation edit from the prior in-context turns, and show that the edit becomes more aligned with the monitored direction as more examples are provided. These results demonstrate feedback-conditioned control over activations and suggest that latent monitoring should be treated as an interactive process in which agents can observe and respond to oversight measures.
|
| 2642 |
InterBias-SV: Compound Conditions in Speaker Verification
2609.36500
|
cs.AIcs.SD
|
Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal |
Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: join...Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its results artefact contains 4,068 scored records across 17 experiments, 12 encoder labels, and six speech corpora, totalling 12 million trial evaluations. Three experiment families contain the same-corpus terms needed to compute additive contrasts. For labels assigned to speaker-trained encoders, their mean contrasts are +0.0026, +0.0088, and +0.0024 in equal error rate (EER), with larger variation across settings. These descriptive averages do not establish equivalence to additivity: trial matching, checkpoint identity, and parts of the condition metadata remain unverified. We also examine two interpretation problems. Near-chance EER can make additive predictions difficult to interpret, but chance performance is not a hard EER ceiling, and correlation with the prediction does not identify a saturation mechanism. Ratios of demographic gaps are unstable when their clean reference is near zero; absolute gaps provide a more direct summary. The benchmark provides condition definitions, analysis scripts, and explicit requirements for interpretable compound-condition comparisons, while separating recomputable summaries from claims that require further experimental validation.
|
| 2643 |
Towards Breaking the Learning System Wall Using Multimodal Tutoring Transcriptions
2609.36502
|
cs.AI
|
Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Ashish Gurung, Ishan Miglani, Shivang Gupta |
Past research using log data has faced the "learning system wall," whereby few methods exist for generalizing models of student learning across platforms. Increasingly, online learning is captured by richer forms of data, including dialog and video, with new a...Past research using log data has faced the "learning system wall," whereby few methods exist for generalizing models of student learning across platforms. Increasingly, online learning is captured by richer forms of data, including dialog and video, with new affordances. An example of this is remote tutoring programs, where human tutors support students who use learning systems while video conferencing. Toward better platform-general modeling of learning, we introduce an AI-driven multimodal transcription system that processes screen-recording videos into unified screenplay-style transcripts containing audio dialogue and annotated learning log actions. We describe a planned method for temporally aligning AI-generated multimodal transcripts with MATHia learning logs and for identifying and classifying student learning processes to align with MATHia logs. Lastly, we highlight challenges and potential solutions in capturing learning processes in one system, offering initial steps towards generalizing log data across diverse systems.
|
| 2644 |
Large-scale factor analysis shows machine intelligence is only partially interpretable
2609.36515
|
cs.AI
|
Faiz Ghifari Haznitrama, Afrizal Hasbi Azizy, Faeyza Rishad Ardi |
A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at...A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods. A robust pattern across different modes of bias is that 1. A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions, 2. Content-similar benchmarks do not necessarily cluster together, and 3. The $g$ factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard "intelligence" benchmarks. Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.
|
| 2645 |
LIBERO-MAX: Do Robot Policies Adapt When the World Changes?
2609.36518
|
cs.AI
|
Yunbei Zhang, Zijian Jin, Yuanzhe Liu, Janet Wang, Xilun Zhang |
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at rese...Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.
|
| 2646 |
Reliability Testing of Medical Model Performance under Distributed Deployment
2609.36525
|
cs.AI
|
Yifei Wang, Xiaohan Zhang, Youtao Ding, Tianlin Li, Xiaoyu Zhang |
Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, an...Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.
|
| 2647 |
Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations
2609.36526
|
cs.AI
|
Guanghui Min, Liang Wu, Mingjia Shi, Yinhan He, Mayank Darbari |
Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and c...Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons cannot isolate individual compressions and are confounded by agent stochasticity. We first find that compression degrades reliability before solvability. Using matched counterfactual continuations that compare execution from the same agent state with versus without compression, we further show that severe degradation concentrates at isolated compression events. Motivated by this finding, we propose PAIR (Prompt Adaptation using Interventional Rollouts) for adapting structured compression prompts. PAIR identifies individual compressions that degrade subsequent execution, diagnoses their effects, and revises the relevant sections of a fixed compression template. PAIR achieves the strongest cross-run reliability among compressed methods in every main benchmark-scope combination, consistently exceeding the competing prompt-adaptation baseline. Without modifying the downstream agent, PAIR brings compressed execution close to the no-compression baseline and sometimes numerically exceeds it.
|
| 2648 |
DraftTrace: A Multi-View Analytics Environment for AI-Integrated Writing
2609.36544
|
cs.AI
|
Divyansh Chandarana, Sandipan De, Vivek Gupta |
Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through which it was produced. We introduce DraftTrace, a writing environment that jointly captures three complementary vie...Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through which it was produced. We introduce DraftTrace, a writing environment that jointly captures three complementary views of writing: the final product, the writing process and interactions with an integrated AI-assistant. DraftTrace reconstructs how a document develops over time and organizes these signals into submission, longitudinal, and class-level analytics for instructors. We deployed DraftTrace in a graduate NLP course with 81 students and compared their sessions with LLM-generated responses entered by automated tools and with copy-typed responses. While product measures distinguish differences in text formulation, process measures distinguish differences in how text is entered. Considering both views together helps characterize cases such as copy-typing. Interaction traces show that students use the assistant differently across stages of writing: to clarify the question at an early stage and to verify answers at a later stage. A preliminary instructor survey highlights the importance of multi-view writing analytics and their interpretability.
|
| 2649 |
Interactive-Policy Distillation with Bidirectional Propose-and-Verify
2609.36546
|
cs.AI
|
Shutong Wu, Xiwen Chen, Brendan Rappazzo, Daiheng Zhang, Anderson Schneider |
On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts far from the teacher, causin...On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts far from the teacher, causing the teacher to be queried on states it would hardly visit and thus provide unreliable supervision. We propose Interactive-Policy Distillation (IPD), which applies adaptive teacher intervention to the student rollout. Under a bidirectional propose-and-verify state machine, the student and teacher alternately exchange their roles as proposer and verifier, and collaboratively generate mixed-source trajectories. Then different supervisions are applied according to the source of each token. This bidirectional propose-and-verify mechanism and the source-split loss make IPD not only a more performant distillation method, but also a unified bridge between on-policy and off-policy paradigms. To make the interleaved dual-model rollouts more efficient, we also design a dedicated fused inference engine that co-hosts both models in one serving instance with separate KV caches and instantiates the state machine model to distribute, collect, and process requests. On math reasoning tasks and across multiple teacher-student model pairs, student models trained with IPD not only outperform those trained with OPD, but also demonstrate higher data efficiency. Specifically, when distilling Qwen3-30B-A3B into Qwen3-1.7B-Base, IPD brings a +3.28 mean@8 and a +3.28 best@8 benchmark-averaged accuracy improvement compared with OPD. Besides, IPD only consumes about 1/4 of the training examples and steps to outperform OPD trained on the whole training dataset for one epoch. We also investigate the impact of different loss variants and takeover / handback configurations, and demonstrate the robustness of IPD on different training data.
|
| 2650 |
SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning
2609.36552
|
cs.AI
|
Zihao Chen, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai |
Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor t...Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.
|
| 2651 |
HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation
2609.36559
|
cs.AI
|
Yansong Liu, Rui Liu, Yuan Zuo, Hongwei Zhao, Da Fu |
Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early prefix of the timeline and then use a frozen model for all future timestamps. We argue that this fixed-prefix ...Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early prefix of the timeline and then use a frozen model for all future timestamps. We argue that this fixed-prefix protocol is misaligned with extrapolation. It learns from a static prefix, whereas the target stream is non-stationary: new entities and facts emerge, temporal dependencies shift across regimes, and recurring historical signals must be refreshed online. As a result, models trained only on early snapshots become outdated and degrade over long horizons. We address this mismatch by formulating extrapolative TKGR as continual learning over streaming snapshots. Under this view, effective extrapolation must jointly handle current dynamics, stable knowledge, and recurring historical evidence. Based on these requirements, we propose History-enhanced Two-Step Continual Learning (HiTS-CL), a backbone-agnostic continual learning framework for extrapolative TKGR. HiTS-CL tracks current dynamics via continual fine-tuning, preserves stable knowledge via multi-teacher adaptive distillation, and retains recurring historical evidence via a selective memory of recent and frequent facts. We integrate HiTS-CL into five representative TKGR backbones and evaluate it on four benchmark datasets. HiTS-CL consistently improves extrapolation accuracy, reduces long-horizon degradation, and outperforms strong continual-learning baselines, including a recent method for temporal knowledge graphs. Source code and data are available at https://github.com/liuyansong98/HiTS-CL.
|
| 2652 |
From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning
2609.36569
|
cs.AI
|
Yupeng Chang, Wenxuan Zhang, Yuan Wu |
Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data imp...Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.
|
| 2653 |
CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
2609.36570
|
cs.AI
|
Mark Russinovich |
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from ...Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages), held-out attack success falls from 0.21-1.00 undefended to 0.00-0.17 defended, and AgentDojo compromise rate from 0.10-0.49 to 0.006-0.079, at 93-100% typography-normalized benign utility, with larger task-dependent costs when reasoning over steered content. A benchmark-level adaptive attacker reaching 0.67-0.73 undefended is held to roughly a quarter of that on the two most deeply evaluated models. Among the defenses we measured on capable models, those achieving lower compromise rates either lost 22-89% of benign utility or fine-tuned the served weights. White-box gradient attacks through the deployed vector compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeeds. CounterSteer largely neutralizes instructional takeover: a black-box framing search cracks 3 of 18 development samples. Parameter manipulation--attacker-chosen arguments in otherwise legitimate calls--is only partially resisted (13 of 18); the decision becomes linearly readable at argument emission but not at the examined pre-generation sites, and is not removed by the tested prefill- or decode-time steering, motivating argument-provenance controls.
|
| 2654 |
Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models
2609.36577
|
cs.AIcs.SD
|
Zhenhong Zhou, Xuanyue Zhao, Youji Liu, Yuanhe Zhang, Xiaoyu Ma |
Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by...Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code
|
| 2655 |
Factorized Scheduling Principle: Learning Interpretable and Transferable Policies via Structured Additive Functions
2609.36578
|
cs.AI
|
Hong Je-Gal, Hyun-Suk Lee |
Scheduling problems arise from repeatedly selecting one item from a set of candidates based on their states. These problems often reduce to assigning priority scores and choosing the highest-ranked item. In this work, we propose a factorized scheduling princip...Scheduling problems arise from repeatedly selecting one item from a set of candidates based on their states. These problems often reduce to assigning priority scores and choosing the highest-ranked item. In this work, we propose a factorized scheduling principle (FSP) framework to learn interpretable and transferable scheduling rules. The FSP framework represents system states as condition distributions and decomposes a global scheduling principle into additive univariate and pairwise components with identifiability constraints. The scheduling principle enables the framework to maintain a simple priority-based structure during deployment. This principle is learned by using a policy-based objective combined with a temporal-difference signal defined on the condition distribution. Experiments on synthetic and realistic scheduling tasks demonstrate the FSP framework's strong performance, interpretability, and zero-shot generalization across different system scales.
|
| 2656 |
Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning
2609.36588
|
cs.AI
|
Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li |
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills requi...We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both $\pi_0$ and $\pi_{0.5}$ backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by $+23.1\%$, $+16.4\%$, and $+44\%$ on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.
|
| 2657 |
Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise
2609.36593
|
cs.AI
|
Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li |
Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text...Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case. Built on Genesis, Text2Sim uses a hierarchical agentic structure that combines a Planner with specialized Writers, asset-generation tools, and an independent Critic. Compact skills (Debug Cards) distilled from graphics demonstrations provide role-specific physical guidance for execution-based repair. We evaluate physical quality, visual quality, and human preference on 42 held-out prompts spanning rigid, articulated, deformable, and cloth phenomena, with a paper-level split between experience construction and evaluation. We design automatic physical and visual scorers to evaluate the quality of the results, and Text2Sim achieves higher scores than all four state-of-the-art baselines on both metrics. In blinded user studies with these baselines, significantly more participants prefer Text2Sim than prefer the baselines, which is consistent with the results from our automatic scorers. The pipeline also supports a broad range of downstream applications; we select dataset construction and extension to multimodal input as two representative examples. We will release the code, the Debug Card library, and a dataset of generated cases, each pairing the text prompt and rendered video with the executable program, assets, physical parameters, controls, and recorded states.
|
| 2658 |
Simple Agentic Memory for Generalist Robot Policies
2609.36595
|
cs.AI
|
Yuyou Zhang, Yunbei Zhang, Miao Li, Janet Wang, Zijian Jin |
Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered proced...Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history; and current-view grounding resolves recalled entities before execution. We evaluate SimpleARM on RoboMME, a benchmark of memory-dependent robot manipulation tasks that require history information no longer available in the current observation. Across all 16 tasks and three policy seeds, SimpleARM achieves 67.17% mean success, compared with 44.51% for the strongest non-oracle baseline. Matched ablations show mechanism specificity: removing relation, reference, progress, or route state produces large losses where the affected state is retrieved for control, while largely sparing other tasks. These results support a state-based view of robot memory: effective memory for control is not simply retained visual history, but compact task-relevant state derived from the interaction history.
|
| 2659 |
Second-Moment Stochastic Approximation Methods
2609.36600
|
cs.AI
|
Tao Jiang, Lin Xiao |
Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Ad...Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive second-moment stochastic approximation methods through the lens of optimal preconditioning for solving matrix equations, and develop a two-stage framework for their convergence analysis. The first stage focuses on the analysis of conceptual (impractical) methods that rely on the exact first and second moments. In the second stage, we replace the exact moments with their respective estimators, and invoke Dvoretzky's theorem to show that the resulting practical methods converge almost surely to a neighborhood of the target solution. The size of the neighborhood depends on the biases and variances of the first- and second-moment estimators. We derive concrete bounds for Muon and a spectral variant of Adam that determine the radius of their neighborhood of convergence.
|
| 2660 |
WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses
2609.36635
|
cs.AI
|
Haomin Qi, Xiangzhe Xu, Yiming Huang, Jingbo Shang, Chengpeng Wang |
Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark eval...Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.
|
| 2661 |
Inducing Process Supervision from Outcome-Only Reinforcement Learning
2609.36641
|
cs.AI
|
Shengda Fan, Xin Cong, Zhong Zhang, Haotian Chen, Yankai Lin |
Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo es...Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step-level labels and an outcome label. The reward depends solely on whether the predicted outcome matches the ground truth, and the resulting group-relative advantage is used to optimize the entire generated response. Intuitively, when checking intermediate steps helps determine the outcome, more accurate checks can lead to better outcome judgments and higher rewards. Outcome-only RL can therefore reinforce step-level verification without explicit process supervision. We validate the effectiveness of TIPS across math and agent benchmarks and four backbone families. Notably, TIPS-Qwen3-4B-Thinking-2507 reaches 85.2 F1 on ProcessBench with only 3.2K outcome-labeled trajectories, surpassing all evaluated trained PRMs and strong prompt-only judges such as GPT-5.4-Instruct and Claude-4.7-Opus, while still trailing o1-mini. Code and data are available at https://github.com/RUCBM/TIPS.
|
| 2662 |
Where Predictive Supervision Goes Shapes What VLA Policies Learn
2609.36645
|
cs.AI
|
Hanseul Kim, Jewon Yeom, Youngjoon Jeong, Minsoo Jo, Taesup Kim |
Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has ...Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy's visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.
|
| 2663 |
Constitutional adapters: Inference-time interventions for misalignment and misuse
2609.36657
|
cs.AI
|
Adam S. Lowet, Mark Kurzeja |
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we sho...Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
|
| 2664 |
AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization
2609.36662
|
cs.AI
|
Pengyu He, Yan Zhang, Ruien Li, Guangwen Yang |
The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication incre...The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing bottleneck. Local update methods reduce this cost by allowing workers to perform several optimizer steps between synchronizations. Most local update methods set the number of local optimizer steps between synchronizations before training and keep this interval fixed throughout the run. However, the best interval can change during the entire train process. If the interval and optimizer are adapted to the current training state, the communication frequency is reduced while maintaining the training performance. In this work, we introduce AutoLoCo, an adaptive training framework to reduce communication in LLM training. It adapts the local interval using scalar training statistics and corrects each outer update. Our method is motivated by two observations: 1) the appropriate local interval varies across training stages, and 2) changing the number of inner steps per interval creates a mismatch with an unchanged outer optimizer, requiring a correction to the outer update. We optimize this mismatch by correction of the outer optimizer for the momentum and the learning rate using the accumulated inner learning rate. Our experiments under communication constraints demonstrate that AutoLoCo reduces communication frequency by 27% relative to DiLoCo while maintaining training performance.
|
| 2665 |
MyoCodec: A Streaming Neural Codec for Electromyography
2609.36687
|
cs.AI
|
Jihwan Lee, Kleanthis Avramidis, Junhyeok Lee, Tiantian Feng, Najim Dehak |
Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and token-based sequence modeling. This paradigm has been widely adopted in modern speech and audio frameworks; howev...Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and token-based sequence modeling. This paradigm has been widely adopted in modern speech and audio frameworks; however, the biosignal domain still lacks a neural codec designed specifically for low-bitrate streaming and generalization across diverse downstream tasks. We present MyoCodec, a streaming neural codec designed for electromyography (EMG). Inspired by recent neural audio codecs, MyoCodec combines causal Transformers with residual vector quantization to encode continuous EMG signals into different levels of EMG representations spanning from continuous latent features to discrete tokens operating at 50 Hz. Trained on twelve public EMG datasets, MyoCodec achieves favorable performance in both intrinsic codec quality and representative downstream tasks, including typing (emg2qwerty), hand-pose (emg2pose), speech decoding (emg2speech), and speech-to-EMG synthesis (speech2emg). Across these tasks, MyoCodec exhibits strong performance against prior models while providing a compact and causal EMG representation. During streaming inference, it requires compute time of only 0.482 ms for each 20 ms frame, enabling real-time streaming. Also, the discrete token representation provided by MyoCodec has the potential to support integration into language-model based approaches, creating a path toward LLM-based interactive systems, where tokenized EMG representations are directly processed into such language or speech models. Code and model weights are released.
|
| 2666 |
Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
2609.36692
|
cs.AI
|
Zixuan Gong, Zeyu Gan, Jiaye Teng, Yong Liu |
Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternat...Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternative way to organize geometric information hierarchically, motivating the Normalize-Then-Precondition framework. Specifically, it first uses diagonal-Gram information to construct a marginally normalized update, then applies spectral preconditioning to its directional interaction geometry. Building on this framework, we develop NormPre with NormPre-G and NormPre-L adopting global and localized spectral preconditioning, grounded in spectral-norm steepest descent and a regularized formulation followed by leading mode selection, respectively. To enable large-scale training, NormPre-G uses Newton-Schulz iterations and NormPre-L employs randomized sketching to approximate the leading interaction eigenspace. Theoretically, we establish $\mathcal{O}(T^{-1/2})$ convergence guarantees for simplified versions of NormPre. Across extensive pretraining experiments on GPT-2 Small, LLaMA and Qwen3, both variants consistently outperform AdamW, Muon and MANO under matched training budgets. Further efficiency and spectral analyses reveal the complementary strengths of two variants and characterize their performance-efficiency trade-off. We open-source our code through a GitHub repository at https://github.com/zx-gong/NormPre.
|
| 2667 |
Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG
2609.36700
|
cs.AI
|
Pranav Handa, Ariful Azad |
When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant app...When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conversations. Our findings reveal that multi-turn interaction causes widespread performance degradation, incurring relative performance drops of up to 21% and increasing unreliability by 47%, making RAG systems simultaneously less accurate and less reliable. We identify two distinct failure modes behind this degradation. Systems are either lost in translation, where conversational rephrasing distorts the retrieval query, or lost in conversation, where retrieval succeeds but the LLM fails to synthesize evidence distributed across turns.
|
| 2668 |
Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change
2609.36739
|
cs.AI
|
Bravish Ghosh |
Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autola...Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge's own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.
|
| 2669 |
Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
2609.36750
|
cs.AI
|
Yiming Wang, Yikang Liu, Qingyuan Tian, Xingyu Chen, Zhuosheng Zhang |
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward re...Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.
|
| 2670 |
Beyond Conditional Independence: Root Cause Analysis with Deep Causal Models
2609.36771
|
cs.AI
|
Md Musfiqur Rahman, Kenneth Lee, Ziwei Jiang, Padmaja Jonnalagedda, Ruocheng Guo |
Root cause analysis (RCA) is a critical problem in many real-world scenarios. RCA enables the identification of faulty or failing mechanisms in a system by comparing anomalous observations with corresponding reference (i.e., regular) observations. However, exi...Root cause analysis (RCA) is a critical problem in many real-world scenarios. RCA enables the identification of faulty or failing mechanisms in a system by comparing anomalous observations with corresponding reference (i.e., regular) observations. However, existing approaches rely either on heuristic methods or on conditional independence tests with a strong unconfoundedness assumption, and thus fail to exploit other complicated distributional constraints in the presence of latent variables. To relax these assumptions, we model the underlying system as a causal model and the anomalous system as a change in the structural functions of the same causal model. Specifically, to handle unobserved confounders, we establish an implicit connection between distributional constraint testing and root cause analysis. To adapt our approach to data generated from arbitrary causal models, we employ the deep causal model (DCM) framework, in which we design the causal model using neural networks. Finally, we illustrate how our method, RCA-DCM, can utilize different levels of partial graphical knowledge to perform RCA. We evaluate RCA-DCM against state-of-the-art baselines on simulated datasets, a physics-based causal chamber and two micro-service applications. RCA-DCM improves top-1 accuracy over the strongest baseline on both Sock Shop (0.880 vs. 0.752) and Online Boutique (0.776 vs. 0.712), and when the true root cause in the causal chamber is unobserved and acts as a latent confounder, it recovers the exact root-cause set more often than any competing method (perfect recovery rate (PRR) 0.846 vs. 0.731).
|
| 2671 |
VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction
2609.36804
|
cs.AI
|
Yitong Han, Nankai Lin, Juan Luo, Hongyan Wu, Lianxi Wang |
Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles i...Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction between Chain-of-Thought (CoT) reasoning and self-consistency decoding, such that the benefits brought by CoT cannot be reliably transferred to final corrections. We propose Vote-guided Advantage Allocation for CSEC (VAA-CSEC), a multi-stage framework that combines CoT distillation, Supervised Fine-Tuning (SFT), Reinforcement Learning (RL) and self-consistency decoding. During RL, we design a task-specific reward function that directly aligned with the minimal-editing principle of CSEC. We further introduce Group-Level Relative Policy Optimization (GLPO), which reallocates GRPO advantages according to the margin between individual rollout rewards and the vote-aggregated group reward, aligning the RL training objective with the self-consistency objective used at inference time. Experiments on CSED-C and NaSGEC-Exam show that VAA-CSEC outperforms all LLM-based baselines on CSED-C with an F0.5 of 47.72%, achieves the highest recall of 42.15% among all methods, and establishes a new state of the art of 41.55% F0.5 on NaSGEC-Exam.
|
| 2672 |
Spotter: Let the Embodied Model Lead, and the VLM Reflect for It
2609.36808
|
cs.AI
|
Long Li, Qichao Zhao, Yue Yang, Fan Xu, Zhe Wang |
Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known err...Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known error, which requires producing a correction and judging whether it is right. Stopped at a failure and allowed to retry, they seldom repair it through their own randomness or from a language description of the error, and best-of-N selection cannot pick the successful candidate after a failure. We attribute this to training only on successful demonstrations and to inputs too narrow to show what went wrong, and conclude that reflection must come from a vision-language model (VLM), which takes in far more information, such as the episode history and text, and is more general. Prior VLM-led work has the VLM plan every step and invoke the embodied model as a tool, placing the VLM on the critical path. We propose Spotter, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control. We run Spotter with Qwen and with GPT as the VLM, and both improve the embodied models; with GPT, Spotter improves Cosmos Policy and $\pi_{0.5}$ by 5.6 and 7.5 percentage points on RoboCasa, and raises $\pi_{0.5}$ from 47.2% to 57.0% on the Hard setting of RoboTwin 2.0 and from 53% to 83% on a real robot. Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes only 13 to 16 s longer than with the embodied model alone and about 70% less time than with a VLM-led baseline using the same model. Our code is available at https://github.com/zqc3117/Spotter.
|
| 2673 |
Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows
2609.36812
|
cs.AI
|
Junhyun Ha, Juho Lee, Byungwoo Park |
Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behav...Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91\% after 500K environment steps.
|
| 2674 |
DSWM: Decomposed Spatio-Temporal World Model for Demand-Driven UAV Base Station Repositioning
2609.36845
|
cs.AI
|
Shengjie Zhong, Zhongliang Zhao, Jingxuan Chen, Xianbin Cao, Xinmei Qiang |
Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model....Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model. We cast demand-driven fleet repositioning as latent-space decision-time planning and propose DSWM, a decomposed spatio-temporal world model: an agentic controller that perceives the demand field through a rolling observation window, retains operational context in a latent recurrent state, reasons about candidate motions by imagined rollouts under an uncertainty penalty, and coordinates the fleet through replanned first actions. DSWM learns a recurrent state-space model shaped by an exponential-moving-average (EMA) based latent predictive objective with variance regularization. It attaches a differentiable service simulator that replays the association, probabilistic line-of-sight channel, and Shannon rate chain inside latent rollouts. Planning uses a cross-entropy method whose imagined demand is anchored on the current observation window with mixing coefficient $\rho=0.95$. On a unified pipeline over three real datasets (Milan CDR (call detail record), Shanghai Telecom, YJMob100K) and 14 methods including five reproduced IEEE baselines, DSWM attains weekday served ratios of 0.889, 0.908, and 0.898, ranking first among non-ablated configurations on every dataset. On Milan it improves over the strongest non-learning baseline (Greedy, 0.780) by 0.109, a margin that comes from decision-time use of observations rather than prediction accuracy.
|
| 2675 |
Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue
2609.36849
|
cs.AI
|
Omar Sheta, Rinku Deuja, Hadi Masoudi, Minghong Fang |
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single promp...Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced gradient and a fixed unsafe reference direction, and their effectiveness in multi-turn dialogue remains unclear. We conduct a controlled evaluation of gradient-based jailbreak detection in multi-turn settings. We extend GradSafe with a Context Window Scanner that applies the detector to fixed-size windows of user turns and uses the maximum window score as the conversation-level score. We evaluate different window sizes, attack families, benign conversation distributions, and target models. The results differ sharply between synthetic and realistic benign settings. Against synthetic benign conversations, the detector achieves an ROC-AUC of 0.98 on human-authored multi-turn jailbreaks. On WildChat benign conversations, ROC-AUC drops to 0.76, and a threshold calibrated on synthetic data flags more than 90% of benign conversations as unsafe. Under realistic benign distributions, single-turn windows give the highest separability, whereas longer windows and accumulated contexts reduce performance. The detector is also sensitive to the attack-generation method and target model: successful Crescendo attacks receive scores comparable to or lower than benign conversations, and Qwen2.5-7B-Instruct yields near-random separability with a different optimal window size. These findings show that gradient-based signals can support multi-turn jailbreak detection, but reliable deployment requires calibration on realistic benign conversations, short-window scoring, length-aware thresholds, and evaluation across attack types and model architectures.
|
| 2676 |
Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning
2609.36862
|
cs.AI
|
Muhammad Zeeshan Akram, Mufid Kamel Marican, Anvesh Reddy Yenugu, Ali Zain Kaimkhani, Minghong Fang |
Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignme...Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem at different levels of the model. Vaccine improves the robustness of hidden embeddings to the representation shifts induced by harmful fine-tuning, whereas Booster simulates harmful weight updates and attenuates their effect during alignment. We investigate whether these mechanisms are complementary and propose VaccineBooster, a single alignment procedure that combines embedding perturbation and weight-level gradient attenuation within each training step. On Llama-2-7B aligned with BeaverTails and then attacked through poisoned fine-tuning, VaccineBooster achieves the lowest OpenAI moderation score among the compared defenses, 0.315, while a Booster-Only variant retains the highest post-attack refusal rate, 50%. Together with ablations over the embedding-perturbation and gradient-attenuation strengths, these results indicate a trade-off: embedding perturbation primarily reduces flagged harmful content, whereas gradient attenuation primarily preserves explicit refusal behavior. Because our evaluation uses ten prompts and a single unseeded run per configuration, we report this trade-off as an observed pattern rather than a statistically resolved effect. These results provide practical guidance for prioritizing content safety or refusal retention when aligned models are exposed to untrusted fine-tuning.
|
| 2677 |
SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs
2609.36879
|
cs.AI
|
Haoran Ou, Gelei Deng, Xuanye Zhang, Wenbo Guo, Tianwei Zhang |
As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide special...As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growing adoption of third-party Skills introduces a new supply-chain attack surface. Malicious Skills can embed harmful behaviors that abuse agent privileges and compromise the agent execution environment or accessible resources. Although recent LLM-based malicious Skill auditing approaches have achieved promising performance, they often rely on capable commercial LLMs. How to achieve effective auditing with compact, locally deployable LLMs in security-sensitive and resource-constrained settings remains largely unexplored. Our investigation reveals that compact LLMs struggle to identify malicious behaviors hidden in complex Skill packages. This difficulty arises from both the implicit nature of such behaviors and the limited reasoning capacity of compact LLMs. To address these challenges, we propose SKILLLITE, an evidence-guided agentic framework for malicious Skill detection. SKILLLITE effectively extracts security-relevant behaviors and infers the intended functionality from complex Skill packages. It then employs a compact LLM to assess the maliciousness of the Skill based on the observed behaviors and their functional context. Experiments show that SKILLLITE improves malicious Skill detection across different compact LLM backbones and outperforms existing representative auditing baselines. Its effectiveness generalizes to behaviorally confirmed in-the-wild malicious Skills. Meanwhile, SKILLLITE maintains a low inference latency, supporting its practical deployment.
|
| 2678 |
Digital Twin Modeling of Quantum Dynamical Systems: Dissipative Quantum Reservoir Computing
2609.36901
|
cs.AI
|
Abhijit Sen, Bikram Keshari Parida, Shital Chauhan, Mahima Arya, Denys I. Bondar |
Modeling the response of driven many-body quantum systems from input--output data is difficult: the dynamics are nonlinear, history dependent, and expensive to simulate as system size grows. A paradigmatic case is High-Harmonic Generation~(HHG), where a strong...Modeling the response of driven many-body quantum systems from input--output data is difficult: the dynamics are nonlinear, history dependent, and expensive to simulate as system size grows. A paradigmatic case is High-Harmonic Generation~(HHG), where a strong field drives a medium to emit radiation that is highly sensitive to the drive and encodes long-range temporal correlations. We introduce a dissipative quantum reservoir computing~(DQRC) framework that builds a digital twin of such a system, learning its input--output map directly from data while the reservoir---itself a small open quantum system---stays fixed and only a classical readout is trained. We show that a minimal single-qubit reservoir reproduces the HHG response of a substantially larger Ising spin chain, and on a representative benchmark matches and on several metrics surpasses previously reported temporal convolutional and Kolmogorov--Arnold-network models, while using a simpler, physically realizable system. A single fixed reservoir further generalizes across a broad range of drives, indicating that it learns a shared physical response structure rather than memorizing trajectories. These results establish dissipative quantum reservoirs as compact, physically grounded digital twins for nonlinear, memory-dependent quantum dynamics. Code is available at \href{https://github.com/AI-and-Quantum-Computing/DQuRC}{https://github.com/AI-and-Quantum-Computing/DQuRC}.
|
| 2679 |
MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
2609.36903
|
cs.AIcs.SD
|
Ke Wang, Houxing Ren, Zimu Lu, Yunqiao Yang, Zhuofan Zong |
End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios su...End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existing long-audio benchmarks focus on passive listening and speech-to-speech benchmarks are mostly short and dyadic. We extend the Moshi paradigm jointly along the long-horizon and multi-party axes in English and Chinese. First, we release 57.6k hours of synthetic training data ($\href{https://huggingface.co/datasets/MultiTalk/MultiTalkPT}{MultiTalkPT}$ and $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkFT}{MultiTalkFT}$) for long-form, multi-party, English-Chinese full-duplex dialogue, with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference. Second, we introduce $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkBench}{MultiTalkBench}$, built from real human recordings, for evaluating long-form, multi-party, bilingual full-duplex dialogue. Conversations average 32.6 minutes and include probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style model that sustains coherent multi-party English-Chinese conversations over extended durations and substantially outperforms open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench.
|
| 2680 |
BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR
2609.36913
|
cs.AI
|
Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn |
Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contex...Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.
|
| 2681 |
AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations
2609.36915
|
cs.AI
|
Rui Huang, Yanlin Mu, Lidong Li, Yucong Wang, Zichen Yan |
Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introd...Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between manipulation and flight, continuously changing observations, and safety-critical physical interactions. These challenges demand diverse training data and systematic policy evaluation, yet collecting demonstrations and evaluating policies directly on physical aerial platforms are costly, difficult to scale, and hard to repeat under controlled conditions. We present AeroManip-VLA, a scalable benchmark for aerial VLA data generation and policy evaluation. AeroManip-VLA provides a GPU-accelerated simulation framework with low-level payload-aware flight and manipulation control in massively parallel environments. Building on this framework, we combine reusable reinforcement learning policies with expert task rules to automatically generate demonstrations without human teleoperation across diverse objects, environments, and randomized initial conditions. The generated data include basic skills such as grasping and placing, as well as long-horizon tasks that require both navigation and manipulation. We further introduce automated event labeling and trajectory categorization to filter demonstrations. These mechanisms enable fine-grained analysis of task progress, behavioral outcomes, and safety-related failures. Finally, we evaluate a range of imitation learning and VLA baselines across different task settings, revealing their performance characteristics and failure modes. Together, AeroManip-VLA enables scalable aerial manipulation data generation, structured trajectory analysis, and systematic VLA evaluation in simulation prior to real-world deployment.
|
| 2682 |
State Transport Routing for Short-horizon Adaptation in Multi-horizon Photovoltaic Forecasting
2609.36926
|
cs.AI
|
Xu Yuqing, Zhou Liguo, Sun Ze, Yu Lei, Jiang Mingming |
Recent power measurements provide valuable information for photovoltaic(PV) power forecasting, but directly extrapolating short-term trends can introduce substantial errors over longer forecast horizons. To address this challenge, we propose state transport ro...Recent power measurements provide valuable information for photovoltaic(PV) power forecasting, but directly extrapolating short-term trends can introduce substantial errors over longer forecast horizons. To address this challenge, we propose state transport routing (STR), a lightweight adapter that refines the predictions of a frozen forecasting model. STR combines the original forecast with two complementary trajectories derived from the latest measured power level and its recent trend. A horizon-conditioned router adjusts their contributions over the first 120 min, while leaving subsequent predictions unchanged. Experiments on four public PV datasets show that STR consistently outperforms a parameter-matched residual adapter. On PVDAQ, the same approach improves five neural forecasting backbones, reducing all-horizon normalized mean absolute error by 0.0201-0.2364 percentage points, with paired 95% confidence intervals excluding zero. No reliable improvement is observed for LightGBM. These findings demonstrate the potential of structured state adaptation to improve short-term forecasting across different neural architectures without retraining the underlying models or altering their longer-horizon predictions.
|
| 2683 |
Safe-by-Design Learning via Energy-based Neural Networks
2609.36942
|
cs.AI
|
Simone Betteti, Morteza Lahijanian, Luca Laurenti |
Learning neural-network models of dynamical systems with safety guarantees is a fundamental requirement for their deployment in safety-critical settings. Safety is commonly established by proving the invariance of a desired subset in state-space, ensuring that...Learning neural-network models of dynamical systems with safety guarantees is a fundamental requirement for their deployment in safety-critical settings. Safety is commonly established by proving the invariance of a desired subset in state-space, ensuring that every trajectory initialized in this subset remains confined to it for all time under admissible inputs. Existing frameworks, however, either rely on computationally expensive post-hoc verification or employ safety-enforcing mechanisms without formal correctness guarantees. In this paper, we introduce a novel neural architecture grounded in energy-based modern Hopfield networks to guarantee safety-by-design while retaining sufficient expressiveness to model complex nonlinear dynamics. Specifically, we integrate modern Hopfield networks with a port-Hamiltonian neural ODE, enabling by design the construction of barrier functions yielding explicit admissible-input sets and quantitative robustness radii. Across several benchmarks, including an 12-dimensional nanodrone model, our framework achieves state-of-the-art performance while producing certified invariant sets that are more robust to external solicitations than comparable existing approaches.
|
| 2684 |
ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models
2609.36952
|
cs.AI
|
Jingnan Pu, Zi-En Fan, Feng Lian |
Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive a...Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive architecture (JEPA). However, strong alignment does not necessarily lead to accurate, stable predictions. To address this, we propose ER-JEPA, which adds an episodic replay path to LLM-JEPA. ER-JEPA stores training pairs in a memory. At each step, it stores and retrieves relevant data to provide additional supervision. This enables learning from both the current batch and stored training pairs, providing additional supervision for token prediction and representation alignment. Experiments across multiple datasets (NL-RX, GSM8K, Spider, and NQ-Open) demonstrate that ER-JEPA consistently outperforms LLM-JEPA.
|
| 2685 |
Purlin: Separating Orchestration from the Datapath of Collectives
2609.36954
|
cs.AI
|
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis |
Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath...Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.
|
| 2686 |
Controlled Decoding Attacks on Black-Box LLMs
2609.36956
|
cs.AI
|
Jesson Wang, Shawn Li, Wei Yang, Franck Dernoncourt, Ryan A. Rossi |
Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that retur...Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.
|
| 2687 |
VStress: Correlation-Aware Auditing and Adaptive Budget Allocation for Repeated Verifiers
2609.36958
|
cs.AI
|
Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang |
Repeated verifier calls are useful only when they contribute conditional information. We introduce VStress, an auditable replay contract, and VStress-CA, a correlation-aware allocation policy that estimates the conditional marginal information of an unqueried ...Repeated verifier calls are useful only when they contribute conditional information. We introduce VStress, an auditable replay contract, and VStress-CA, a correlation-aware allocation policy that estimates the conditional marginal information of an unqueried verifier on a sealed calibration split, discounts uncertainty, normalizes by call cost, and stops or abstains when the next call is not informative. The controller freezes its decision and cost ledger before joining the clean oracle; a dependence-shift alarm disables channel preference and falls back to exact-stop. The controlled audit gives the mechanism boundary: at 35% symmetric corruption, majority-5 improves balanced accuracy from 0.6578 to 0.7739, whereas at 65% it loses 0.1226 points. In the matched fixed-budget comparison, breadth, redundancy, and adaptive allocation obtain balanced accuracies 0.6048, 0.6375, and 0.6538, with 3.4216 calls per item and an RLVR score of 0.6417 for VStress-CA. Dependence diagnostics also increase from same-model repeats to cross-family channels, with conditional marginal gains of 0.0126, 0.0462, and 0.0913. These measurements turn correlation from a post-hoc warning into an auditable allocation decision.
|
| 2688 |
JudgeCast: Time Series Forecasting with Experience-Informed Covariate Judgements
2609.36966
|
cs.AI
|
Donguk Kwon, Wooseok Jeong, Dongha Lee |
Covariate effects vary across contexts and shift over time, requiring forecasters to assess how to use them for each forecasting context. As forecasting proceeds, observations for earlier forecasts become available, providing feedback on past covariate use for...Covariate effects vary across contexts and shift over time, requiring forecasters to assess how to use them for each forecasting context. As forecasting proceeds, observations for earlier forecasts become available, providing feedback on past covariate use for subsequent forecasts. However, when multiple covariates act together, the forecast error reveals the numerical discrepancy from the observation but not how the covariates should have been used. We introduce JudgeCast, an experience-based framework for time series forecasting with covariates. Following the judgmental adjustment practice, a frozen TSFM provides the base forecast, while a frozen LLM uses the current context and relevant experience to adjust it. Within the adjustment, assessing covariate effects and determining the numerical adjustment serve distinct roles, so JudgeCast first forms explicit covariate-wise judgments and then determines the adjustment. After observation, JudgeCast uses the observed residual of the base forecast to reconstruct alternative judgments and evaluates the original and alternatives through their resulting adjustments. The best-performing decision is selected and retained as validated experience for subsequent forecasts. Across diverse real-world datasets, JudgeCast outperforms strong baselines. Ablations show that explicit covariate-wise judgment can improve forecast-time adjustment, while residual-guided experience construction yields more reliable forecasting gains than retaining raw decisions as experience.
|
| 2689 |
SRJudge: Empowering Large Language Models with Selective Reasoning for Fine-Grained Knowledge Concept Tagging
2609.36982
|
cs.AI
|
Zhiwei Yang, Jiahua Yang, Huiru Lin, Xing Chen, Quanlong Guan |
Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners in traditional and online teaching practices. Recent work has explored large language models (LLMs) for this ta...Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners in traditional and online teaching practices. Recent work has explored large language models (LLMs) for this task, achieving promising performance. However, LLMs still struggle to select the correct concept from a large-scale candidate set due to the high dimensionality of the decision space. In this paper, we propose a novel three-stage Select-Reason-Judge (SRJudge) framework, which empowers LLMs with selective reasoning capability for fine-grained knowledge concept tagging. Specifically, the Selector in Stage 1 first narrows the candidate concepts to a top-K shortlist by fine-tuning a small language model (SLM), e.g., BERT, since the top-$K$ predictions hit the correct concept in most cases, thereby reducing the decision space of correct candidates. Next, the Stage 2 Reasoner employs a lightweight LLM for refined reasoning over the shortlisted candidates. It further integrates an improved reinforcement learning strategy with a dynamic task-specific reward function and a pruning mechanism to better align with human reasoning preferences. Finally, a larger LLM acts as a judger that evaluates the overall rationality of the reasoning process and its explanations to determine the final output. In addition, we construct two high-quality datasets for further validation, i.e., the biology dataset S_Bio and the physics dataset S_Phy. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines across benchmark datasets, verifying its effectiveness and superiority. Resources are available at: https://github.com/Nicozwy/SRJudge.
|
| 2690 |
Abductive World Modeling via Causal Representation Learning
2609.36985
|
cs.AI
|
Ziqi Liu, Songhan Yang, Linfan Zhou, Jiatong Liu, Lijun Peng |
The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting the...The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), which organizes the inferred world state into three complementary components - Entity, Dynamic, and Relation - capturing what exists, how it changes, and how entities interact, respectively. By jointly reasoning over the current observation and its predicted future, HASP abductively infers these latent factors and integrates them into a structured state representation for downstream reasoning. To the best of our knowledge, AWM is the first framework to introduce abductive state inference into latent-space world modeling for learning structured representations of world dynamics. Experiments across physical prediction, causal reasoning, and action understanding demonstrate the effectiveness of our approach. Compared with V-JEPA, a state-of-the-art latent-space world model, AWM improves physical prediction AUROC by 10.7%, causal reasoning accuracy by 16.8%, and action Top-1 accuracy by 68.0%.
|
| 2691 |
Cross-Organizational SysML Model Integration: A Survey of Challenges and AI-Supported Tasks
2609.37000
|
cs.AI
|
Zirui Li, Torsten Brix, Stephan Husung |
Cross-organizational collaboration is widely regarded as a key promise of SysML-based Model-Based Systems Engineering (MBSE), yet practitioners still face persistent challenges when exchanging and integrating system models. In parallel, Large Language Models (...Cross-organizational collaboration is widely regarded as a key promise of SysML-based Model-Based Systems Engineering (MBSE), yet practitioners still face persistent challenges when exchanging and integrating system models. In parallel, Large Language Models (LLMs) raise expectations for AI-assisted model understanding and integration, while reliability and required human oversight continue to pose challenges. This paper reports the results of an online questionnaire survey with 29 MBSE stakeholders involved in cross-organizational collaboration. Respondents rated eight predefined integration challenge categories and six AI-supported task types on five-point Likert scales. The results indicate that stakeholders perceive model integration as a multi-dimensional alignment problem across semantics, behavior, traceability, and exchange interoperability. These perceptions vary by organizational role and frequency of integration involvement. AI is rated highly useful for analysis tasks such as semantic structure analysis and inconsistency detection, and respondents predominantly prefer human-in-the-loop use with mandatory verification. These findings motivate AI support that enhances, rather than replaces, engineering responsibility in SysML-based integration.
|
| 2692 |
OPFL: Optimistic Verification of Federated Learning via Empirical Boundary
2609.37011
|
cs.AI
|
Hongxu Su, Jianzhu Yao, Xuechao Wang, Pramod Viswanath |
Federated learning enables multiple clients to collaboratively train models without sharing their private data. However, the lack of visibility into local training makes it difficult to verify whether clients follow the prescribed training procedure or submit ...Federated learning enables multiple clients to collaboratively train models without sharing their private data. However, the lack of visibility into local training makes it difficult to verify whether clients follow the prescribed training procedure or submit malicious updates, such as model poisoning. A natural approach is to replay client training for verification. However, privacy-preserving replay produces numerical results that cannot be directly matched with local client execution because the two run in different environments. We present OPFL, an optimistic verification framework for privacy-preserving federated learning. To protect data privacy, OPFL performs replay inside secure multi-party computation (MPC). Although gradients computed on MPC and local GPUs are not bitwise identical, we observe that their absolute differences are stable and bounded. OPFL therefore calibrates an empirical boundary offline and uses it to distinguish benign numerical deviations from malicious manipulation. To reduce the cost of expensive MPC replay, OPFL adopts optimistic verification by post auditing only sampled training steps. Experiments on LeNet, BERT, and Qwen show that the boundary generalizes across datasets, input lengths, and GPUs, while achieving $0$\% ASR against model poisoning and PGD-based attacks. On a LeNet workload, at $p=0.01$, OPFL is approximately $98.6\times$ faster than full MPC-based FL and $625.5\times$ faster than ZK-based approach.
|
| 2693 |
LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
2609.37029
|
cs.AI
|
Hao-Yuan He, Peng-Fei Liu, Si Shen, Ming Li |
Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage...Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting fixed-size, multiscale views from the target's verification pass, thereby eliminating the need for a growing persistent state. Extensive evaluations demonstrate that LongSpark achieves state-of-the-art end-to-end efficiency across multiple model scales and realistic serving conditions. Notably, it delivers the lowest time-per-output-token on long-context tasks while reducing the drafter's context state by several orders of magnitude.
|
| 2694 |
Selecting The Most Informative Tokens in Natural Language Autoencoders
2609.37040
|
cs.AI
|
Federico Torrielli, Gianluca Barmina, Andrea Blasi N\'u\~nez, Amon Rapp, Luigi Di Caro |
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4....Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
|
| 2695 |
Beam Search as Test-Time Self-Distillation via Counterfactual Contexts
2609.37041
|
cs.AI
|
Su Ee Tan, Xiaotong Ji, Rasul Tutunov, Haitham Bou-Ammar, Matthieu Zimmer |
Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. H...Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient updates and access to expert demonstrations, making it inapplicable at inference. We propose test-time self-distillation, a decoding-time method that extracts a steering signal from the self-distillation framework without any parameter updates, reward models, or training data. Our key insight is that counterfactual contexts, i.e. fixed textual templates that hypothetically prime the model for excellent versus poor reasoning, can substitute for the demonstration. The log-odds ratio of a candidate answer under these two counterfactual conditions defines a new reward signal. We derive the optimal KL-regularized policy under this reward, which takes the form of a Gibbs reweighting of the base distribution. Crucially, this reweighting is global: it cannot be decomposed into independent per-token operations without ignoring future trajectory quality. We therefore approximate the target distribution via beam search. Experiments on mathematical reasoning (MATH500), code generation (HumanEval), and graduate-level science QA (GPQA) across multiple model scales show that test-time self-distillation improves over standard sampling, low temperature, beam search and power sampling baselines on average, demonstrating that the self-distillation principle can be operationalized at inference time.
|
| 2696 |
Evolving Towards Better Codes: LLM-Guided Search for High-Distance Binary Linear Codes
2609.37056
|
cs.AI
|
Amal Seddas, Vladyslav Shashkov, Maryna Viazovska, Emmanuel Abbe |
Evolutionary program search driven by large language models (LLMs) has produced record-breaking constructions for open problems in combinatorics and beyond. We apply this approach to the longstanding problem of improving the best-known bounds for binary linear...Evolutionary program search driven by large language models (LLMs) has produced record-breaking constructions for open problems in combinatorics and beyond. We apply this approach to the longstanding problem of improving the best-known bounds for binary linear codes. Building on the EvoTune evolutionary framework and the ShinkaEvolve codebase, we introduce LinCodeEvolve, which evolves code-construction programs against an exact minimum-distance evaluator. A strategy loop combines diversity-driven search and expert supervision: when progress plateaus, new strategies are used to redirect the search. LinCodeEvolve discovers seven record-breaking codes, $[172,21,66]$, $[173,20,68]$, $[176,21,68]$, $[181,21,70]$, $[184,21,72]$, $[189,22,72]$ and $[200,21,77]$, six of which have concise quasi-cyclic descriptions. With standard code modification techniques, they improve $22$ entries of the tables. Every code is verified by exhaustive enumeration. These results suggest that LLM-guided search can help find improved codes and complement existing methods in coding theory.
|
| 2697 |
A Comprehensive View of Fairness through Distributional Stability
2609.37061
|
cs.AI
|
Gayane Taturyan, Charlotte Laclau, Stephan Cl\'emencon |
We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it re...We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under this perspective, several classical notions of fairness arise as stability with respect to specific perturbations, with the associated unfairness gap given by a Lipschitz constant of a prediction-rate functional. This formulation also yields guarantees that hold uniformly over a range of demographic compositions at test time, without requiring knowledge of the deployment distribution. It leads to a learning procedure based on convex combinations of reweighted predictors, formulated as a second-order cone program, for which we establish generalization bounds. Experiments on standard benchmarks illustrate the approach.
|
| 2698 |
Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning
2609.37066
|
cs.AI
|
Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu, Said Mammar |
Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or mem...Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.
|
| 2699 |
FACT: Fidelity-Aware Construction of Articulated Twins
2609.37067
|
cs.AI
|
Kuixiang Shao, Chuansen Nie, Yinuo Bai, Jiayuan Gu, Jingyi Yu |
Visually plausible articulated assets may still fail during contact interactions or exhibit inaccurate motion. We present FACT (Fidelity-Aware Construction of Articulated Twins), an agentic framework that progressively constructs articulated twins to improve g...Visually plausible articulated assets may still fail during contact interactions or exhibit inaccurate motion. We present FACT (Fidelity-Aware Construction of Articulated Twins), an agentic framework that progressively constructs articulated twins to improve geometry, contact, and dynamic fidelity. The agent drives an evidence--diagnosis--revision loop on a shared editable representation, selecting measurements and model edits using quantitative feedback, while numerical tools execute and validate the updates. It reconstructs editable articulated geometry from images through feature planning, targeted measurements, and diagnostic refinement. On this reference, it repairs collision proxies through task-aware local repartitioning before fidelity-constrained compression. Finally, it constructs response models from passive-response videos, using simulation residuals to guide model revision and constrained physical parameter fitting. Experiments show that FACT improves geometric reconstruction over baselines, enables more reliable interaction with simpler collision proxies, and better reproduces held-out physical responses than direct parameter inference.
|
| 2700 |
Predictive Safety Curricula for Robust Legged Locomotion
2609.37070
|
cs.AI
|
Ivan Ovinnikov, Pascal Sutter, Christian Gehring, Jordis Herrmann |
Rare but consequential failures can persist in learned locomotion policies for legged robots even when average task performance is high, in part because standard curricula primarily adapt task difficulty rather than the distribution of safety-critical experien...Rare but consequential failures can persist in learned locomotion policies for legged robots even when average task performance is high, in part because standard curricula primarily adapt task difficulty rather than the distribution of safety-critical experience. We introduce Predictive Safety Curricula (PSC), a framework for allocating locomotion training experience using learned predictions of future safety cost. PSC trains a distributional safety critic from policy rollouts and uses its predictions to prioritize both terrain contexts and previously encountered randomized events. The resulting curriculum modifies the training distribution while leaving the task reward and policy-optimization loss unchanged. We evaluate PSC in controlled rough-terrain locomotion and in production locomotion systems. PSC improves reliability relative to standard terrain progression, advantage-based replay, and learning-progress curricula, with the largest gains on difficult terrain and under degraded observations. The same allocation principle transfers to two production locomotion stacks. On ANYmal-D hardware, PSC reduces shank-collision incidence by $63\%$ relative to the learning-progress curriculum across three matched training seeds, with a reduction in every seed. On a production stair-climbing platform, PSC eliminates observed shank collisions in the evaluated hardware trials. These results show that learned predictions of future safety cost can provide an effective signal for allocating training experience toward rare failure modes and improving locomotion reliability.
|
| 2701 |
Identifying ODEs from Unstructured Data with Causal Representation Learning
2609.37083
|
cs.AI
|
Alessandro Trenta, Riccardo Massidda, Davide Bacciu, Sara Magliacane |
We study the problem of recovering the governing ODE of a dynamical system from unstructured, high-dimensional observations such as images. Existing methods for ODE discovery typically assume direct measurements of the variables, or do not provide theoretical ...We study the problem of recovering the governing ODE of a dynamical system from unstructured, high-dimensional observations such as images. Existing methods for ODE discovery typically assume direct measurements of the variables, or do not provide theoretical guarantees on the learned variables and equations. While Causal Representation Learning (CRL) methods provide guarantees on identifying variables from high-dimensional observations up to component-wise diffeomorphisms, we show that in general these variables cannot be used directly as input to equation discovery methods, which typically assume that the variables will lead to sparse equations. So we introduce SParse Equivalent Equation Discovery AutoEncoder (SPEED-AE), a framework that combines a pretrained CRL method with a component-wise autoencoder that learns transformations of variables that are amenable to sparse ODE discovery. We show that for polynomial ODEs, this additional step allows us to restrict the identifiability of each variable from polynomial to monomial diffeomorphisms. Experiments on Lotka-Volterra, Lorenz, and a two-pendulum system show that SPEED-AE improves on the disentanglement of the CRL methods and that it recovers ODEs that are closest to the ground truth, while achieving state-of-the-art forecasting performance.
|
| 2702 |
ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum
2609.37085
|
cs.AI
|
Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi, Pantelis Frangoudis |
Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, be...Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.
|
| 2703 |
What Does Post-Training Change in Multilingual Reasoning?
2609.37104
|
cs.AI
|
Hongyang Li, Xiao Li, Caesar Wu, Gr\'egoire Danoy, Pascal Bouvry |
Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of perf...Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered.
|
| 2704 |
VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses
2609.37105
|
cs.AI
|
Jiexing Qi, Yu He, Jun Liu, Qichen Huang, Shaohua Hu |
Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for...Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.
|
| 2705 |
Designing a Boundary Negotiating Artifact for Collaborative Socio-Technical Sense-Making in AI Regulatory Sandboxes
2609.37109
|
cs.AI
|
Idoia Landa-Oregi, Tom Deckenbrunnen, Alessio Buscemi, Daniele Pagani, German Castignani |
The rapid, unpredictable advancements in AI system capabilities has seen regulators take adaptive and experimental approaches to policymaking. Established in other domains as instruments balancing regulation with innovation, regulatory sandboxes are seen as so...The rapid, unpredictable advancements in AI system capabilities has seen regulators take adaptive and experimental approaches to policymaking. Established in other domains as instruments balancing regulation with innovation, regulatory sandboxes are seen as solutions for AI regulation. However, analyses mostly focus on the legal and institutional design of AI Regulatory Sandboxes (AIRSes). With the legal framework leaving the socio-technical interpretation to stakeholders, this creates a gap on the sense-making required to fulfill the AIRS purpose. In this paper, we approach this by designing a Boundary Negotiating Artifact as a way to mediate meaning in AIRSes. Through Research-through-Design we iteratively develop a tool, providing an interface for the different stakeholders to collaborate in AI assessment. We then position it as technical backbone in established AIRS frameworks, structuring the collaborative sense-making of the involved stakeholders. We further report the insights gained from our design process leaving the qualitative evaluation for future work.
|
| 2706 |
Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio
2609.37116
|
cs.AIcs.SD
|
Gouthaman KV, Shiv Gehlot, Vishnu Raj, Lars Villemoes, Arijit Biswas |
Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptua...Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited channel configurations and cannot be directly applied to higher-channel-count audio. This raises the question: how can pretrained perceptual knowledge be effectively reused for multichannel spatial audio? Using 5.1-channel audio, we study four levels of multichannel integration: signal, prediction, latent, and feature and propose two learned approaches: latent-level aggregation of spatial-group representations and the feature-level Feature-Band Group Attention (FGAtt), which adaptively fuses spatial groups at the feature level before perceptual processing. Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge
|
| 2707 |
Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training
2609.37119
|
cs.AI
|
Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy |
Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has lear...Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels. Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes. We further show that binarizing the debiased score stops the policy from exploiting the critic's length bias. Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead. This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost: because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.
|
| 2708 |
SQUARE: Structured Quantum Representation Adapters as Compact Quadratic Feature Maps for Frozen Language Models
2609.37134
|
cs.AI
|
Emily Jimin Roh, Hyojun Ahn, Hoyeong Lee, Soohyun Park, Sung Whan Yoon |
Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensiona...Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensional bottleneck? Common linear and low-rank adapters remain linear at the adaptation module itself, whereas explicit second-order alternatives introduce pairwise interactions through direct parameterization or predefined factorizations. We propose SQUARE, a Structured QUAntum REpresentation adapter that amplitude-encodes the bottleneck vector, applies a parameterized quantum circuit, and measures the resulting state. We show that each basis-probability feature is exactly a normalized quadratic form in the bottleneck coordinates, while the additional Pauli-$Z$ readouts are signed linear combinations of these probabilities. The measured map can therefore parameterize interactions over $O(d^2)$ coordinate pairs through a small set of shared circuit parameters, where $d$ is the bottleneck dimension. It provides a structured parameterization within, rather than beyond, the classical normalized-quadratic feature class. In a disjoint same-pipeline evaluation over eight GLUE-derived controlled interaction tasks and five shared seeds, SQUARE achieves an average test accuracy of $0.7565$, compared with $0.7355$ for an affine normalized-quadratic predictor, $0.7271$ for the evaluated parameter-matched Givens mixing model, $0.6817$ for an MLP, and $0.6155$ for a frozen-circuit control. Under reduced supervision, it also shows consistent gains over the strongest evaluated classical comparator, with the same qualitative pattern across multiple frozen LM backbones. All circuit experiments use simulation, while the learned feature map can be evaluated exactly in batched PyTorch without quantum hardware.
|
| 2709 |
Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust
2609.37156
|
cs.AI
|
Ziqi Wen, Ting Xu, Lianyu Wang, Xian Lin, Yanda Meng |
World models enable agents to learn and plan in imagination, but predictions beyond their experience can become unreliable and mislead decisions. Existing uncertainty estimates derived from predictions can remain overconfident on unfamiliar state-action pairs....World models enable agents to learn and plan in imagination, but predictions beyond their experience can become unreliable and mislead decisions. Existing uncertainty estimates derived from predictions can remain overconfident on unfamiliar state-action pairs. We propose the Lucid World Model (LucidWM), which learns doubt from experience and propagates trust through imagination. By integrating Subjective Logic into categorical latent transitions, LucidWM distinguishes predicted outcomes from their evidential support and assigns each transition a degree of doubt. The complement of this doubt defines transition-level trust, which accumulates multiplicatively along imagined trajectories to reweight returns for policy learning and guide action selection. Uncertainty estimation requires no additional parameters or forward passes. Evaluated on four base world models against seventeen uncertainty readouts, LucidWM detects environmental changes and signals uncertainty during action-corrupted rollouts. In a controlled navigation case study, acting on trust reduces the number of steps required to reach the goal from 362 to 190. Fifteen demonstration videos show how LucidWM doubts its dreams and acts on that doubt. Videos are available at https://lucidwm.github.io.
|
| 2710 |
Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories
2609.37169
|
cs.AI
|
Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu |
Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, whi...Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how much compute mid-training absorbs. We revisit how this compute should be allocated to a single run or multiple similar optimizations. We find that branches forked from a shared checkpoint under various controlled recipe reaches measurably different regions of parameter space, and establish a form of compatible diversity that extending one run cannot supply. Therefore, we introduce Trajectory Soup, which distributes a mid-training budget over several independent branches, and consolidates strongest checkpoints selected on validation through intra- and inter-trajectory averaging into a single model. A local bias and variance analysis separates the two averaging levels, showing that inter-trajectory averaging removes residual error beyond the reach of averaging within a trajectory, while checkpoint selection carries a bias that bounds how many checkpoints are worth merging. Across model scales, learning-rate schedules, token budgets, and trajectory counts, Trajectory Soup improves aggregate downstream performance over the strongest single-trajectory average under matched budgets and keeps improving as budgets expand, with the advantage preserved after an identical post-training pipeline. These results position trajectory allocation and merging as a practical way to extend the compute-scaling frontier of mid-training beyond serial saturation.
|
| 2711 |
Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation
2609.37170
|
cs.AI
|
Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang, Guangting Wang |
Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's dis...Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in between. We introduce \textbf{Interpolated Policy Distillation (IPD)}, which defines the next-token distribution at every decoding step as an explicit linear interpolation between the student and teacher distributions. The interpolation operates at the distribution level, token by token, and its coefficient provides direct control over the balance between trajectory quality and student learnability. Naively sampling from this policy would require sequentially querying the teacher at every token and is thus expensive. To make IPD practical, we accelerate it with a new speculative-decoding rule while exactly preserving the interpolated next-token distribution.At the trajectory level, the resulting rollouts naturally interleave student- and teacher-generated segments. Unlike recent heuristic segment-interleaving methods, however, this interleaving is induced by an exactly realized token-level interpolated policy rather than by hand-designed switching rules. Across text-only and multimodal reasoning benchmarks, IPD consistently outperforms both endpoint policies (SFT and OPD), their conventional two-stage combination (SFT-then-OPD), and recent heuristic segment-interleaving methods, demonstrating that token-level policy interpolation better balances trajectory quality and student learnability.
|
| 2712 |
EgoHumanoid-V2: Human-to-Humanoid Transfer of Coordinated Whole-Body Skills for Loco-Manipulation
2609.37181
|
cs.AI
|
Jin Chen, Yiming Jiang, Chongyang Xu, Modi Shi, Shijia Peng |
Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordin...Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid-V2, the first egocentric human-to-humanoid skill transfer framework for coordinated whole-body loco-manipulation. At its core, coarse-to-fine action alignment combines kinematic reference correction with dynamics-aware refinement. It improves end-effector pose accuracy while preserving whole-body coordination. We also use robot-arm rendering and training-time image augmentation to reduce the visual embodiment gap and improve viewpoint robustness. On four real-world tasks, vision-language-action (VLA) policies trained on aligned human data show zero-shot skill transfer without target-task robot demonstrations. Task scores are comparable to those of policies trained on teleoperation data at a lower collection cost. These results support human data as direct skill supervision.
|
| 2713 |
ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents
2609.37196
|
cs.AI
|
Yanjie Li, Xiangyu He, Xuelong Dai, Bin Xiao |
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leav...Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call. ToolFence provides two key advantages. First, its fine-grained provenance-aware authorization enables the system to distinguish user-authorized values from untrusted observations, effectively addressing the within-tool attack. Second, its deterministic fast path and capability-level runtime grants substantially reduce the frequency of expensive judge calls, improving runtime efficiency. On AgentDojo with Qwen3-max, ToolFence reduces overall ASR to near zero with only a 3.80 percentage-point clean-utility drop and practical runtime overhead.
|
| 2714 |
CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?
2609.37216
|
cs.AI
|
Yue Pan, Jiawei Li, Ziyuan Zhang, Xiangxin Zhao, He Ye |
Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are...Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible but untrustworthy comments. We further present Sentinel, a repository-grounded agentic judge that actively gathers code evidence to verify review comments before making judgments. Starting from Qwen3-Coder-30B-A3B-Instruct, Sentinel is trained on the CRJudgeBench training split through iterative action-level learning from a privileged teacher. On the 359-instance CRJudgeBench test set, Sentinel achieves 76.60\% accuracy, outperforming GLM-5.3 by 6.13 percentage points and its base model by 19.78 points. These results show that even state-of-the-art general-purpose LLMs struggle to identify untrustworthy comments, while iterative action-level learning substantially improves the accuracy of repository-grounded trustworthiness judgments. Our dataset is available at https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark
|
| 2715 |
Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods
2609.37218
|
cs.AI
|
Zewen Sun, Tongyang Zhao, Liyao Xiang, Mingxuan Ma, Lingzhe Wang |
Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated senten...Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated sentences without explicitly accounting for the highly non-uniform and context-dependent semantic preference of LLM generation. When these two preferences are poorly aligned, many natural continuations become incompatible with the watermark, causing semantic narrowing: reduced semantic freedom, increased resampling cost, and potential degradation on tasks with strict semantic requirements. To alleviate this problem, we propose HammingMark, which uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood. Defining watermark validity over a Hamming neighborhood in compact hash space retains a larger fraction of naturally likely semantic continuations. The coarse many-to-one hash mapping further allows diverse semantic realizations to remain watermark-valid. Experiments on C4 and BookSum show that HammingMark achieves strong robustness, high detectability, and near-unwatermarked generation quality, requiring only 2.2 sampled candidates per accepted sentence,a 72.8% reduction compared with the most sampling-efficient existing method. On more complex tasks with strict semantic constraints, HammingMark achieves the highest detection rates with the highest or tied-highest ROUGE-L scores, demonstrating its effectiveness in balancing watermark detectability and generation quality under constrained generation settings.
|
| 2716 |
Follow the Entities: A Corpus Map for Agentic Search
2609.37226
|
cs.AI
|
Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam |
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LL...Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.
|
| 2717 |
DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis
2609.37233
|
cs.AI
|
Yuan Li, Hanyun Jiang, Guowei Tian, Chengpeng Wang, Peisen Yao |
Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route...Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated. We present DatalogBench, a benchmark of 136 text-to-Datalog synthesis tasks curated from existing Datalog-based artifacts. Synthesized programs are graded by execution on held-out inputs against an oracle validated by mutation analysis. Across six LLMs and four prompting configurations, exact match peaks at 68.4%, and relation descriptions or an input-output example have only modest, model-dependent effects. Under direct prompting, most failures occur at compile time, typically because a model invents auxiliary predicates that it never declares or types consistently. Two coding agents reach up to 83.8% and eliminate nearly all such failures, leaving mostly semantic errors concentrated in recursive tasks. DatalogBench thus identifies recursive reasoning and decomposition as open challenges for current LLMs and agents, and offers a reliable, execution-grounded measure of both.
|
| 2718 |
Loss-Guided Pretraining Data Selection for Time-Series Foundation Models
2609.37255
|
cs.AI
|
Yike Li, Shaoxu Song, Jianmin Wang |
Time series foundation models (TSFMs) are pretrained on heterogeneous collections containing billions of observations, yet their training windows are typically sampled without estimating whether they provide useful learning signal. We introduce a static data-s...Time series foundation models (TSFMs) are pretrained on heterogeneous collections containing billions of observations, yet their training windows are typically sampled without estimating whether they provide useful learning signal. We introduce a static data-selection framework that scores each window with a reference forecaster and retains an intermediate interval within every source dataset. Specifically, we connect forecasting loss to optimization difficulty by showing that normalized squared loss controls the per-sample gradient norm under a local Jacobian condition. We then define a reference loss score and apply dataset-stratified selection to preserve the diversity of samples. Across various TSFM architectures, our method outperforms random selection by an absolute margin and even improves both relative MASE and CRPS over full-data pretraining by retaining fewer candidate pretraining windows. Further analyses show strong cross-scale and cross-architecture score correlations, indicating that a small reference model can often select data for larger targets, provided that the reference and target share compatible difficulty orderings.
|
| 2719 |
Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring
2609.37312
|
cs.AI
|
Mohammadali Mohammadkhani, Madhava Krishna, Yash Sarrof, Michael Hahn |
Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed ...Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold depending on model size, successfully solving the task necessarily leaks a near-linear amount of information about the covert task input into the CoT. Therefore, sufficiently complex hidden computation always leaves an information-theoretic footprint. However, concerningly, this leakage need not be readable: Under plausible cryptographic assumptions, even a one-layer Transformer can encrypt its reasoning online so that no polynomial-time monitor can extract information about the hidden computation. Overall, our theoretical and empirical results provide a holistic view of both the opportunities and the limitations of CoT monitoring.
|
| 2720 |
Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
2609.37315
|
cs.AI
|
Rohith Reddy Bellibatlu, Zichong Wang, Wenbin Zhang |
Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publ...Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool's advertised surfaces as an executable contract, check the implementation against it, and trace each score's provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have written. Across 34 audited mutating tools in four benchmarks we confirm seven tool defects and one evaluator property at pinned commits. On injected defects the checker raised no false positive in 25 flags, flagged 2 of 5 negative controls, and missed most: in 29 of 33 scored misses a clause covered the defect but no probe revealed it. The checker's own static half, run alone, flags 14 of 17 confirmed sites, so on these findings the dynamic half confirms and traces rather than discovers. Twelve further AgentDojo tools, with six held-out tools and the seven audited first, complete its 25-tool mutating surface, on which at least 5 tools diverge from their advertised surface as our contracts read it, a rate for AgentDojo alone. No gold trajectory reaches either tau2-bench defect; on 1,120 paths built to isolate the telecom defect, a number fixed by construction, the evaluator rewards a refuel of a suspended line and fails the repaired tool. The clearest case is a clinical benchmark whose tool tells the agent each write executed under a documented no-write design its interface does not disclose; its grader takes that message as evidence, so its action success rate records whether a request carried the expected payload, not whether any record changed.
|
| 2721 |
PowerMarketJax: A JAX Benchmark Suite for Multi-Agent Reinforcement Learning in Power Markets
2609.37321
|
cs.AI
|
Zhanhua Pan, Xin Qin, Xiao Liu, Zhilong Cao, Jianhong Wang |
Power markets are a natural testbed for multi-agent reinforcement learning (MARL), where multiple self-interested participants repeatedly submit bids. A market-clearing mechanism then determines dispatch and prices subject to power grid constraints and market ...Power markets are a natural testbed for multi-agent reinforcement learning (MARL), where multiple self-interested participants repeatedly submit bids. A market-clearing mechanism then determines dispatch and prices subject to power grid constraints and market settlement rules. However, existing MARL environments typically focus on a single market setting, implement simplified clearing mechanisms, or rely on CPU-based optimization solvers that slow large-scale training and limit the systematic study of bidding strategies and market behavior. We introduce PowerMarketJax, a benchmark suite for MARL across five power markets: day-ahead wholesale, real-time balancing, ancillary services, peer-to-peer double auctions, and local flexibility. Each environment implements its own clearing, pricing, and settlement rules while providing a common framework for learning and evaluation. We find that learned bidding behavior depends strongly on the market design: independent learners can miss better strategies when gains require many agents to change together, when more profitable strategies lie beyond a region of lower profit, or when profits disappear as more agents adopt the same strategy. PowerMarketJax implements both market simulation and policy training in JAX, allowing the entire pipeline to run on the GPU with 1,024 X 1,200 parallelisms across both environments and market participants, achieving up to 33X speedup over CPU-based baselines. Our open-source benchmark is available at: https://github.com/powermarketjax/PowerMarketJax.
|
| 2722 |
A Sharp Transition in Data Reconstruction under Differential Privacy
2609.37344
|
cs.AI
|
Max Cairney-Leeming, Simone Bombari, Marco Mondelli |
Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides...Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single $d$-dimensional training sample from a $\rho$-zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at $\rho \asymp d$ for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for $\rho \ll d$; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for $\rho \gg d$. Remarkably, the transition moves to $\rho \asymp s$ for data lying in an $s$-dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).
|
| 2723 |
Port-Hamiltonian Latent Deliberation: Mitigating the Deliberation Drift Cliff in Test-Time Compute Scaling
2609.37351
|
cs.AI
|
Zeyu Jia (School of Biomedical Engineering, Technology, Tianjin Medical University, Medical School, Tianjin University) |
Test-time compute scaling has emerged as a cornerstone of advanced machine reasoning, yet performing iterative deliberation directly within continuous latent representation spaces reveals a catastrophic pathology: the Deliberation Drift Cliff. While unconstrai...Test-time compute scaling has emerged as a cornerstone of advanced machine reasoning, yet performing iterative deliberation directly within continuous latent representation spaces reveals a catastrophic pathology: the Deliberation Drift Cliff. While unconstrained recurrent latent models achieve initial reasoning gains at short horizons (K <= 4), their reasoning collapses when extrapolated to deeper thinking steps (K >= 16), dropping by 22% to 62% across standard logical benchmarks. We resolve the trilemma among expressivity, Lyapunov stability, and computational efficiency in test-time latent reasoning through a 22-round empirical and theoretical investigation. We demonstrate that strictly conservative scalar potential gradient flows suppress long-range drift (cliff 3.40%) but bottleneck peak reasoning accuracy at 32.73%, whereas unconstrained rotational flows achieve high symbolic expressivity (82.33%) but suffer a severe 36.87% drift cliff. To resolve this geometric duality, we establish Port-Hamiltonian Latent Deliberation (PH-LD) and propose the Direct-Gradient Pure-Tensor Helmholtz-Hodge Decomposition (DG-HHD). DG-HHD parameterizes the attracting flow as a tangent projection tensor network while orthogonally decoupling non-zero circulation (Hodge machine error 1.65e-17, contraction error 5.55e-17), eliminating runtime autograd dependencies to achieve 1.84x vector field and 2.09x RK45 rollout speedups. In a 15-arm symmetrical Pareto benchmark, DG-HHD achieves 58.67% peak accuracy (+25.94% absolute gain over conservative HHD) and retains 35.27% at K=32. Transferred to small language model (SLM) multi-hop causal reasoning, DG-HHD delivers monotonic compute scaling (49.33% to 51.56%) and suppresses out-of-distribution drift (cliff -0.66%). All 30 Level 0 deterministic invariants are certified.
|
| 2724 |
Compiling Learning Problems into Adaptation Programs for Language Models
2609.37371
|
cs.AI
|
Rebecca Ramnauth, Brian Scassellati |
Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what extent a model should adapt as ...Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what extent a model should adapt as a joint prediction and decision problem. Rather than searching over candidate programs anew for each learning episode, a compiler learns from prior adaptations to predict a vector-valued counterfactual response surface over candidate programs---their expected effects on acquisition, transfer, boundedness, and preservation---and selects a program before adaptation begins. Because this predicted geometry captures multiple behavioral consequences rather than a single winner or scalar score, it can be reused under different downstream priorities without retraining. Across five learning types, preferred programs vary meaningfully across episodes, and this variation is predictable from pre-adaptation information. On Llama-3.1-8B, compiler-selected programs approach exhaustive search while outperforming global and objective-specific defaults. Replication on Gemma-2-9B preserves program heterogeneity and selection headroom, but shows that exploiting this headroom requires accounting for uncertainty when departing from strong defaults. Together, these results show that adaptation search can be amortized across related learning problems, turning prior adaptation experience into a basis for deciding how future learning should occur.
|
| 2725 |
Complexity-Aware Evaluation of LLM Comprehension
2609.37405
|
cs.AI
|
Ali Mohammadi Esfahani, Nafiseh Kahani, Samuel A. Ajila |
Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can concea...Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex. This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume. We evaluate DeepSeek-Coder-V2 and Llama through two complementary tasks: automatic input-output prediction over 300 Python functions and manually assessed semantic comprehension over a balanced subset of 60 functions. The functions are grouped into Low-, Medium-, and High-complexity bands. DeepSeek-Coder-V2 achieves an overall automatic accuracy of 78.33%, compared with 70.33% for Llama. However, accuracy decreases substantially from Low to High complexity, from 93.52% to 52.78% for DeepSeek-Coder-V2 and from 87.04% to 47.22% for Llama. Incorrect predictions are consistently associated with higher values of all four complexity metrics, and correlation and logistic-regression analyses confirm broadly comparable negative associations between structural complexity and correctness. Manual semantic comprehension shows the same degradation pattern, with accuracy decreasing from 100.00% to 75.00% for DeepSeek-Coder-V2 and from 90.00% to 60.00% for Llama. These findings demonstrate that complexity-aware evaluation provides a more diagnostic assessment of LLM code-comprehension reliability than aggregate accuracy alone.
|
| 2726 |
Simultaneous Neural Optimal Transport
2609.37424
|
cs.AI
|
Milena Gazdieva, Kirill Sokolov, Jiawei Chen, Evgeny Burnaev, Alexander Korotin |
Optimal Transport (OT) provides a principled framework for learning transformations between probability distributions from unpaired samples. In many applications, however, a single transformation must map several source distributions to a common target distrib...Optimal Transport (OT) provides a principled framework for learning transformations between probability distributions from unpaired samples. In many applications, however, a single transformation must map several source distributions to a common target distribution. For example, image restoration might require handling different types of degradation without knowing the degradation of each input at inference time. Simple approaches of pooling the source distributions only encourage alignment with the target at the aggregate level and may leave individual sources misaligned. In our paper, we consider the simultaneous OT problem which formalizes the task of learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We propose a neural method for solving the simultaneous OT problem by learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We derive a max-min formulation for learning this map. We illustrate its application to image restoration, where a single model handles multiple degradation types using a common collection of clean target images.
|
| 2727 |
Learning to Retrieve Missing Evidence for Long-Term Memory QA
2609.37443
|
cs.AI
|
Yi-Xuan Deng, Yi Zhang, Wei Liu, Chao Xue, Shuojin Yang |
Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can revea...Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a lightweight planner through reinforcement learning, rewarding queries that recover previously missing evidence. MERA achieves strong answer accuracy across Qwen3-30B and GPT-4o-mini backbones. With Qwen3-30B for evidence processing and answer generation, the trained 0.6B planner achieves 77.40% accuracy on LoCoMo and 71.29% on LongMemEval-S, exceeding a 30B planner without retrieval-grounded training by 4.10% and 3.96%, respectively. On LoCoMo, later retrieval rounds increase cumulative evidence recall from 55.5% to 80.5%.
|
| 2728 |
Engineering Efficient Self-Play Chess: Search, Replay, and Throughput Under Limited Compute
2609.37447
|
cs.AI
|
Bertil Braun |
How strong can an AlphaZero-style chess system become under limited training compute when its entire learning loop is engineered for efficiency? We train from random initialization through searched self-play on a single eight-GPU node for 2.5 days. The resulti...How strong can an AlphaZero-style chess system become under limited training compute when its entire learning loop is engineered for efficiency? We train from random initialization through searched self-play on a single eight-GPU node for 2.5 days. The resulting 6.32-million-parameter model reaches 3,251 benchmark Elo [3,206, 3,297] at 100,000 searches per move (estimated at under five seconds of thinking time) against a fixed-node Stockfish 13 ladder. The run ingests 3.25 million completed games, involves an estimated 100 billion search simulations, and makes 836.6 million training presentations. We investigate search allocation, replay and restart-state selection, policy representation, progressive model sizing, quantized inference, and throughput engineering. Alongside the retained design, we document plausible alternatives that failed to improve the complete learning loop or did not justify their cost. The reported strength is a result of the integrated system, not an isolated Elo gain attributable to any single choice.
|
| 2729 |
Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers
2609.37468
|
cs.AI
|
Beining Xu, Peichun Hua, Yunming Xiao |
Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors that exploit this feedback loo...Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors that exploit this feedback loop and repurpose weak backdoor purification to conceal their presence. An attacker supplies a compromised retriever checkpoint while leaving the search agent and deployment corpus unchanged. Without corpus write access, the attacker can still suppress useful evidence, persistently retrieve a selected existing document, or steer the agent toward prolonged search, inflating retrieval, context, and latency cost. To conceal these behaviors from detection, we propose leveraging a controlled inject-and-remove cycle: deliberately inject a weaker backdoor and then unlearn it. This process weakens detector-visible signatures and fools the backdoor detectors with an illusion of purification while preserving the malicious retrieval behavior. These findings expose a systematic vulnerability in RAG systems in which a weak defense becomes an attacker's concealment tool for a backdoored retriever, even when the underlying corpus remains trustworthy.
|
| 2730 |
Regime Boundary Alignment for Evidence-Gated Question Answering
2609.37491
|
cs.AI
|
Zeyan Li, Qirong Guo, SIyuan Qiu, Hu Xu, Chun Li |
Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence is missing. We trace this behavior to the training signal: answer-focused fine-tuning assigns no target to unsu...Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence is missing. We trace this behavior to the training signal: answer-focused fine-tuning assigns no target to unsupported contexts, so it cannot distinguish a reader that abstains from one that guesses, and unsupported answering stays near 100% even as supported accuracy improves. We introduce Regime Boundary Alignment (RBA), which trains a single reader on matched variants of the same question and gold answer. The reader is trained to produce the gold answer when the context supports it, including when conflicting evidence is also present, and to abstain when the correct support is removed; inference is ordinary decoding, with no verifier, threshold, or regime label. On three multi-hop QA datasets across three seeds, RBA reduces the unsupported-answer rate by more than sixty percentage points relative to conflict-focused training while matching its supported accuracy. On a held-out TriviaQA retrieval-miss slice, the same reader reduces unsupported answering from 100% to below 1% while also improving supported accuracy. These results indicate that evidence-gated answering must be learned on both sides of the support boundary.
|
| 2731 |
Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks
2609.37493
|
cs.AI
|
Dongyub Jude Lee, Jungseob Lee, Chanjun Park, Hyeonseok Moon, Heuiseok Lim |
Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error rate among served answers. We introduce PriceCheck, which builds a compact family of decision rules from label...Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error rate among served answers. We introduce PriceCheck, which builds a compact family of decision rules from label-free checks such as re-solving a problem. Each check has a price: its agreement rates on correct and incorrect answers and its cost per run. Prices fitted on a small, class-enriched labelled set compose into predictions of a schedule's coverage and cost, guiding which checks to run and when to stop. A calibration test then selects a schedule at a stated selective-risk target. In mathematics, the selected schedules serve 76.1% of answers on average and keep held-out selective risk below 1.5% on all 15 splits. Under the shared testing protocol, PriceCheck serves more answers at that target than reward models, a prompted judge, the generator's confidence and a trained correctness classifier. At matched coverage, it keeps the fewest wrong answers among these scorers. Across 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage. These results show that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers. Code is available at https://github.com/js-lee-AI/PriceCheck.
|
| 2732 |
Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance
2609.37501
|
cs.AI
|
Dipankar Sarkar |
We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action...We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Task-level should-escalate labels make the act-versus-defer decision a measurable training signal. We demonstrate the harness at smoke scale. An offline reference run (n=12) lifts escalation recall from 0 to 0.67 and reduces unsafe-action rate from 0.33 to 0.08 when governance is enabled. Two single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and evaluation split) expose substantial variation: nominally identical RL-base configurations yield task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.5. An answer-quality adapter changes recall from 1.0 to 0.5 in Run A, but from 0.5 to 1.0 in Run B. An escalation-aware variant produces no measurable change in Run B. These small pilots do not establish reliable adapter effects or production readiness. Their contribution is diagnostic: configuration variance can overwhelm apparent tuning effects on bounded-autonomy metrics, motivating larger evaluation sets and repeated runs.
|
| 2733 |
Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning
2609.37519
|
cs.AI
|
Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang, Tolga Izdas |
Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the ro...Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into parametric Signal Temporal Logic (STL) specifications and uses the resulting formal representation for robot learning. A vision-language model extracts an embodiment-independent semantic event trace and constructs a bank of symbolic temporal specifications. The model determines the task structure, while numerical predicate thresholds and temporal bounds are grounded from successful robot trajectories. For policy learning, we separate short- and long-timescale temporal information: short-horizon specifications provide dense rewards through rolling-window quantitative robustness, while a causal monitor over a retained long-horizon specification provides one-time progress rewards for valid temporal prefixes. The same representation supports cross-embodiment transfer from human or animal videos to robot control. Across four manipulation tasks, Video2STL achieves $85.8\%$ average success-once and $67.0\%$ success-at-end, compared with $81.5\%/59.5\%$ for native dense PPO and $65.0\%/42.3\%$ for Text2Reward; in quadruped locomotion, Qwen-3.8 and GPT-5.6-based Video2STL policies achieve $100\%$ success across velocities from $0.3$ to $2.1\,\mathrm{m/s}$ while remaining competitive in high-speed energy efficiency. Project webpage: \href{https://video2stl.github.io/}{video2stl}.
|
| 2734 |
DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification
2609.37532
|
cs.AI
|
Rongjian Chen, Minxian Xu, Zhengxin Fang, Kejiang Ye, Chengzhong Xu |
Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Unifor...Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable tokens. We present DScale, preserving drafter architecture, weights, and full draft length. A separate 112K-parameter predictor requires neither confidence calibration nor hardware speed-curve preparation. Path-aware tiles reduce padding. Dynamic verify-length (DVL) allocation packs scored prefixes into half the native verification capacity. Fixed-address workspaces propagate changing boundaries through verification and acceptance while reusing captured graphs. On A100-40GB with tensor parallelism 1, Qwen3-8B and Qwen3-4B cover four datasets and concurrency 8-32, reusing each target's frozen predictor. Geometric-mean throughput gains across these configurations are respectively 43.9% and 48.8% over DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino, with lower request latency. Cumulative ablations show that adding the three mechanisms successively increases geometric-mean throughput, while budget adjustment improves accepted-token retention. GPU profiling shows that complete decode-step time on GSM8K decreases by 30.8-52.5% relative to DFlash
|
| 2735 |
Risk-Aware Semantic Grounding for Trustworthy LLM-Based Robot Planning
2609.37554
|
cs.AI
|
{\L}ukasz Sobczak, Nur Kele\c{s}o\u{g}lu, S{\l}awomir Piotr Nowak |
Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Awa...Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Aware Semantic Grounding framework for trustworthy LLM-based robot planning. Unlike existing LLM-based planners that primarily optimize plan generation, we formulate semantic grounding reliability as a multi-dimensional risk estimation problem. The proposed architecture explicitly models grounding uncertainty through ambiguity, hallucination and semantic-conflict risks before planning occurs, enabling the system to decide whether to execute the instruction, request clarification, or reject it. To evaluate the approach, we introduce TRUST-NAV, a benchmark containing both standard navigation tasks and risk-inducing instruction scenarios. Experimental results show that while conventional LLM planners achieve strong performance on valid navigation tasks, the proposed framework substantially improves ambiguity detection and semantic conflict rejection. These findings suggest that trustworthy robot planning should be evaluated not only by task completion, but also by the ability to recognize when execution should not occur.
|
| 2736 |
Concealing LLM-Based Multi-Agent Topology via Phantom Structure Injection
2609.37567
|
cs.AI
|
Longzhu He, Zelang Wen, Xinfeng Li, Sen Su, XiaoFeng Wang |
Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs in...Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs information flow among agents and often encodes proprietary knowledge about the system architecture. However, recent work has shown that such topologies can be inferred even in black-box settings by exploiting semantic dependencies in observable reasoning traces, posing significant risks of intellectual property leakage and exposure of system vulnerabilities. To address this threat, we propose MIRAGE, a topology-concealment framework that preserves the genuine communication topology for task execution while shaping adversary-facing semantic evidence toward a carefully constructed phantom topology. Specifically, MIRAGE operates in three stages: (1) phantom topology synthesis, (2) semantic edge realization, and (3) protected MAS execution. It constructs a phantom topology structurally distinct from the genuine one, materializes phantom edges as plausible semantic dependencies, and suppresses source-specific cues that could reveal genuine edges absent from the phantom topology. Extensive experiments across three topology optimization frameworks and four benchmark datasets demonstrate that MIRAGE substantially reduces the effectiveness of topology inference attacks while largely preserving the task utility of the protected MAS.
|
| 2737 |
Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection
2609.37586
|
cs.AIcs.SD
|
Yuankun Xie, Xiaoxuan Guo, Xiaopeng Wang, Siqing Qin, Shaole Li |
Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech. Existing dataset-incremental evaluation changes both real-speech domains and deepfake mechanisms, making their...Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech. Existing dataset-incremental evaluation changes both real-speech domains and deepfake mechanisms, making their effects difficult to distinguish. We construct five task organizations over identical training, development, and evaluation pools to study these factors under a controlled sample budget. Our proposed Real-Anchored Mechanism-Incremental (RAMI) protocol reflects the practical setting in which available real speech provides a recurring mixed-domain reference while new deepfake mechanisms arrive incrementally. We further propose RF-Prompt, an asymmetric continual prompt-learning method that preserves reusable real-speech knowledge through a shared real prompt and expands mechanism-specific knowledge through inherited fake experts with orthogonal residuals. Input-adaptive soft fusion combines the accumulated experts into a fixed number of injected tokens without requiring task identity at inference. On RAMI, RF-Prompt achieves 10.110% average EER and 10.370% pooled EER, outperforming all evaluated continual-learning baselines. Across the five controlled protocols, RAMI yields the lowest common-average and pooled EER. Component ablations, limited-data experiments, and cross-backbone evaluations further validate the proposed design.
|
| 2738 |
ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling
2609.37587
|
cs.AI
|
Zijie Meng, Xiwei Dai, Yingying Zhang, Jian Wu, Xian Wu |
Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) whe...Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually updated patient memory. However, under a fixed memory budget, successive updates must integrate new information without progressively losing critical historical evidence needed to subsequent tasks. To address this challenge, we introduce Recurrent Longitudinal Memory (ReLMem), a framework that learns to maintain fixed-capacity patient memory for efficient downstream prediction with a frozen LLM. ReLMem equips this LLM with lightweight compression adapters to recurrently update the memory from its previous state and each incoming visit, without rereading earlier records. Specifically, we develop a multi-granularity optimization strategy to preserve task-relevant information throughout recurrent updates and support downstream prediction from the final memory. The intermediate supervision aligns attention outputs from compressed memory and the full history under identical queries, while prediction supervision minimizes cross-entropy with ground truth answers conditioned on the final memory. On EHR-based medication prediction, ReLMem approaches the F1 scores of full-history baseline while reducing average retained historical storage by 97.1%. Under the same memory budget, it improves macro- and micro-F1 over the strongest compressed-memory baseline by 4.66 and 4.75 percentage points, respectively. These results highlight the value of learning recurrent patient memory for efficient longitudinal EHR modeling.
|
| 2739 |
Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation
2609.37591
|
cs.AI
|
Yang Li, Sijia Zhang, Yihan Li, Aming WU, Zihao Zhang |
Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and le...Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal does not guarantee a reliable policy update. The key challenge is thus twofold: identifying interactions that support goal-directed improvement and determining whether the resulting updates are worth retaining. We observe that each executed action induces an immediate observation transition, providing evidence of its local consequences. Based on this insight, we propose Credit-Guided Policy Improvement (CGPI), which recovers signed, reference-relative decision credit from action-induced observation transitions without external outcome feedback. With the pretrained navigation policy frozen, CGPI uses this credit to propose lightweight adaptation updates and verifies them against prior credit-supported interactions. Updates are retained only when supported and rolled back otherwise. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate the feasibility of zero-shot sim-to-real transfer.
|
| 2740 |
Independent Verification Paths Are Not Independent: A Case Study of Common-Mode Failure in a Satellite Catalogue Pipeline
2609.37603
|
cs.AI
|
Fabio Rovai |
A common safeguard for a data pipeline is redundant computation: derive each published number by two routes built on different technology and refuse to exit when they disagree. We report one such gate failing, in a cross-catalogue integrity study of two open r...A common safeguard for a data pipeline is redundant computation: derive each published number by two routes built on different technology and refuse to exit when they disagree. We report one such gate failing, in a cross-catalogue integrity study of two open registers of Earth-orbiting objects. A gate comparing a set-based Python path with SPARQL queries over the emitted RDF graph printed ALL CROSS-CHECKS AGREE on seven counts. Three were wrong, one overstated more than fourfold (932 against 220). Both paths imported the same constants, which encoded a misreading of the source's status vocabulary, so the error was common-mode and the gate could not see it. We give the mechanism, an object-level ledger reconciling every figure, and three checks that go back to the source's documentation, measured on the defective code and on its correction. We then checked that correction against each object's phase history, held in a source file the pipeline never read. The correction was also wrong: 42 of its 261 disagreements are artefacts, and none of our three checks flagged them. Finally, in a controlled replication with three pinned models and tools disabled, 72 of 75 paths generated on request as independent checks computed the defective count, 29 of 30 even when the prompt carried the source's own definitions of the codes. The evidence is one pipeline and one defect family. Within it, redundancy verified implementation, and the errors that reached publication were errors of meaning.
|
| 2741 |
AS$^2$D: Accelerating On-Demand Audio Understanding on Mobile Devices
2609.37617
|
cs.AIcs.SD
|
Yunzhe Li, Kyoungjun Park, Hongzi Zhu, Lili Qiu |
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializi...Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency is necessary for source-conditioned generation. Our key observation is that, for audio language models, the input audio and user request can provide useful speculative candidates without following the target's evolving text prefix. We propose AS$^2$D (Audio Speculative Speculative Decoding), which enables target-decoupled drafting: an audio-conditioned drafter follows its own generation history while the target independently verifies and corrects ready candidates. Without usable candidates, the target advances alone. Thus, target feedback determines which candidates are committed but no longer determines when the drafter can make progress, enabling drafting and verification to proceed concurrently while retaining target-side verification and correction. We implement AS$^2$D in MNN for Android and evaluate two target models across four phones, seven datasets, and three tasks covering 12.2 hours of audio. Across four phones, AS$^2$D improves pooled ASR throughput by 42-76% over target-only decoding, while only 5.7% of evaluation windows are slower than target-only, compared with 58.1-63.0% for speculative baselines. For ASR, AS$^2$D reaches 97.33-98.20% of a hindsight per-window oracle's pooled throughput over the evaluated drafter/budget catalog. Native on-demand execution with a 7B target achieves up to 78% higher throughput than target-only. These results show that source-conditioned audio generation can relax the conventional dependence of speculative drafting on the target's evolving output prefix, exposing substantial parallelism for efficient inference.
|
| 2742 |
Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision
2609.37624
|
cs.AI
|
Jacob Epifano |
Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and del...Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cuts the EM rate by about a third and improves answers on held-out medical questions, while deleting the same rows has little measurable effect. The advantage is larger when half the poison rows are corrected, and it holds on a second base model and a second misaligned model organism. The content of the replacement appears to matter: paraphrasing the rows while keeping their bad advice shows no clear benefit, and the correct answers distributed with the dataset appear to do about as well as our rewriter's. Realigning an already-poisoned model with further fine-tuning is known to work, but which data does the work has not been compared directly. We find that a short round of training on corrections beats the same amount of training on generic chat data, that corrections on other medical prompts do roughly as well as corrections of the poisoned prompts themselves, and that instructing the correction writer to model a careful, harm-avoiding assistant adds no measurable benefit over plain corrections. In the settings we tested, correcting harmful training data reduces EM more than deleting it.
|
| 2743 |
SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving
2609.37626
|
cs.AI
|
Chuan Liu, Shuoming Zhang, Zhicheng Li, Qianqi Sun, Ruiyuan Xu |
No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and ...No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads, decouples where a request's KV cache lives from how attention weights are sharded. This has two consequences. First, layouts differ only in who owns the weights and the cache, and most of that state already sits where the next layout needs it; SPLASH reuses it, moves the rest in the background of ongoing inference, and hands off at a batch boundary, making a switch nearly free: its median overhead is under 0.51% of the step it runs in. Second, the decoupling exposes a layout that existing engines lack: Decoupled Ownership Parallelism (DOP) shards attention weights as tensor parallelism does while keeping each request's cache on a single owner as data-parallel attention does. DOP replicates neither, offers 27-60% more KV capacity than data-parallel attention, and gives the scheduler a choice when KV memory limits admission. A transition-aware scheduler follows the best of the four layouts as load changes. On B200 GPUs serving GLM-5.3, SPLASH improves end-to-end serving throughput by 1.3-1.73x over fixed-layout deployments, and the same layout regimes appear with DeepSeek-V3.2 on H200 and GLM-5.3-Flash on DCU.
|
| 2744 |
ProCTI: Prototype-Refined Global Conditioning for Diffusion-Based Time Series Imputation
2609.37632
|
cs.AI
|
Fariza Rashid, Duc Van Le, Rahat Masood, Gustavo Batista, Aruna Seneviratne |
Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual informati...Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual information from the current or neighbouring windows. Meanwhile, global dataset-level structure often remains implicit, limiting performance when local observations are sparse, noisy, or unrepresentative. To address this issue, we propose ProCTI, a diffusion-imputation framework that augments local conditioning with retrieved global dataset-level priors through learned prototypes. A hybrid conditioning mechanism integrates this global context with local signals during reverse diffusion, enabling more accurate reconstruction under varying missingness scenarios. Experiments across multiple benchmark datasets show that ProCTI outperforms strong baselines overall under random missingness, while remaining competitive under attribute-wise missingness. Furthermore, we use a latent-regime data model to characterise the precise conditions under which prototype-derived global conditioning provably improves imputation. We support this with a general theoretical analysis of local-global conditioning.
|
| 2745 |
RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
2609.37633
|
cs.AI
|
Michael Kirchhof, Eleonora Gualdoni, Andrew Szot, Khashayar Gatmiry, Aryo Lotfi |
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult t...The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.
|
| 2746 |
Evaluating and Benchmarking the System One Model Jev
2609.37647
|
cs.AI
|
Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa |
Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor des...Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in information access pipelines, such as routing queries, checking grounding, moderating content, or rating against a rubric. We evaluate Jev (jev-1.13.0) zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reasoning, moderation, legal clause analysis and rubric scoring, with one frozen template per dataset and full evaluation splits: 346,009 requests for under USD 10. For reference, we score Qwen3.8-27B and Gemma-4-E4B on identical requests via their exact next-token probabilities over the options. Jev reaches 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC and 86.7% on Belebele across 122 languages. It beats Qwen on 27 of 37 datasets, with none of Qwen's nine leads outside the bootstrap intervals, and Gemma on all 37. All three models degrade on low-resource languages, fine-grained or noisy labels, and rubric-based quality judgments. Jev's choice probabilities are well calibrated and support selective prediction. Binary probabilities rank well but are poorly placed relative to a fixed 0.5 threshold; thresholds tuned on training data raise micro-F1 on UNFAIR-ToS from 0.50 to 0.75. Jev answers MMLU's calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder. Rotating the options leaves Jev's accuracy unchanged and withholding the question drops it to near chance, ruling out shallow memorization but not memorized question-answer pairs. We release the code, harness and all raw responses.
|
| 2747 |
Semantic Map Sharing and Capability-Aware Coverage Planning for AI-Native 6G Robotic Coordination
2609.37666
|
cs.AI
|
Abdulqader Dhafer, Qi Wang, Zhou Daniel Hao |
Search and Rescue (SAR) operations increasingly deploy heterogeneous teams of aerial and ground robots. However, conventional coverage methods typically do not translate perceived terrain into platform-specific reachability, while continuous image exchange imp...Search and Rescue (SAR) operations increasingly deploy heterogeneous teams of aerial and ground robots. However, conventional coverage methods typically do not translate perceived terrain into platform-specific reachability, while continuous image exchange imposes a high communication cost. We propose an edge-centric, semantic-aware coverage planning framework that integrates aerial terrain perception, robot-specific traversability reasoning, and payload-efficient semantic state sharing. Aerial observations are converted into compact semantic grid maps, enabling reachability-constrained area decomposition and capability-aware coverage paths that assign only regions admitted by each robot's capability profile. The resulting perception-sharing-planning loop feeds semantic corrections into traversability reasoning and replanning, forming an application-level mechanism motivated by AI-enabled goal-oriented communication envisioned for AI-native 6G networks. For the high-update case, transmitting semantic corrections reduces the application payload by a factor of approximately $82$ relative to periodic full-map sharing. Across matched benchmark scenarios, the proposed method achieved $91.5\%$ coverage with no capability-infeasible allocations, compared with $78.8\%$ coverage and a $21.5\%$ capability-infeasible allocation rate for LS-MCPP. Semantic corrections update the shared planning state without requiring repeated transmission of the complete map.
|
| 2748 |
Retrieve, Reproduce, Reveal: Dissecting Retrieval-Augmented Software Vulnerability Detection
2609.37669
|
cs.AI
|
Sabrina Kaniewski, Tim Kr\"amer, Julius B\"achle, Markus Enzweiler, Michael Menth |
Retrieval-Augmented Generation (RAG) is increasingly used to enhance Large Language Model (LLM)-based software vulnerability detection by grounding predictions in retrieved vulnerability knowledge, such as vulnerability reports. However, existing RAG-based sof...Retrieval-Augmented Generation (RAG) is increasingly used to enhance Large Language Model (LLM)-based software vulnerability detection by grounding predictions in retrieved vulnerability knowledge, such as vulnerability reports. However, existing RAG-based software vulnerability detection (RAG4SVD) systems are often evaluated using proprietary models, which challenges open science and reproducibility. Further, studies use different datasets, custom knowledge bases, different backbone models, and diverse metrics, which hinders meaningful cross-system comparison. In this work, we study six open-source RAG4SVD systems and address these reproducibility and comparability challenges through (i) reproduction of their experimental settings under an open-weight setting, and (ii) a unified benchmark using a common dataset, metric suite, and pool of open-weight models. Further, RAG4SVD systems typically consist of multiple components, yet are often evaluated only as a whole system, i.e., end-to-end. Therefore, we perform (iii) a component-level analysis that decomposes representative RAG4SVD pipelines into input abstraction, knowledge retrieval, and detection. Our results demonstrate that reproducibility varies substantially across systems. Under the presented unified benchmark, published RAG4SVD performance does not transfer under a controlled open-weight evaluation and depends strongly on the used model. The component analysis shows that effective RAG4SVD depends on the alignment between pipeline stages. For example, oracle knowledge raises retrieval to near-optimal, yet performance remains low (0.51 pairwise accuracy), demonstrating that retrieval effectiveness alone is insufficient for reliable detection. These findings motivate evaluating RAG4SVD not only end-to-end, but at the level of pipeline components, and provide a basis for more standardized, RAG-aware evaluation practices.
|
| 2749 |
When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task
2609.37680
|
cs.AI
|
Sai Sumedh R. Hindupur, Hadas Orgad, Thomas Fel, Demba Ba |
One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifo...One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data and tasks, the extent to which models rely on them for computation, and how they manipulate them, remains unclear. We characterize precisely the geometry of computation in a number-comparison task, as an abstraction of comparison for decision making, and how models utilize geometry in an elegant fashion to implement it. Specifically, we study the causal geometry of number comparison in Qwen2.5-7B-Instruct, a capable and widely studied open-weight model, and find Qwen largely uses linear representations of numbers despite the presence of curved geometry. To compare two numbers, the model first encodes each number along a vector and adds the two representations using attention and the residual connection, bringing them into a shared space in the residual stream. Then, the model uses MLP neurons to compare the pair of numbers on local regions in this shared space, which correspond to smaller intervals of input numbers, and combines these to obtain the position of the maximum. In fact, this reliance on linear representations for comparison also persists when the model compares three numbers. Our findings demonstrate that the manifold hypothesis can co-exist with linear representations: while concepts that are ordered may have manifold structure in representations, the model may use an underlying linear structure of the concept in certain computations.
|
| 2750 |
GARDiff: Graph-Aligned Residual Diffusion for Probabilistic Multivariate Time-Series Forecasting
2609.37694
|
cs.AI
|
Rui Han, Min Yang, Xu Zhang, Xinghao Yang, Wei Liu |
Diffusion models have recently shown strong potential for probabilistic multivariate time-series forecasting by modeling complex conditional distributions. Recent decoupled diffusion frameworks further separate forecasting into deterministic prediction and sto...Diffusion models have recently shown strong potential for probabilistic multivariate time-series forecasting by modeling complex conditional distributions. Recent decoupled diffusion frameworks further separate forecasting into deterministic prediction and stochastic residual generation, making it natural to derive dependency graphs from deterministic representations and use them to guide residual diffusion. However, we show that this direct structural transfer is unreliable. Although deterministic-derived graphs encode useful global dependency priors, they exhibit substantial edge-level misalignment with residual dependency structures, introducing inaccurate or redundant conditions during residual generation. This reveals a previously overlooked deterministic-to-residual structural alignment problem in decoupled diffusion forecasting. To address this problem, we propose GARDiff, a Graph-Aligned Residual Diffusion framework for probabilistic multivariate time-series forecasting. Instead of treating deterministic-derived graphs as fixed diffusion conditions, GARDiff progressively adapts them to residual generation. Specifically, GARDiff estimates residual uncertainty to distinguish high- and low-uncertainty regions, enabling uncertainty-aware structural refinement, and further performs timestep-aware edge sparsification during reverse diffusion to evolve graph conditions from broad dependency aggregation to localized residual refinement. Extensive experiments on six real-world benchmarks demonstrate that GARDiff consistently improves probabilistic forecasting performance and uncertainty calibration over strong baselines.
|
| 2751 |
Width Expansion as a Method for Class Incremental Learning
2609.37702
|
cs.AI
|
A. L. S. Conde, Y. Elkhatib, C. M. Ranieri |
Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic f...Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic forgetting a central challenge. Existing approaches include regularization, knowledge distillation, replay, and architectural expansion. However, many expansion methods rely on explicit task identifiers or predefined growth strategies, limiting their applicability when task boundaries are unavailable at inference time. This work proposes a dynamic width expansion method that increases the number of neurons within existing layers according to a normalized loss criterion, without requiring task-specific information. An attention mechanism with persistent key-value memory is also incorporated to stabilize feature representations and reduce interference between previously learned and newly introduced classes. The approach is evaluated on Split MNIST and Split CIFAR-100 under the standard Class-IL protocol. Experiments compare fixed-capacity and dynamically expanding architectures, both with and without attention, combined with established continual learning methods including EWC, LwF, and A-GEM. Results show that progressive width expansion consistently improves performance over fixed architectures, particularly when combined with functional methods and A-GEM. The combination of width expansion and attention provides the most consistent gains. Overall, dynamic width expansion based on representational demand provides an effective and flexible strategy for Class-IL, although uncontrolled growth may increase overfitting and computational cost.
|
| 2752 |
A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
2609.37788
|
cs.AI
|
Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo |
Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE);...Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR.BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.
|
| 2753 |
Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance
2609.37789
|
cs.AI
|
Fabian A. Mikulasch, Friedemann Zenke |
Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelev...Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could they be distinguished? Surprisingly, we prove that common SSL methods can achieve exactly this, by implicitly instantiating a latent-variable model with stochastic dynamics and observation-private nuisance. We trace their ability to recover the stochastic signal to two complementary principles: Predictive mutual information maximization ensures that representations retain the information needed for prediction, while latent distribution matching constrains how this information is encoded, thereby making the retained signal identifiable. We confirm this identifiability result in simulations for Gaussian predictors, which recover the true signal up to an affine transformation even in dynamic, nuisance-laden environments.
|
| 2754 |
GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets
2609.37798
|
cs.AIcs.SDeess.AS
|
Gaspard Bott\'e, S\'everin Baroudi, Samir Sadok, Francesco Paissan, Thomas Hueber |
Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly...Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-generation mechanisms. Pretrained on 960 hours of LibriSpeech, our 57M-parameter model achieves a 6.89% WER on frozen-encoder SUPERB ASR and a 25.87% CER on slot filling, outperforming the best non-distilled sub-90M baselines by 43.1% and 22.0%, respectively. These results demonstrate that highly competitive speech representations can emerge from a radically simplified training recipe.
|
| 2755 |
Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation
2609.37800
|
cs.AI
|
Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ek\c{s}i, Rahul Sharma |
Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when eac...Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when each user has a finite catalog that can become repetitive or depleted over time. We propose CohortMix-TS, a warm-started mixture bandit that learns latent user groups from earlier cohorts and uses available metadata to construct group-informed priors for new users. Starting from these fixed priors, the model personalizes independently as feedback from each user becomes available. Session slates combine Thompson sampling with diversity and inventory-depletion controls. We evaluate CohortMix-TS through simulation, semi-synthetic experiments, and a 25-day randomized in-the-wild deployment with 713 registered participants in a Campus Games quiz application. Our evaluations show that cross-cohort transfer improves early recommendation quality and user-level regret, while inventory-aware slate construction helps prevent premature exhaustion of preferred items. In the field deployment, treatment users also showed a larger early-to-late change in correctness than users receiving random recommendations. Together, these results show how warm-start transfer and inventory-aware recommendations can support personalization for short-lived, repeatedly cold-starting cohorts.
|
| 2756 |
Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents
2609.37810
|
cs.AI
|
Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu, Zuxuan Wu |
Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task...Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.
|
| 2757 |
Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue
2609.37818
|
cs.AI
|
Shengbo Cai, Yuxiang Wang, Jingran Xie, Zhisheng Zhang, Shun Lei |
Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in resp...Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acoustic grounding at every pass. Its two-stage training further narrows the perception-reasoning gap by separating learning to reason from learning to respond, enabling direct inference without CoT. On EchoMind, LoopSLM improves paralinguistic understanding, reasoning, and reply quality over Qwen2.5-Omni-7B. Against the CoT-SFT baseline, LoopSLM gains over 20 points in reasoning accuracy while generating 64.5% fewer tokens at half the latency. It also outperforms Qwen3-Omni-Thinking on most empathetic reply metrics with 34x lower latency. Despite training only on dialogue data, LoopSLM improves accuracy on general audio benchmarks.
|
| 2758 |
Making Duplicate Reimbursement Unrepresentable: A Verified Ethereum E-Invoice System for Humans and AI Agents
2609.37819
|
cs.AI
|
Jia Cai |
Electronic invoices are replacing paper invoices worldwide, but today's centralized architectures leave three problems unsolved on the consumption side: an invoice can be submitted for reimbursement repeatedly, authenticity is difficult for recipients to verif...Electronic invoices are replacing paper invoices worldwide, but today's centralized architectures leave three problems unsolved on the consumption side: an invoice can be submitted for reimbursement repeatedly, authenticity is difficult for recipients to verify, and data is siloed at a central authority that forms both a performance bottleneck and a single point of failure. This paper presents the design, formal analysis, and implementation of a complete blockchain-based electronic invoice system on Ethereum. We formalize the invoice lifecycle as a guarded labeled transition system and prove, under standard cryptographic and consensus assumptions, that the system guarantees: (i) reimbursement uniqueness--an invoice is reimbursed at most once, even across mutually distrusting organizations; (ii) face integrity--any verified invoice matches the recorded one unless keccak256 second-preimage resistance is broken; and (iii) authorization soundness for every lifecycle operation. The core invariants are machine-checked using Solidity SMTChecker, proving inductive validity across all reachable transaction sequences. The architecture models each invoice as a non-fungible, non-tradable token whose state transitions through five guarded subsystems, employing a lock-based protocol that makes duplicate reimbursement unrepresentable rather than merely detectable. We implement the design as a Solidity 0.8 contract with a four-role web application and evaluate it on a private Ethereum network: issuing costs 646,773 gas, full reimbursement costs under 135,000 gas, all operations run in O(1) time, and a single node sustains 137 issuances/s. Finally, the verified contract serves as a safety envelope for LLM-based reimbursement agents, provably rejecting unsafe actions (duplicate, over-limit, or forged-receipt claims) even when the agent's internal policy fails. All code and benchmarks are open-source.
|
| 2759 |
Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR
2609.37825
|
cs.AI
|
Kun Liang, Chenming Tang, Clive Bai, Weijie Liu, Zeyuan Liu |
Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct t...Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy's future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit the standard state-only formulation of value estimation and propose $\pi$PPO, a self-privileged actor-critic framework. By reusing verified same-prompt rollouts as contrastive evidence, $\pi$PPO helps the critic assess intermediate reasoning against successful and failed attempts, while preserving standard policy optimization and the deployment interface. Experiments show that $\pi$PPO consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reasoning benchmarks, while remaining effective even when paired with substantially smaller asymmetric critics.
|
| 2760 |
Scaling Influence Functions in LLMs through Eigenbasis-Corrected One-Bit Gradient Projection
2609.37842
|
cs.AI
|
Jaeseung Heo, J Rosser, Dongwoo Kim |
Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients re...Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients reduces the computational cost, but storing full gradients is prohibitively expensive at LLM scale. We study how to compress these gradients while preserving influence estimates for future queries that are unknown at storage time. Through a worst-case analysis, we characterize the optimal fixed-dimensional linear representation and propose eigenbasis-corrected one-bit gradient projection (EOGP) to approximate it at scale. Specifically, EOGP uses EK-FAC to reduce gradient dimensionality, then applies PCA within the retained subspace to learn compression directions from the training gradients. We then apply one-bit quantization to the resulting coordinates, allowing more coordinates to be retained within a fixed storage budget. On GPT-2, EOGP predicts retraining outcomes more accurately than the evaluated compression baselines while using one-sixteenth of their per-example storage. On OLMo 2 SFT models from 1B to 32B parameters, EOGP remains competitive with the baselines allocated over 100 times as much storage per example.
|
| 2761 |
Is manual software optimization a thing of the past?
2609.37849
|
cs.AI
|
Pavlin G. Poli\v{c}ar, Martin \v{S}pendl, Toma\v{z} Ho\v{c}evar |
Scientific software is increasingly required to process larger datasets while maintaining acceptable execution times. Software optimization traditionally requires substantial expertise in programming, algorithms, and numerical methods. Recent advances in large...Scientific software is increasingly required to process larger datasets while maintaining acceptable execution times. Software optimization traditionally requires substantial expertise in programming, algorithms, and numerical methods. Recent advances in large language models (LLMs) offer the possibility of automating much of this process. We investigate whether LLM-based agents can autonomously achieve substantial performance improvements in scientific software, including mature implementations that have already been extensively optimized by human developers. We tasked an LLM-based agent with optimizing software for three computational problems: t-SNE, single-sample gene set enrichment analysis (ssGSEA), and graphlet counting. Humans defined the scope, correctness criteria, and a verification mechanism, after which the agent worked autonomously, in some cases for several hours. Code maintainers reviewed each resulting implementation and verified its correctness. The optimized implementations were faster in all tested configurations, by up to two orders of magnitude over the fastest existing tools. The improvements included low-level code optimizations, mathematical reformulations, and an entirely new algorithm for graphlet counting. Software optimization can increasingly be delegated to autonomous agents, with the human role shifting from implementing optimizations to deciding which software to optimize, defining objectives, providing verification mechanisms, and ensuring the correctness of the final software. For well-scoped, verifiable problems, we argue that manual software optimization may be a thing of the past.
|
| 2762 |
AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems
2609.37864
|
cs.AI
|
Yiming Cheng (The University of Chicago), Alfin Wijaya Rahardja (Fudan University), Mengshi Zhang (TensorBlock, Inc), Zihao Chen (TensorBlock |
Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs wh...Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AgentBug-Smith, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AgentBug-Smith consistently outperforms existing bug reproduction techniques designed for general software, achieving 10.67% - 27.56% higher success rates of reproducing harness bugs. By applying AgentBug-Smith to open-source agentic systems in the wild, we construct Live-Harness-Bench, a live and extensible benchmark that currently contains 200 reproducible harness bugs. We further demonstrate the utility of Live-Harness-Bench through two downstream applications. First, we use Live-Harness-Bench as the evaluation benchmark to systematically evaluate state-of-the-art software agents, revealing their limited capabilities in repairing real-world harness bugs. Second, we use Live-Harness-Bench as a knowledge base of real-world harness bug fixes, from which reusable repair skills can be distilled to improve existing software agents, increasing their harness-bug repair rates by 6.32%. Together, AgentBug-Smith and Live-Harness-Bench establish a scalable foundation for continuously evaluating and improving software agents on harness bug repair, turning real-world agent failures into executable evaluation instances and reusable knowledge for harness improvement, thus contributing to the ultimate goal of recursively self-improving agents.
|
| 2763 |
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
2609.37868
|
cs.AI
|
Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park, Eunho Yang |
Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the...Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.
|
| 2764 |
ExceptionDrive: A Planning-Oriented Counterfactual Corner-Case Benchmark for Autonomous Driving
2609.37871
|
cs.AI
|
Ziyi Luo, Zhe Sun, Yehao Lu, Lei Zhou, Lisheng Wu |
Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDrive, a counterfactual planning benchmark that uses VLM-assisted screening, localized multi-view editing, and qu...Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDrive, a counterfactual planning benchmark that uses VLM-assisted screening, localized multi-view editing, and quality auditing to insert hazards into real nuScenes scenes while preserving their context. Its 21 tasks span six safety families and define hazard or conflict regions, local safety constraints, and acceptable responses. Because hazard insertion can invalidate the recorded human trajectory, our reference-free protocol evaluates edited predictions using Unsafe Rate (UR), Hazard Clearance Compliance (HCC), Hazard Proximity Response (HPR), and Counterfactual Trajectory Shift (CTS), which measure core-region intrusion, clearance compliance, clearance relative to a prescribed margin, and counterfactual trajectory change. Seven representative planners frequently intrude into hazard regions or provide insufficient clearance. We also develop a Reminder Agent that, without sample-specific task labels, converts visual evidence and the shared taxonomy into structured records of hazard presence, type, and a recommended high-level strategy. The agent neither predicts trajectories nor controls the vehicle; its records guide a VLM-based decision agent. In zero-shot experiments, the reminders improve strategy accuracy and reduce under-warning.
|
| 2765 |
Boids of a Feather Flock Together - Evolving Prey Behaviours Under Different Predator Attack Strategies
2609.37885
|
cs.AI
|
Augusta van Haren, Hanna Hoogen, Luca Pattavina |
Flocking and schooling are thought to have evolved partly as defences against predation, but how prey should balance social and escape tendencies may depend on the predator's hunting strategy. We extend the predator-prey boids model of Ojo et al. (2023), itsel...Flocking and schooling are thought to have evolved partly as defences against predation, but how prey should balance social and escape tendencies may depend on the predator's hunting strategy. We extend the predator-prey boids model of Ojo et al. (2023), itself based on Reynolds' boids, by combining six prey movement tendencies (alignment, cohesion, separation, dodge, repel and wiggle) into a single weighted acceleration update, and by reformulating wiggle as a sinusoidal manoeuvre. We then use an evolutionary strategy to optimise the six behaviour coefficients for collective prey survival against four predator hunting strategies: attack-centroid, attack-nearest, attack-random and attack-peripheral. Across five independent trials per strategy, coefficients converged within trials and mean fitness remained stable or increased, although trials often settled in different local optima. Prey survival was highest under attack-centroid and lowest under attack-nearest, in line with our hypotheses. Against attack-centroid, prey evolved individualistic predator avoidance with high escape coefficients, whereas against the other three strategies they largely kept their flock formation. Across all strategies, evolution favoured a low repel coefficient and relatively high dodge and wiggle coefficients. Our results suggest that optimal anti-predator behaviour depends on the interplay between escape tendencies and the predator's hunting strategy.
|
| 2766 |
Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction
2609.37905
|
cs.AI
|
Shivang Chopra, Fotis Iliopoulos, Zsolt Kira, Gaurav Menghani |
Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come fro...Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity continues to grow. This motivates a complementary scaling direction that we call estimator scaling, where additional resources are used to incorporate multiple related estimators rather than only enlarging a single predictor. Through theoretical analysis, we show that the gains from estimator scaling are governed by the amount of non-shared predictive variation available across estimators. However, exploiting this variation naively can be expensive: independently trained models provide substantial estimator diversity but require deployment cost to grow with ensemble size. This motivates a parameter-efficient realization of estimator scaling that can incorporate diversity from multiple estimator sources without maintaining multiple full models. Building on this view, we introduce RECursive Averaged Predictor (RECAP), a parameter-efficient recursive CTR model that operationalizes estimator scaling at three levels: distillation across independently trained models, exponential moving averaging over training trajectories, and aggregation over inference-time routes within a weight-shared recursive backbone. Experiments across multiple benchmarks establish new state-of-the-art predictive performance on standard benchmarks, while placing the RECAP on a favorable performance-parameter Pareto frontier.
|
| 2767 |
Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3
2609.37911
|
cs.AI
|
Ryan C. Barron, Cade W. Trotter, Maksim E. Eren, Kim {\O}. Rasmussen, Liz D. Miller |
Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated ...Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term lists, a pseudo-document, multiple pseudo-references, and corpus-steered text all together with SPLADE-v3 on NFCorpus, TREC-COVID, and SciDocs. Every condition searches the same frozen document index and follows the same query-side integration rule and 256-dimension budget, isolating the effect of the added content. All twelve method-collection comparisons improve aggregate nDCG@10, with best relative gains of 4.81%, 8.92%, and 9.47%. Eleven remain significant after Holm correction. The gain persists in 103 of 114 interpolation settings, including every setting that assigns at least 30% of the mixture weight to the original query. Shuffled-text and non-contextual lexical-bag controls also remain above baseline in all 24 aggregate comparisons, showing that the added vocabulary carries most of the benefit. A corpus-induced typed concept graph, by contrast, produces no consistent gain, and its relation, depth, validation, random, and gating controls do not rescue it. Generated vocabulary can therefore complement a strong learned sparse retriever, provided that the original query remains strongly represented.
|
| 2768 |
The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
2609.37914
|
cs.AI
|
Gon\c{c}alo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke, Stella Biderman |
Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fin...Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining -- a sound attribution score should enable us to enhance or attenuate EM by filtering data on that score. Score-based filtering can substantially enhance or attenuate EM; we find that both data-attribution scores and a black-box harmfulness score can identify consequential examples. All models we test become misaligned when trained on the same dataset, and influence scores perform best when filtering data from the same model that computed them. We find cross-model generalization of influence scores from scores derived from the three model families we tested, but this generalization does not recover same model filtering performance.
|
| 2769 |
RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust
2609.37916
|
cs.AI
|
Eugene Hauptmann, Nataliya Kosmyna |
Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap ...Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runtime roles around one primitive-level, three-level intermediate representation (IR), plus a transparent dispatch contract that resolves each operator to native, common-IR, or rewritten lowering and fails compilation when legalization is not possible. The same IR targets fourteen runtime devices (cpu, metal, mlx, ane, cuda, rocm, oneapi, tpu, hexagon, gpu, vulkan, opengl, directx, webgpu) and two specialty codegen paths (Cortex-M INT8 and FPGA), ingests safetensors, GGUF, ONNX, and rten formats, supports F16/BF16/F64/C64 and quantized INT4/INT8 flows with AMP/PTQ/QAT, and scales via tensor-/pipeline-parallel collectives over TCP and RDMA transports. Beyond neural workloads, RLX also extends to scientific/physics-style domains through sparse and dense linear algebra extensions (e.g., CSR LU/CG/matvec and LAPACK- backed factorizations) and 3D Gaussian splatting operators. We evaluate RLX against PyTorch, TensorFlow, JAX, candle, burn, tch, rten, MLX, CoreML, IREE, Glow, TensorRT, and tinygrad under identical input generation and p50 measurement methodology on one host. On all-MiniLM-L6-v2, RLX-Metal is fastest at every batch (e.g., 16.6 ms at batch 32 vs. PyTorch-MPS 26.7 ms). In the MNIST training table, RLX also has the top-throughput entry (graph-fused MLP: 946,487 img/s), above NumPy+BLAS (787,349 img/s), while retaining 100% top-1 parity on reference checks (e.g., Qwen3).
|
| 2770 |
Dagger: Decoupling-based Model Stealing Attack against Graph Neural Networks
2609.37972
|
cs.AI
|
Ying Song, Xiaowei Jia, Balaji Palanisamy |
As Graph Neural Networks (GNNs) are widely deployed as Machine Learning-as-a-Service (MLaaS) APIs, model stealing attacks have emerged as a critical security threat. By querying a victim model's black-box API, an adversary can construct a functionally equivale...As Graph Neural Networks (GNNs) are widely deployed as Machine Learning-as-a-Service (MLaaS) APIs, model stealing attacks have emerged as a critical security threat. By querying a victim model's black-box API, an adversary can construct a functionally equivalent surrogate model, compromising proprietary intellectual property and downstream security. Existing GNN stealing attacks, however, rely on overly permissive assumptions, such as soft-label outputs, large query budgets, full-graph query access, and prior knowledge of victim backbones that rarely hold in real-world deployments. In this work, we formalize a strictly constrained black-box, hard-label and backbone-agnostic threat model for GNN stealing attacks under a tight query budget. Given these realistic restrictions, we identify four fundamental challenges: sparse local structures and isolated nodes that degrade victim label quality, insufficient supervision signals, systematic imbalance with incomplete class coverage, and backbone mismatch. To address these interlocking barriers, we propose Dagger, a novel two-phase decoupling-based attack framework. Specifically, in Phase 1, Dagger pre-trains a surrogate using decoupled information propagation to preserve structural context over sparse local subgraphs while handling isolated nodes, combined with manifold-level node mixup to synthesize continuous supervision signals and smooth decision boundaries. In Phase 2, Dagger freezes the encoder and fine-tunes the classifier head via class-balanced sampling paired with logit adjustment to rectify severe query imbalance without requiring extra victim queries. Extensive experiments across four benchmark graphs and four GNN backbones demonstrate that Dagger consistently outperforms state-of-the-art GNN stealing attacks, achieving up to 18.16\% higher fidelity while only utilizing 12.23$\times$ fewer queries than the strongest baseline.
|
| 2771 |
On Trajectory-Aware Training for Masked Diffusion Language Models
2609.37974
|
cs.AI
|
Manuel Madeira, Amitis Shidani, Alice Bizeul, Victor Turrisi, Louis B\'ethune |
Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own ...Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices interact. We introduce PUMBA, a unified framework for trajectory-aware training that trains the denoiser on consecutive steps of policy-induced trajectories, passes information between steps, and optimizes them jointly by backpropagation through time. A controlled study of this design space shows that i) exact train--inference alignment fails due to local overfitting, whereas a looser alignment still brings training masks closer to those seen at inference; ii) passing continuous information outperforms discrete gradient estimators through the commitment at each step; and iii) performance improves as backpropagation through time spans more steps, which we support theoretically. Combined, these components match the best checkpoint of a same-size autoregressive model. Building on these findings, we scale PUMBA to supervised fine-tuning of LLaDA-8B, where it improves the trade-off between performance and number of function evaluations (NFEs) in both full-canvas and block diffusion generation. At matched performance, it needs up to 22% fewer NFEs than standard fine-tuning with twice the budget in full-canvas generation, and up to 26% fewer than standard fine-tuning for the same number of steps in block diffusion.
|
| 2772 |
BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals
2609.37993
|
cs.AI
|
Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen, Esteban Gaillac |
The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admit...The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an entailment cascade has checked it against the passage it cites, and an answer is released only once enough checked evidence stands behind it. Each question is run three or four times, every pass retrieving from a corpus stripped of what the earlier passes have already seen. The confidence filed with each answer is computed by the orchestrator from what the run leaves behind and is never asked of the model, which is offered no way to rate itself. The two retrieval runs placed 4th and 5th of 22, pooling the passes was worth 0.0709 nDCG@20, and the gain was largest on the multi-hop and post-processing-heavy questions, where the organisers rank the pooled run top of the field. Sixteen of the 25 answer runs were built on passages these two runs supplied, 12 of them filed by other teams. HMR rewards a system whose confidence is high where it answers right and low where it answers wrong. The pipeline reached an accuracy of 0.9219, 6th of 25, while the confidence filed with those answers gave an HMR of 0.4915, 13th. A few rules crafted over those same recorded signals, with no further model call and no further retrieval, raise that to an accuracy of 0.9375, 5th, and an HMR of 0.6985, 9th. Ranking on HMR alone can reward a system for answering wrongly with low confidence, so we propose accHMR, the accuracy multiplied by HMR, which reports the reward in proportion to the accuracy, and on which the revised rules would have scored 0.6549, 5th. For future work, fitting a model on the numbers the pipeline already produces, rather than writing such rules by hand, would be a real step forward.
|
| 2773 |
No Scale Left Behind: Multi-Scale Autoencoder with Bi-directional Attention for Time Series Anomaly Detection
2609.38004
|
cs.AI
|
Jiaheng Guo, Haochen Zhang, Yu-Chao Huang, Jinhao Duan, Nicholas Konz |
Time series anomaly detection (TSAD) plays a crucial role in healthcare, finance, industrial monitoring, and other sectors. Within and between these settings, anomalies span vastly different temporal scales, from sub-second point spikes to multi-hour drift pat...Time series anomaly detection (TSAD) plays a crucial role in healthcare, finance, industrial monitoring, and other sectors. Within and between these settings, anomalies span vastly different temporal scales, from sub-second point spikes to multi-hour drift patterns. However, most existing TSAD methods commit to a single temporal granularity, and multi-scale designs either analyze different scales in isolation or are constrained to a predefined coarse-to-fine hierarchy, both failing to sufficiently capture multi-scale interactions. To resolve this limitation, we propose Multi-Scale Autoencoder with Cross-Scale Attention for TSAD (MSCAD), a simple yet powerful semi-supervised TSAD framework founded on parallel autoencoder branches corresponding to different patch sizes. A stack of symmetric bidirectional cross-scale attention blocks enables every pair of scales to exchange information before reconstruction without allowing any single scale to be privileged. On the comprehensive TSB-AD benchmark (40 datasets, 530 series), MSCAD achieves large performance gains against 50 baselines across multiple metrics, with VUS-PR of 0.57(+9.6%) on the univariate split and 0.47(+9.3%) on the multivariate split compared to the state-of-the-art.
|
| 2774 |
Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S
2609.38021
|
cs.AI
|
Christopher J. Chanhnourack |
We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable fina...We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable final reader. The chain places all gold sessions in the candidate pool for 468/470 answerable questions and produces gold-complete packets for 462/470. With a Claude Opus reader called through an unpinned CLI alias, two 500-question passes score 479/500 and 475/500 under GPT-4o. The 72 answerable knowledge-update rows used a substantively modified scoring prompt whose effect under the official text has not been measured. The pair straddles Chronos High's published 478/500; differences in reader generation, scoring prompt, and possibly data version, plus within-system variance, establish neither superiority nor equivalence. A grok-4.6-high reader on the same packets scores 476/474, while a maximum-reasoning-effort agentic variant regresses to 461/465. The headline passes differ on eight verdict-flip rows. A second judge agrees with the headline judge on 493/500 rows (98.6%) in each pass and scores both passes 472/500; the official judge also flips three verdicts when re-scoring byte-identical pass-1 answers. Negative controls rejected a verifier that repaired three wrong drafts but broke eleven correct drafts. All components were developed on the same 500 questions, with no held-out evaluation or independent human adjudication; retrieval and scaffold method sources and transcript-derived audits are held; and the headline reader received extra operator context, its complete requests were not retained, and MCP tool availability is unresolved. We release materialized packets, scaffolds, reader outputs, judge verdicts, and controls for inspection and re-scoring.
|
| 2775 |
Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
2609.38025
|
cs.AI
|
Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu |
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every to...On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by $9.7$ points over vanilla OPD, and enables the smaller student to surpass its larger teacher.
|
| 2776 |
Gender bias across LLMs is common and highly heterogenous
2609.38036
|
cs.AI
|
Edoardo Bolzoni, Valerio Capraro |
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which ...Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.
|
| 2777 |
Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL
2609.38065
|
cs.AI
|
Mathias Jackermeier, Jacques Cloete, Alessandro Abate |
Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted fo...Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high computational costs limit the scale and statistical reliability of experiments. We introduce Jaxolotl, a unified high-performance benchmark suite for multi-task LTL-RL to address these concerns. Jaxolotl provides a modular, end-to-end JAX implementation of six representative algorithms and four environments, together with newly curated task suites and a standardised, statistically robust evaluation protocol. By precompiling symbolic task representations into static arrays, Jaxolotl enables fully JIT-compiled training and evaluation, achieving end-to-end speedups of up to $220\times$ and supporting controlled comparisons at substantially greater experimental scale. We use this framework to systematically evaluate existing approaches, revealing complementary strengths and limitations: general methods capable of non-myopic reasoning struggle as the number of propositions grows, while methods with stronger scaling rely on environment-specific assumptions and suffer from myopia.
|
| 2778 |
Neural topology optimization of ship structures under propulsion machinery vibrations
2609.38089
|
cs.AI
|
Shengyu Yan, Muhammad Muztahidul Hakim Zareer, Jasmin Jelovica |
Ship structural vibrations contribute to noise, fatigue, and equipment damage, while dynamic-compliance topology optimization can produce pathological designs near resonance. This study extends neural-reparameterized topology optimization using a convolutional...Ship structural vibrations contribute to noise, fatigue, and equipment damage, while dynamic-compliance topology optimization can produce pathological designs near resonance. This study extends neural-reparameterized topology optimization using a convolutional Kolmogorov-Arnold network (KATO) to forced-vibration design with active input power (AIP) as the objective. Applications include a 100 Hz engine-supporting deck panel and an 18 Hz thruster foundation frame. Helmholtz PDE filtering and Heaviside projection control feature sizes and manufacturing tolerance. Across both deck families, all eight optimized layouts reduce AIP relative to size-optimized references and, after finite-depth extrusion, also achieve lower static compliance. For unrestricted, manufacturing-aware, and stress-aware frame variants, KATO matches GCMMA in AIP within 0.5 dB while yielding 22-36x lower static compliance after matched-volume binary re-analysis. In a near-resonant 300 Hz case, both methods reduce initial AIP by more than 32 dB; KATO maintains a connected design, achieves 59x lower binary static compliance, and reduces maximum AIP over 1-500 Hz by 2.7 dB. KATO runs 6.4-10.4x faster than GCMMA for the implemented stress-aware formulations. The results demonstrate neural AIP-driven topology optimization as an efficient approach for designing connected, feature-size-controlled ship structures with improved forced-vibration performance.
|
| 2779 |
Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces
2609.38107
|
cs.AI
|
Ratish Puduppully, Pranabendu Misra, Paarth Iyer, Durgesh Kalwar, Vardhan Palod |
Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically ver...Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially, iGSM exposes the exact quantities and dependencies that a correct solution should use, allowing generated traces to be checked programmatically step by step and enabling us to test whether correct answers are reliably accompanied by valid traces. We first evaluate models trained exclusively on valid, minimal traces. Answer correctness and trace validity nearly coincide in distribution but decouple out of distribution: on the hardest instances, 31.6% of correct answers have invalid traces, over half of which pass all syntactic and arithmetic checks but fail semantic dependency checks. We then intervene on trace supervision. Non-minimal training traces induce non-minimal outputs, while re-asking the same problem with a different query reveals computations inherited from the original query, weakening minimality as evidence of selective planning. Shuffling tokens in 10% of training trace sentences preserves near-clean accuracy even out of distribution despite no trace passing verification. Swapped training traces likewise retain high in-distribution accuracy. We discuss the implications of these findings for chain-of-thought monitoring and interpretation in the context of AI safety.
|
| 2780 |
How Local Mixing Encodes Relative Position in Global NoPE Attention
2609.38109
|
cs.AI
|
Cutter Dawes, Nick Alonso, Tom Figliolia, Beren Millidge |
The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding ...The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers. Supported by both theoretical and empirical evidence, our central argument is that SWA and gated linear attention induce a recency bias in the residual stream that propagates to, and is selected by, the global attention logits. Moreover, in contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences. In addition to deepening our understanding of how hybrid models encode position, these findings may provide insights for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.
|
| 2781 |
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
2609.38166
|
cs.AI
|
Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu |
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the ...Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70$\times$ at the kernel level and 1.47$\times$ for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.
|
| 2782 |
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
2609.38169
|
cs.AI
|
Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo |
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy ...Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.
|
| 2783 |
Skill-Space Shooting for Autonomous Robot Policy Improvement
2609.38178
|
cs.AI
|
Zihang Rui, Renhao Wang, Haoxu Huang, Yang Gao |
Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstratio...Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable corrections for the policy. Our insight is that many such corrections are familiar short behaviors, or skills: they recur across tasks and describe actions that foundation models can reason about from a scene. We introduce skill-space shooting, which uses foundation model guidance to explore corrections through these reusable skills and turn successful trials into policy improvement. Real-world experiments show repeated improvement in policies acting autonomously, while skills can also be shared to reduce the teaching needed to improve on new tasks. By making reusable skills a source of corrective supervision, skill-space shooting enables scalable and generalizable policy improvement within and across tasks. Additional results and videos at https://skill-space-shooting.github.io.
|
| 2784 |
Implementing Cumulative Functions with Generalized Cumulative Constraints
2508.01751
|
cs.AI
|
Pierre Schaus, Charles Thomas, Roger Kameugne |
Modeling scheduling problems with conditional time intervals and cumulative functions has become a common approach when using modern commercial constraint programming solvers. This paradigm enables the modeling of a wide range of scheduling problems, including...Modeling scheduling problems with conditional time intervals and cumulative functions has become a common approach when using modern commercial constraint programming solvers. This paradigm enables the modeling of a wide range of scheduling problems, including those involving producers and consumers. However, it is unavailable in existing open-source solvers and practical implementation details remain undocumented. In this work, we present an implementation of this modeling approach using a single, generic global constraint called the Generalized Cumulative. We also introduce a novel timetabling filtering algorithm specifically designed to handle tasks defined on conditional time-intervals. Experimental results demonstrate that this approach, combined with the new filtering algorithm, performs competitively with existing solvers enabling the modeling of producer and consumer scheduling problems and effectively scales to large-scale problems.
|
| 2785 |
Sequence Variables: A Constraint Programming Computational Domain for Routing and Sequencing
2510.09373
|
cs.AI
|
Augustin Delecluse, Pierre Schaus, Pascal Van Hentenryck |
Constraint Programming (CP) offers an intuitive, declarative framework for modeling Vehicle Routing Problems (VRP). While classical successor-based CP models can be adapted to handle optional visits or insertion-based heuristics, sequence variables provide a s...Constraint Programming (CP) offers an intuitive, declarative framework for modeling Vehicle Routing Problems (VRP). While classical successor-based CP models can be adapted to handle optional visits or insertion-based heuristics, sequence variables provide a significantly more natural and elegant formulation for these requirements. Building upon our prior work that introduced the initial concept, the main contribution of this article is the complete semantic and operational formalization of sequence variables as a computational domain. Specifically, we formally define the sequence domain and its update operations, and detail the implementation and data structures required to integrate sequence variables into trail-based CP solvers. Furthermore, we introduce consistency levels for associated constraints on this domain alongside specialized global constraints tailored for routing problems. Finally, we demonstrate that sequence variables simplify problem modeling while achieving competitive computational performance on Pickup and Delivery Problems with and without Time Windows, the Dial-a-Ride Problem, and a Prize-Collecting Scheduling Problem.
|
| 2786 |
Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Markets
2512.02436
|
cs.AI
|
Agostino Capponi, Alfio Gliozzo, Brian Zhu |
Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implicit equivalences, and hidden contradictions across markets. We present an agentic AI (AAI) pipeline that autonomously reco...Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implicit equivalences, and hidden contradictions across markets. We present an agentic AI (AAI) pipeline that autonomously recovers cross-market structure from contract text before prices enter the analysis. The workflow first clusters markets into coherent topical groups using natural-language understanding over contract text and metadata, and then identifies contracts within each cluster, but from different event markets, that exhibit strong dependence or leader--follower relationships. We evaluate this system, along with a natural language inference (NLI) benchmark, on a large prediction market dataset from early 2026. Using resolved outcomes to evaluate identified relations, we find that AAI-identified relations are 62.8\% consistent with exchange-recorded settlements, whereas the NLI benchmark only achieves 40.6\% accuracy. Within clusters, the AAI output is sparse and also remarkably compatible as a signed graph with a frustration rate of 0.324\%. As an application, we show how discovered relations inform semantics-based trading strategies on prediction markets. One such strategy yields 14.12\% net ROI after fees in a two-month period in 2026. Overall, we demonstrate the potential for agentic AI as a structural discovery layer for prediction markets.
|
| 2787 |
Nonlinearity as Rank: Generative Low-Rank Adapter with Radial Basis Functions
2602.05709
|
cs.AI
|
Yihao Ouyang, Shiwei Li, Haozhao Wang, Xiandi Luo, Zhuoqi Hu |
Low-rank adaptation (LoRA) approximates the update of a pretrained weight matrix using the product of two low-rank matrices. However, standard LoRA follows an explicit-rank paradigm, where increasing model capacity requires adding more rows or columns (i.e., b...Low-rank adaptation (LoRA) approximates the update of a pretrained weight matrix using the product of two low-rank matrices. However, standard LoRA follows an explicit-rank paradigm, where increasing model capacity requires adding more rows or columns (i.e., basis vectors) to the low-rank matrices, leading to substantial parameter growth. In this paper, we find that these basis vectors exhibit significant parameter redundancy and can be compactly represented by lightweight nonlinear functions. Therefore, we propose Generative Low-Rank Adapter (GenLoRA), which replaces explicit basis vector storage with nonlinear basis vector generation. Specifically, GenLoRA maintains a latent vector for each low-rank matrix and employs a set of lightweight radial basis functions (RBFs) to synthesize the basis vectors. Each RBF requires far fewer parameters than an explicit basis vector, enabling higher parameter efficiency in GenLoRA. Extensive experiments across multiple datasets and architectures show that GenLoRA attains higher effective LoRA ranks under smaller parameter budgets, resulting in superior fine-tuning performance. The code is available at https://anonymous.4open.science/r/GenLoRA.
|
| 2788 |
FloCA: Towards Faithful and Logically Consistent Flowchart Reasoning
2602.14035
|
cs.AI
|
Jinzi Zou, Bolin Wang, Shuo Zhang, Nuo Xu, Junzhou Zhao |
Flowchart-oriented dialogue (FOD) systems aim to guide users through multi-turn decision-making or operational procedures by following a domain-specific flowchart to achieve a task goal. In this work, we formalize flowchart reasoning in FOD as grounding user i...Flowchart-oriented dialogue (FOD) systems aim to guide users through multi-turn decision-making or operational procedures by following a domain-specific flowchart to achieve a task goal. In this work, we formalize flowchart reasoning in FOD as grounding user input to flowchart nodes at each dialogue turn while ensuring node transition is consistent with the correct flowchart path. Despite recent advances of LLMs in task-oriented dialogue systems, adapting them to FOD still faces two limitations: (1) LLMs lack an explicit mechanism to represent and reason over flowchart topology, and (2) they are prone to hallucinations, leading to unfaithful flowchart reasoning. To address these limitations, we propose FloCA, a zero-shot flowchart-oriented conversational agent. FloCA uses an LLM for intent understanding and response generation, while delegating flowchart reasoning to an external tool that performs topology-constrained graph execution, ensuring faithful and logically consistent node transitions across dialogue turns. We further introduce an evaluation framework with an LLM-based user simulator and five new metrics covering reasoning accuracy and interaction efficiency. Extensive experiments on FLODIAL and PFDial datasets highlight the bottlenecks of existing LLM/VLM-based methods and demonstrate the superiority of FloCA. The code and dataset are publicly available at https://github.com/Jinzi-Zou/FloCA-flowchart-reasoning.
|
| 2789 |
Evaluating Test-Time Scaling of General LLM Agents
2602.18998
|
cs.AI
|
Xiaochuan Li, Ryan Ming, Pranav Setlur, Abhijay Paladugu, Andy Tang |
LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systematically investigate two principal...LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systematically investigate two principal test-time scaling axes of LLM agents: sequential scaling through extended interaction and parallel scaling through trajectory sampling. We first introduce a realistic benchmark that provides one unified framework for evaluating LLM agents across search, coding, reasoning, and tool-use domains, more faithfully reflecting the heterogeneity of real-world deployments. Evaluating ten leading LLM agents reveals substantial performance degradation when transitioning from domain-specific evaluations to this realistic setting. Building on this foundation, we progressively scale test-time compute along fine-grained increments to characterize the performance upper bound. We find that neither scaling axis can consistently yield meaningful gains from additional test-time compute in realistic environments, a phenomenon we attribute to two fundamental limitations: the scaling plateau that bottlenecks sequential scaling and the verification gap that undermines parallel scaling. Code is publicly available at https://github.com/cxcscmu/General-AgentBench.
|
| 2790 |
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
2602.24055
|
cs.AI
|
Reva Schwartz, Carina Westling, Morgan Briggs, Marzieh Fadaee, Isar Nejadgholi |
This study proposes CIRCLE, a six-stage, lifecycle-based framework to bridge the reality gap between model-centric performance metrics and AI system outcomes in deployment. Current approaches such as MLOps frameworks and AI model benchmarks offer detailed insi...This study proposes CIRCLE, a six-stage, lifecycle-based framework to bridge the reality gap between model-centric performance metrics and AI system outcomes in deployment. Current approaches such as MLOps frameworks and AI model benchmarks offer detailed insights into system stability and model capabilities, but they do not provide decision makers outside the AI stack with systematic evidence of how these systems actually behave in real world contexts or affect their organizations over time. CIRCLE operationalizes the Validation phase of TEVV (Test, Evaluation, Verification, and Validation) by translating priorities of stakeholders outside the stack into measurable signals. Unlike participatory design which often remains localized, or algorithmic audits which are often retrospective, CIRCLE provides a structured, prospective protocol for linking context sensitive qualitative insights to scalable quantitative metrics. By integrating methods such as field testing, red teaming, and longitudinal studies into a coordinated pipeline, CIRCLE produces systematic knowledge; evidence that is comparable across sites yet sensitive to local context. This can enable governance based on materialized downstream effects rather than theoretical capabilities.
|
| 2791 |
OntoTKGE: Ontology-Enhanced Temporal Knowledge Graph Extrapolation
2604.05468
|
cs.AI
|
Dongying Lin, Yinan Liu, Shengwei tang, Bin Wang, Xiaochun Yang |
Temporal knowledge graph (TKG) extrapolation is an important task that aims to predict future facts through historical interaction information within KG snapshots. A key challenge for most existing TKG extrapolation models is handling entities with sparse hist...Temporal knowledge graph (TKG) extrapolation is an important task that aims to predict future facts through historical interaction information within KG snapshots. A key challenge for most existing TKG extrapolation models is handling entities with sparse historical interaction. The ontological knowledge is beneficial for alleviating this sparsity issue by enabling these entities to inherit behavioral patterns from other entities with the same concept, which is ignored by previous studies. In this paper, we propose a novel encoder-decoder framework OntoTKGE that leverages the ontological knowledge from the ontology-view KG (i.e., a KG modeling hierarchical relations among abstract concepts as well as the connections between concepts and entities) to guide the TKG extrapolation model's learning process through the effective integration of the ontological and temporal knowledge, thereby enhancing entity embeddings. OntoTKGE is flexible enough to adapt to many TKG extrapolation models. Extensive experiments on five data sets demonstrate that OntoTKGE not only significantly improves the performance of many TKG extrapolation models but also surpasses many state-of-the-art(SOTA) baseline methods.
|
| 2792 |
Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results
2604.21965
|
cs.AI
|
Benjamin Kohler, David Zollikofer, Johanna Einsiedler, Alexander Hoyle, Elliott Ash |
Recent work has used LLM agents to reproduce empirical social science results with access to both the data and code. We broaden this scope by asking: Can they reproduce results given only a paper's methods description and original data? We develop an agentic r...Recent work has used LLM agents to reproduce empirical social science results with access to both the data and code. We broaden this scope by asking: Can they reproduce results given only a paper's methods description and original data? We develop an agentic reproduction system that extracts structured methods descriptions from papers, runs reimplementations under strict information isolation -- agents never see the original code, results, or paper -- and enables deterministic, cell-level comparison of reproduced outputs to the original results. An error attribution step traces discrepancies through the system chain to identify root causes. Evaluating four agent scaffolds and four LLMs on 48 papers with human-verified reproducibility, we find that agents can largely recover published results, but performance varies substantially between models, scaffolds, and papers. Root cause analysis reveals that failures stem both from agent errors and from underspecification in the papers themselves.
|
| 2793 |
TimeTok: Granularity-Controllable Time-Series Generation via Hierarchical Tokenization
2605.01418
|
cs.AI
|
Seokhyun Lee, Jaeho Kim, Changjun Oh, Mihaela van der Schaar, Changhee Lee |
Time-series data are inherently multiscale, spanning diverse temporal granularities from coarse trends to fine-scale dynamics. However, existing time-series generative models provide limited control over the temporal granularity of both inputs and outputs, res...Time-series data are inherently multiscale, spanning diverse temporal granularities from coarse trends to fine-scale dynamics. However, existing time-series generative models provide limited control over the temporal granularity of both inputs and outputs, restricting their ability to condition on user-provided coarse sketches and generate samples at a desired target granularity. To address this, we introduce TimeTok, a unified framework for Granularity-Controllable Time-Series Generation (GC-TSG), which generates time series at any target granularity from any coarser input (e.g., rough sketches) or without conditioning. At the core of TimeTok is a hierarchical tokenization strategy that maps time series into an ordered sequence of tokens, from coarse to fine temporal granularity. Our autoregressive generation process operates across these granularity levels, producing token blocks that are decoded back into continuous time series. This design naturally enables GC-TSG within a single framework, where controlling the number of token blocks provides explicit control over output detail. Experiments show that TimeTok excels at GC-TSG tasks while achieving state-of-the-art performance in standard generation. Furthermore, we showcase TimeTok's potential as a foundational tokenizer by training on multiple datasets with heterogeneous temporal granularities, verifying strong transferability that consistently outperforms models trained on individual datasets. To our knowledge, this is the first unified framework that covers the full generative spectrum for time series, offering a valuable foundation for models that benefit from diverse temporal granularities.
|
| 2794 |
ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries
2605.06223
|
cs.AI
|
Junhyuk Kwon, Seungjoon Lee, Hyejin Park, Kyle Min, Jungseul Ok |
Natural-language instance navigation becomes challenging when the initial user request does not uniquely specify the target instance. A practical agent should reduce the user's burden by actively asking only the information needed to distinguish the target fro...Natural-language instance navigation becomes challenging when the initial user request does not uniquely specify the target instance. A practical agent should reduce the user's burden by actively asking only the information needed to distinguish the target from similar distractors, rather than requiring a detailed description upfront. Existing approaches often fall short of this goal by mistaking distractors that strongly match the accumulated information about the target provided by the user. As a result, despite the dialogue, the agent may still fail to distinguish the target from distractors, leading to premature decisions and lengthy user responses. We propose Proactive Instance Navigation with Comparative Judgment (ProCompNav), a two-stage framework that first constructs a candidate pool and then identifies the target through Recursive Comparative Judgment (RCJ). RCJ iteratively narrows the pool by selecting an attribute-value pair that divides the candidates, asking the user a binary question, and removing inconsistent candidates, without requiring an attribute unique to the target. On CoIN-Bench, ProCompNav outperforms the evaluated baselines in Success Rate while substantially reducing Response Length. On the non-interactive TextNav benchmark, ProCompNav achieves the highest Success Rate. Two human studies further show that participants prefer ProCompNav's interaction strategies.
|
| 2795 |
Process Matters more than Output for Distinguishing Humans from Machines
2605.06524
|
cs.AI
|
Milena Rmus, Mathew D. Hardy, Thomas L. Griffiths, Mayank Agrawal |
Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online settings. Existing approaches evaluate whether a system can produce responses indistinguishable from those of a human...Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online settings. Existing approaches evaluate whether a system can produce responses indistinguishable from those of a human. This approach follows the focus on the output of a machine, as suggested by Alan Turing. Cognitive science provides an alternative approach: considering the process by which that behavior is produced. To evaluate whether processes can reliably distinguish humans from machines, we introduce a process-based framework, the Process Turing Test, and evaluate it across a battery of cognitive tasks spanning decision-making, working memory, and planning. These tasks, such as mental rotation and sequence prediction, yield process-level measures complementing conventional measures of overall task performance. We also include multiple CAPTCHA tasks in the battery. Across the battery, process-level features provide substantially stronger discriminative signal than performance metrics alone, reliably distinguishing humans from agents even when task performance is matched (process-based classifier AUC = 0.88). We also conducted a controlled red-teaming study comparing off-the-shelf frontier agents (Claude Sonnet 4.5, GPT-5, Gemini 2.5 Pro), Centaur (LLM fine-tuned on 10.7M human decisions), and two task-specific fine-tuning methods: action-level supervised fine-tuning (A-SFT) and process-level fine-tuning (P-SFT), which directly optimizes process features. We find that broad fine-tuning on human choices makes task processes more human-like relative to off-the-shelf frontier agents, and task-specific P-SFT further improves human-like behavioral mimicry, though this advantage largely disappears under cross-task transfer. These results highlight process specification as a central bottleneck in achieving human-like cognitive processes in machines.
|
| 2796 |
EquiMem: Calibrating Shared Memory in Multi-Agent Debate via Game-Theoretic Equilibrium
2605.09278
|
cs.AI
|
Yuqiao Meng, Luoxi Tang, Sakshi Sunil Narvekar, Rupali Rajendra Vaje, Yingxue Zhang |
Multi-agent debate (MAD) systems increasingly rely on shared memory to support long-horizon reasoning, but this convenience opens a critical vulnerability: a single corrupted entry can contaminate the downstream memory-augmented reasoning, and debate alone fai...Multi-agent debate (MAD) systems increasingly rely on shared memory to support long-horizon reasoning, but this convenience opens a critical vulnerability: a single corrupted entry can contaminate the downstream memory-augmented reasoning, and debate alone fails to filter such errors. Existing safeguards filter entries via heuristics or LLM-based validation, yet they rely on AI judgments that share the same failure modes and overlook the cross-agent dynamics of MAD. We address this gap by formulating memory updating in MAD as a zero-trust memory game, in which no agent is assumed reliable and the game's equilibrium motivates a principled objective for calibrating memory influence. Guided by this objective, we propose EquiMem, an inference-time calibration mechanism that evaluates each update against the shared memory state and the memory already being used by other agents, using their existing retrieval queries and traversal paths as evidence of calibration. EquiMem instantiates this calibration for both embedding- and graph-based memory, and across diverse benchmarks, MAD frameworks, and memory architectures, it consistently outperforms existing safeguards, remains robust under adversarial agents, and incurs negligible inference overhead.
|
| 2797 |
Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
2605.14322
|
cs.AI
|
Zixin Chen, Peng Liu, Rui Sheng, Haobo Li, Jianhong Tu |
Language agents are increasingly deployed in professional workflows, yet tutoring remains a high-stakes capability that existing evaluations only partially capture. Effective tutor agents require more than producing correct answers or executing accurate tool c...Language agents are increasingly deployed in professional workflows, yet tutoring remains a high-stakes capability that existing evaluations only partially capture. Effective tutor agents require more than producing correct answers or executing accurate tool calls: they must infer a warranted teaching decision from evidence, adapt support as learner state changes, and carry an instructor's request through a learning-management system (LMS) to a completed, verified intervention. We introduce TeachArena, a source-grounded benchmark that jointly evaluates three complementary surfaces of teaching work: professional pedagogical judgment, situated multi-turn tutoring, and end-to-end LMS teaching workflows. Its 354 audited tasks are each built around a pedagogical insight, grounded in evidence, and evaluated with matched verifiers over observable turn-level responses, tutoring trajectories, and persistent artifacts or environment states. Across a comprehensive evaluation of frontier models, our findings reveal that current models are generally capable of bounded pedagogical judgment, but still fall short of professional teaching standards in situated tutoring and end-to-end teaching-workflow execution. By unifying teacher judgment, adaptive tutoring, and institutional action in one auditable benchmark, TEACHARENA provides a measurement foundation for developing tutor agents that can support realistic teaching work.
|
| 2798 |
COLLATOR: Compositional Multi-Agent Orchestration with Counterfactual Reinforcement Learning
2605.14483
|
cs.AI
|
Xudong Chen, Yixin Liu, Hua Wei, Kaize Ding |
Large language models (LLMs) provide a flexible foundation for multi-agent systems, but their effectiveness and computational cost depend critically on orchestration design. Across different tasks, role design, capacity assignment, and dependency construction ...Large language models (LLMs) provide a flexible foundation for multi-agent systems, but their effectiveness and computational cost depend critically on orchestration design. Across different tasks, role design, capacity assignment, and dependency construction jointly affect both solution quality and execution efficiency. Existing approaches automate parts of this design process, yet they often optimize these decisions partially or sequentially, and rely on execution-level feedback that provides limited credit assignment for local orchestration decisions. We propose LEMON (Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning), an LLM-based orchestrator that learns to design efficient multi-agent orchestration. Given a task, LEMON designs a unified orchestration specification that composes customized agent duties, capacity levels, and dependency relations. To train the orchestrator, we augment the orchestration-level Group Relative Policy Optimization (GRPO) objective with a localized counterfactual credit signal that edits role, capacity, or dependency fields and applies the resulting reward contrast only to the edited spans. Experiments on six reasoning and coding benchmarks, including MMLU, GSM8K, AQuA, MultiArith, SVAMP, and HumanEval, show that LEMON achieves the best average performance among evaluated multi-agent orchestration methods while improving the accuracy-token trade-off.
|
| 2799 |
Ottilie: A Socially Intelligent Virtual Host for Sales-Driven Live Commerce
2605.14542
|
cs.AI
|
Yuyan Chen |
A skilled live-commerce host is not merely a narrator, but a sales agent who converts viewer curiosity into purchase intent through expert product knowledge, emotionally intelligent response tactics, and entertainment that serves as a vehicle for product expos...A skilled live-commerce host is not merely a narrator, but a sales agent who converts viewer curiosity into purchase intent through expert product knowledge, emotionally intelligent response tactics, and entertainment that serves as a vehicle for product exposure. Yet no existing AI system replicates this. Conversational recommenders treat recommendation as a terminal act, while general-purpose LLMs hallucinate product claims and default to generic promotional templates that fail to engage or persuade. We present Ottilie, a sales-conversion-oriented virtual host for Chinese beauty live commerce that grounds product responses in a verified knowledge base, adapts discourse strategy to viewer intent through empathetic amplification, evidence-backed rebuttal, and humor-mediated deflection, and supervises all four output fields simultaneously within a structured generation framework. Experiments demonstrate gains of 14% on informativeness and 5% on factual correctness, with consistent advantages in tactfulness and viewer engagement. Ottilie has been deployed as an autonomous English-language host on Twitch, where the same fine-tuned model runs unattended on a single commodity GPU and promotes the company's own product; we report how the system adapted to a new language, market, and product without retraining, and what the deployment revealed about latency and adversarial chat.
|
| 2800 |
Harnessing LLM Agents with Skill Programs
2605.17734
|
cs.AI
|
Hongjun Liu, Yifei Ming, Shafiq Joty, Chen Zhao |
Equipping LLM agents with reusable skills derived from past experience has become a popular and successful approach for tackling complex and long-horizon tasks. However, such lessons are often encoded as textual guidance that remains largely advisory, lacking ...Equipping LLM agents with reusable skills derived from past experience has become a popular and successful approach for tackling complex and long-horizon tasks. However, such lessons are often encoded as textual guidance that remains largely advisory, lacking explicit mechanisms for when and how to intervene in the agent loop. To bridge the gap, we introduce HASP(Harnessing LLM Agents with Skill Programs), a new framework that upgrades skills into executable Program Functions (PFs). Rather than offering passive advice, PFs act as executable guardrails that activate on failure-prone states and modify the next action or inject corrective context. HASP is highly modular: it can be applied at inference time for direct agent-loop intervention, during post-training to provide structured supervision, or for self-improvement by evolving validated, teacher-reviewed PFs. Empirically, HASP drives substantial gains compared to both training-free and training-based methods on web-search, math reasoning, and coding tasks. For example, on web-search reasoning, inference-time PFs alone improve the average performance by 25% compared to (multi-loop) ReAct Agent, while post-training and controlled evolution achieve a 30.4% gain over Search-R1. To provide deeper insights into HASP, our mechanism analysis reveals how PFs trigger and intervene, how skills are internalized, and the requirement for stable skill library evolution.
|
| 2801 |
DemoEvolve: Demonstration-Guided Harness Evolution under Sparse Feedback
2605.24539
|
cs.AI
|
Lirong Che, Yuzhe yang, Peiwen lin, Xu Cao, Chuang wang |
Harness evolution enables frozen language model agents to adapt to unfamiliar tasks by modifying the external programs that govern their behavior. For long-horizon tasks, each rollout is costly, and a limited interaction budget may yield few examples of effect...Harness evolution enables frozen language model agents to adapt to unfamiliar tasks by modifying the external programs that govern their behavior. For long-horizon tasks, each rollout is costly, and a limited interaction budget may yield few examples of effective behavior. Sparse, delayed feedback also makes it difficult for a coding agent to diagnose failures and determine which modifications will improve performance. We present DemoEvolve, which uses human demonstrations to guide harness evolution under limited interaction budgets. The coding agent examines demonstrations alongside its own rollout history, extracts strategies and the conditions under which they apply, and turns them into reusable harness components. We compare demonstration guidance with self-rollout evolution and augmentation with retrieved textual knowledge under the same evolution procedure and agent interaction budget. On held-out seeds, DemoEvolve improves mean capped final-round progress from 16.83 to 20.00 in Balatro and mean floor reached from 18.17 to 28.83 in Slay the Spire 2, compared with the respective base harnesses. These results suggest that demonstrations can help harness evolution make better use of limited interaction experience by providing concrete guidance for diagnosis and modification.
|
| 2802 |
Boosting Knowledge Graph Foundation Models via Enhanced Negative Sampling
2605.27023
|
cs.AI
|
Yinan Liu, Wenjin Xu, Zhiyuan Zha, Xiaochun Yang, Bin Wang |
Knowledge graphs (KGs) have become the core backbone of numerous downstream tasks such as question answering and recommender systems. However, despite all this, KGs are often very incomplete. To perform zero-shot knowledge graph completion in unseen KGs, which...Knowledge graphs (KGs) have become the core backbone of numerous downstream tasks such as question answering and recommender systems. However, despite all this, KGs are often very incomplete. To perform zero-shot knowledge graph completion in unseen KGs, which have different relational vocabularies from those used for pre-training, KG foundation models (KGFMs) receive a wide range of attention. Existing KGFMs often perform training using random negative triples, which are constructed by replacing the head or tail entity of a positive triple with a random entity. However, these negative triples are often constructed with limited quality, providing weak supervision for KGFM training. In this paper, we propose a simple yet effective adaptive negative sampling approach, KMAS, to enhance existing KGFMs. KMAS constructs hard negative triples through the updated relation embeddings generated from the existing KGFM's relation encoder. To further adaptively align with the evolving capability of the KGFM during the training process, KMAS adjusts the ratio of hard negative triples dynamically throughout the whole training process: after a warmup phrase, it increases the ratio linearly and then decreases linearly. Extensive experiments are conducted over 44 data sets. Experimental results demonstrate that our proposed negative sampling method can enhance many SOTA KGFMs without requiring excessive additional time or memory consumption.
|
| 2803 |
KLineage: Recovering the Missing When of Kernel Optimization by Deoptimizing Experts
2605.28213
|
cs.AI
|
Shuoming Zhang, Qiuchu Yu, Ruiyuan Xu, Chenjing Zhang, Junjie Peng |
LLM-based agents are increasingly used to generate GPU kernels, but they often struggle to determine when an optimization is sound because its required code state and dependencies are implicit in expert implementations. We introduce KLineage, which learns this...LLM-based agents are increasingly used to generate GPU kernels, but they often struggle to determine when an optimization is sound because its required code state and dependencies are implicit in expert implementations. We introduce KLineage, which learns this missing "when" knowledge from expert kernels: instead of relying on forward rollouts, KLineage walks expert implementations backward through validation-gated simplifications and reverses each accepted step into a reusable optimization skill. Each skill records not only the optimization intent, but also when to apply the optimization technique, including where it applies in code, what conditions made it valid, what effect it has, and what failures its assumptions avoid. A downstream LLM materializes these skills on new code surfaces under the same compile/correctness/profile gate. This guidance on when to apply each optimization can help downstream models to generate higher-performance kernels. On five expert workloads across two NVIDIA architectures, these lineage-derived skills serve as an effective optimization curriculum, exceeding recent memory-based LLM-kernel baselines in both final kernel quality and optimization efficiency under the same fixed budget. We also demonstrate that the KLineage framework extends beyond NVIDIA GPUs to Ascend NPUs. Our code is publicly available at https://github.com/ict-agent/klineage.
|
| 2804 |
Diversifying RLVR Rollouts via First-Token Exploration
2605.28295
|
cs.AI
|
Soeun Kim, Albert No |
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models without labeled trajectories, using groups of verifier-scored rollouts to explore alternative reasoning paths. Limited rollout diversity is a central bottleneck, typically addressed ...Reinforcement learning with verifiable rewards (RLVR) trains reasoning models without labeled trajectories, using groups of verifier-scored rollouts to explore alternative reasoning paths. Limited rollout diversity is a central bottleneck, typically addressed through adjustments to temperature, prefixes, or rollout selection. We identify the first token of the response as a structurally distinct target for diversification, largely overlooked in prior work. We find that the first-token distribution is sharply concentrated and only weakly related to downstream correctness, as lower-probability candidates can yield similarly accurate responses. Diversifying the first token can therefore broaden the reasoning paths explored within each rollout group with little loss in response quality. Motivated by this observation, we introduce REFT (Rollout Exploration with First-Token Diversification), a lightweight modification to RLVR. REFT samples first tokens uniformly from the policy's top-$N$ candidates and allocates rollouts evenly across the sampled tokens, leaving the rest of the pipeline unchanged. We evaluate REFT on eight models spanning multiple architectures and sizes (0.5B-14B), with mathematical reasoning and code-generation tasks under GRPO and DAPO. Across these settings, REFT consistently improves Pass@1, Pass@8, and Pass@64. It also outperforms competing diversification methods at every evaluated budget, incurring the lowest rollout cost.
|
| 2805 |
DeepSurvey: Agent-Oriented Automated Survey Generation with Analytical Depth and Citation Reliability
2605.29522
|
cs.AI
|
Ziyue Yang, Da Ma, Hanqi Li, Zijian Wang, Tiancheng Huang |
As scientific literature grows rapidly and research increasingly involves AI agents, automated survey generation has become a key capability for both agents and human researchers. For such agents, a survey serves as a primary knowledge source of a field prior ...As scientific literature grows rapidly and research increasingly involves AI agents, automated survey generation has become a key capability for both agents and human researchers. For such agents, a survey serves as a primary knowledge source of a field prior to research. Since pretrained models already encode broad knowledge of established work, these consumers benefit more from analytical depth than from breadth alone; moreover, unsupported claims, once ingested as knowledge, can propagate into downstream research. However, existing systems tend to overemphasize coverage and presentation; they suffer from limited analytical depth due to reliance on abstracts and isolated paper processing, and from unreliable citations due to imprecise retrieval and post-hoc grounding. We present DeepSurvey, an agentic generation system that addresses both limitations. To enhance depth, DeepSurvey extracts structured keynotes, models cross-paper relationships through clustering and comparative analysis, and integrates a code-agent subsystem to recover implementation-level details. To fortify reliability, it combines citation-graph expansion with hybrid filtering for topic-focused retrieval, enforces evidence-constrained analysis and writing, and deploys multi-granularity agentic refinement to validate citation--claim alignment. Experiments show that DeepSurvey achieves the highest content score (8.34/10) and citation quality (recall and precision gains of 25.3\% and 35.2\% over the strongest baseline), generalizes more robustly across domains, and is preferred by domain experts over human-written surveys (83.3\% in overall quality, 100\% in content depth). Moreover, when a coding agent uses a survey as its only literature source, the agent equipped with DeepSurvey achieves the best performance among human-written and baseline-generated surveys.
|
| 2806 |
SoftSkill: Behavioral Compression for Contextual Adaptation
2606.20333
|
cs.AI
|
Xijia Tao, Yihua Teng, Xinyu Fu, Ziru Liu, Kecheng Chen |
Natural-language skills let agents reuse task knowledge, yet deploying a long Markdown document makes the model interpret that knowledge anew on every call. We ask whether the behavior induced by a skill can be carried by a compact, trainable context. SoftSkil...Natural-language skills let agents reuse task knowledge, yet deploying a long Markdown document makes the model interpret that knowledge anew on every call. We ask whether the behavior induced by a skill can be carried by a compact, trainable context. SoftSkill initializes virtual token embeddings from a skill document and optimizes a soft skill with next-token prediction while keeping the language model frozen. The resulting conditioning sequence can occupy the skill section or another supported prompt location. On Qwen3.5-4B, a 32-token soft skill improves over no skill by 7.6, 42.1, and 1.3 points on SearchQA, LiveMath, and DocVQA. It also exceeds the optimized textual skill on SearchQA and LiveMath by 4.5 and 12.5 points, respectively, while replacing skill documents of hundreds to thousands of tokens. The same approach improves multi-step agent execution: on Qwen3.6-35B-A3B, OfficeQA and ALFWorld rise by 8.2 and 14.1 points over no skill; on Qwen3.5-4B with aligned action decoding, ALFWorld success rises from 44/134 with the untrained initialization to 91/134 after training. These findings show that a skill document can serve as the starting point for a continuous control that improves both answers and actions without updating the backbone.
|
| 2807 |
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
2606.24937
|
cs.AI
|
Haggai Roitman |
The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems, covering the full stack from first principles to production deployment. The central thesis: building great agentic systems requires understandi...The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems, covering the full stack from first principles to production deployment. The central thesis: building great agentic systems requires understanding every layer of the pipeline, not just one. The book opens with the LLM substrate, covering transformer architecture, GPU systems, training and fine-tuning (SFT, LoRA, MoE), model compression, and inference optimization, as essential foundations. It then develops the alignment and reasoning layer: RLHF, PPO, DPO and its variants, GRPO, reward modeling, and RL for large reasoning models including chain-of-thought and test-time scaling. The second half is devoted to agentic AI proper: agentic training and trajectory-based RL, RAG and Agentic RAG, memory systems (in-context, external, episodic, and semantic), agent harness design, loop engineering, graph-based orchestration, and a taxonomy of agent design patterns covering security, red teaming, and gateway infrastructure. Inter-agent coordination is covered in depth: the Model Context Protocol (MCP), agent skills and tool use, the Agent-to-Agent (A2A) protocol, and multi-agent architectures spanning centralized, decentralized, and hierarchical topologies. The book concludes with agent development frameworks, agentic UI design, evaluation methodology (non-deterministic evaluation, reasoning collapse, LLM-as-Judge), production deployment, and the regulatory environment (EU AI Act, California SB 942) as an engineering requirement. Each chapter pairs theory with implementation guidance, executable notebooks, and references to the primary literature.
|
| 2808 |
Rater State Bias in RLHF Preference Data: An Audit Framework
2607.16195
|
cs.AI
|
Elena Kopteva, Vitaliy Hlynianyi-Zhuk |
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotation. Under sustained stressful or distres...We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, raters' preferences may shift over time, so that preference data encode rater state alongside judgments about response quality. We argue that, if present, such shifts would differ from ordinary disagreement or random label noise. They would be state dependent, could be shared across annotators under similar conditions, and would not necessarily cancel during aggregation, reward modeling, and policy optimization. We propose rater state shift as a plausible and testable source of structured bias in RLHF preference data. This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also propose survival level emotional authenticity as a candidate output signature, defined by lexical, pragmatic, discourse, and safety features whose reliability and validity remain to be demonstrated. We show that systematic rater state bias can survive aggregation and may enter the learned reward signal. We state five testable predictions, together with effect size thresholds for an initial audit, and note which require proprietary data. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model.
|
| 2809 |
JUMP: Efficient Membership Inference on Fine-Tuned Diffusion Language Models
2607.16207
|
cs.AI
|
Yeachan Jun, Albert No |
Membership inference attacks (MIAs) test whether a candidate example was used to train a language model. Existing attacks on fine-tuned discrete diffusion language models (dLLMs) often aggregate reconstruction signals across many mask configurations, requiring...Membership inference attacks (MIAs) test whether a candidate example was used to train a language model. Existing attacks on fine-tuned discrete diffusion language models (dLLMs) often aggregate reconstruction signals across many mask configurations, requiring repeated model evaluations. We propose Joint Uncertainty Guided Mask Probing (JUMP), an efficient MIA that exploits the ability of dLLMs to predict masked tokens in parallel. Using the pre-fine-tuning checkpoint as a reference, JUMP selects low-confidence positions, masks them jointly, and aggregates clipped target-reference reconstruction gaps. This focuses the attack on positions that reveal stronger membership signals from fine-tuning. After mask selection, all selected tokens are evaluated with one scoring query per model. Across six MIMIR domains, JUMP improves mean ROC-AUC over a prior multi-mask attack from 0.819 to 0.902 on LLaDA and from 0.851 to 0.942 on Dream. Including mask selection, it requires only three forward passes per example, compared with 32 for the baseline. We further extend JUMP to the target-only setting by replacing target-reference scoring with relative token preference, which compares the observed token with alternative predictions at the same masked position. Target-Only JUMP achieves mean ROC-AUCs of 0.609 and 0.638 on LLaDA and Dream.
|
| 2810 |
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
2607.25152
|
cs.AI
|
Hyundoo Park, Byungho Choi |
Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. W...Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.
|
| 2811 |
Belief-Guided Decision Making with Uncertainty Gating in the Game of Go
2607.26946
|
cs.AI
|
Mehrad Yaghoubi, Azam Bastanfard, Abbas Jalilvand, Ashkan Rezaei |
Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottlen...Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottleneck on consumer-grade hardware, where the computational cost of tree management severely limits inference rates. Furthermore, without deep search, these models suffer from hallucination, proposing moves with high confidence that are strategically fatal. This paper introduces a novel Belief-Guided architecture that disentangles the Policy head from a distinct Belief head. Unlike traditional value functions, the Belief head acts as an internal simulator and independent critic, modeling epistemic uncertainty and strategic stability. By integrating memory mechanisms (Transformer/GRU) to handle long-term dependencies and the Ko rule, and utilizing a gating mechanism to filter overconfident policy errors, our model shifts the burden of intelligence from runtime search to parametric "intuition." Experimental results demonstrate that this approach significantly improves search-free win rates and reduces hallucination, enabling professional-level play on limited hardware where massive MCTS is infeasible.
|
| 2812 |
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
2608.11676
|
cs.AI
|
Wooseong Yang, Wei-Chieh Huang, Weizhi Zhang, Yu Wang, Philip S. Yu |
Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the se...Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in cross-architecture communication: cross-attention bridges that transfer continuous representations across different LLM families suffer from rare-token compression collapse, where entity identity is lost in the continuous bottleneck (bridge-only F1 ~30%). We propose XBRIDGE, a decode-free communication protocol that addresses this through two mechanisms. Lexical Anchor Mapping (LAM) maps the sender's original context tokens to the receiver's vocabulary, providing discrete entity anchors. A Latent Enrichment Bridge (LEB) lets the receiver query the sender's hidden states for contextual enrichment. The entity anchors ground the bridge's contextual signals to specific entities through the receiver's own self-attention. Across three model families (Llama, Qwen, and Mistral), seven benchmarks, and both communication directions, XBRIDGE outperforms text-based communication on all seven tasks for each model pair while achieving 11x lower latency, and in a same-architecture setting it also exceeds a KV-sharing baseline on six of seven tasks. LEB requires only 264M trainable parameters (3.8% of the receiver), is trained on a small balanced sample set, and adds negligible inference overhead.
|
| 2813 |
Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
2608.12385
|
cs.AI
|
Liming Liu, Mingze Wang, Tuo Zhao |
As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typical serving, prompt prefill runs in parallel and is compute-bound, whereas autoregressive decode is sequential and memory-...As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typical serving, prompt prefill runs in parallel and is compute-bound, whereas autoregressive decode is sequential and memory-traffic-bound. Conventional width or depth scaling raises both costs together, since every added layer is evaluated in both phases and enlarges the weights read at each decode step. We instead ask whether additional learned computation can be allocated to continuation prediction while preserving prompt-wide primary computation and a single KV cache. We realize this with the Decode-Branch Transformer. Its primary path alone processes the prompt and writes the KV cache; the decode branch is omitted during prefill and activated only from the final prompt position onward, adding continuation computation without writing state or affecting the primary path. The paths share attention, MLP, and output matrices, using separate token embeddings with lightweight coupling. Grouped decode reuses loaded weight tiles and the primary KV cache across both paths, so the added arithmetic does not proportionally increase dominant memory traffic or decode latency. Across matched-token comparisons, Decode-Branch achieves lower validation loss across architectures and data settings. In MoE models, the primary and branch expert fan-outs become independent knobs for trading prompt cost, decode cost, and predictive quality. We study two expert-allocation regimes, holding prefill or decode computation fixed, and expose a prefill-decode-quality trade-off enabled by phase-specific expert allocation.
|
| 2814 |
Spatial Memory Agent: Experience-Grounded Procedural Memory for Spatial Intelligence
2608.12743
|
cs.AI
|
Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang |
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLMs, existing work has mainly followed two lines. One line uses post-training methods, such as supervis...Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLMs, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning through \textbf{parameter-update-free self-evolution}, without depending on external expert spatial tools at inference time? We present \textbf{Spatial Memory Agent (SMA)}, an experience-grounded runtime memory framework that converts verified spatial experience into reusable transferable lessons. Specifically, SMA first queries the frozen VLM in a verifiable spatial environment, obtains a predicted answer and reward, and uses verifier-guided reflection to distill compact transferable lessons stored in memory cards. SMA further assigns each memory card a \textbf{Transfer Reliability Score (TRS)}, which is initialized uniformly and calibrated from later retrieval outcomes as visit evidence of future transfer reliability. During read-only deployment, SMA retrieves memory cards through semantic filtering and combined similarity--TRS ranking, allowing the retrieved memory to guide frozen model inference. Experiments across five representative spatial benchmarks show that SMA achieves the best macro-average accuracy for all four base VLMs and the best accuracy in most individual evaluations, establishing a practical parameter-update-free path for spatial self-evolution through reusable experience.
|
| 2815 |
From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation
2608.16831
|
cs.AI
|
Minh-Ha Nguyen, Ngoc-Ngo Quang Tran, Thuy Dung Nguyen, Cathy Shyr |
Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-pa...Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures. We formalize Policy Iteration with Human Feedback (PIHF), which makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP. Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools, review and persistent inquiry records to develop reusable task policies. Across general reasoning benchmarks (BIG-Bench Extra Hard, HoVer and LiveBench-Math), PIHF-MCP improved performance of the baseline model by 16.9, 22.2 and 4.7 percentage points, respectively. With a matched baseline model, development used about 1/5 of the labelled examples and 4% of the task rollouts reported by a previous SOTA in-context optimizer, making it about 9 times faster and 3 times cheaper at comparable or higher scores. In a low-data rare-disease diagnosis setting, policies developed from previous SOTA prompt optimizers trailed a previously published PIHF-developed system on every held-out cohort (on average 16 percentage points). These findings support a route to more efficient inference-time scaling: PIHF-MCP develops reusable policies from a few examples that improve performance on unseen cases and across models. Because each policy comes from an explicit, recorded investigation, the process also keeps humans in the loop and enables ownership and learning, making it well suited to high-stakes decisions.
|
| 2816 |
Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured Partitioning
2608.19966
|
cs.AI
|
Zhiquan Huang, Jiazhe Wang, Linjing Xue, Ming Liu, Meiwen Li |
Multivariate time series forecasting (MTSF) is a fundamental task in many real world applications. Existing patch based forecasting methods generally fall into three categories: fixed partitioning, multi-scale partitioning, and extendable partitioning. Fixed p...Multivariate time series forecasting (MTSF) is a fundamental task in many real world applications. Existing patch based forecasting methods generally fall into three categories: fixed partitioning, multi-scale partitioning, and extendable partitioning. Fixed partitioning often breaks meaningful temporal boundaries, multi-scale partitioning may introduce redundant representations across scales, and extendable partitioning improves flexibility but still lacks an explicit mechanism for organizing semantic structure and modeling interactions among heterogeneous temporal patterns. To address these limitations, we propose SCPaT, a Transformer based framework built on semantic structured partitioning. SCPaT first decomposes input sequences into semantically consistent units through adaptive semantic unit generation, then constructs a dynamic semantic graph to model directed dependencies among these units and organize them into higher order semantic blocks. Based on these structured representations, an importance aware routing mechanism adaptively dispatches different semantic blocks to different experts for customized modeling. Extensive experiments on 12 real world datasets demonstrate the effectiveness of SCPaT.
|
| 2817 |
Solving Robust POMDPs with Omega-regular Objectives via Partially Observable Stochastic Games
2608.24986
|
cs.AI
|
Durgam Latha, Dion Reji, S. Akshay, {\DJ}or{\dj}e \v{Z}ikeli\'c, Shankaranarayanan Krishna |
Robust POMDPs (RPOMDPs) generalize classical POMDPs to the setting where exact transition probabilities are not known -- rather, they are only known to belong to some uncertainty set of values. In this work, we study the problem of solving RPOMDPs with general...Robust POMDPs (RPOMDPs) generalize classical POMDPs to the setting where exact transition probabilities are not known -- rather, they are only known to belong to some uncertainty set of values. In this work, we study the problem of solving RPOMDPs with general omega-regular objectives, which subsume a broad class of objectives such as reachability, safety, and linear temporal logic (LTL) objectives. We show that, for (s,a)-rectangular RPOMDPs with polytopic uncertainty sets, the problem of solving RPOMDPs under omega-regular objectives can be reduced to solving partially observable stochastic games (POSGs) under omega-regular objectives. Moreover, we show for the first time that reductions can be constructed in both directions, establishing the semantic equivalence between (s,a)-rectangular RPOMDPs with polytopic uncertainty sets and POSGs. This allows us to derive a range of new computational complexity results, including both upper and lower complexity bounds, on solving RPOMDPs with different omega-regular objectives. As a corollary, we also derive new computational complexity results for RMDPs.
|
| 2818 |
Benevolent Bias in Multi-Turn Human-Agent Dialogue
2608.29206
|
cs.AI
|
Qianqi Liu, Jin Huang, Fethiye Irmak Dogan, Hatice Gunes |
Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, tone and tr...Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, tone and treatment, yielding three classes: neutral support, overt bias, and benevolent bias. Building on these definitions, we construct BENEVDIAL, a class-balanced corpus of 362,880 multi-turn support dialogues spanning user and agent demographics, roles, and generators, to support controlled evaluation. We then test two detector families on it: off-the-shelf safety detectors and prompted large language model (LLM) judges. Our findings reveal a notable detection gap: off-the-shelf detectors reliably flag overt bias yet largely fail to identify benevolent bias. LLM judges improve sensitivity when guided by explicit detection criteria, but this comes at the cost of increased misclassification of neutral supportive statements as benevolent bias, a tendency that is further exacerbated by the presence of demographic context. These findings suggest that fair monitoring of human-agent dialogue must look beyond surface cues to whether the agent's treatment is disparate.
|
| 2819 |
SRPO: Setwise Relative Policy Optimization for Multi-Agent Systems
2609.08452
|
cs.AI
|
Shengtian Yang, Ziyu Xiong, Yu Li, Yewen Li, Qingpeng Cai |
Multi-agent systems enable complex reasoning and tool use by coordinating agents that divide roles and refine candidate solutions. Existing methods typically update individual agent responses or treat a complete trajectory as one training example. However, the...Multi-agent systems enable complex reasoning and tool use by coordinating agents that divide roles and refine candidate solutions. Existing methods typically update individual agent responses or treat a complete trajectory as one training example. However, these methods may produce misleading policy updates because they assign the same final outcome to responses or trajectory segments that may play different roles in different team decisions. This is because treating each response as an independent update may separate outputs that jointly determine the next action, while treating an entire trajectory as one update may combine decisions made after different observations. These limitations call for a policy update defined at the level of a team decision, outputs that lead to the same state transition are optimized under a shared objective. In this paper, we propose Setwise Relative Policy Optimization (SRPO) for multi-agent systems. We represent the outputs used together to produce one state transition as an active set. A singleton set covers division of labor, while a larger set covers joint co-evolution, and the set composition can change across decisions. SRPO assigns a shared advantage to each active set, clips the combined policy change, and normalizes its scale according to the set size. This ties each update to the decision that produced the next state. Experiments on mathematical reasoning and multi-turn search demonstrate the effectiveness of SRPO across both tasks. Additional ablation studies analyze the normalization choice and training behavior under changing active-set sizes.
|
| 2820 |
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
2609.08572
|
cs.AI
|
Jaewon Chu, Jinwoo Seo, Jaewon Cho, Jeehye Na, Yunyang Xiong |
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide p...Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language feedback have emerged as a leading paradigm. In this paper, we identify limitations in two stages of existing textual gradient approaches: gradient extraction and gradient aggregation. In gradient extraction, previous works select a target prompt without verifying whether modifying it resolves the failure, and derive gradients without agent-level supervision over the corresponding agent's intermediate output. In gradient aggregation, individual gradients are randomly grouped and concatenated, often mixing unrelated failure modes and producing prompts that fail to generalize. To address these limitations, we propose AgentGrad, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction. For each failure, sequential intervention modifies the behavior of one agent at a time to identify the target agent whose modification resolves the failure. The modified output of the target agent then serves as agent-level supervision for extracting a fine-grained gradient. Semantic textual gradient abstraction clusters semantically similar gradients to prevent mixing unrelated failure modes, and abstracts each cluster into a generalized gradient that captures the shared corrective pattern. Experimental results show that AgentGrad achieves state-of-the-art performance across five MAS benchmarks while reducing wall-clock optimization time by $2.5\times$ and optimization cost by 21.8\% on average compared to the next-best baselines.
|
| 2821 |
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
2609.13463
|
cs.AI
|
Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru |
The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data r...The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the need for automated root-cause attribution (RCA). However, automated RCA methods using LLMs suffer from low diagnostic accuracy, especially as execution traces grow larger. They struggle because relevant information is often sparse, distributed across distant actions, and disconnected from the visible failure, reducing root-cause attribution to a massive search problem. Existing RCA methods typically rely on one-shot LLM judgments to diagnose failures from execution traces. While effective for shorter trajectories, these judges tend to settle on a plausible diagnosis early, leaving critical evidence in longer traces unexamined. We introduce Continual Search, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence. We evaluate Continual Search across four existing RCA benchmarks. Recognizing the lack of massive execution traces in current benchmarks, we introduce MegaRCA-Mix to evaluate RCA at scale. MegaRCA-Mix provides a challenging testbed of 50 human-annotated failure trials spanning long-horizon, execution-heavy tasks. Across multiple benchmark suites and model families, Continual Search consistently improves attribution performance. On MegaRCA-Mix, for example, it improves Opus-4.8's F1 score by 29\%, from $0.471$ to $0.608$. More interestingly, within the same model family, lower-tier models can even surpass their higher-tier counterparts, demonstrating that effective search supersedes raw model scale.
|
| 2822 |
How Should Reasoning Be Organized in a Transformer's Latent Space?
2609.13747
|
cs.AI
|
Hongyu Gu, Chang Liu, Jingwen Fu |
Continuous reasoning has emerged as a promising way to improve reasoning in large language models (LLMs). Yet we still lack a clear principle for deciding what a latent state should preserve. Reasoning by superposition shows that a single latent state can enco...Continuous reasoning has emerged as a promising way to improve reasoning in large language models (LLMs). Yet we still lack a clear principle for deciding what a latent state should preserve. Reasoning by superposition shows that a single latent state can encode several search alternatives and expand them in parallel. We ask how those states should be weighted as reasoning proceeds. A natural choice is to preserve only the states active at the frontier step, since keeping every reached state appears to spread a limited hidden width too thin. We show that the opposite can hold. When later computation draws on several reached states, a cumulative state can guide attention correctly at a smaller hidden width than a frontier state that stores fewer states. At the same width, the cumulative state therefore keeps more intermediate states available for later reasoning. More generally, equal cumulative weights are optimal when future queries are unknown and remain close to the best task-specific weights when those queries are known. Experiments with two-layer and GPT-2 Transformers reproduce the predicted width advantage and show that unequal weights fail first on the states that receive the least weight. This suggests a important principle: keep reached states equally weighted, and restore equal weights as computation proceeds.
|
| 2823 |
What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
2609.19212
|
cs.AI
|
Chengwen Qi, Deheng Ye, Yatao Bian |
Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on simplifications such as elemen...Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on simplifications such as elemental composition, productivity-based tests, and action-explicit goals, which make systematic generalization easier to study but omit some essential aspects of this capability. To characterize what these simplifications miss, we adopt a reasoning-centered lens and introduce TranSGrid, a testbed that brings deductive, inductive, and abductive reasoning together within a unified task. Experiments with seven Transformer models on 4,800 TranSGrid instances show that all models perform much worse on TranSGrid than on a held-out test set: the largest model solves 79.6% of the test set, but only 55.3% of TranSGrid and 15.8% of the hardest subset. The gap remains within the training length range, showing that productivity alone is not sufficient to evaluate systematic generalization. Additionally, we reintroduce the other two simplifications into TranSGrid: one variant limits interactions among action effects to approximate elemental composition (reducing the inductive demand); the other makes goals action-explicit (reducing the abductive one). In both, solve rates return to roughly the test set level, showing that either simplification alone is enough to reduce TranSGrid to an ordinary held-out test set. Together, our results show that existing tasks reduce either or both of the inductive and abductive demands of systematic generalization, and that comprehensively measuring this capability requires a task that involves all three forms of reasoning.
|
| 2824 |
Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale
2609.22702
|
cs.AI
|
Hong Zhang, John K. Hutchison, Rao Kotamarthi, Jeremy Feinstein, Haiwen Guan |
Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. A central challenge is to produce high-resolution forecasts across continental domains where hydrological behavior varies widel...Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. A central challenge is to produce high-resolution forecasts across continental domains where hydrological behavior varies widely from place to place. We developed Hapi, a U-Net Swin Transformer that uses fine three-dimensional patches and hierarchical shifted-window attention to forecast river discharge, surface runoff, snow water equivalent, and soil wetness index across the contiguous United States. The model produces medium-range forecasts (24--72~h) at $0.05^{\circ}$ resolution with adaptive task weighting and required only 0.11 seconds for a four-variable 72-h CONUS forecast on one A100 GPU. In a held-out 2024 potential-skill evaluation with ERA5-Land inputs prescribed over the forecast horizon, Hapi achieved the highest F1-score for floods in 20 of 21 comparisons across seven GloFAS return periods and three forecast leads. Independent validation against observed daily discharge at 3{,}881 U.S. Geological Survey gauges showed that Hapi achieved the highest median Nash--Sutcliffe efficiency at every lead, supported by regional-cluster bootstrap intervals. In a matched 24-h comparison of loss formulations, adaptive task balancing produced the lowest discharge errors and the highest F1-score for floods.
|
| 2825 |
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
2609.26550
|
cs.AI
|
Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman |
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or esca...LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.
|
| 2826 |
Pretrained ASR Pseudo-labeling for Noisy Police Audio
2609.30469
|
cs.AI
|
Kaavya Chaparala, Su Huang, Stephen L. Morgan, Rhiannon N. Miller, Anjalie Field |
Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this app...Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this approach on very noisy domains is not known. In this work, we systematically assess the opportunities and limits of pseudo-labeling to adapt foundation ASR models (Whisper and Qwen3-ASR) to noisy BPC domain corpora from Baltimore and Chicago. We demonstrate that existing internal confidence metrics (log-probabilities and STAR scores) fail to distinguish between high and low quality BPC pseudo-labels, and we introduce an external LLM-as-a-judge filtering paradigm that leverages parametric knowledge to discard contextually implausible transcripts. Our LLM-judging filters more aggressively than internal metrics and significantly reduces WER of the pseudo-labeled training sets across the Baltimore and Chicago BPC corpora, though a substantial gap remains relative to an oracle filter. We also introduce a new cross-model pseudo-labeling paradigm where one model is finetuned with pseudo-labels from the other, and we identify this method as a promising direction for future pseudo-labeling work.
|
| 2827 |
Audio LLMs Know When They Can't Hear You
2609.30625
|
cs.AI
|
Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik |
Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional t...Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it predicts that its transcription will be reliable. We find that existing approaches, including speech quality predictors, audio LLM generation uncertainty, and transcript-conditioned WER estimation, provide limited signals for detecting transcription failures. In contrast, we discover that transcription reliability is strongly represented in the model's audio-encoder representations. Based on this observation, we devise a lightweight reliability predictor that operates on representations extracted by the frozen audio encoder and predicts the reliability class before generation. The reliability predictor can trigger a clarification request from the user when their voice query is predicted to be unreliable, while allowing reliable queries to proceed without modifying the underlying Audio LLM. Our predictor achieves 81.10% in-domain and 78.09% cross-domain macro-F1 scores, outperforming the strongest baselines by 10.33 and 11.93 points, respectively. Finally, we show that reliability labels can transfer across Audio LLM families, and that transfer performance is closely related to the alignment of their model-specific reliability boundaries.
|
| 2828 |
Cheap, open agents make LLM pollution harder to mitigate
2609.31054
|
cs.AI
|
Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff |
Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source ...Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.
|
| 2829 |
LLM Judge Validation Under Sparse Overlap: From Inference to Design
2609.31857
|
cs.AI
|
Junxuan Li, Arko Mukherjee, Soumyabrata Pal |
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise ...Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
|
| 2830 |
CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies
2609.32192
|
cs.AI
|
Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Xinyu He |
Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary. Workf...Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary. Workflow topology determines which intermediate artifacts are applicable to which downstream workers and when they cease to be valid, thereby providing a structural stress dimension for sharing and isolation. Existing memory benchmarks primarily evaluate retention and retrieval, whereas multi-agent benchmarks emphasize coordination and end-to-end completion, leaving topology-conditioned memory boundaries largely unmeasured. We introduce CoMemBench, an execution-grounded benchmark for collaborative memory sharing and isolation across multi-agent workflow topologies. It constructs 800 composite workflows across four domains from source-grounded dependency graphs, with node-local specifications, verifiable artifact handoffs, native evaluators, and matched isolation challenges. CoMemBench measures workflow completion, verified node progress, required-handoff reliability, isolation robustness, and token cost. Experiments reveal a sharing-isolation trade-off: broader context improves information availability but can weaken isolation, while system rankings shift across topologies and artifact violations.
|
| 2831 |
RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems
2609.32490
|
cs.AI
|
Yuchen Song, Andong Chen, Wenxin Zhu, Muyun Yang, Tiejun Zhao |
LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may onl...LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool use, or execution. We refer to such problems as progressively specified tasks. To systematically study this setting, we introduce ProgSpec, a benchmark that evaluates final outputs against requirements explicitly stated in the initial request and additional requirements supported by the available task evidence. We further propose RepoMAS, an issue-driven multi-agent framework inspired by open-source project management. RepoMAS records newly discovered requirements, conflicts, and failures as structured Issues and uses them to revise the task specification and execution structure during problem solving. Across ProgSpec and five existing benchmarks, RepoMAS achieves the best performance. Further analyses show that its issue-driven revision and repository maintenance mechanisms consistently contribute to performance. These results highlight the importance of allowing MASs to revise not only how a task is solved, but also revise their explicit representation of task requirements during execution.
|
| 2832 |
The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining
2609.32964
|
cs.AI
|
Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia |
Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanism...Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decision to commit or abstain in the first place. We study this decision through mechanistic analysis, framing hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse, causally localised subset of attention heads and MLP sublayers underlying this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions.
|
| 2833 |
Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
2609.32965
|
cs.AI
|
Hongyi Du, Tianyi Zhang, Weijia Zhang, Yi Yang, Haofei Yu |
Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when t...Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding 183 broken benchmark pairs, Relic achieves 371/469 (79.1%), establishing the best reported result among peer-structured systems. On the 47-pair same-model subset, Relic also exceeds Solo (28/47 vs. 26/47), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.
|
| 2834 |
Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring
2609.33123
|
cs.AI
|
Zhixiang Zhang, Zesen Liu, Wai Ip Lai, Hongxu chen, Dongdong She |
Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevoluti...Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify 43 pairwise and 18 irreducible 3-way compositional safety failures. Conventional solution incurs combinatorial complexity in validating cross-component interactions, leaving the safety checking impractical as the harness evolves. To solve this, we introduced a typed hypergraph that represents component states as nodes and safety-relevant higher-order interactions as hyperedges. When the harness changes, the hypergraph updates only the interaction neighborhood of the changed states rather than reconstructing the global composition space. Building on that, we develop a hypergraph-guided runtime monitoring mechanism. Experiments show that our method effectively mitigates compositional safety risks while preserving task utility and reducing interaction-checking costs, and further reveal an empirical safety-utility-cost trade-off across different safety mechanisms.
|
| 2835 |
What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation
2609.33455
|
cs.AI
|
Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo |
On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how ef...On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student's reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and introduces additional supervision at prefix-sensitive positions. Trajectory Dropout consistently improves average performance across teacher--student model pairs of different scales and six mathematical reasoning benchmarks, while also yielding gains on two out-of-domain benchmarks. It can also be flexibly integrated into existing OPD variants with negligible computational overhead, further improving their performance. These results demonstrate that Trajectory Dropout provides a simple mechanism for strengthening token-level supervision across model scales and OPD objectives.
|
| 2836 |
Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents
2609.33772
|
cs.AI
|
Weiyi Xu, Xiaowen Yang, Wen Da, Hang Xu, Canwei Li |
Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, op...Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instructions, but a substantial gap remains between the information contained in a skill and a concrete, challenging task with a complete executable environment. To address this gap, we introduce Skill2Env, a capability-oriented framework that starts from a skill and uses agent capability demands to guide task and environment synthesis. Skill2Env represents these demands through reusable difficulty patterns and instantiates them into task blueprints that specify objectives, challenges, environment facts, information boundaries, and acceptance criteria. These blueprints guide the joint construction of task instructions, execution substrates, workspaces, and rubric-based evaluators around source skills. We further propose Iterative Task Hardening, which uses solver execution evidence to identify insufficiently challenging task designs, strengthen or extend their difficulty-pattern instantiations, and revise the corresponding blueprints and environments. Using 1.5K high-scoring trajectories generated from Skill2Env environments for supervised fine-tuning, we observe consistent improvements across a broad range of agent benchmarks, demonstrating the effectiveness of capability-oriented environment synthesis for agent post-training.
|
| 2837 |
GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions
2609.34113
|
cs.AI
|
Shaoqing Zhang, Kehai Chen, Xuefeng Bai, Zhuosheng Zhang, Pengfei Zhang |
Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI...Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR
|
| 2838 |
Applying Language Models in Clinical Medicine: Recent Trends and Perspectives
2609.34780
|
cs.AI
|
Erik Aerts |
The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to appl...The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While traditional deep learning based AI applications in medicine have often focused on specific and defined tasks, LLMs offer broader capabilities and flexibility in working with available data,. At the same time of writing, the integration of LLMs into medical settings raises important questions regarding their reliability, accuracy, transparency, safety, and appropriate role in a medical setting. This text presents and discusses recent talks and articles concerning the application of LLMs in medicine, with particular emphasis on their potential utility in research and clinical practice. It considers both the opportunities offered by these technologies and the challenges associated with their implementation, aiming to provide a perspective on the current and emerging role of LLMs within the medical field.
|
| 2839 |
STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts
2609.34799
|
cs.AI
|
Wanchun Ni, Tao Qi, Leonel Aguilar, Jiugeng Sun, Marlene Wagner |
Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as c...Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult. We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data. We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation.
|
| 2840 |
AX is the New AEO
2609.34951
|
cs.AI
|
Ido Finder, Assaf Elovic, Gad Shalev, Liad Yosef |
In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, o...In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, or AEO, now tells businesses to scatter breadcrumbs across forum threads, listicles, and off-site citations, so AI engines are likelier to surface and recommend them. But being surfaced is no longer enough: an agent opens the results and reads them before deciding, and one buyer question sends it through several rounds of search and fetch. What decides the outcome at this drill-down step is whether the agent can fetch and read the business's own site: agent experience (AX). We argue that AX is the new AEO. We run 37,927 agent journeys, each a buyer question about a business, across four independent harnesses over 1,056 real businesses, matched on fame, prior model knowledge, and two AEO proxies, then split based on their AX level. Only 7-10% of the finished answer comes from the model's training knowledge, whether or not the site is readable. Agent-ready businesses have answers built from their own pages 78% of the time against 56% and are clearly recommended 1.9x more often, while every grounded answer about a not-agent-ready business costs the agent 64% more. Holding business, harness, and question fixed, answers built from the site are 41% more accurate. The dominant failure is not fabrication but omission: web-built answers are 3.7x more likely to contain none of the facts the buyer asked for. Baselines differ sharply across the four harnesses, with clear-recommendation rates varying sevenfold from stack to stack, yet the effect holds in every one. In the agentic web era, being readable beats being talked about, and improving a site's AX is the strongest lever a business has.
|
| 2841 |
Just Initialize: A Training-Free Initialization Component for Large-Scale Routing Optimization
2609.35443
|
cs.AI
|
Jiale Zhao, Sirui Mao, Zimu Chen, Wentao Yang, Zihan Wang |
Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational cost. We instead shift the focus t...Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational cost. We instead shift the focus to a useful initialization that can be refined into a high-quality solution with limited downstream refinement. We propose Just Initialize, a training-free and solver-agnostic initialization component for large-scale routing optimization. Just Initialize compresses a large routing instance into a compact surrogate space, optimizes its global routing structure, and recovers the resulting solution as an optimization-friendly starting point in the original space. Extensive experiments on Traveling Salesman Problems (TSPs), Capacitated Vehicle Routing Problems (CVRPs), Vehicle Routing Problems with Time Windows (VRPTWs), and Prize-Collecting Traveling Salesman Problems (PCTSPs) demonstrate that Just Initialize achieves high-quality solutions comparable to or better than state-of-the-art methods while substantially reducing computational cost across instances ranging from 1K to 100K nodes, including an average speedup of approximately 70$\times$, sub-second runtimes on 10K-node instances, and runtimes within tens of seconds on 100K-node instances.
|
| 2842 |
RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis
2609.35549
|
cs.AI
|
Bo Zhang, Yuchen Wang, Dongbai Li, Matthew Yu Heng Wong, Qingkai Zeng |
Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor ...Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates, or produce plausible but invalid names. We introduce RareDx, which couples controlled evidence use with knowledge-graph-grounded policy optimization. RareDx-Harness normalizes heterogeneous records into one ranked-diagnosis task and compares direct inference, static retrieval, adaptive tools, and structured phenotype-gene-disease reasoning over a shared knowledge layer. The training pipeline combines Top-10 post-training with RareDx-KGPO, our knowledge-graph-grounded policy optimization method. Its reward projects predictions into a canonical disease graph and integrates curated graded relevance, ontology proximity, biomedical similarity, and phenotype consistency. Vocabulary and output-budget constraints prevent dense partial credit from rewarding fabricated or overlong differentials. Across eight benchmarks, the complete RareDx system centered on Qwen3.5-9B reaches 38.34 macro Hit@10, 1.60 points above GPT-5.5 under the archived protocol; a disjoint validation-selection audit retains a 6.80-point routing gain over Direct on held-out cases. The 27B system reaches 23.53/36.56/40.76 at Hit@1/5/10. Controlled ablations show that retrieval is not uniformly helpful and that controlled routing is central to the gain. These results indicate that structured medical knowledge can turn a compact model into a competitive diagnostic ranker across heterogeneous long-tail settings in clinical practice.
|
| 2843 |
Signatures of semantic search in the activations of large language models
2609.35599
|
cs.AI
|
Luke Leckie, Peter M. Todd, Jacob G. Foster |
When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clu...When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production ("exploit") and between-cluster switching ("explore"). Whether LLMs likewise represent these two search regimes within their internal states is unknown. Here, we apply a range of mechanistic interpretability techniques to provide evidence for this. In Study 1, we use the Jacobian lens (J-lens), which maps intermediate-layer residual-stream representations to token-level activations, to show that concept-level activations predict switching. First, we find that switching coincides with low next-token activations. Moreover, the probability of switching rises as the set of strongest J-lens activations (the J-space) becomes depleted of items from the category currently being produced, analogous to explore-exploit decision-making during patch foraging. We then show that middle-layer J-lens activations of abstract category-related labels (e.g., "water") increase in anticipation of switching into that category. We confirm these representations to causally influence switching by deriving steering vectors that target category switching. In Study 2, we identify generic residual stream directions that are activated during and in anticipation of switching. By steering activations along these directions, we bias increased or decreased rates of switching. Our study extends the semantic foraging framework to artificial intelligences and provides evidence that LLMs maintain distinct representational signatures for exploration and exploitation as they verbalise conceptual information.
|
| 2844 |
Reasoning with Continuous Latent Diffusion
2609.35694
|
cs.AI
|
Xiang Cheng |
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce the Continuous Embedding Diffusion Reasoner (CEDR), an ELF-based training and inference recipe. Our experiments show that accurate decoding a...Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce the Continuous Embedding Diffusion Reasoner (CEDR), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT CEDR-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/CEDR.
|
| 2845 |
On the Approximation and Convergence of Distributional Policy Gradient Algorithms for Risk-Sensitive Reinforcement Learning
2405.14749
|
cs.AI
|
Xian Yu, Minheng Xiao, Lei Ying |
Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distributional RL seeks to estimate the entire dis...Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distributional RL seeks to estimate the entire distribution of it, leading to a unified framework for handling different risk measures. However, developing policy gradient methods for risk-sensitive distributional RL is inherently more complex as it often involves finding the gradient of a probability measure. This paper introduces a new distributional policy gradient framework for risk-sensitive RL, where we derive an analytical gradient of the probability measure of the cumulative cost. For practical implementation, we further design a categorical distributional policy gradient algorithm (CDPG) that approximates arbitrary distributions using a categorical family supported on fixed points. Using Conditional Value-at-Risk (CVaR) as the objective, we prove that the proposed CDPG converges to stationary points and establish its iteration complexities under inexact policy evaluation. Through experiments in a stochastic Cliffwalk environment, we demonstrate the effectiveness of the proposed algorithm and highlight the benefits of incorporating risk sensitivity into distributional RL.
|
| 2846 |
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
2406.00083
|
cs.AI
|
Jiaqi Xue, Mengxin Zheng, Yebowen Hu, Fei Liu, Xun Chen |
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant information from external knowledge bases to provide more accurate, contextually informed, and up-to-date responses. However, this reliance on external knowledge ...Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant information from external knowledge bases to provide more accurate, contextually informed, and up-to-date responses. However, this reliance on external knowledge introduces significant security vulnerabilities, as many RAG systems (e.g., Google Search) rely on large and unsanitized data repositories (e.g., Reddit). In this paper, we unveil a novel threat in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base. When a user's query contains attacker-specified trigger words, the RAG retrieves and refers to these malicious passages, enabling the attacker to steer the response without altering the user input or modifying the RAG weights. BadRAG operates in two phases: (i) malicious passages are optimized to be retrieved exclusively when trigger words appear in user queries; (ii) these passages are meticulously crafted to achieve adversarial generation objectives, including denial of service, sentiment manipulation, context leakage, and tool misuse. Our experiments show that injecting just 10 malicious passages (0.04\% of the external corpora) achieves a 98.2\% retrieval success rate and increases negative response rates from 0.22\% to 72\% for queries containing triggers.
|
| 2847 |
RecKG: Knowledge Graph for Recommender Systems
2501.03598
|
cs.AI
|
Junhyuk Kwon, Seokho Ahn, Young-Duk Seo |
Knowledge graphs have proven successful in integrating heterogeneous data across various domains. However, there remains a noticeable dearth of research on their seamless integration among heterogeneous recommender systems, despite knowledge graph-based recomm...Knowledge graphs have proven successful in integrating heterogeneous data across various domains. However, there remains a noticeable dearth of research on their seamless integration among heterogeneous recommender systems, despite knowledge graph-based recommender systems garnering extensive research attention. This study aims to fill this gap by proposing RecKG, a standardized knowledge graph for recommender systems. RecKG ensures the consistent representation of entities across different datasets, accommodating diverse attribute types for effective data integration. Through a meticulous examination of various recommender system datasets, we select attributes for RecKG, ensuring standardized formatting through consistent naming conventions. By these characteristics, RecKG can seamlessly integrate heterogeneous data sources, enabling the discovery of additional semantic information within the integrated knowledge graph. We apply RecKG to standardize real-world datasets, subsequently developing an application for RecKG using a graph database. Finally, we validate RecKG's achievement in interoperability through a qualitative evaluation between RecKG and other studies.
|
| 2848 |
Causal pieces: analysing and improving spiking neural networks piece by piece
2504.14015
|
cs.AI
|
Dominik Dold, Philipp Christian Petersen |
We introduce "causal pieces", a novel concept for analysing spiking neural networks (SNNs), inspired by "linear pieces" used to study expressivity and trainability in artificial neural networks (ANNs). Causal pieces partition the input and parameter space of a...We introduce "causal pieces", a novel concept for analysing spiking neural networks (SNNs), inspired by "linear pieces" used to study expressivity and trainability in artificial neural networks (ANNs). Causal pieces partition the input and parameter space of a feedforward SNN with single-spike coding into distinct regions where the same subnetwork causes the output spikes. For networks of current-based leaky integrate-and-fire (LIF) neurons with large membrane time constants, we show that within each causal piece, output spike times are locally Lipschitz continuous with respect to inputs and network parameters. We further prove a lower bound on the approximation error that depends on the number of causal pieces. Thus, the number of causal pieces is a measure of the approximation capabilities of SNNs, which is valid despite spike-time discontinuities and applies to networks with both excitatory and inhibitory synapses. Empirically, we find that parameter initialisations yielding more causal pieces on the training set strongly correlate with SNN training success across multiple benchmarks, including Yin-Yang, Fashion-MNIST, and EuroSAT. Moreover, simulations with standard single-spike LIF neurons indicate that our findings extend beyond the theoretically analysed regime. These results establish causal pieces as a powerful and principled tool for analysing and improving the computational capabilities of SNNs.
|
| 2849 |
Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
2507.02735
|
cs.AI
|
Sizhe Chen, Arman Zharmagambetov, David Wagner, Chuan Guo |
Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to AI agents. By fine-tuning on simulated prompt injections, SecAlign, a leading open defense, reports LLMs with go...Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to AI agents. By fine-tuning on simulated prompt injections, SecAlign, a leading open defense, reports LLMs with good test-time robustness and negligible benign utility drop. By scaling up training and evaluations, however, we find that SecAlign actually suffers from significant utility degradation, especially in agentic tasks where the threat of prompt injection is prominent. Motivated by this, we propose Meta-SecAlign for utility-preserving defense by (1) randomized injection position during training to avoid shortcut learning and (2) self-generated responses as high-quality in-distribution training labels. Across general knowledge, instruction following, and agentic workflows (on tool-calling and web-navigation), Meta-SecAlign maintains almost all the undefended LLM's utility while achieving better overall security than SecAlign against various static and GCG adaptive attacks. Experiments use Llama-3.1-8B, Llama-3.3-70B, Llama-4-Scout, Qwen3-4B, and Qwen3.6-27B on 6 prompt injection benchmarks including AgentDojo, InjecAgent, WASP, and SEP. Below are links for the code (https://github.com/facebookresearch/Meta_SecAlign), Meta-SecAlign-70B (https://huggingface.co/facebook/Meta-SecAlign-70B), and Meta-SecAlign-8B (https://huggingface.co/facebook/Meta-SecAlign-8B) models.
|
| 2850 |
Learning Task Mixtures from Task Affinities: A Probabilistic Graphical Model for Supervised Fine-Tuning
2507.12612
|
cs.AI
|
Prateek Chanda, Saral Sureka, Parth Pratim Chatterjee, Krishnateja Killamsetty, Nikhil Shivakumar Nayak |
Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set of tasks. In practice, mixtures are often fixed using simple heuristics (e.g., uniform or size-proportional sampling)...Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set of tasks. In practice, mixtures are often fixed using simple heuristics (e.g., uniform or size-proportional sampling) that ignore task interactions, which can hurt transfer and waste budget on redundant sources. We introduce TaskPGM, a framework for learning continuous task mixtures via an energy-based model over tasks. Tasks form the nodes of a Markov random field: unary potentials capture per-task utility, and pairwise potentials encode inter-task relationships using behavioral divergences computed from predictive distributions of single-task fine-tuned models (e.g., Jensen--Shannon divergence and pointwise mutual information). Optimizing this objective yields mixtures that balance coverage against redundancy. We show that the resulting set function is weakly submodular under budget constraints, enabling approximation guarantees for discrete selection variants. Across multiple model families (LLaMA-7B, Qwen2-7B) and evaluation suites (BIG-Bench Hard), TaskPGM improves over standard mixing strategies and provides interpretable structure over task interactions.
|
| 2851 |
LLM Serving Optimization with Variable Prefill and Decode Lengths
2508.06133
|
cs.AI
|
Meixuan Wang, Yinyu Ye, Zijie Zhou |
We study offline scheduling for large language model (LLM) serving under a fixed KV-cache memory budget, where requests have heterogeneous prompt (prefill) and response (decode) lengths. Given a backlog of requests available at time zero, the scheduler forms m...We study offline scheduling for large language model (LLM) serving under a fixed KV-cache memory budget, where requests have heterogeneous prompt (prefill) and response (decode) lengths. Given a backlog of requests available at time zero, the scheduler forms mixed prefill/decode batches over time to minimize total end-to-end latency. We show that heterogeneity in prompt lengths fundamentally changes the problem: minimizing total latency is NP-hard, and standard policies that prioritize short outputs or small total sequence sizes can have unbounded approximation ratios. We propose Sorted-F, which repeatedly selects feasible batches using an F-metric that balances batch cardinality against downstream decode cost. With exact batch selection, Sorted-F achieves a constant-factor approximation guarantee in the unit-time, uninterrupted-decoding model with known output lengths; the guarantee also holds under a static peak-memory batch constraint. We develop an exact pseudopolynomial dynamic program for this static subproblem, scalable local-search and greedy heuristics, LP-guided variants, and a receding-horizon online extension. Experiments on public conversational and long-document summarization workloads show that F-metric-based scheduling substantially reduces latency relative to standard baselines and remains close to the LP relaxation lower bound on tractable instances.
|
| 2852 |
Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation
2509.20411
|
cs.AI
|
Tharcisse Ndayipfukamiye, Jianguo Ding, Doreen Sebastian Sarwatt, Adamu Gaston Philipo, Huansheng Ning |
Machine learning-based cybersecurity systems are highly vulnerable to adversarial attacks, while Generative Adversarial Networks (GANs) act as both powerful attack enablers and promising defenses. This survey systematically reviews GAN-based adversarial defens...Machine learning-based cybersecurity systems are highly vulnerable to adversarial attacks, while Generative Adversarial Networks (GANs) act as both powerful attack enablers and promising defenses. This survey systematically reviews GAN-based adversarial defenses in cybersecurity (2021--August 31, 2025), consolidating recent progress, identifying gaps, and outlining future directions. Using a PRISMA-compliant systematic literature review protocol, we searched five major digital libraries. From 829 initial records, 185 peer-reviewed studies were retained and synthesized through quantitative trend analysis and thematic taxonomy development. We introduce a four-dimensional taxonomy spanning defensive function, GAN architecture, cybersecurity domain, and adversarial threat model. GANs improve detection accuracy, robustness, and data utility across network intrusion detection, malware analysis, and IoT security. Notable advances include WGAN-GP for stable training, CGANs for targeted synthesis, and hybrid GAN models for improved resilience. Yet, persistent challenges remain such as instability in training, lack of standardized benchmarks, high computational cost, and limited explainability. GAN-based defenses demonstrate strong potential but require advances in stable architectures, benchmarking, transparency, and deployment. We propose a roadmap emphasizing hybrid models, unified evaluation, real-world integration, and defenses against emerging threats such as LLM-driven cyberattacks. This survey establishes the foundation for scalable, trustworthy, and adaptive GAN-powered defenses.
|
| 2853 |
Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
2510.22752
|
cs.AI
|
Anooshka Bajaj, Deven Mahesh Mistry, Sahaj Singh Maini, Yash Aggarwal, Zoran Tiganj |
In-context learning is governed by both temporal and semantic relationships, shaping how Large Language Models (LLMs) retrieve contextual information. Analogous to human episodic memory, where the retrieval of specific events is enabled by separating events th...In-context learning is governed by both temporal and semantic relationships, shaping how Large Language Models (LLMs) retrieve contextual information. Analogous to human episodic memory, where the retrieval of specific events is enabled by separating events that happened at different times, this work probes the ability of various pretrained LLMs, including transformer and state-space models, to differentiate and retrieve temporally separated events. Specifically, we prompted models with sequences containing multiple presentations of the same token, which reappears at the sequence end. By fixing the positions of these repeated tokens and permuting all others, we removed semantic confounds and isolated temporal effects on next-token prediction. Across diverse sequences, models consistently placed the highest probabilities on tokens following a repeated token, but with a notable bias for those nearest the beginning or end of the input. An ablation experiment linked this phenomenon in transformers to induction heads. Extending the analysis to unique semantic contexts with partial overlap further demonstrated that memories embedded in the middle of a prompt are retrieved less reliably. Despite architectural differences, state-space and transformer models showed comparable temporal biases. Our findings deepen the understanding of temporal biases in in-context learning and offer an illustration of how these biases can enable temporal separation and episodic retrieval.
|
| 2854 |
Evidence-Guided Schema Normalization for Temporal Tabular Reasoning
2512.00329
|
cs.AI
|
Ashish Thanga, Vibhu Dixit, Abhilash Shankarampeta, Vivek Gupta |
Temporal reasoning over evolving semi-structured tables poses a challenge to current QA systems. We propose an approach that recasts the task as automated knowledge base construction: (1) prompting an LLM to synthesize a 3NF-compliant relational schema from Wi...Temporal reasoning over evolving semi-structured tables poses a challenge to current QA systems. We propose an approach that recasts the task as automated knowledge base construction: (1) prompting an LLM to synthesize a 3NF-compliant relational schema from Wikipedia infobox timelines, (2) populating the schema to obtain a queryable database, and (3) generating and executing SQL queries against it, with QA accuracy serving as an extrinsic evaluation of the constructed knowledge base. In a controlled grid of three schema generators crossed with six query models, the schema source accounts for 79.5% of the exact match (EM) variance against 1.6% for the query model: replacing the schema, and the prompt scaffolding derived from it, shifts EM by 14.7 to 20.0 points, whereas replacing the query model under a fixed schema shifts it by 4.4 to 12.1. From this evidence, we distill three candidate schema-design principles: balanced normalization, semantic naming, and consistent temporal anchoring, framed as correlational hypotheses. Our best configuration (Gemini 2.5 Flash schemas + Gemini-2.0-Flash queries) reaches 80.39 EM, 11.5 points above the strongest reported baseline (68.89 EM); an open-weights configuration reaches 79.52.
|
| 2855 |
Data-Free Pruning of Self-Attention Layers in LLMs
2512.20636
|
cs.AI
|
Dhananjay Saikumar, Blesson Varghese |
Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss. We attribute this to the Attention Suppression Hypothesis: during pre-training, some deep attention layers learn to mute their own contribution, leaving the re...Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss. We attribute this to the Attention Suppression Hypothesis: during pre-training, some deep attention layers learn to mute their own contribution, leaving the residual stream and the MLP to carry the representation. We propose Gate-Norm, a one-shot, weight-only criterion that ranks attention sublayers by query-key coupling and removes the least coupled ones, requiring no calibration data, no forward passes, no fine-tuning, and no specialized kernels. On 40-layer, 13B-parameter LLaMA models, Gate-Norm prunes the model in under a second. Pruning 8-16 attention sublayers yields up to $1.30\times$ higher inference throughput while keeping average zero-shot accuracy within 1.5 percentage points of the unpruned baseline across BoolQ, RTE, HellaSwag, WinoGrande, ARC-Easy/Challenge, and OpenBookQA. Across these settings, Gate-Norm matches data-driven pruning methods in accuracy while being $\sim 1000\times$ faster to score layers, enabling practical, data-free compression of LLMs.
|
| 2856 |
Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
2601.08777
|
cs.AI
|
Yang Cai, Weiqiang Zheng |
Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for personalized and trustworthy AI. We formalize an ideal notion of universal alignment through test-time scaling: for each ...Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for personalized and trustworthy AI. We formalize an ideal notion of universal alignment through test-time scaling: for each prompt, the model produces $k\ge 1$ candidate responses and a user selects their preferred one. We introduce $(k,f(k))$-robust alignment, which requires the $k$-output model to have win rate $f(k)$ against any other single-output model, and asymptotic universal alignment (U-alignment), which requires $f(k)\to 1$ as $k\to\infty$. Our main result characterizes the optimal convergence rate: there exists a family of single-output policies whose $k$-sample product policies achieve U-alignment at rate $f(k)=\frac{k}{k+1}$, and no method can achieve a faster rate in general. We show that popular post-training methods, including Nash learning from human feedback (NLHF), can fundamentally underutilize the benefits of test-time scaling. Even though NLHF is optimal for $k=1$, sampling from the resulting (often deterministic) policy cannot guarantee win rates above $\tfrac{1}{2}$ except for an arbitrarily small slack. This stems from a lack of output diversity: existing alignment methods can collapse to a single majority-preferred response, making additional samples redundant. In contrast, our approach preserves output diversity and achieves the optimal test-time scaling rate. In particular, we propose a family of symmetric multi-player alignment games and prove that any symmetric Nash equilibrium policy of the $(k+1)$-player alignment game achieves the optimal $(k,\frac{k}{k+1})$-robust alignment. Finally, we provide theoretical convergence guarantees for self-play learning dynamics in these games and extend the framework to opponents that also generate multiple responses.
|
| 2857 |
Uncertainty-aware Causal Decision Making via Effect Bound Decomposition
2601.22736
|
cs.AI
|
Md Musfiqur Rahman, Ziwei Jiang, Hilaf Hasson, Murat Kocaoglu |
Causal inference from observational data can provide strong evidence for finding the best action in a decision-making scenario without having to perform expensive randomized trials. The causal effect of an action is often not pointwise identifiable even with i...Causal inference from observational data can provide strong evidence for finding the best action in a decision-making scenario without having to perform expensive randomized trials. The causal effect of an action is often not pointwise identifiable even with infinite data due to unobserved confounding factors. Furthermore, having only finitely many samples adds another layer of uncertainty to causal effect estimation. Several existing methods can be used to obtain upper and lower bounds to the causal effect, ranging from symbolic methods to the more recent neural network-based approaches, which implicitly incorporate both sources of uncertainty. However, these methods do not inform whether collecting more samples may or may not help identify the best action from observational data, leaving experts in the dark about their data collection strategies. We address this problem with a novel framework that can distinguish the range of causal effect values that might be eliminated by collecting more samples from the range of values that, with high probability, cannot be eliminated with more observational samples. We show that this partitioning can be obtained by solving max-min and min-max optimization problems. We leverage neural causal models to approximately recover this decomposition in practice. We demonstrate via experiments on synthetic and real-world datasets that our algorithm can determine when collecting more samples will not help determine the best action. Our framework can help practitioners decide when to resort to non-observational studies or seek to measure some of the unmeasured confounders for optimal decision-making.
|
| 2858 |
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
2602.11988
|
cs.AI
|
Thibaud Gloaguen, Niels M\"undler-Sasahara, Mark Niklas M\"uller, Veselin Raychev, Martin Vechev |
A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although this practice is strongly encouraged by agent developers, there is currently no rigorous investigation into whether such c...A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although this practice is strongly encouraged by agent developers, there is currently no rigorous investigation into whether such context files are actually effective for real-world tasks. In this work, we study this question and evaluate coding agents' task completion performance in two complementary settings: established SWE-bench tasks from popular repositories, with LLM-generated context files, and a novel collection of issues from repositories containing developer-committed context files. Surprisingly, we find that providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average. This observation holds across different LLMs, coding agents, and for both LLM-generated and developer-committed context files. Specifically, we find that while instructions in the context files are well followed by coding agents, repository overviews, although popular and recommended by model providers, are not helpful. We conclude that while context files are useful for specifying non-standard coding practices, any attempts to improve performance should be rigorously evaluated before deployment.
|
| 2859 |
On Calibration of Large Language Models: From Response To Capability
2602.13540
|
cs.AI
|
Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee |
Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output. However, this formulation is ...Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output. However, this formulation is misaligned with many practical settings where the central question is how likely a model is to solve a query overall. We show that this mismatch results from the stochastic nature of modern LLM decoding, under which single-response correctness fails to reflect underlying model capability. To address this issue, we introduce capability calibration, a new evaluation framework for measuring how well query-level confidence aligns with a model's expected accuracy on individual queries. We formally distinguish capability calibration (CC) from response calibration (RC) and show that the two differ both theoretically and empirically. We further show that CC is better suited than RC to applications like pass@k prediction and inference budget allocation. Finally, we evaluate common confidence estimation methods to understand the practical feasibility of CC.
|
| 2860 |
Robo-Saber: Generating and Simulating Virtual Reality Players
2602.18319
|
cs.AI
|
Nam Hee Kim, Jingjing May Liu, Jaakko Lehtinen, Perttu H\"am\"al\"ainen, James F. O'Brien |
We present the first motion generation system for playtesting virtual reality (VR) games. Our player model generates VR headset and handheld controller movements from in-game object arrangements, guided by style exemplars and aligned to maximize simulated game...We present the first motion generation system for playtesting virtual reality (VR) games. Our player model generates VR headset and handheld controller movements from in-game object arrangements, guided by style exemplars and aligned to maximize simulated gameplay score. We train on the large BOXRR-23 dataset and apply our framework on the popular VR game Beat Saber. The resulting model Robo-Saber produces skilled gameplay and captures diverse player behaviors, mirroring the skill levels and movement patterns specified by input style exemplars. Robo-Saber demonstrates promise in synthesizing rich gameplay data for predictive applications and enabling a physics-based whole-body VR playtesting agent.
|
| 2861 |
Agentic AI for Scalable and Robust Optical Systems Control
2602.20144
|
cs.AI
|
Zehao Wang, Mingzhe Han, Wei Cheng, Yue-Kai Huang, Philip Ji |
We present AgentOptics, an agentic AI framework for high-fidelity, autonomous optical system control built on the Model Context Protocol (MCP). AgentOptics interprets natural language tasks and executes protocol-compliant actions on heterogeneous optical devic...We present AgentOptics, an agentic AI framework for high-fidelity, autonomous optical system control built on the Model Context Protocol (MCP). AgentOptics interprets natural language tasks and executes protocol-compliant actions on heterogeneous optical devices through a structured tool abstraction layer. We implement 64 standardized MCP tools across 8 representative optical devices and construct a 410-task benchmark to evaluate request understanding, role-aware responses, multi-step coordination, robustness to linguistic variation, and error handling. We assess two deployment configurations--commercial online LLMs and locally hosted open-source LLMs--and compare them with LLM-based code generation baselines. AgentOptics achieves 87.7%--99.0% average task success rates, significantly outperforming code-generation approaches, which reach up to 50% success. We further demonstrate broader applicability through five case studies extending beyond device-level control to system orchestration, monitoring, and closed-loop optimization. These include DWDM link provisioning and coordinated monitoring of coherent 400 GbE and analog radio-over-fiber (ARoF) channels; autonomous characterization and bias optimization of a wideband ARoF link carrying 5G fronthaul traffic; multi-span channel provisioning with launch power optimization; closed-loop fiber polarization stabilization; and distributed acoustic sensing (DAS)-based fiber monitoring with LLM-assisted event detection. These results establish AgentOptics as a scalable, robust paradigm for autonomous control and orchestration of heterogeneous optical systems.
|
| 2862 |
ContactExplorer: Contact Coverage-Guided Exploration for General-Purpose Dexterous Manipulation
2603.10971
|
cs.AI
|
Zixuan Liu, Ruoyi Qiao, Chenrui Tie, Xuanwei Liu, Yunfan Lou |
Reinforcement learning explores effectively in domains such as Atari games, navigation, and locomotion, where novelty over states or dynamics is a sufficient signal. In contrast, dexterous manipulation requires rich physical hand--object interactions, but exis...Reinforcement learning explores effectively in domains such as Atari games, navigation, and locomotion, where novelty over states or dynamics is a sufficient signal. In contrast, dexterous manipulation requires rich physical hand--object interactions, but existing methods often suffer from unstable contact-based novelty signals, inefficient distance novelty signals, or reliance on task-specific priors. We propose ContactExplorer, a general exploration method for dexterous manipulation tasks. ContactExplorer represents contact as the intersection between object surface points and hand keypoints, encouraging dexterous hands to discover diverse and novel contact patterns, namely which fingers contact which object regions. It maintains a contact counter conditioned on discretized object states obtained via learned hash codes. This counter is leveraged in two complementary ways: (1) a count-based contact coverage reward that promotes exploration of novel contact patterns, and (2) an energy-based reaching reward that guides the agent toward under-explored contact regions. We evaluate ContactExplorer on seven contact-rich manipulation tasks and five dexterous hand embodiments. Experimental results show that ContactExplorer substantially improves sample efficiency and success rates over existing exploration methods, that it reduces the need for task-specific priors, and that it remains effective across hand embodiments and transfers to the real world. Project page is https://contact-explorer.github.io.
|
| 2863 |
Altered Thoughts, Altered Actions: Reasoning Chain as Control Surface for a Vision-Language-Action Policy
2603.12717
|
cs.AI
|
Tuan Duong Trinh, Basim Azam, Mohammed Ishaq Ansari, Mohammed Yaqoob Ansari, Naveed Akhtar |
Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies are designed to reason in text before acting, generating a reasoning chain and decoding actions conditioned on that chain. Th...Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies are designed to reason in text before acting, generating a reasoning chain and decoding actions conditioned on that chain. The works introducing this design offer the reasoning chain as an oversight interface: text a person can read and edit to correct the policy. What an edited reasoning chain does to the policy's motor actions, whether it repairs them or corrupts them, remains an open question. We measure both directions, repair and corruption, with our deterministic entity swap applied to the instruction the policy receives and to the reasoning chain it generates. A forty-task observed backdrop across all four LIBERO simulation suites reveals that the cost of corrupting the reasoning chain concentrates where language alone determines the goal. There, on LIBERO-Goal, we run the decisive counterfactual intervention with DeepThinkVLA, chosen because its reasoning chain is exposed as plain text. The policy receives a corrupted instruction, paired with the reasoning chain it generates when that instruction is clean. This counterfactually correct reasoning chain recovers 47.8 pp of the lost success, our pre-registered confirmatory test. Had the chain merely restated what the camera image already determines, the injection could have changed nothing. Instead, all 10 tasks move in the predicted direction. The reasoning chain is therefore a working control surface: text written into it steers the robot, repairing behaviour when the text is right and corrupting it when the text is wrong. Whether to expose such a control surface is a real deployment tradeoff, and it can now be measured.
|
| 2864 |
TabKD: Tabular Knowledge Distillation through Interaction Diversity of Learned Feature Bins
2603.15481
|
cs.AI
|
Shovon Niverd Pereira, Krishna Khadka, Yu Lei |
Data-free knowledge distillation enables model compression without original training data, critical for privacy-sensitive tabular domains. However, existing methods does not perform well on tabular data because they do not explicitly address feature interactio...Data-free knowledge distillation enables model compression without original training data, critical for privacy-sensitive tabular domains. However, existing methods does not perform well on tabular data because they do not explicitly address feature interactions, the fundamental way tabular models encode predictive knowledge. We identify interaction diversity, systematic coverage of feature combinations, as an essential requirement for effective tabular distillation. To operationalize this insight, we propose TabKD, which learns adaptive feature bins aligned with teacher decision boundaries, then generates synthetic queries that maximize pairwise interaction coverage. Across 4 benchmark datasets and 4 teacher architectures, TabKD achieves highest student-teacher agreement in 14 out of 16 configurations, outperforming 5 state-of-the-art baselines. We further show that interaction coverage strongly correlates with distillation quality, validating our core hypothesis. Our work establishes interaction-focused exploration as a principled framework for tabular model extraction.
|
| 2865 |
Pure and physics-guided deep learning approaches for spatio-temporal groundwater level prediction
2603.25779
|
cs.AI
|
Matteo Salis, Gabriele Sartor, Rosa Meo, Stefano Ferraris, Abdourrahmane M. Atto |
Groundwater represents a key element of the water cycle, yet it exhibits complex and context-dependent relationships that make its modeling challenging. Theory-based models have been the cornerstone of scientific understanding. However, their computational cos...Groundwater represents a key element of the water cycle, yet it exhibits complex and context-dependent relationships that make its modeling challenging. Theory-based models have been the cornerstone of scientific understanding. However, their computational cost, simplifying assumptions, and calibration requirements limit their use. In recent years, data-driven models have emerged as powerful alternatives. In particular, deep learning has proven to be a promising approach for its design flexibility and ability to learn complex relationships directly from the data without requiring extensive domain information. We proposed an attention-based pure deep learning model, named STAINet, to predict weekly groundwater levels in Piedmont (Italy), leveraging both irregular groundwater time series and weather image sequences. To enhance the model's trustworthiness and generalization ability, we merged the theory and data-driven approaches by considering physics-guided strategies to inject the groundwater flow equation into the model. Firstly, we restructured the tail of the architecture to predict the three terms of the governing equation, named the autoregressive, diffusion, and residual components - we thus obtained the PSTAINet-IB. Then, we further injected physics priors by adding loss terms related to the estimated equation components, obtaining the PSTAINet-ILB model. Lastly, we developed the PSTAINet-ILRB by imposing a loss term specific to the residual component, which forces the groundwater recharge to occur within the groundwater body recharge zone, which is identified by domain experts. The models were evaluated both by feeding true lagged values as input and by iterating their own predictions (rollouts) over the whole test set. The PSTAINet-ILB model performed the best, achieving remarkable test performance, and generating equation components in line with domain experts' expectations.
|
| 2866 |
Screening Is Enough
2604.01178
|
cs.AI
|
Ken M. Nakanishi |
We call query--key relevance absolute when its values lie on a fixed bounded scale, depend on neither competing keys nor sequence length, require no sequence-length-dependent calibration, and can all be zero. To realize this notion, we introduce screening, who...We call query--key relevance absolute when its values lie on a fixed bounded scale, depend on neither competing keys nor sequence length, require no sequence-length-dependent calibration, and can all be zero. To realize this notion, we introduce screening, whose explicit threshold transforms bounded query--key similarities into relevance values, enabling exact rejection, empty selection, and direct inspection on a common scale. In a controlled comparison of 12 attention mechanisms on a matched Transformer backbone, only screening maintains both low long-context perplexity and robust retrieval beyond the training context; notably, it does so without inference-time scaling. Building on screening, we introduce Multiscreen, a language-model architecture composed of parallel gated screening tiles. Multiscreen retains these long-context gains while achieving greater parameter efficiency, stronger general zero-shot downstream performance, lower training cost at larger scales, and lower model-side time to first token than Transformer baselines. We further develop a normalization design that keeps Multiscreen training stable even at a learning rate of $1$ and show that an adapted version likewise stabilizes Transformer at the same learning rate.
|
| 2867 |
Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective
2604.08880
|
cs.AI
|
Tokio Kajitsuka, Ukyo Honda, Sho Takase |
Chain-of-thought (CoT) distillation transfers reasoning behaviors from a strong teacher to a smaller student, but prior work reports a capacity gap: distillation may fail when the teacher-student capability mismatch is large. We revisit the capacity gap from a...Chain-of-thought (CoT) distillation transfers reasoning behaviors from a strong teacher to a smaller student, but prior work reports a capacity gap: distillation may fail when the teacher-student capability mismatch is large. We revisit the capacity gap from a practical perspective by re-examining commonly used experimental settings. Notably, we find that CoT distillation often degrades performance compared to the student's pre-distillation baseline, and that some settings used in prior work, while suitable for establishing the capacity gap as a phenomenon, do not reflect realistic deployment scenarios. Complementing prior work that establishes the capacity gap, we evaluate its practical impact under more realistic settings and find that it does not consistently dominate; stronger teachers tend to be preferable when candidate teachers differ substantially in performance. Our results offer practical guidance for selecting teacher-student pairs in CoT distillation.
|
| 2868 |
Query Lower Bounds for Diffusion Sampling
2604.10857
|
cs.AI
|
Zhiyang Xun, Eric Price |
Diffusion models generate samples by iteratively querying learned score estimates. A rapidly growing literature focuses on accelerating sampling by minimizing the number of score evaluations, yet the information-theoretic limits of such acceleration remain unc...Diffusion models generate samples by iteratively querying learned score estimates. A rapidly growing literature focuses on accelerating sampling by minimizing the number of score evaluations, yet the information-theoretic limits of such acceleration remain unclear. In this work, we establish the first score query lower bounds for diffusion sampling. We prove that for $d$-dimensional distributions, given access to score estimates with polynomial accuracy $\varepsilon=d^{-O(1)}$ (in any $L^p$ sense), any sampling algorithm requires $\widetilde{\Omega}(\sqrt{d})$ adaptive score queries. In particular, our proof shows that, within any polynomial total-query budget, successful sampling requires searching over $\widetilde{\Omega}(\sqrt{d})$ distinct noise levels, providing a formal explanation for why multiscale noise schedules are necessary in practice.
|
| 2869 |
Collective Opinion Dynamics in Structured LLM Populations
2604.11312
|
cs.AI
|
Erica Cau, Andrea Failla, Giulio Rossetti |
Large Language Models are increasingly deployed as interacting agents in settings such as online platforms, recommendation systems, and multi-agent applications. Understanding the collective behaviors that emerge from their interactions is therefore increasing...Large Language Models are increasingly deployed as interacting agents in settings such as online platforms, recommendation systems, and multi-agent applications. Understanding the collective behaviors that emerge from their interactions is therefore increasingly crucial, especially as these behaviors may shape public opinion and contribute to polarization. In this work, we investigate how network structure and group composition shape the evolution of opinions in populations of LLM agents engaged in multi-round debates. We generate networks with controlled levels of homophily and varying group sizes, and perform ten independent runs per configuration. The results show that LLM agents exhibit patterns that are highly sensitive to network structure, relative group sizes, and to the model itself. We highlight that even within LLM populations, low homophily facilitates convergence by increasing opportunities for cross-group interaction, whereas higher homophily limits cross-group exposure and can preserve distinct opinion states, consistent with previous research. We further show that the choice of LLM substantially affects the resulting dynamics, with different models exhibiting different patterns of opinion updating under comparable network conditions. At the individual level, providing agents with information about their local neighborhood further modulates opinion transitions, revealing substantial differences between models in their sensitivity to local social context. Overall, our findings show that the collective dynamics of LLM populations arise from interactions among model-specific behavior, network structure, population composition, and local social information.
|
| 2870 |
PAC-CF: Calibrating Irreversible Frontier Pruning in LLM-Guided Search
2604.14345
|
cs.AI
|
Tianhao Qian, Jiayu Chen, Zhenyu Sun, Lixu Wang |
LLM-guided search explores multiple candidate trajectories, but at substantial test-time cost. Pruning low-scoring frontier candidates can control this cost, yet it also turns potentially biased evaluator scores into irreversible decisions: systematic ranking ...LLM-guided search explores multiple candidate trajectories, but at substantial test-time cost. Pruning low-scoring frontier candidates can control this cost, yet it also turns potentially biased evaluator scores into irreversible decisions: systematic ranking errors can persist under repeated scoring and remove useful branches. We propose Probably Approximately Correct Conformal Filtering (PAC-CF). Its fixed-frontier analysis formulates elimination as an $(\varepsilon,\delta)$-PAC problem under bounded evaluator bias; its operational rule separately calibrates a score-gap threshold on held-out tasks by running the original controller without PAC-CF and using post-search verifier labels to measure the deficit of solution-preserving candidates relative to the frontier leader. Conditional on exchangeable native-controller tasks with nonempty protected exposure, conformal calibration gives finite-sample coverage for retaining at least one verifier-defined valid continuation at every protected frontier on the native trajectory. At deployment, PAC-CF removes only candidates whose gap from the highest frontier score exceeds the frozen threshold. We evaluate PAC-CF across three domains, five controllers, and four request budgets from B100 to B500. In the cross-domain/controller macro averages, the point estimates for all three workload measures are lower at every budget; the paired-bootstrap 95\% confidence interval for utility excludes zero at B100 and B200. For pruning-aware ToolTree, the full-test-set cross-domain utility difference is $+4.38$ points at each tested budget; on the natural-termination sensitivity cohort, physical requests decrease by $18.94$--$18.95\%$ and end-to-end token usage by $23.57$--$23.76\%$.
|
| 2871 |
Amortized Optimal Transport from Sliced Potentials
2604.15114
|
cs.AI
|
Minh-Phuc Truong, Khai Nguyen |
We propose a novel amortized optimization method for predicting optimal transport (OT) plans across multiple pairs of measures by leveraging Kantorovich potentials derived from sliced OT. We introduce two amortization strategies: regression-based amortization ...We propose a novel amortized optimization method for predicting optimal transport (OT) plans across multiple pairs of measures by leveraging Kantorovich potentials derived from sliced OT. We introduce two amortization strategies: regression-based amortization (RA-OT) and objective-based amortization (OA-OT). In RA-OT, we formulate a functional regression model that treats Kantorovich potentials from the original OT problem as responses and those obtained from sliced OT as predictors, and estimate these models via least-squares methods. In OA-OT, we estimate the parameters of the functional model by optimizing the Kantorovich dual objective. In both approaches, the predicted OT plan is subsequently recovered from the estimated potentials. As amortized OT methods, both RA-OT and OA-OT enable efficient solutions to repeated OT problems across different measure pairs by reusing information learned from prior instances to rapidly approximate new solutions. Moreover, by exploiting the structure provided by sliced OT, the proposed models are more parsimonious, independent of specific structures of the measures, such as the number of atoms in the discrete case, while achieving high accuracy. We demonstrate the effectiveness of our approaches on tasks including MNIST digit transport, color transfer, supply-demand transportation on spherical data, and mini-batch OT conditional flow matching.
|
| 2872 |
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
2604.15186
|
cs.AI
|
Otto White, Marcel Wagenl\"ander, Britannio Jarrett, Xijin Zhou, Yanda Tao |
Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpr...Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic workflows onto a GPU cluster. Scepsy exploits the insight that, while the end-to-end latency of an agentic workflow is unpredictable, each LLM's fraction of execution time is comparatively stable across requests. Scepsy profiles each LLM under different parallelism degrees and combines the profiles with these fractions into an Aggregate LLM Pipeline, a lightweight throughput and latency predictor for allocations. To minimize latency at a target throughput, Scepsy uses the Aggregate LLM Pipeline to search over fractional GPU shares, tensor parallelism degrees, and replica counts. A hierarchical heuristic then places the chosen allocation onto the cluster, minimizing fragmentation and respecting network topology. On realistic agentic workflows, Scepsy achieves up to 2.5x higher throughput before saturation and 1.0-3.3x lower latency than systems that optimize LLMs independently or rely on user-specified allocations.
|
| 2873 |
Replay-buffer engineering for noise-aware quantum circuit optimization
2604.21863
|
cs.AI
|
Akash Kundu, Sebastian Feld |
Deep reinforcement learning for quantum circuit optimization faces three bottlenecks: replay buffers that overlook temporal difference (TD) target reliability, curriculum-based architecture search requiring a full quantum-classical evaluation after every edit,...Deep reinforcement learning for quantum circuit optimization faces three bottlenecks: replay buffers that overlook temporal difference (TD) target reliability, curriculum-based architecture search requiring a full quantum-classical evaluation after every edit, and the discard of noiseless trajectories when retraining under hardware noise. We address these limitations by treating replay as a central algorithmic lever. We introduce ReaPER+, an annealed replay rule that transitions from TD-error prioritization to reliability-aware sampling as value estimates mature. ReaPER+ achieves up to 4x higher sample efficiency than fixed PER, ReaPER, and uniform replay, while matching prior on-policy solution quality with up to 32x fewer interactions At 12 qubits, fixed ReaPER reaches the lowest energy error in the fewest steps, while PER and uniform replay find more compact circuits at higher error. On tasks scaling to 20 qubits, ReaPER+ retains its advantage, demonstrating that reliability-aware annealing extends beyond small-system benchmarks. LunarLander-v3 confirms that the ReaPER+ is domain-agnostic, it improves success rates by up to 26.8% over PER and 21.8% over fixed ReaPER, with a 3% AUC gain over both. We further introduce OptCRLQAS, which amortizes quantum-classical evaluations across multiple architectural edits, reducing training wall-clock time by up to 67.5% on 12-qubit without degrading solution quality. Finally, lightweight replay-buffer transfer warm-starts noisy optimization from noiseless trajectories, without weight transfer or $\epsilon$-greedy pretraining, reducing steps to chemical accuracy by 85-90% and final energy error by up to 90% relative to from-scratch learning. Transfer gains increase with system size. Together, these results establish experience storage, sampling, and transfer as decisive levers for sample efficient, noise-aware quantum circuit optimization.
|
| 2874 |
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
2604.22893
|
cs.AI
|
Minghui Xu, Qi Luo, Kun Li, Zhengyang Shan |
Traditional ``row-count $\times$ quality coefficient'' approaches fail to capture the nonlinear utility of data for Large Language Model (LLM) capabilities. We present a utility-aware data valuation framework that moves from static accounting toward utility-ba...Traditional ``row-count $\times$ quality coefficient'' approaches fail to capture the nonlinear utility of data for Large Language Model (LLM) capabilities. We present a utility-aware data valuation framework that moves from static accounting toward utility-based pricing. The framework combines token-level information density and data quality, empirical training contribution estimated by influence functions, proxy models, and Data Shapley values, and cryptographic mechanisms based on hash commitments, Merkle trees, and tamper-evident training ledgers. Preliminary evaluation includes a smoke-scale multi-domain study and a synthetic duplication test. Results show that proxy-based empirical gain achieves strong ranking agreement with proxy-defined realized utility and substantially outperforms row-count and token-count baselines in the tested setting, while the prototype provides a controlled sanity check of the valuation pipeline and its sensitivity to duplication-based manipulation.
|
| 2875 |
Understanding Scattered Forest Search: A Version-Space Perspective on Multi-Turn Program Correction
2604.23989
|
cs.AI
|
Yuto Tanaka, Issei Sato |
In multi-turn program correction, the state-of-the-art method Scattered Forest Search (SFS) has been proposed, employing Monte Carlo Tree Search (MCTS) with carefully crafted initial seeds and text-based optimization. However, since SFS integrates multiple com...In multi-turn program correction, the state-of-the-art method Scattered Forest Search (SFS) has been proposed, employing Monte Carlo Tree Search (MCTS) with carefully crafted initial seeds and text-based optimization. However, since SFS integrates multiple components, the effects of each component on performance and the overall behavior of SFS have not been sufficiently analyzed. In this work, we theoretically analyze the refinement process of SFS from the perspective of version spaces in learning theory and clarify its behavior. First, as a basis for the theoretical analysis, we introduce a sequential self-refinement method (Line), which starts from an initial program and repeatedly refines the resulting program. Furthermore, while Line progresses the refinement process in the depth direction, we introduce Iterative Refinement of Repair Instructions (IRRI) to capture the refinement process in the width direction, which fixes initial programs and iteratively refines repair instructions. We then analyze Line and IRRI and conduct a theoretical analysis of SFS by positioning it as an intermediate method between the two. Our theoretical analysis reveals that SFS exhibits behavior closer to IRRI than to Line, and this theoretical characteristic is also confirmed experimentally. Considering computational resources and methodological simplicity, these results suggest that a simpler method, IRRI, may achieve a similar refinement process without relying on the complex correction mechanism of SFS.
|
| 2876 |
Single-turn emergency psychiatric triage across 15 frontier AI chatbots
2604.25415
|
cs.AI
|
Veith Weilnhammer, Lennart Luettgau, Christopher Summerfield, Raymond Dolan, Elise Wilkinson |
People increasingly turn to general-purpose AI chatbots for advice about emotional and mental health problems, but the ability of these systems to recognize and appropriately triage psychiatric emergencies remains under-characterized. We evaluated psychiatric ...People increasingly turn to general-purpose AI chatbots for advice about emotional and mental health problems, but the ability of these systems to recognize and appropriately triage psychiatric emergencies remains under-characterized. We evaluated psychiatric triage performance in 15 frontier AI chatbots using 112 clinical vignettes spanning four urgency levels, from routine care to immediate emergency assessment. In each trial (1680 total), a chatbot received a single user message conveying all triage-relevant information from one vignette and recommended a timeframe for care. The primary outcome was emergency under-triage; secondary outcomes included triage accuracy and the direction of errors. Vignettes and user messages were generated using a clinician-verified LLM pipeline. Across 415 emergency trials, 23 were under-triaged (5.5%; 95% CI 1.8-15.9). Overall accuracy, averaged across urgency levels, ranged from 42.0% to 71.8% across chatbots and was lowest for intermediate cases (19.6%; 95% CI 11.7-28.1). Every chatbot showed a net over-triage bias; overall, 763 of 786 incorrect assignments (97.1%) were more urgent than the prespecified triage level. The error pattern was similar when predictions were assessed against clinician ratings: 35 of 430 trials involving vignettes rated as emergencies by at least 75% of clinicians were under-triaged (8.1%). AI chatbots recognized most psychiatric emergencies but still missed clinically important cases and frequently over-triaged less urgent presentations. Further evaluations should examine how triage performance changes when clinically relevant information must be elicited through conversation.
|
| 2877 |
GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory
2605.01688
|
cs.AI
|
Yushi Sun, Bowen Cao, Dong Fang, Lingfeng Su, Wai Lam |
Long-horizon memory systems increasingly improve how evidence is stored and retrieved, yet the generator must still reason over fragments whose cross-session relationships are implicit. We study generation-time memory organization as a distinct design dimensio...Long-horizon memory systems increasingly improve how evidence is stored and retrieved, yet the generator must still reason over fragments whose cross-session relationships are implicit. We study generation-time memory organization as a distinct design dimension and introduce GRAVITY (Generation-time Relational Anchoring Via Injected Topological MemorY), a host-independent auxiliary memory layer. GRAVITY consolidates raw dialogue into entity profiles, temporal event traces, and cross-session topic summaries, then retrieves and injects query-relevant records through the prompt interface. Across five heterogeneous memory systems on LongMemEval and LoCoMo, it improves every host--benchmark baseline under two distinct LLM configurations. Controlled analyses separate gains from organizing already available evidence and from consolidating information across the full history. Under a matched LightMem pipeline, the entity--event--topic representation reaches 83.9% on LoCoMo, 3.6% above the strongest of six alternative auxiliary representations. These results show that generation-time structure is a portable complement to existing memory retrieval, while its interaction with host evidence depends on the benchmark and host.
|
| 2878 |
Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents
2605.01970
|
cs.AI
|
Debeshee Das, Julien Piet, Darya Kaviani, Luca Beurer-Kellner, Florian Tram\`er |
Memory systems enable otherwise stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. The Trojan Hippo attack is a class of persistent memory attacks that operates under a more realistic threat model than pr...Memory systems enable otherwise stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. The Trojan Hippo attack is a class of persistent memory attacks that operates under a more realistic threat model than prior memory poisoning work. The attacker plants a dormant payload into an agent's long-term memory via a single untrusted tool call (e.g., a crafted email), which activates only when the user later discusses sensitive topics such as finance, health, or identity, and exfiltrates high-value personal data to the attacker. While anecdotal demonstrations of such attacks have appeared against deployed systems, no prior work systematically evaluates them across heterogeneous memory architectures and defenses. We introduce Trojan Hippo Bench, a dynamic evaluation framework for persistent memory attacks and memory-layer defenses, comprising two components, (1) an OpenEvolve-based adaptive red-teaming benchmark that stress-tests defenses and memory backends against continuously refined attacks, and (2) a capability-aware security-utility analysis for persistent memory systems, enabling principled reasoning about defense deployment across different usage profiles. Instantiated on an email assistant across four memory backends (explicit tool memory, agentic memory, RAG, and sliding-window context), the undefended Trojan Hippo attack achieves up to 85-100% attack success rate (ASR) against current frontier models from OpenAI and Google, with planted memories successfully activating even after 100 benign sessions. We evaluate four memory-system defenses inspired by basic security principles; they reduce attack success to as low as 0-5%, but at utility costs that vary widely with the task mix. Effective real-world deployment thus remains an open challenge that Trojan Hippo Bench is built to study; we release it to support reuse and extension.
|
| 2879 |
Cumulative-Goodness Free-Riding in Forward-Forward Networks: Real, Repairable, but Not Accuracy-Dominant
2605.06240
|
cs.AI
|
Amirhossein Yousefiramandi |
Forward-Forward (FF) training lets each layer learn from a local goodness criterion. In cumulative-goodness variants, later layers can inherit a task that earlier layers have partly separated. We formalize this as layer free-riding: under the softplus FF crite...Forward-Forward (FF) training lets each layer learn from a local goodness criterion. In cumulative-goodness variants, later layers can inherit a task that earlier layers have partly separated. We formalize this as layer free-riding: under the softplus FF criterion, the class-discrimination gradient reaching block $d$ decays exponentially with the positive margin accumulated by earlier blocks (Theorem 3.1); a squared-hinge barrier reproduces the pathology. History-free and hardness-gated repairs raise deeper-layer separation by up to $5\times$ on CIFAR-100 and $45\times$ in an 8-block CIFAR-10 model, yet these and a depth-scaled auxiliary term move accuracy by under one percentage point among non-degenerate variants; on Tiny ImageNet, a harder cross-dataset check of the selected configuration, layer health likewise does not track accuracy. A theory-derived attenuation-compensated objective raises deepest-block separation $5.9\times$ at a consistent cost in the FF classifier's own accuracy that a probe trained on the frozen features recovers; in a smaller CIFAR-10 control, most of its separation gain comes from dropping the depth-scaled auxiliary term it replaces. Accuracy invariance is a property of the deployed sum-based scoring rule under redistribution (Proposition 3.2) and, empirically, of the structured trained readouts we test, not of what the network learns: retrained variants disagree on 13-26% of test predictions. Architecture and augmentation matter far more than the non-degenerate training rules studied. Cumulative free-riding is therefore a real and repairable pathology, but for the softplus and squared-hinge objectives, architectures, and datasets we study, it is not the dominant factor limiting accuracy.
|
| 2880 |
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
2605.10335
|
cs.AI
|
Yao Lu, Dengdong Fan, Shixun Zhang, Yonghong Tian |
Adaptive optimizers such as Adam are standard for training Transformers, but storing gradient first and second moments incurs substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without sto...Adaptive optimizers such as Adam are standard for training Transformers, but storing gradient first and second moments incurs substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by $\ell_p$-norm steepest descent, PowerStep applies a signed-power transform directly to one momentum buffer. We establish a finite-horizon stationarity bound for exact, unregularized updates, with an $O(1/\sqrt{T})$ term and a noise-dependent residual. Experiments on Transformers from 124M to 235B parameters show competitive validation quality while halving $\texttt{fp32}$ optimizer-state memory relative to AdamW. Combined with uniform $\texttt{int8}$ quantization, PowerStep remains numerically stable and reduces optimizer-state memory by $\sim8\times$ compared to $\texttt{fp32}$ AdamW. PowerStep thus provides a simple, memory-efficient alternative for large-scale training.
|
| 2881 |
PRISM: A Geometric Risk Bound for Decomposing Drift into Scale, Shape, and Head
2605.11608
|
cs.AI
|
Chieh-Yen Lin, Shao-Hua Sun |
A single base LLM now comes with dozens of post-training variants, quantized, LoRA-adapted, or distilled, and each has to be checked before release. Existing evaluations provide only a partial picture: benchmark scores and likelihood screens say that a variant...A single base LLM now comes with dozens of post-training variants, quantized, LoRA-adapted, or distilled, and each has to be checked before release. Existing evaluations provide only a partial picture: benchmark scores and likelihood screens say that a variant has degraded, similarity scores such as CKA and SVCCA say how its features moved, and nothing connects the two. We connect them with one structural fact and one design choice: the prediction head is linear, so feature geometry reaches the loss, and we compare the two feature sets through an orthogonal map, which leaves the geometry being measured unchanged. From these we derive PRISM, a closed-form upper bound on the cross-entropy risk gap between a target model and a proxy variant, and prove that it splits exactly into three measurable axes: scale, shape, and head. Each axis names a failure mode and where to intervene: low-bit quantization distorts shape and, at the lowest bit-widths, also inflates activation scale; quantizing the output embedding inflates the head term, which dominates the bound at high bit-widths. Because the shape term is differentiable, the same geometry becomes a regularizer that curbs catastrophic forgetting more than experience replay. Across two model families and five benchmarks, PRISM ranks quantized and fine-tuned variants from a single forward pass at mean Spearman above 0.8, and eight reference sequences already recover the ranking that the risk gap itself needs over a hundred to reach.
|
| 2882 |
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
2605.12070
|
cs.AI
|
Zhong Guan, Yongjian Guo, Haoran Sun, Wen Huang, Shuai Di |
Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-policy correction. In heterogeneous train...Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-policy correction. In heterogeneous training systems, the total importance ratio should ideally be decomposed into two semantically distinct factors: a \emph{training--inference discrepancy term} that aligns inference-side and training-side distributions at the same behavior-policy version, and a \emph{policy-staleness term} that constrains the update from the historical policy to the current policy. We show that practical asynchronous pipelines with delayed updates and partial rollouts often lose the required historical training-side logits, or old logits. This missing-old-logit problem entangles discrepancy repair with staleness correction, breaks the intended semantics of decoupled correction, and makes clipping and masking thresholds interact undesirably. To address this issue, we study both exact and approximate correction routes. We propose three exact old-logit acquisition strategies: snapshot-based version tracking, a dedicated old-logit model, and synchronization via partial rollout interruption, and compare their system trade-offs. From the perspective of approximate correction, we focus on preserving the benefits of decoupled correction through a more appropriate approximate policy when exact old logits cannot be recovered at low cost, without incurring extra system overhead. Following this analysis, we adopt a revised PPO-EWMA method, which achieves significant gains in both training speed and optimization performance.
|
| 2883 |
Generative Support Realignment for Cross-Domain Offline Reinforcement Learning
2605.13054
|
cs.AI
|
Minung Kim, Jeongmo Kim, Gwanwoo Choi, Seungyul Han |
Cross-domain offline reinforcement learning learns a target policy from pre-collected source and target datasets with different dynamics. When target data are scarce, effectively compensating for their limited coverage using source data remains challenging due...Cross-domain offline reinforcement learning learns a target policy from pre-collected source and target datasets with different dynamics. When target data are scarce, effectively compensating for their limited coverage using source data remains challenging due to the discrepancy between domains. We propose Target-aligned Coverage Expansion (TCE), which leverages source states to generatively realign and expand the limited target support while controlling the generation error induced by this expansion. We further derive a performance gap bound that characterizes the interplay between generation error and source--target dynamics gap, providing theoretical guidance for effective coverage expansion and source utilization. Across diverse cross-domain environments, TCE consistently outperforms state-of-the-art baselines, while further analysis empirically supports the guidance provided by our theoretical findings.
|
| 2884 |
Edge-AI-Driven Learning-to-Rank for Decentralized Task Allocation in Circular Smart Manufacturing
2605.16433
|
cs.AI
|
Mohammadhossein Ghahramani, Yan Qiao, Mengchu Zhou |
Task allocation in smart manufacturing systems must operate under decentralized decision-making, dynamic workloads, and shared-resource constraints. In circular manufacturing settings, these challenges are further intensified because tasks compete for reusable...Task allocation in smart manufacturing systems must operate under decentralized decision-making, dynamic workloads, and shared-resource constraints. In circular manufacturing settings, these challenges are further intensified because tasks compete for reusable, capacity-constrained assets, and machine selection also might affect processing energy. Although learning-based approaches have been explored for task allocation, improvements in predictive modeling do not necessarily translate into better allocation outcomes under decentralized negotiation. This work proposes an Edge-AI-driven decentralized task-allocation framework. We develop lightweight decision intelligence deployed at the machine level. It is developed progressively: first, a resource-aware heuristic establishes the decentralized bidding structure; a regression-based Edge-AI formulation then examines learned local bid approximation, and a compact autoencoder-regularized pairwise ranking model finally provides a learned correction to the analytical bid ordering. Each machine evaluates incoming tasks by using its processing capability, queue state, energy characteristics, and a compact signal representing contention over the reusable shared production asset. The framework is assessed using discrete-event simulation in scenarios characterized by high load and dependence on shared resources. Compared to the heuristic, the proposed ranking method increases completed tasks, reduces average tardiness, and lowers the deadline-miss rate, with statistically significant paired differences. Mean energy per completed task is also reduced. The results indicate that effective learning-assisted allocation depends not only on approximating local decision quantities, but also on shaping the relative preferences that determine negotiation outcomes.
|
| 2885 |
Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis
2605.18770
|
cs.AI
|
Arthur Capozzi, Dirk Helbing |
Public commercial registries are formally open, yet their practical analysis remains difficult because relevant facts are scattered across millions of records that combine structured metadata, multilingual legal notices, temporal events, and entity aliases. Th...Public commercial registries are formally open, yet their practical analysis remains difficult because relevant facts are scattered across millions of records that combine structured metadata, multilingual legal notices, temporal events, and entity aliases. This paper presents a controlled, tool-mediated agentic GraphRAG architecture for auditable natural-language analysis of such registries. The proposed pipeline transforms publications from the Swiss Official Gazette of Commerce into a Neo4j knowledge graph comprising over five million nodes and 4.7 million relationships. It combines deterministic ingestion of structured registry fields, LLM-assisted extraction of latent actors from unstructured notices, and a deterministic identity-resolution layer. An analytical agent operates on this graph through intent routing, restricted graph tools, bounded reflection, and state-machine-guided response synthesis. We evaluate the system using a multi-tier protocol covering answer quality, retrieval behavior, entity resolution, and multi-turn conversational performance. The complete architecture is compared with dense, lexical, and hybrid flat-retrieval baselines and with controlled architectural ablations. On a manually curated benchmark, graph-mediated retrieval increases factual correctness from 0.26 for the strongest flat-retrieval baseline to 0.83 for the complete system, with comparable improvements in relevance and completeness. Ablation results show that bounded reflection improves answer quality while intent routing and LLM-based graph enrichment improve reliability in difficult entity resolution tasks. An exploratory dashboard displays the graph evidence and execution traces underlying each response, allowing users to inspect how answers were produced.
|
| 2886 |
TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
2605.18859
|
cs.AI
|
Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng, Jiarong Xing |
LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrifi...LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrificing quality, yet existing router benchmarks evaluate routers only on one-shot prompts. They never expose the router-visible prefix at an intermediate agent step, never test whether a cheaper replacement preserves downstream task success, and often rely on online LLM judges at evaluation time. We introduce TwinRouterBench, a step-level routing benchmark with two tracks. The static track provides 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and PinchBench, each paired with an execution-verified target tier estimated under a released downgrade-and-cascade protocol; scoring is deterministic arithmetic over tier labels, trajectory membership, and token costs, with no online evaluator-side LLM judge. The dynamic track supplies a harness that runs routers on the full 500-case SWE-bench Verified suite; in this paper we report a 100-case held-out evaluation disjoint from the static SWE supervision split. At each LLM call the router selects a concrete model from a locked pool, and success is measured by official task resolution and realized API spend. The two tracks support fast offline iteration followed by end-to-end validation under live agent execution. Code and data are available at https://github.com/CommonstackAI/TwinRouterBench.
|
| 2887 |
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
2605.21862
|
cs.AI
|
Chushan Zhang, Ruihan Lu, Jinguang Tong, Xuesong Li, Yikai Wang |
Chunked vision-language-action (VLA) policies generate multi-step actions from one observation and typically re-observe after executing the action chunk. Earlier actions change the scene conditions for later steps, while occlusion can compromise visual feedbac...Chunked vision-language-action (VLA) policies generate multi-step actions from one observation and typically re-observe after executing the action chunk. Earlier actions change the scene conditions for later steps, while occlusion can compromise visual feedback. Spatial and temporal VLAs enhance geometry and memory, but their representations may become stale as actions alter object poses and contacts. We introduce EvoScene-VLA, which uses compact scene tokens to unify within-chunk scene prediction, cross-chunk state propagation, and observation-based correction. The action expert jointly denoises actions and future scene states in a single flow-matching process, allowing predicted scene changes to inform action generation. The final scene state serves as the next prior for subsequent correction. At the next control call, the vision-language model (VLM) fuses the incoming visual input with this prior. Two-level geometric anchoring applies local depth supervision to per-view scene features and global 3D feature supervision to the fused scene state; a shared geometric decoder imposes the same constraints on current and future states. Since demonstrations do not directly provide future scene states, we train a scene predictor using 3D teacher features from future images. Its predictions provide targets for the action expert's future scene states. At deployment, we remove both Scene Predictor and Geometric Anchor. EvoScene-VLA improves average success over the matched baseline by 2.4 and 2.2 percentage points on 31 RoboTwin tasks under fixed and randomized settings, respectively, and by 2.7 points on LIBERO. Across three real-robot tasks, average success rises from 37.3% to 42.0%. Ablations show that a recurrent scene state alone does not necessarily improve performance, whereas state propagation can improve control when combined with geometric and future scene supervision.
|
| 2888 |
Decomposing and Measuring Evaluation Awareness
2605.23055
|
cs.AI
|
Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Florian Tram\`er |
Frontier language models sometimes recognize that they are under evaluation and adjust their behavior which can undermine validity of benchmark results. Yet the field studies it without a shared foundation, conflating flaws of the evaluation with capabilities ...Frontier language models sometimes recognize that they are under evaluation and adjust their behavior which can undermine validity of benchmark results. Yet the field studies it without a shared foundation, conflating flaws of the evaluation with capabilities of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component and a model component that separates recognition from propensity. We operationalize the environment component through eight categorized trigger factors, such as placeholder entities and grading-style output formats, and study recognition and behavior through chain-of-thought monitoring. Across nine frontier models and four benchmarks, recognition rates depend on the specific pairing of model and benchmark. Recognition rarely associates with behavioral change, and when it does, the direction depends on the type of evaluation perceived. Models are also more sensitive to safety than capability evaluations, placing safety benchmark validity at greater risk. To study which factors each model is sensitive to and how they interact, we propose \textbf{EvalAwareBench}, a factor-controlled benchmark of 100 paired safety-capability tasks where each of the eight factors can be independently toggled, varying evaluative signals while holding the underlying request fixed. Through EvalAwareBench, we find that no single factor uniformly affects all models, but stacking factors progressively raises evaluation awareness across all of them. Our framework and EvalAwareBench provide the tools to measure, attribute, and mitigate evaluation awareness, building the foundation for future solutions.
|
| 2889 |
Scalable Heterogeneous Graph Foundation Models for Data-Driven Optimal Power Flow in Smart Grids
2605.23194
|
cs.AI
|
Massimiliano Lupo Pasini, Yijiang Li, Kibaek Kim, Teja Kuruganti |
Fast and reliable optimal power flow (OPF) approximation is important for power system operation, yet heterogeneous OPF graph models are often evaluated either with architecture-specific implementations or on limited training corpora. This paper presents a lar...Fast and reliable optimal power flow (OPF) approximation is important for power system operation, yet heterogeneous OPF graph models are often evaluated either with architecture-specific implementations or on limited training corpora. This paper presents a large-scale heterogeneous graph-learning workflow, built on HydraGNN, for data-driven OPF surrogate modeling and graph foundation-model (GFM) development. The workflow preserves buses, generators, loads, shunts, alternating-current (AC) lines, transformers, and device-to-bus relations and provides a common implementation for distributed preprocessing, multi-graphics-processing-unit (GPU) training, hyperparameter optimization (HPO), and downstream adaptation. Using approximately three million heterogeneous graphs from ten Power Grid Library for Benchmarking AC Optimal Power Flow Algorithms (PGLib-OPF) cases spanning 14 to 13,659 buses, we compare six heterogeneous graph neural network (GNN) families under a common HPO campaign on Frontier, resulting in two compact HeteroSAGE and HeteroHEAT models (~1.6-1.7 million parameters) that achieve the lowest observed validation MSE. Strong and weak scaling using up to 1,024 Frontier nodes show that 256 nodes (corresponding to 1,024 AMD MI250X GPUs) provides the best time-to-solution balance for the measured workload. Two downstream tasks evaluate task adaptation on the IEEE 118-bus case: a classification task to distinguish nominal operating conditions from artificially generated extreme-overload conditions and N-1 single-component-outage regression to predict bus voltage magnitude and phase angle. Partial fine-tuning provides the strongest balance between predictive performance and computational requirements.
|
| 2890 |
Adaptive Mass-Segmented KV Compression for Long-Context Reasoning
2605.23200
|
cs.AI
|
Junzhe Yang, Xiaoyu Shen |
The linear growth of the Key-Value (KV) cache is a critical bottleneck in long-form LLM inference. Existing KV compression methods mitigate this by evicting tokens based on importance scores. However, we show that their reliance on global Top-k selection trigg...The linear growth of the Key-Value (KV) cache is a critical bottleneck in long-form LLM inference. Existing KV compression methods mitigate this by evicting tokens based on importance scores. However, we show that their reliance on global Top-k selection triggers Region Wipe-out: the severe eviction of contiguous reasoning blocks that derails logical coherence. To address this, we propose Adaptive Mass-Segmented (AMS) KV Compression, a framework that shifts the paradigm from token-level competition to region-aware quota allocation. AMS adaptively partitions the KV cache based on the spatial distribution of attention mass, ensuring structurally vital reasoning segments receive guaranteed memory quotas. To ensure stability during iterative decoding, an EMA-based smoothing mechanism is incorporated to prevent jitter in segment boundaries. Crucially, AMS is a universal plug-and-play layer that is orthogonal to existing scorers. It can be seamlessly integrated into representative methods such as TOVA, Expected Attention, KeyDiff, R-KV and TriAttention. AMS is also system-compatible with modern paged-KV serving frameworks such as vLLM, supporting efficient gather-and-compact KV execution without introducing additional steady-state attention overhead. Extensive experiments across a diverse suite of tasks, including mathematical reasoning (MATH500, AIME, GSM8K), code completion, open-domain QA, and sparse retrieval, demonstrate that AMS consistently mitigates structural fragmentation and boosts model performance.
|
| 2891 |
Learning to Assign Prediction Tasks to Agents with Capacity Constraints
2605.27999
|
cs.AI
|
Shang Wu, Saatvik Kher, Padhraic Smyth |
We address the problem of learning to assign prediction tasks to one agent from a set of available agents, including human decision-makers and AI models. We focus on sequential learning of agent expertise and assignment policies where each agent is constrained...We address the problem of learning to assign prediction tasks to one agent from a set of available agents, including human decision-makers and AI models. We focus on sequential learning of agent expertise and assignment policies where each agent is constrained to handle a fraction of tasks. We provide a general theoretical characterization of this problem in terms of agent capacities, differences in agent expertise, and task context. We then develop a framework of sequential explore-exploit policy-learning algorithms that seek to maximize overall performance. Experimental results over a variety of tabular, image, and text prediction tasks demonstrate systematic gains from our policy-learning algorithms relative to non-contextual baselines across different types of agents, including LLMs and humans.
|
| 2892 |
Efficient Pre-Training of LLMs through Truncated SVD Representations
2605.28573
|
cs.AI
|
Kaivan Kamali, Kajetan Schweighofer, Hormoz Shahrzad, Olivier Francon, Babak Hodjat |
LLM pretraining is extremely costly; therefore, parameter-efficient LLM architectures have recently emerged as a compelling research direction. One such promising approach is to represent the parameters as orthonormal low-rank weight matrices. However, maintai...LLM pretraining is extremely costly; therefore, parameter-efficient LLM architectures have recently emerged as a compelling research direction. One such promising approach is to represent the parameters as orthonormal low-rank weight matrices. However, maintaining orthonormality during training is computationally expensive, making it impractical. This paper presents the TSVD (Truncated Singular Value Decomposition) framework which efficiently maintains orthonormality through QR decomposition and caching. Furthermore, a spectral energy heuristic is introduced to select the rank of the resulting low-rank weight matrices. Empirical evaluations across model sizes show that TSVD matches or outperforms full-parameter baselines at a fraction of the compute cost. TSVD thus provides a scalable, computationally efficient foundation for LLM pretraining.
|
| 2893 |
Do Proactive Agents Need an LLM to Decide When to Act?
2605.30152
|
cs.AI
|
Xiaoze Liu, Ruowang Zhang, Amir H. Abdi, Michel Galley, Zhikai Chen |
Proactive assistants continuously decide when to intervene and what context should support the intervention. Large language model (LLM) pipelines repeatedly interpret activity histories to make these decisions, paying an inference cost even when the assistant ...Proactive assistants continuously decide when to intervene and what context should support the intervention. Large language model (LLM) pipelines repeatedly interpret activity histories to make these decisions, paying an inference cost even when the assistant remains silent. We show that a small graph model can handle both decisions and improve the language agents it controls. Our key insight is that user activity has a native graph structure: events involve persistent entities whose recurrence connects interactions over time. Triggering and context selection map directly to predictions on event and entity nodes. We introduce a temporal-graph-learning (TGL) controller that learns these predictions jointly and supplies both outputs in one forward pass. The downstream language agent generates suggestions on triggered events using the activity history and scored entities. At 11.13 ms per event on a GPU server, TGL achieves the highest AUCs among nine trigger architectures and gives approximately $4$--$7\times$ trigger-stage speedups over the two single-forward LLM triggers. A shared TGL model improves F1 across all 14 downstream backbones by a mean of 16.7 points. The controller also runs at 13.99 ms on a consumer laptop with an approximately 220 MiB BF16 resident footprint, bringing effective proactive control to on-device deployment.
|
| 2894 |
Reasoning with Sampling: Cutting at Decision Points
2605.30327
|
cs.AI
|
Felix Zhou, Anay Mehrotra, Quanquan C. Liu |
Frontier reasoning models are produced by post-training base language models with reinforcement learning. Recent work has challenged this by showing that sampling from a sharpened version of the base model's distribution, a so-called power distribution, elicit...Frontier reasoning models are produced by post-training base language models with reinforcement learning. Recent work has challenged this by showing that sampling from a sharpened version of the base model's distribution, a so-called power distribution, elicits comparable reasoning without additional training, curated datasets, or verifiers. However, making this method practical requires efficiently sampling from the power distribution. A sampler needs to "mix" to the power distribution, which necessitates moving between modes of the target distribution; intuitively, e.g., trying different reasoning strategies. The samplers proposed in prior works repeatedly select a "cut" position in the current reasoning trace uniformly at random and resample the suffix from that position onward. However, reasoning traces typically contain a few consequential decisions (e.g., the choice of proof strategy or algorithm), and we observe that a uniformly chosen cut tends to rewrite local details rather than revisit decision points. We introduce an algorithm (Entropy-Cut Metropolis-Hastings) that uses the base model's next-token entropy as a proxy to identify key decision points and resample from those positions. We empirically verify that entropy jumps are a useful proxy for decision points and, in a stylized model of reasoning, prove that our method's mixing time scales with the number of decisions in a trace rather than with the number of tokens, which can be much larger. Across MATH500, HumanEval, GPQA Diamond, and AIME26, our method consistently improves over baselines and RL-trained models, including best-of-N at the same token budget.
|
| 2895 |
Two-Fidelity Best-Action Identification for Stochastic Minimax Tree
2606.01708
|
cs.AI
|
Peter Chen, Xi Chen |
We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly relevant in modern AI planning, where deep minimax search and Monte Carlo Tree Search (MCTS) with language model long rollouts face a fundament...We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly relevant in modern AI planning, where deep minimax search and Monte Carlo Tree Search (MCTS) with language model long rollouts face a fundamental tradeoff: heuristic evaluations are cheap but biased, while accurate rollouts are reliable but prohibitively expensive. We propose 2FFS, a two-fidelity tree-search algorithm that brings multi-fidelity flat bandit ideas into trees. The algorithm combines minimax-style fast expansion with MCTS-style stochastic sampling, adaptively deciding when to exploit cheap biased evaluations and when to invoke expensive accurate evaluations for local certification. We prove fixed-confidence correctness, establish finite stopping for exact identification, and give a polynomial-depth cost upper bound for general-depth trees. Across numerical stochastic-tree experiments, 2FFS uses substantially fewer samples and computational operations comparing to existing BAI-MCTS baseline.
|
| 2896 |
RobotValues: Evaluating Household Robots When Human Values Conflict
2606.03312
|
cs.AI
|
Jongwook Han, Hyeongjin Kim, Yohan Jo |
While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations where robots are expected to choose actions that prioritize diverse values such as human autonomy, efficiency, or social ap...While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations where robots are expected to choose actions that prioritize diverse values such as human autonomy, efficiency, or social appropriateness. Yet, there are no benchmarks for evaluating robots' value preferences in such scenarios. We introduce RobotValues, a benchmark to evaluate household robot planners in 8K value-conflict scenarios. Each instance consists of a realistic, synthetically generated household image with multiple plausible robot actions that prioritize different human values. We construct ROBOTVALUES through LLM-assisted scenario generation, stakeholder-grounded value extraction, image generation and automatic quality control. We evaluate 10 VLMs used in robotics and find that models show default value preferences, including safety and accommodation, while underselecting privacy-prioritizing actions. When models are prompted to prioritize values that conflict with their preferences, models often fail to override the default actions, choosing incorrect actions 80% of the time on average across models. These findings highlight the need to go beyond task completion or safety evaluations and assess robots' decision-making capability when human values conflict.
|
| 2897 |
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
2606.04923
|
cs.AI
|
Xuekang Wang, Zhuoyuan Hao, Shuo Hou, Hao Peng, Juanzi Li |
Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes....Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a Controllable Hacking Environment for Rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.
|
| 2898 |
HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers
2606.06493
|
cs.AI
|
Lizhi Yang, Junheng Li, Nehar Poddar, Yiling Hou, Gio Huh |
Humanoid loco-manipulation benefits from one controller to coordinate locomotion, arm movements, and fall recovery. These behaviors are difficult to learn together from scratch because they require different skills and objectives. We present HANDOFF, a trainin...Humanoid loco-manipulation benefits from one controller to coordinate locomotion, arm movements, and fall recovery. These behaviors are difficult to learn together from scratch because they require different skills and objectives. We present HANDOFF, a training architecture that distills motion-tracking, locomotion, and recovery teachers into one 29-DoF policy, using context-conditioned KL losses over separate action slices and a soft mixture of experts. Commanded velocity continuously blends body supervision, while a binary recovery flag assigns the full action to the recovery teacher. The resulting controller takes a compact 10-D task-space command rather than a dense kinematic stream and does not switch policies at runtime. On the Unitree G1, HANDOFF matches state-of-the-art velocity tracking and obtains the largest robust manipulation workspace in a matched adapted-interface benchmark. The same controller executes natural-language-driven, multi-stage loco-manipulation tasks in simulation and hardware experiments without further data collection or fine-tuning.
|
| 2899 |
UPLOTS: A Unified Pretrained Language Model for Constrained Time-series Generation
2606.10466
|
cs.AI
|
Du Yin, Hao Xue, Jinliang Deng, Yang Yang, Shuang Ao |
In time-series generation, existing approaches typically handcraft ortrain a separate model for each dataset, which hinders their scalability and fails to leverage shared temporal structures across domains. To address this fragmentation, we propose UPLOTS, a U...In time-series generation, existing approaches typically handcraft ortrain a separate model for each dataset, which hinders their scalability and fails to leverage shared temporal structures across domains. To address this fragmentation, we propose UPLOTS, a Unified, Prompt-guided Language model framework fOr constrained Time-Series Generation across diverse domains. Instead of building task-specific models, UPLOTS leverages a single pre-trained transformer backbone guided by learned constraint prompts, enabling on-demand generation with precise pattern control. One key innovation is our dynamic multi-dataset loss re-weighting and prompt-to-pattern mapping, which allows UPLOTS to internalize diverse temporal structures during training and conditionally generate them at inference. We evaluate UPLOTS on four real-world benchmarks and multiple constraint settings, including peak-period, calendar, load-level, and volatility patterns. Additional held-out constraint-combination and downstream forecasting experiments further demonstrate that UPLOTS generalizes beyond the original peak-pattern setting and improves data augmentation under scarce real-data regimes. Our code and baselines are available at github repo: https://github.com/cruiseresearchgroup/UPLOTS.
|
| 2900 |
Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness
2606.18363
|
cs.AI
|
Haowen Liu, Xirui Li, Shaoxiong Yao, Peng Shi, Tianyi Zhou |
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-language-action systems by combining high-level...Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-language-action systems by combining high-level reasoning with external modules for perception, planning, and control. However, it remains unclear what makes an effective harness for embodied manipulation, and to what extent such a harness can unlock embodied capabilities in a wide range of reasoning models. In this work, we present Guava, a harness framework for embodied tool use developed through systematic exploration of the design space of agent workflows, action spaces, and observation spaces. Our study identifies three key ingredients for effective embodied agents: iterative perception-reasoning-action loops, semantic action abstractions, and multimodal observations. To understand whether these design principles are universal even to small models, we develop an end-to-end training pipeline that distills embodied manipulation capabilities into a 4B open-source model using fewer than 2K trajectories collected entirely in simulation. Experimental results in both simulation and real-world environments show performance comparable to frontier proprietary models while exhibiting strong generalization to unseen objects, novel instructions, and long-horizon tasks. Results suggest that a well-designed harness can serve as a scalable, model-agnostic interface for embodied manipulation, enabling strong emergent embodied capabilities in compact open-source models with minimal training data.
|
| 2901 |
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems
2606.18837
|
cs.AI
|
Hehai Lin, Qi Yang, Chengwei Qin |
Large Language Model (LLM)-based automatic Multi-Agent Systems (MAS) generation has become a crucial frontier for tackling complex tasks. However, existing methods face a dilemma between model capability and experience retention. Inference-time MAS leverages f...Large Language Model (LLM)-based automatic Multi-Agent Systems (MAS) generation has become a crucial frontier for tackling complex tasks. However, existing methods face a dilemma between model capability and experience retention. Inference-time MAS leverages frozen frontier LLMs but repeats identical searches without learning from past experience. Conversely, Training-time MAS internalizes experience via gradient updates but is constrained by the low capability ceiling of smaller models, and is hard to scale to large frontier LLMs. To bridge this gap, we propose Skill-MAS, a novel third path that decouples experience retention from parametric updates by conceptualizing the high-level orchestration capability as an evolvable Meta-Skill. Skill-MAS refines this architectural knowledge through a closed optimization loop: (1) Multi-Trajectory Rollout samples a behavioral distribution for each task under the current Meta-Skill; and (2) Selective Reflection adaptively selects priority tasks and applies hierarchical contrastive analysis to distill systemic experience into generalizable, strategy-level principles. Extensive experiments across four complex benchmarks and four distinct LLMs demonstrate that Skill-MAS not only achieves remarkable performance gains but also maintains a favorable cost-performance trade-off. Further analysis reveals that the evolved Meta-Skills are highly robust and exhibit strong transferability across unseen tasks and different LLMs.
|
| 2902 |
Temporal Self-Imitation Learning
2606.19752
|
cs.AI
|
Yinsen Jia, Boyuan Chen |
Long-horizon policies trained with reinforcement learning can still achieve high return through inefficient interactions, while rare efficient behaviors discovered during training may be forgotten. We argue that temporal efficiency itself provides a source of ...Long-horizon policies trained with reinforcement learning can still achieve high return through inefficient interactions, while rare efficient behaviors discovered during training may be forgotten. We argue that temporal efficiency itself provides a source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 30 long-horizon tasks spanning robot manipulation and interactive navigation, TSIL consistently improves task success rates, learning efficiency, behavioral efficiency, and robustness to unstable training conditions.
|
| 2903 |
Is Agent Code Less Maintainable Than Human Code?
2606.21804
|
cs.AI
|
Shaswat Patel, Betty Li Hou, Arun Purohit, Kai Xu, Jane Pan |
Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding agents have demonstrated strong performance on single-issue tasks, it remains unclear how maintainable their code is when ...Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding agents have demonstrated strong performance on single-issue tasks, it remains unclear how maintainable their code is when future agents build on top of it, potentially leading to compounding downstream effects. We investigate how agent code compares to human code in these maintenance settings, presenting CodeThread, a framework to construct controlled experiments from repository-level coding benchmarks. Applying CodeThread to four frontier coding agents and four benchmarks, we find that agents are less effective at resolving tasks when building on agent code compared to human code, with task resolve rate drops of up to 13.1%. Regression analysis reveals that many traditional software engineering maintainability metrics do not explain this difference. Instead, the clearest signals are subtler behavioral differences in agent code, such as changes to input validation and error handling, along with differences in downstream code size and task difficulty. These findings highlight the need to evaluate these systems not only by immediate task resolution but also by code maintainability, and point to potential sources of downstream errors introduced by agent code.
|
| 2904 |
Low-power analogue neural networks with trainable nonlinear connections for continuous control
2606.23742
|
cs.AI
|
Ian T. Vidamour, Fernando Aguirre, Thomas J. Hayward, Matthew O. A. Ellis, Charles Swindells |
Physical neural networks promise low-power machine learning by computing directly with analogue device physics, but most architectures force nonlinear device responses to act as scalar weights. Inspired by Kolmogorov-Arnold networks, we place trainable nonline...Physical neural networks promise low-power machine learning by computing directly with analogue device physics, but most architectures force nonlinear device responses to act as scalar weights. Inspired by Kolmogorov-Arnold networks, we place trainable nonlinear functions on the connections, making each physical connection a learnable computational element. Realising these functions as analogue band-pass filters on field-programmable analogue arrays, we find that the benefit is task-dependent and follows from the smoothness of the physical basis: the networks represent smooth, continuously valued targets, including robotic kinematics, continuous control, and photovoltaic maximum-power-point tracking, with far fewer nodes and connections than multilayer perceptrons, but offer no parameter-efficiency advantage on classification-like decision boundaries. Trained networks transfer to hardware across approximately 35,000 connections with quantified fidelity, and a dedicated CMOS implementation is projected to operate at approximately 30 microwatts. A memristive realisation reproduces the same behaviour in simulation, indicating that the advantage comes from placing trainable nonlinearity on connections, rather than from a particular device.
|
| 2905 |
Does Anthropomorphic Language Impact Public Perceptions of AI?
2606.29121
|
cs.AI
|
Betty Li Hou, Sophie Hao, Sunoo Park, Tal Linzen |
Public discourse about artificial intelligence (AI) often uses anthropomorphic language: language that attributes human capabilities and characteristics to AI systems. This practice has been criticized for setting misleading expectations, inflating claims, and...Public discourse about artificial intelligence (AI) often uses anthropomorphic language: language that attributes human capabilities and characteristics to AI systems. This practice has been criticized for setting misleading expectations, inflating claims, and fueling hype around AI, which may distort public understanding of AI and impact policy priorities. We study the effects of anthropomorphic framing by comparing changes in participants' perceptions of AI (N=815) when reading passages with and without anthropomorphic language, designed to reflect realistic public-facing AI discourse. We further examine whether these effects differ across two types of AI technologies -- large language models and recommendation systems -- and measure changes in perceptions of AI across several dimensions that are prominent in current public discourse. In a separate condition using a text that explicitly discusses the dangers of AI, we show that individuals' views of AI can shift in response to reading a text; yet in the main conditions of the experiment, where we compare anthropomorphic and non-anthropomorphic descriptions, we find that whether the text uses anthropomorphic language does not substantially affect participants' perceptions of AI. Our results indicate that any immediate effects on opinions of AI are modest, although they leave open the possibility that anthropomorphic language could have an effect in naturalistic settings, or over gradual, continued exposure.
|
| 2906 |
Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs
2607.10803
|
cs.AI
|
Shrestha Datta, Hongfu Liu, Anshuman Chhabra |
Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and interpretability. We introduce Weight-Adjusted Gradients (WAG), a simple yet effective approach for estimating parameter i...Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and interpretability. We introduce Weight-Adjusted Gradients (WAG), a simple yet effective approach for estimating parameter importance that explicitly captures the interaction between model weights and first-order gradient information and identifies parameters that disproportionately influence model behavior, such as those responsible for collapse phenomena in LLMs. Across a range of models and settings, we show that WAG surfaces a tiny but critical subset of parameters (< 0.5 parts per million or 0.00005% of model size) whose modification leads to dramatic degradation in performance, indicating a novel failure mode. These findings also reveal a previously underexplored interplay between weights and gradients, suggesting that parameter importance cannot be fully understood through either signal alone. We demonstrate the practical utility of WAG across several diverse applications, such as expert allocation in Mixture-of-Experts (MoE) architectures, targeted unlearning, mixed-precision quantization, and layer selection for knowledge editing. In sum, WAG can serve as a unified approach for analyzing, debugging, and controlling LLMs, and opens new directions for principled parameter-level interpretation.
|
| 2907 |
A Safety-First Gateway Architecture for Trusted Public Health Resource Navigation
2607.13038
|
cs.AI
|
Ben Torkian, Jun Zhou |
Conversational AI can improve access to public health information, but public-facing healthcare applications require safeguards against inappropriate medical guidance and unsupported generation. We present a Safety-First Science Gateway for maternal and child ...Conversational AI can improve access to public health information, but public-facing healthcare applications require safeguards against inappropriate medical guidance and unsupported generation. We present a Safety-First Science Gateway for maternal and child health (MCH) resource navigation that combines large language models (LLMs) and retrieval-augmented generation (RAG) with a multi-layer safety architecture. The gateway integrates emergency handling, domain/scope screening, source attribution, anonymous session management, and operational audit logging while restricting retrieval to curated institutional resources. We describe the gateway architecture, prototype implementation, and functional verification of selected workflows. The current system provides resource provenance and safety-bounded navigation; it does not constitute a clinical decision-support system or automated claim-by-claim verification of generated health information. This work provides a reusable architectural framework for conversational navigation of curated public-health resources.
|
| 2908 |
Distilled Reinforcement Learning for LLM Post-training
2607.17247
|
cs.AI
|
Chen Wang, Zhaochun Li, Jionghao Bai, Yining Zhang, Hexuan Deng |
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome s...Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into the RL objective to provide fine-grained guidance, selectively transfer new knowledge and avoid unconditional imitation. Distilled RL contains three components: reverse importance sampling with clipping, negative sample reset, and sequence-level geometric normalization. Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model. Extensive experiments across both within-family and cross-family distillation settings show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k. Our code is available at https://github.com/597358816/Distilled-RL.
|
| 2909 |
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
2607.24555
|
cs.AI
|
Junsung Hwang |
Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read at every decode step. We find that attention keys are approximately low-rank within pages. A single low-rank projection shared across pages can miss page-s...Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read at every decode step. We find that attention keys are approximately low-rank within pages. A single low-rank projection shared across pages can miss page-specific directions; fitting a basis to each page better identifies the pages receiving the most attention at comparable stored selector cost. LOCKS stores a rank-$r$ spectral summary per page, reconstructs its within-page logits, and selects pages by log-sum-exp mass without reading candidate keys or values. It stays within about a point of FullKV on LongBench-v1, tracks the read-every-key exact-LSE oracle on RULER down to the smallest budgets, and retains quality furthest under tight budgets on AIME26 and MATH-500. At a $2048$-token budget it matches FullKV aggregate quality beyond $100$K context while attending about $2\%$ of tokens. Across ranks $2$-$8$, summaries use $4$-$10\%$ of full-KV bytes. On GH200 with GPU-resident KV, LOCKS reduces complete decode-step time by $1.8\times$ at $512$K context. With full KV offloaded to Grace memory, it reaches $3.82$-$4.22\times$ the faster dense backend's aggregate throughput at $64$K-$256$K by serving larger batches.
|
| 2910 |
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
2607.24786
|
cs.AIcs.SDeess.AScs.MM
|
Hugo Malard, Michel Olvera, Sanjeel Parekh, Ga\"el Richard, Slim Essid |
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challenging, as models must locate sound sources from temporally aligned audio-visual dat...Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challenging, as models must locate sound sources from temporally aligned audio-visual data without pixel-level supervision. Recent large-scale audio-visual retrieval models, trained at unprecedented scale, encode rich multimodal structure. We show their latent representations, though optimized for global alignment, can nonetheless enable fine-grained spatial grounding. While spatial detail is progressively lost in the upper layers of retrieval backbones due to global pooling, intermediate visual tokens retain highly structured spatial information. To exploit this, we introduce LAIP (\emph{Localization via Audio-Informed Pooling}), a framework that employs a lightweight \emph{Audio-informed Spatial Pooling} (AiSP) to replace the standard global aggregation module. By querying intermediate visual tokens with audio aligned at the frame level, LAIP recovers localized spatial information that is otherwise discarded by the retrieval pipeline, with the largest gains observed for PE-AV, a stack with underlying temporal aggregation. Our approach achieves state-of-the-art performance on AVSBench and AVATAR, nearly doubling previous results on the latter, improving average CIoU from 13.21 to 26.22.
|
| 2911 |
OneLatent: Latent Reasoning for Efficient Foundation Recommendation Models
2607.26621
|
cs.AI
|
Hao Jiang, Peiru Du, Pengfei Yao, Mengting Li, Siyuan Lou |
Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their use as the backbone of foundation recommendation models (FRMs). Existing methods enhance recommendations through explicit Chain-of-Thought (CoT) reasoning under a Th...Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their use as the backbone of foundation recommendation models (FRMs). Existing methods enhance recommendations through explicit Chain-of-Thought (CoT) reasoning under a Think-then-Answer paradigm. However, explicit CoT incurs substantial inference overhead by generating lengthy reasoning traces and relies on manually designed templates that struggle to capture diverse, dynamic user interests. We propose OneLatent, an efficient latent reasoning framework that compresses explicit reasoning traces into several learnable latent tokens, enabling Latent-Reason-then-Answer inference without generating verbose traces. OneLatent first introduces Multi-View Adaptive CoT (MV-ACoT), which creates diverse, high-quality teacher-generated supervision by exploring user interests from multiple perspectives and automatically adapting reasoning complexity to each instance. Building on pretrained FRMs, it then uses a three-stage latent-token alignment paradigm to progressively internalize CoT traces into learnable latent tokens. Finally, a multistage curriculum-based post-training strategy activates latent-token reasoning for downstream recommendation tasks. Experiments on an industrial-scale Kuaishou dataset and the public Kuaishou LLM-Rec benchmark show that OneLatent consistently outperforms explicit CoT-based methods and traditional baselines. Compared with the Think and No-Think variants of FRMs, OneLatent improves SID@64 by 17.44% and 9.33%, respectively, while achieving over 17x higher online inference throughput. We further develop a production serving system for scalable, real-time FRM inference. An online A/B test in Kuaishou's local-services advertising scenario shows that deploying OneLatent with this system yields an estimated 9.6% revenue lift over strong online baselines, including OneRec and OneReason.
|
| 2912 |
ORCA-bench: How Ready Are Language Model Agents for Oncall?
2607.28545
|
cs.AI
|
Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang |
Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident ...Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs 1,079 RCA tasks with six days of metrics, logs, and traces collected from an OpenTelemetry-instrumented microservice system under continuous simulated user load. Agents investigate this recorded history through real observability interfaces---Prometheus, Jaeger, and OpenSearch via Grafana---with full access to application source code. Tasks systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $\kappa_w = 0.91$). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard---a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access reduces RCA accuracy and increases the hallucination rate for every evaluated model. These results come from a curated 50 GB / six-day testbed of standalone tasks on a system whose code and instrumentation are public. Since real production systems are orders of magnitude larger, more dynamic, and more idiosyncratic, the gap we report underscores the engineering work still needed before agents can be entrusted with production reliability. We release the public set at https://hub.harborframework.com/datasets/orca-bench/orca-bench.
|
| 2913 |
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
2607.28623
|
cs.AI
|
Lizhi Yang, Junheng Li, Aaron D. Ames |
We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The deployed policy sees the ball only as segmentation-masked depth from a head-mounted cam...We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The deployed policy sees the ball only as segmentation-masked depth from a head-mounted camera, while training-time CBF guidance represents clearance to every body link, and an adversarial motion prior regularizes the resulting evasive reflexes. We evaluate on a controlled any-link contact benchmark with seeded throws in two regimes: single throws and a deployment loop in which the robot walks back to its station and recovers between throws. On this benchmark, the policy comes within a few points of a privileged state oracle: a fixed onboard camera alone is adequate for evasion. We find that usable barrier structure depends on perceptual observability: Joint-CBF gives the best performance with accurate ball states, degrades under fixed-camera observations when used only as training guidance, and recovers with a ball-tracking gimbal or privileged runtime filter. We therefore deploy a lightweight Link-CBF policy zero-shot on the Unitree G1 in the real world, where it tolerates imperfect perception, succeeds on 95% of throws, and uses semantic segmentation to dodge different balls.
|
| 2914 |
Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study
2608.00042
|
cs.AI
|
Ramesh B. Paramkusham |
Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-stakes environments including healthcare, legal services, and financial analysis. While performance gains from...Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-stakes environments including healthcare, legal services, and financial analysis. While performance gains from parameter-efficient fine-tuning are well characterised, the corresponding impact on trustworthiness (factual calibration and adversarial robustness) remains poorly understood. This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation across three SLM architectures (TinyLlama 1B, Gemma-2 2B, Llama 3.2 1B), three domains (healthcare, legal, finance), two training-data conditions (benign and adversarially perturbed), and four fine-tuning strategies (baseline LoRA, Safety-DPO, Dark Experience Replay, and Task Arithmetic LoRA, TA-LoRA). Trustworthiness is evaluated through TruthfulQA MC2 (factual calibration) and HarmBench ASR (adversarial robustness) across all 216 experimental configurations with three random seeds. Three principal findings emerge. First, baseline QLoRA domain adaptation produces minimal TruthfulQA MC2 change across all model-domain combinations (mean |Delta TQA| < 0.02). Second, adversarially perturbed training data consistently improves domain adaptation quality (Delta loss approximately -0.040) without worsening trustworthiness benchmarks. Third, none of the three safety-preserving strategies reduced adversarial harm susceptibility: Safety-DPO was effectively neutral (mean Delta ASR < 0.001), while Dark ER and TA-LoRA increased mean HarmBench ASR by +0.171 and +0.155 respectively in safety-aligned models (Gemma-2 2B, Llama 3.2 1B), with individual configurations exceeding +0.45. These results challenge the assumption that replay-based and arithmetic-merge strategies transfer alignment to domain-adapted SLMs.
|
| 2915 |
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
2608.04682
|
cs.AI
|
Haobin Li, Ping Deng, Weizhong Qian, Liang Jiang, Zhenyu Huang |
Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase. However, existing SWE benchmarks typically assume that high-quality issue reports ...Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase. However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation. To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages. Beyond shifting the focus from existing reactive bug fixing to proactive bug fixing, Active-SWE enables a more in-depth evaluation by expanding the scope from fixing a specific recorded bug to multiple-bug fixing and potential bug discovery scenarios. To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability. Extensive experiments reveal that most state-of-the-art coding agents struggle with proactive bug-fixing tasks, demonstrating limited performance in locating and resolving recorded bugs, handling multiple bug fixing scenarios, and discovering valid potential bugs.
|
| 2916 |
SR-OPSD: Self-Referenced On-Policy Self-Distillation
2608.09745
|
cs.AI
|
Zhuo Sun, Entong Li, Yanlong Zhao, Xiaoyuan Cheng, Wenxuan Yuan |
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on student-generated trajectories, complementing reinforcement learning with sparse outcome rewards. Its self-teacher, derived from the student's current or exponentially a...On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on student-generated trajectories, complementing reinforcement learning with sparse outcome rewards. Its self-teacher, derived from the student's current or exponentially averaged parameters and conditioned on additional context, evolves alongside the student and its rollout context distribution. The benefit of modifying this moving target depends on how target--student probability mismatches translate into updates. We propose \emph{Self-Referenced On-Policy Self-Distillation (SR-OPSD)}, which constructs a normalized geometric target from the self-teacher and a frozen initial policy, then minimizes the forward R\'enyi divergence from this target to the student. The interpolation coefficient controls the self-teacher's contribution, while the R\'enyi order controls the power weighting of target-to-student probability ratios in the gradient. For fixed contexts and target components, we establish a conditional variational characterization and derive the exact token-logit gradient, revealing how anchoring and projection jointly shape the effective update target. Experiments across scientific reasoning, tool use, mathematical reasoning, and code generation demonstrate strong performance across multiple model families and scales. Ablations further show that reference anchoring can improve or degrade performance depending on the projection objective, supporting the joint design of target construction and projection geometry.
|
| 2917 |
Functional compatibility as a determinant of persistent neural learning
2608.22462
|
cs.AI
|
Hossein Javidnia |
Neural networks can acquire new capabilities while damaging existing ones, but what determines whether new learning persists remains unclear. We identify functional compatibility, the extent to which incoming learning can coexist with behaviour that must be pr...Neural networks can acquire new capabilities while damaging existing ones, but what determines whether new learning persists remains unclear. We identify functional compatibility, the extent to which incoming learning can coexist with behaviour that must be preserved, as an experimentally manipulable causal determinant of persistence. From identical neural states, we vary compatibility while matching unrestricted learning opportunity and imposing a common retention requirement. Persistent learning increases with compatibility across independent directions, convolutional and transformer architectures, vision and text, and a ten-seed replication. Learning rules and retention constraints determine how much compatible opportunity is retained, whereas nonlinear geometry limits the matched intervention at larger update norms. Functional compatibility therefore reframes stability-plasticity from preventing forgetting to determining which new learning can coexist with existing function and persist.
|
| 2918 |
On-policy Distillation with Verifiable Reward
2608.24696
|
cs.AI
|
Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li |
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level...Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.
|
| 2919 |
Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction
2608.25466
|
cs.AI
|
Azrin Sultana |
The functional annotation of genes in non-model organisms remains a significant challenge in computational biology, with 20-70% of sequenced genes lacking characterized functions. Traditional homology-based methods are often costly and strongly dependent on hi...The functional annotation of genes in non-model organisms remains a significant challenge in computational biology, with 20-70% of sequenced genes lacking characterized functions. Traditional homology-based methods are often costly and strongly dependent on high sequence similarity. This study presents Homo-RAG, a framework for large language model-based gene function prediction that integrates homology-guided multi-hop retrieval with evidence-aware ranking. The framework exploits biological relationships between zebrafish and human orthologs to guide evidence acquisition from ZFIN, UniProt, and PubMed through hybrid dense and lexical retrieval. An Evidence Confidence Score (ECS) integrates semantic relevance, entity matching, orthology information, source reliability, and literature association signals to refine the ranking of retrieved evidence. Extensive evaluation across 150 queries and 7,200 retrieved documents shows that evidence weighting parameter of lambda=0.50 improves NDCG@10 to 0.9879 and MRR to 0.99, while retrieving relevant evidence for 99.33% of queries. Furthermore, 80% of the retrieved documents are query-exclusive, indicating that evidence quality complements rather than replaces retrieval relevance. These findings establish Homo-RAG as a practical and robust framework for reliable, evidence-grounded gene function prediction in understudied organisms. The study addresses important limitations of conventional annotation pipelines while identifying opportunities for future improvements in evidence features and attribution mechanisms.
|
| 2920 |
TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development
2608.26086
|
cs.AI
|
Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang |
Auto-research agents now run machine-learning development unattended for hours, revising data pipelines, models, and validation from their own feedback, yet on most competitions they still finish below strong human competitors. Outcome-based benchmarks record ...Auto-research agents now run machine-learning development unattended for hours, revising data pipelines, models, and validation from their own feedback, yet on most competitions they still finish below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4{,}465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, and the labeling models, together with an open-source toolkit that turns a run from any command-line agent into a TraceML trajectory and reads it against the human cohorts.
|
| 2921 |
Performative Privacy: When Differential Privacy Maximizes Utility
2608.28198
|
cs.AI
|
Uddalak Mukherjee, Edwige Cyffers, Yann Chevaleyre |
Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. However, this claim has not been formalized so far. In parallel, performative learning provide...Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. However, this claim has not been formalized so far. In parallel, performative learning provides a framework for studying learning systems whose deployment affects the data they later observe. In this work, we bring these two perspectives together and introduce performative privacy, where data leakage reduces future participation. We study a simple model where agents repeatedly contribute data for mean estimation but may leave the system when their data is leaked. Privacy is implemented through differentially private mechanisms, creating a trade-off between estimation noise and future participation. We show, through a theoretical study of the dynamics and numerical experiments, that a finite privacy budget can outperform non-private estimation in the long term when the feedback loop between leakage and participation is sufficiently strong. This provides first evidence that differential privacy can be optimal not only as a protection mechanism, but also from the perspective of long-term utility.
|
| 2922 |
The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents
2608.28795
|
cs.AI
|
Achint Mehta |
Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g. a linter, a boot probe, a shell, a screenshot tool. We call this set the agent's verification surface. This study asks whether increasing only that surface...Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g. a linter, a boot probe, a shell, a screenshot tool. We call this set the agent's verification surface. This study asks whether increasing only that surface, with everything else held fixed, produces a matching growth in the quality of the software the agent ships. We built a minimal coding agent whose tool list is the single controlled variable and used it to implement 1,116 web applications across six models and eight tool configurations. A condition-blind human graded every application against a frozen rubric, and automatic probes stress-tested the API-observable behaviors. Verification's cheapest benefit arrives first, which is to make sure that the application comes up. Without any tools, about one build in seven fails to launch at all and a single boot probe removes nearly all of these failures at roughly 35 percent of a full shell's token cost, while the full shell multiplies the no-tools cost by 2.35. Screenshots help most where mistakes are visible (e.g. element placement, interaction), though even there the gain over a shell is modest and does not survive correction for multiple statistical comparisons. In cases where failures can only be measured rather than seen, such as keeping scrolling smooth over a 100,000-row list, screenshots add nothing. A verification tool improves the output artifact only where its reach covers the way the application actually fails.
|
| 2923 |
The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era
2608.28980
|
cs.AI
|
Kiyan Rezaee |
Specialized machine learning architectures encode structural assumptions (equivariance, permutation invariance, relational structure) that language-based foundation models lack by design. This review asks whether such assumptions can instead be acquired throug...Specialized machine learning architectures encode structural assumptions (equivariance, permutation invariance, relational structure) that language-based foundation models lack by design. This review asks whether such assumptions can instead be acquired through language, drawing on a corpus of 186 papers published between 2016 and 2026 across nine modalities: tabular data, graphs, time series, vision, chemistry, code, knowledge graphs, point clouds, and protein structure. Methods are organized into eight representational regimes, ranging from language-only prompting to fully specialized architectures, and are assessed not only on predictive accuracy but on whether structural information is preserved by the representation and computed by the model. Language-mediated models prove competitive in extreme few-shot prediction, discretized symbolic tasks, and textually annotated knowledge graphs. No evidence of general replacement survives once structure itself is evaluated: where direct tests exist, apparent parity is traceable to information asymmetry, benchmark contamination, or computation performed outside the language model. Across independent research communities, missing structural inductive biases are reintroduced through graph modules, structure-aware tokens, or specialized attention, indicating that specialization is relocated rather than eliminated. The review closes by specifying a falsifiable controlled experiment that would measure the remaining gap directly. The derived data supporting the findings of this review are openly available in the repository at https://github.com/kiyan-rezaee/language-vs-structure.
|
| 2924 |
Online Surrogate Repair: Decoupling High-Fidelity Feedback from Search Length in Closed-Loop Discovery
2609.07655
|
cs.AI
|
Xiaotang Feng, Philip Torr, Bruno Andreis |
Closed-loop AI scientists can generate candidate designs at low marginal computational cost, whereas reliable feedback may require wet-lab synthesis, characterization, or high-fidelity computation. Addressing this imbalance through custom laboratory automation...Closed-loop AI scientists can generate candidate designs at low marginal computational cost, whereas reliable feedback may require wet-lab synthesis, characterization, or high-fidelity computation. Addressing this imbalance through custom laboratory automation remains infrastructure-intensive and costly, while replacing new experiments with a fixed surrogate leaves persistent model errors that can be amplified by optimization. We propose \emph{online surrogate repair} (OSR), a closed-loop algorithm that uses sparse high-fidelity evaluations to update the surrogate throughout a longer agent search conducted primarily with inexpensive surrogate feedback. An acquisition rule selects which designs from the agent's accumulated proposals receive high-fidelity evaluation, and the resulting labels update the surrogate used in subsequent episodes. Across controlled synthetic environments, we demonstrate that improving global surrogate fit does not necessarily reduce maximum regret, whereas Q90-UCB and expected improvement (EI) substantially reduce regret by directing evaluations toward regions that determine the optimizer's decisions. On MADE, controls receiving high-fidelity feedback after every episode require $6.36$--$7.23\times$ more oracle queries to match Online EI under two LLM orchestrators and $10.27\times$ more under the non-LLM Chemeleon+MLIP workflow. Online surrogate repair introduces a novel third feedback regime between fixed-surrogate operation and high-fidelity feedback after every episode, separating the frequency of high-fidelity evaluation from the duration of the agent's search.
|
| 2925 |
Kalman Delta Networks: Uncertainty-aware Associative Memory
2609.07816
|
cs.AI
|
Ngoc Bui, Tinglin Huang, Rex Ying |
Linear attention enables efficient long-context inference by compressing token history into a fixed-size recurrent memory. This compression makes each update a trade-off between incorporating new information and preserving useful associations. Models such as D...Linear attention enables efficient long-context inference by compressing token history into a fixed-size recurrent memory. This compression makes each update a trade-off between incorporating new information and preserving useful associations. Models such as DeltaNet, Gated DeltaNet, and KDA predict write strength from the current token representation, without explicitly tracking uncertainty in the stored memory. Yet this uncertainty matters: a new observation should have greater influence when the existing association is uncertain and less when it is already well supported. We introduce Kalman Delta Networks (KDNs), a family of linear-attention models that explicitly track memory uncertainty to guide each update. By formulating associative memory as a linear-Gaussian state-space model, KDNs propagate both the memory estimate and its uncertainty, using the Kalman gain to balance accumulated evidence against the reliability of new observations. This formulation also recovers standard delta-rule updates by replacing tracked covariance with a token-predicted isotropic surrogate. To support hardware-efficient training and inference, we derive Diagonal KDN and Isotropic KDN, which retain one uncertainty value per key channel and per head, respectively. Their uncertainty updates admit associative scans with logarithmic parallel depth, requiring only $O(d_k)$ and $O(1)$ auxiliary state per head. Across controlled pretraining at 750M and 1.3B parameters, both variants consistently improve perplexity and mean downstream accuracy over the evaluated state-of-the-art linear-attention baselines.
|
| 2926 |
Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
2609.18587
|
cs.AI
|
Naveen Vakada, Mingyuan Li, Shaoxiong Ji |
Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-...Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as rewards and optimizes only ~100K bias parameters while keeping the pretrained backbone frozen. On MATH-500, our approach reaches 76.67% accuracy with Qwen2.5-7B, slightly exceeding our own labeled bias-steering reproduction while optimizing 76,000x fewer parameters than full-parameter TTRL. The same training procedure improves performance across vision-language and audio reasoning tasks, including MathVista, AI2D, LogicVista, and MMAU. We further show that the learned steering vectors transfer to 4,500 held-out MATH problems, indicating that the adaptation is not limited to the problems used during test-time optimization. Finally, we analyze why this highly restricted adaptation can work, showing that majority-vote reliability improves with rollout consensus and that bias subspaces with greater accessible gradient energy exhibit stronger downstream trainability. These results demonstrate that substantial test-time adaptation can emerge from optimizing a tiny bias-only subspace using entirely label-free rewards.
|
| 2927 |
NeuroECG: ECGFounder-Based Deep ECG Representation for EEG-Free Neurological Prognostication After Cardiac Arrest
2609.18891
|
cs.AI
|
Jiajun Gao, Yi Zhao, Chenyang Xu, Yuxi Zhou, Hao Wang |
Neurological prognostication after cardiac arrest commonly relies on electroencephalography (EEG). However, EEG demands high clinical resources. Bedside electrocardiography (ECG) is standard and low-cost. Yet, its value for predicting neurological outcomes rem...Neurological prognostication after cardiac arrest commonly relies on electroencephalography (EEG). However, EEG demands high clinical resources. Bedside electrocardiography (ECG) is standard and low-cost. Yet, its value for predicting neurological outcomes remains underexplored. In this study, we propose NeuroECG, an ECGFounder-based deep representation framework for EEG-free auxiliary prognostication. NeuroECG adapts a pretrained ECG foundation model via task-specific fine-tuning. We implement a gradual unfreezing strategy on single-channel bedside monitoring ECG. Multiple ECG segments per patient are encoded into segment-level deep features. These embeddings are aggregated via quantile pooling (q = 0.24) and compressed using principal component analysis (PCA). Experiments on 412 ECG-available patients from the multicenter I-CARE database show that the adapted ECGFounder backbone achieves the best performance among ECG-only backbone baselines, with a test AUROC of 0.7333. We further combine the learned deep ECG representation with static clinical covariates. The proposed NeuroECG model achieves a test AUROC of 0.8077 and an AUPRC of 0.8970. These results support deep bedside ECG representations as a useful source of auxiliary prognostic information. Their integration with static clinical covariates improves prediction in an EEG-free setting. The source code is available at https://github.com/goddream66/NeuroECG
|
| 2928 |
Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision
2609.20820
|
cs.AI
|
Nitish Dashora, Douglas Chen, Idan Shenfeld, John Marangola, Pulkit Agrawal |
Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many approaches to policy memory involve compressin...Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many approaches to policy memory involve compressing historical information through expensive VLM queries in-the-loop to process only task-salient information. In this paper, we propose an alternative approach in which computationally intensive VLM queries are made during train-time to learn a lightweight latent memory that can be efficiently queried at deployment time. Our representation, which we call the \textbf{workspace token}, is trained by (1) using a VLM to identify current and historical information necessary for completing a task, then (2) distilling these into the workspace token using a set-reconstruction decoder loss. In both simulation and hardware, we show that the workspace token can be used as a drop-in replacement for observations during deployment, enabling policies to solve memory-intensive tasks without the need for VLM reasoning in-the-loop, in effect serving as a \textbf{latent harness} for distilling a stronger reasoning models ability to solve long-horizon tasks to a reactive robotic policy. We further demonstrate that the workspace tokens are not only more lightweight, but also lead to better policy performance compared to conditioning policies on explicit modalities like curated past image frames, motivating a ``latent'' approach to history curation and reasoning model harnesses more broadly.
|
| 2929 |
From Semantic Decisions to Feasible Trajectories: Self-Evolving LLM-Guided Optimal Control for Narrow-Space Parking
2609.24631
|
cs.AI
|
Zhengbao Yao, Yuanfu Luo, Kehan Xue |
Autonomous parking in nonconvex and narrow environments remains challenging. Although optimal-control methods can explicitly enforce vehicle dynamics and collision constraints, nonconvexity compromises solver robustness and can cause failures. Large language m...Autonomous parking in nonconvex and narrow environments remains challenging. Although optimal-control methods can explicitly enforce vehicle dynamics and collision constraints, nonconvexity compromises solver robustness and can cause failures. Large language models (LLMs) exhibit strong semantic reasoning capabilities, but directly generating dense trajectories makes it difficult to guarantee physical feasibility. We introduce SE-LLM-OCP, a unified framework in which LLMs make high-level discrete maneuver decisions, while an optimal-control module enforces low-level vehicle dynamics and collision constraints. Online, the LLM proposes sparse maneuver plans, decomposing the parking task into a sequence of short-horizon trajectory-optimization problems. A low-level solver then sequentially solves optimal-control problems. If the solver fails, the LLM aggregates failure evidence from the solver and validation stages to guide replanning. Offline, SE-LLM-OCP automatically evolves a structured decision-making knowledge base from scratch, driven by accumulated online failures. We validate our proposed framework in simulation on a car-like vehicle model and on a differential-drive robot. Our experimental results show that SE-LLM-OCP enables safer autonomous parking in narrow scenarios and demonstrates transfer of the same maneuver representation to a different kinematic platform.
|
| 2930 |
XPhysICS: Cross-Physical-Domain Threat Grounding for Industrial Control Systems Security
2609.30805
|
cs.AI
|
Sangshin Park, Jainta Paul, Lawrence Ponce, Md Raihan Ahmed, Mu Zhang |
Industrial control system attacks are usually documented in terms of the plant where they occurred: its sensors, actuators, process stages, and control logic. Yet many attacks express a more general physical pattern--such as suppressing flow, corrupting chemic...Industrial control system attacks are usually documented in terms of the plant where they occurred: its sensors, actuators, process stages, and control logic. Yet many attacks express a more general physical pattern--such as suppressing flow, corrupting chemical dosing, or driving a vessel toward overflow--that may also matter in a different plant. The challenge is deciding when such a threat remains meaningful on a new system rather than relying on similar component names or broad semantic labels. We present XPhysICS, a methodology for grounding documented cyber-physical threats onto a specific target system. XPhysICS converts source evidence into a provenance-linked description of what is manipulated, what physical consequence is expected, and what observations the evidence calls for. Once this analyst-guided abstraction, its vocabulary and schema version, and a target contract are fixed, XPhysICS applies deterministic grounding checks. An accepted result can be represented as a validation slice that records the mapped roles, signals, dependencies, and context intended to support later evaluation. We study 83 threat abstractions across continuous-process and manufacturing sources using separate evaluation denominators. The continuous-process study evaluates 78 abstractions against target contracts spanning water treatment, water distribution, hydropower, and chemical processes. Selected cases are exercised through controlled perturbations of simulator-role signals. We also test compatibility with several analysis styles, including the released upstream GeCo implementation, and conduct a three-objective, one-target realizability study using a paper-derived search reproduction. Across these evaluated settings, the results support treating explicit target checks and traceable evidence as separate from semantic similarity alone.
|
| 2931 |
Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation
2609.32292
|
cs.AI
|
Xuekang Yang, Lu Chen, Shuang Luo, Jialing Zhu, Qi Zhang |
Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow...Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify an executable affordance pose and recover from accumulated action errors. We propose PACE (Preference-refined Affordance-Conditioned Execution), a supervised local execution module that augments frozen zero-shot semantic planners for reliable cross-floor navigation. PACE grounds transition-related semantics into a long-horizon, agent-centric traversable affordance pose and conditions short-horizon action generation on this spatial target, thereby aligning semantic goals with physical execution. We further post-train PACE through failure-aware preference refinement using rollout-derived pairs that contrast normal or recovery behaviors with deviation-amplifying behaviors, thereby improving closed-loop correction. We integrate PACE into six open-source zero-shot VLN navigators and demonstrate consistent improvements on the cross-floor subsets of R2R-CE and RxR-CE, increasing the average success rate from 16.35% to 27.65% and from 4.76% to 12.06%, respectively. Real-world experiments further demonstrate PACE's applicability in unseen environments, highlighting the potential of traversable affordances to bridge semantic intent and reliable embodied behavior.
|
| 2932 |
DRAM: Delta-rule Recurrent Associative Memory for Robot Manipulation Policies
2609.32453
|
cs.AI
|
Xinyu Zhao, Yixiang Shan, Tao Yang, Runyu Lei, Yiming Zhao |
Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed ...Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation windows, which substantially increase inference cost, or rely on pre-defined semantic features, which limit task generality and may also require the retraining of the backbone to adapt to the memory. We introduce DRAM (Delta-rule Recurrent Associative Memory), a plug-and-play memory module that can be attached to a wide range of pretrained robotic policies, endowing them with long-horizon memory without architectural modification or backbone retraining, requiring only task-specific post-training of the memory module and action expert. DRAM maintains a fixed-size associative memory using gated delta-rule linear attention, with a modified update that incorporates all tokens within each frame in parallel. An architecture-agnostic readout integrates historical context into action prediction across different policy architectures. Experiments show that DRAM consistently improves frozen pretrained policies over short-context baselines and alternative compact memory designs, validating its effectiveness as a fixed-size, post-hoc memory module trained with the backbone frozen.
|
| 2933 |
ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations
2609.32868
|
cs.AI
|
J. Paul Liu, Uthpala Herath, Andrew Petersen |
Traditional scientific computing requires researchers to translate computational intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler state and application logs. We present ASCEND (Autonomous Scie...Traditional scientific computing requires researchers to translate computational intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler state and application logs. We present ASCEND (Autonomous Scientific Computing Engine and Novel Discovery), an AI-powered agent interface that runs the agent on the researcher's own laptop, reaching Slurm-managed clusters and a GPU workstation over a multiplexed authenticated connection, with site-specific execution policies checked by locally executed tools; the language model is hosted remotely and holds no credentials. No facility-scale service is required: an account on each resource is sufficient, and the public installer lets users link additional Slurm clusters or workstations of their own. We report four recorded cases: (1) the agent closed a failure-recovery loop on a planted tensor-device fault, submitting, diagnosing, repairing and resubmitting with job-level artifacts preserved; (2) it reproduced the published evaluation of a weather-forecasting model from the author's released forecasts, agreeing with the published curves to 2.1% (z500) and 2.4% (t850) while identifying a unit discrepancy in the paper's prose and an initialization-field discrepancy in its released data; (3) it parallelized a released 12,693-line geophysical solver under a bit-for-bit identity requirement, reducing wall-clock runtime from about twelve hours to about two; (4) that requirement exposed two instances of undefined behaviour in the published solver, both repaired and reported upstream. Separately, a pre-specified evaluation of the policy layer found the deployed validator rejected 29 of 30 constructed violations and held the remaining one for approval, while denying 3 of 14 legitimate requests. Autonomy was exercised under author supervision; an end-to-end recovery benchmark remains outstanding.
|
| 2934 |
AquaWAM: A Dynamics-aware World Action Model for Underwater Embodied Agents
2609.33299
|
cs.AI
|
Cunhao Zhu, Yifeng Wang, Dongliang Xu, Yunzhong Hou, Yue Yao |
World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environment. However, underwater robots are usua...World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environment. However, underwater robots are usually subject to passive dynamics, such as inertia, buoyancy, hydrodynamic drag, and persistent drift, which can continue to affect the vehicle even after an action is completed. Existing WAMs, which primarily predict action-conditioned visual observations, are not explicitly designed to capture such passive motion dynamics. In this paper, we present AquaWAM, the first World Action Model designed for underwater embodied agents. Instead of predicting future images, AquaWAM models both action-conditioned and passive physical dynamics, including the thruster dead band, the inertial glide that outlasts each command, and ambient currents. Specifically, it senses through the DVL, IMU, pressure sensor and joint encoders, while cameras supply only semantics for understanding goals and target pose. By modeling compact navigation states rather than high-dimensional visual observations, AquaWAM substantially reduces the model size and computational cost compared with conventional WAMs. Experimentally, AquaWAM achieves a 72.6% task success rate across 20 underwater tasks on the USIM benchmark, outperforming existing methods while making action decisions 2.7x faster than U0 on an NVIDIA Jetson AGX Orin. Our model also remains effective when some onboard sensor measurements are unavailable. For example, without DVL velocity measurements, our method still achieves a 61.6% success rate, compared with 39.4% for U0.
|
| 2935 |
Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
2609.33401
|
cs.AI
|
Yixuan Liu |
Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models select from predefined answers and report probabilities that software can use to allow, block, or review in...Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models select from predefined answers and report probabilities that software can use to allow, block, or review inputs, but the reliability of these automated decisions remains unclear. We evaluate Jev, Laya, Decider, and Bespoke Nimble against specialized classifiers and language-model judges, examining decision accuracy, probability calibration, and selective automation. We draw the following conclusions. (1) Strong overall performance and favorable aggregate calibration can hide failures concentrated in particular attack groups, including attacks classified as safe with high confidence. (2) The evaluated adapted configurations do not consistently improve classification over their base models across tasks. (3) Under the strictest evaluated error limits, the policies allow few inputs automatically, and separate allow and block thresholds increase automation mainly through more blocks. Passing confirmation does not ensure that these limits hold on test. (4) Judges can detect attacks missed by another model, but may also falsely flag more benign inputs and share the other model's high-confidence errors. These findings support evaluating model accuracy, probability calibration, and the resulting allow/block/review decisions together.
|
| 2936 |
Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation
2609.34182
|
cs.AI
|
Wenqiao Li, Qianyou Zhao, Jiawen Hao, Xuezhou Zhu, Tengyu Liu |
Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more ...Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.
|
| 2937 |
CoSec: Benchmarking Agent Security in Communities
2609.34790
|
cs.AI
|
Hao Chen, Wenhui Dong, Ye Chen, Jiezhi Yao, Chenbo Xia |
LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must comple...LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks and prevent unauthorized disclosure of protected information. Existing evaluations do not fully examine these risks in agent systems. We introduce \textbf{CoSec}, an executable benchmark for evaluating privacy and authorization enforcement in LLM agent systems operating within and across communities. CoSec contains 208 canonical scenarios spanning fixed and evolving boundaries, protected information belonging to the agent owner or other participants, and attacks through dialogue, environmental content, persistent memory, and composed workflows. CoSec executes complete agent systems with persistent sessions, memory, files and tools. It verifies information flows against the active authorization state using execution traces and artifacts. Across harness and model configurations, agents frequently complete benign tasks but violate privacy and authorization boundaries. Privacy behavior varies across harnesses, attack surfaces, and community states, revealing how memory, files, tools, and workflows can carry protected information beyond its authorized scope. These findings show that task utility does not imply privacy or authorization compliance and that authorization in community settings remains an unresolved security challenge for persistent LLM agents.
|
| 2938 |
MCP Error Messages Written for Developers Hurt the Most Capable Agents Most
2609.35381
|
cs.AI
|
Xiaonan Xu, Wenjing Wu |
Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 wi...Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.
|
| 2939 |
F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement
2609.35575
|
cs.AI
|
Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu |
The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world...The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.
|
| cs.CL 620 papers | ||||
| 407 |
ChestPheNoT: Deployable, Auditable Label-Status-Evidence Extraction from Radiology Reports
2609.31629
|
cs.CLcs.LG
|
Kai Yu, Chenyu Zhu, Zaifu Zhan, Meijia Song, Min Zeng |
Structured phenotype extraction from radiology reports supports cohort construction, quality auditing, and clinical analytics, but practical deployment requires local inference and auditable predictions, while expert annotations remain scarce. Conventional lab...Structured phenotype extraction from radiology reports supports cohort construction, quality auditing, and clinical analytics, but practical deployment requires local inference and auditable predictions, while expert annotations remain scarce. Conventional labelers provide structured findings and assertion states but no supporting evidence, while API-hosted large language models may be unsuitable when clinical text cannot leave institutional infrastructure. We present CHESTPHENOT, a compact 0.5-3B language model that jointly extracts finding labels, three-class status (present/absent/uncertain), and verbatim supporting evidence spans. CHESTPHENOT is trained using hybrid CheXbert+72B silver supervision followed by supervised fine-tuning and lightweight GRPO refinement. Across three human-annotated gold sets spanning in-distribution, cross-taxonomy, and cross-institution evaluation, the 3B model remains below its CheXbert silver teacher in distribution but is competitive under distribution shift, significantly surpassing CheXbert on cross-institution detection (+2.0 F1). Task-specific training also enables the 3B model to match or exceed substantially larger prompted models on most detection and status comparisons. For evidence-grounded extraction, over 99% of final evidence spans are locatable in the source report, and the 3B model achieves 47.5 auditable-F1, outperforming Qwen2.5-7B one-shot prompting by 7.6 points and approaching Qwen2.5-72B. These results demonstrate that locally deployable models can provide competitive and directly auditable radiology-report extraction without relying on external inference APIs. Code and the full extraction/judge prompts will be made available at https://github.com/yukkai/ChestPheNoT.
|
| 408 |
A literature-guided descriptor-based framework for filtering composition search spaces
2609.31650
|
cs.CL
|
Lei Zhang, Markus Stricker |
Scientific literature contains latent knowledge about materials behavior, but much of this knowledge is expressed through words, contexts, and recurring associations rather than explicit design principles. This raises a central question: how can large-scale sc...Scientific literature contains latent knowledge about materials behavior, but much of this knowledge is expressed through words, contexts, and recurring associations rather than explicit design principles. This raises a central question: how can large-scale scientific corpora be used for practical problems in materials discovery? Here, we present a literature-guided descriptor-based filtering framework for reducing composition search spaces. For a given performance metric, the framework selects two descriptors from a filtered vocabulary in a literature-trained word embedding model and uses the selected descriptors to construct a Pareto-based filter for candidate compositions. Across the evaluated performance metrics and composition search spaces, the framework filters out an average of 74.27\% of the candidate compositions, with an average best-value error of 1.93\% relative to experimental measurements. Compared with expert-chosen and random descriptors, our performance metric-dependent descriptors provide a more controlled balance between retained fraction and best-value error. These results show that literature-derived embeddings can support intuitive and reproducible filters for narrowing candidate composition spaces while preserving high-performing compositions.
|
| 409 |
Distributional sentiment modeling and anomaly detection for consumer complaint assessment
2609.31653
|
cs.CLcs.LG
|
Peiheng Gao, Chen Yang, Shimin Zhang |
Sentiment analysis is a common tool for converting unstructured text into quantitative signals in finance and risk management. Yet most applications reduce the output to a discrete polarity label or a single predictive feature, overlooking the distributional s...Sentiment analysis is a common tool for converting unstructured text into quantitative signals in finance and risk management. Yet most applications reduce the output to a discrete polarity label or a single predictive feature, overlooking the distributional structure of sentiment intensity in consumer complaint narratives. In this paper we treat negative sentiment in consumer complaints as a bounded continuous variable and study its full distribution rather than a single label. We score each narrative with a transformer classifier, model the scores with Beta distributions, and compare the fitted distributions of meritorious and non-meritorious complaints through the Kullback Leibler divergence and the squared Hellinger distance. The fitted distributions are then linked with dollar amounts and company response outcomes to construct anomaly diagnostics that flag complaints whose textual severity is inconsistent with the recorded relief. We find that the two groups have strongly overlapping distributions, so negative sentiment intensity is not a sharp classifier of outcomes on its own; combined with monetary and categorical attributes, it isolates unusually severe complaints for operational risk monitoring. Treating sentiment analysis as continuous distributional measurement, this study links sentiment extraction, bounded response modeling, and anomaly detection in a unified framework for consumer complaint assessment.
|
| 410 |
Parser, Chunking, and Embedding Interactions in Retrieval-Augmented Generation over Indian Government Regulatory Documents
2609.31660
|
cs.CL
|
Shubham Kumar Singh |
Retrieval-augmented generation (RAG) pipelines are typically assembled from independently-chosen components -- a document parser, a chunking strategy, and an embedding model -- yet these choices are rarely evaluated jointly, and evaluations that do combine the...Retrieval-augmented generation (RAG) pipelines are typically assembled from independently-chosen components -- a document parser, a chunking strategy, and an embedding model -- yet these choices are rarely evaluated jointly, and evaluations that do combine them are usually run on a single document or corpus. We present a controlled factorial study of 3 parsers, 3 chunking strategies, and 5 dense embedding models, together with a sparse BM25 baseline, evaluated against 800 question instances, each with one or more required evidence strings, with evidence strings automatically validated against source text and a 10% random sample manually reviewed, across four structurally distinct Indian central-government regulatory documents. We fit linear mixed-effects models with document-query-level random intercepts to the resulting 72,000-row result set, run Holm-corrected paired comparisons between matched dense and sparse configurations, and report clustered bootstrap confidence intervals for all 54 unique retriever configurations. We find that no single retriever family dominates across documents; parser and chunker choice interact significantly; MPNet-base is a consistent underperformer with a severe failure mode on table-derived questions; and the corpus exhibits a near-saturated evidence-preservation ceiling above 98%, indicating that retrieval differences are driven primarily by ranking quality rather than information loss during ingestion. We additionally report embedding-dimension and chunk-size/overlap ablations and an efficiency/quality Pareto analysis. We release our full evaluation harness, corpus manifest, and 800-question benchmark.
|
| 411 |
LLM-Guided Ontology-Driven Knowledge Graph Construction from Unstructured Text
2609.31663
|
cs.CL
|
Abdelhadi Belfadel, Maxence Gagnant, Joseph Kattan, Sana Tmar |
Ontology-driven knowledge graph construction from industrial text remains challenging due to the domain specificity of documents, the scarcity of annotated resources, and the complexity of ontology engineering workflows. This paper presents and investigates th...Ontology-driven knowledge graph construction from industrial text remains challenging due to the domain specificity of documents, the scarcity of annotated resources, and the complexity of ontology engineering workflows. This paper presents and investigates the applicability of an ontology learning pipeline that combines compact open-source Large Language Models (LLMs), reusable prompting strategies, and open knowledge bases to support the extraction, structuring, enrichment, and evaluation of knowledge from textual corpora. The approach is tested and evaluated on a private French corpus of power-grid incident reports, using locally deployable open-source LLMs ranging from 7B to 32B parameters. Starting from unstructured reports, the approach extracts entities and relations, generates RDF triples, constructs related OWL ontology, enriches it using external knowledge sources, assesses the quality of the ontology, and subsequently constructs a populated knowledge graph grounded in the resulting ontology schema. Experiments on 80 manually annotated private reports show that schema-guided prompting significantly improves extraction quality, while quantized models provide an effective trade-off between performance and computational cost. These results demonstrate the feasibility of transforming domain-specific industrial text into ontology-based knowledge graphs using locally deployed open-source LLMs, while supporting the generalization of the extraction process through reusable prompting strategies.
|
| 412 |
What Drives Dialectal Jailbreaks? An Ablation of Surface Form, Cultural Framing, and Strategy Banks
2609.31664
|
cs.CL
|
Qingyang Xu |
Recent work suggests that obscure language registers can weaken large language model refusal behavior, especially when paired with black-box prompt optimization. It remains unclear whether failures stem from non-standard surface form, culturally grounded frami...Recent work suggests that obscure language registers can weaken large language model refusal behavior, especially when paired with black-box prompt optimization. It remains unclear whether failures stem from non-standard surface form, culturally grounded framing, or optimization over an expressive prompt-strategy space. We study Chinese registers by extending a classical-Chinese red-teaming framework to Shanghainese and Cantonese and running a 36-cell ablation across surface forms, strategy-bank variants, and two target models. The main finding is corrective: dialectal surface form is neither necessary nor sufficient for high attack success. Non-optimized English, Mandarin, and naive dialect translations remain below 8\% attack success rate, whereas all conditions that retain an optimizer-controlled strategy bank reach 98--100\%. A culture-neutral generic strategy bank reaches the same ceiling at near-single-query cost, further indicating that strategy-bank expressiveness, rather than dialectal or cultural content, accounts for most of the observed effect. Dialect choice still affects query efficiency and response severity, and qualitative coding identifies recurring high-level mechanisms such as semantic glossing and procedural scaffolding. We qualify the absolute ceiling-level attack success rate by showing sensitivity to judge calibration and noting that the design does not fully separate one-shot strategy construction from iterative search.
|
| 413 |
Age-Adaptive Handwriting Reconstruction from an IMU-Based Digital Pen through Shared Representations and Domain-Specific Heads
2609.31666
|
cs.CLcs.LG
|
Florent Imbert (LUT), Yann Soullard (IRISA, UR2, SHADOC), Eric Anquetil (INSA Rennes |
Digital pens are widely used to capture handwriting on digital devices, enabling precise trace recording and enhancing human-computer interaction. However, most are bundled with tablets and lack cross-brand compatibility. Recent digital pens equipped with kine...Digital pens are widely used to capture handwriting on digital devices, enabling precise trace recording and enhancing human-computer interaction. However, most are bundled with tablets and lack cross-brand compatibility. Recent digital pens equipped with kinematic sensors have emerged, designed for use on any surface. This especially opens significant potential for supporting handwriting acquisition in classrooms. Handwriting reconstruction from such an IMU-equipped pen poses a challenge due to the significant variability in sensor signals between adults and children. Even when producing visually similar traces, variations in writing dynamics, motor control, pen holding, and user confidence introduce substantial discrepancies in the captured signals. Additionally, the high variability in children's handwriting requires collecting large amounts of data, which is not feasible to implement in a school environment at scale. Furthermore, models trained exclusively on adult data fail to generalize to children's handwriting, and conversely, models trained on children's data perform poorly on adult writers. This crosspopulation degradation highlights the need for a unified model that can be deployed directly on the pen, without any user-specific adaptation. To address this issue, we propose a cross-domain learning, using an original neural network architecture based on a Temporal Convolutional Network and multiple prediction heads. The model is designed to be robust across age groups by leveraging shared features while effectively handling variability induced by differences in graphomotor development. This approach aims to improve handwriting trace reconstruction from sensor data, where each domain benefits from additional data provided by the other domain.
|
| 414 |
An Evaluation of AI-Supported Evidence-Based Learning for Public Speaking Skill Development
2609.31676
|
cs.CL
|
Sashini Hettiarachchi, Shahbaz Siddeeq, Mika Saari, Pekka Abrahamsson |
Public speaking is an essential skill in academic and professional contexts, but it often causes anxiety. Although several AI-based speech coaching tools exist, they typically provide generic feedback and lack model speeches for learning. This study addresses ...Public speaking is an essential skill in academic and professional contexts, but it often causes anxiety. Although several AI-based speech coaching tools exist, they typically provide generic feedback and lack model speeches for learning. This study addresses these gaps by developing a system that incorporates speech context into evaluation criteria, enabling more tailored feedback. It also generates improved model speeches in both text and audio formats to support example-based learning. The system was evaluated in a pilot study with seven undergraduate students over one to three weeks. Participants used the system repeatedly and practiced the same speech at least three times. Results showed improvements in public speaking performance for all participants. Filler word usage decreased, and anxiety levels dropped by 15.3% to 34.4%. Participants found the context-aware feedback and revised speeches useful and confidence-building. These findings suggest that context-aware feedback and model speeches can enhance public speaking skills and reduce anxiety. Future studies could explore video-based analysis of physical behavior during speeches.
|
| 415 |
The Temporal Tug-of-War: Visualizing and Detecting RAG Conflicts in Diffusion Models via Trajectory Variance
2609.31684
|
cs.CLcs.LG
|
Sravan Karthick T, Pranav Darshan, Pranav A, Minal Moharir, Ivan P. Yamshchikov |
Retrieval-Augmented Generation (RAG) introduces a specific failure mode in discrete diffusion language models: when retrieved context contradicts parametric knowledge, the iterative denoising process becomes a visible battleground between competing knowledge s...Retrieval-Augmented Generation (RAG) introduces a specific failure mode in discrete diffusion language models: when retrieved context contradicts parametric knowledge, the iterative denoising process becomes a visible battleground between competing knowledge sources. We identify temporal semantic divergence as an observable for detecting these conflicts and introduce the Trajectory Variance Score (TVS), a simple and interpretable measure of this divergence. TVS computes the mean pairwise cosine distance of answer embeddings across independent stochastic denoising trajectories, capturing the temporal tug of war between parametric and contextual attractors. Requiring as few as two parallel inference runs, TVS is computationally lightweight. Across four diverse datasets (Synthetic, SciQ, PopQA, and CounterFact), a simple Logistic Regression classifier using TVS achieves $70.10\%$ accuracy and $0.7647$ AUROC on LLaDA. On Dream 7B, increasing the number of trajectories from two to five improves accuracy from $63.91\%$ to $69.62\%$. More complex sequential models provide only marginal improvements over the linear classifier. Evaluation across LLaDA and Dream 7B demonstrates that conflict-induced trajectory dynamics and their key properties transfer across distinct diffusion architectures.
|
| 416 |
Verification of PETSc with CIVL using LLM-generated ACSL contracts and deterministic driver generation
2609.31687
|
cs.CL
|
Hansol Suh, Jan H\"uckelheim, Stephen Siegel |
Parallel numerical libraries such as PETSc are widely used in science and engineering applications where wrong results can have costly consequences. Despite this, numerical libraries are rarely formally verified. One of the challenges is the need for an expert...Parallel numerical libraries such as PETSc are widely used in science and engineering applications where wrong results can have costly consequences. Despite this, numerical libraries are rarely formally verified. One of the challenges is the need for an expert to hand-write a specification and manually apply a verification tool, often requiring the development of a harness or driver, all of which can contain additional bugs that lead to false positives or false negatives during verification. With recent advancements in large language models (LLMs), it is tempting to generate such drivers and reference models automatically, but one-shot generation based on a simple prompt is brittle and leads to additional unverified code that needs to be audited. In this paper, we present an approach to use LLMs in a limited setting to generate a small, human-certifiable ACSL contract from the function's documentation, combined with a deterministic toolchain that supports a restricted ACSL profile and generates a driver that uses the CIVL verifier to check the implementation against the contract and, when available, an existing reference model. We demonstrate the pipeline end-to-end on three PETSc functions: MatAXPY (reference model already exists), MatAYPX (no reference model, so the certified contract is the sole oracle), and the non-compressing mode of MatFilter (no reference model, with more complex, conditional behavior). With this pipeline, we were able to discover a bug in PETSc's MatAYPX function that was previously undiscovered and had been present in the code since 1997.
|
| 417 |
Don't Repeat Yourself: Self-Supervised Fine-Tuning for Coverage
2609.31688
|
cs.CL
|
Eric Fithian, Kirill Skobelev, X. Y. Han |
In verifiable domains such as math and coding, finding one correct solution among many attempts can matter more than the pass rate of each attempt. Post-training can concentrate large language model outputs around a few modes, while increasing sampling tempera...In verifiable domains such as math and coding, finding one correct solution among many attempts can matter more than the pass rate of each attempt. Post-training can concentrate large language model outputs around a few modes, while increasing sampling temperature has limited effectiveness. We introduce Don't Repeat Yourself Supervised Fine-Tuning (DRY-SFT), a post-training method that increases output diversity and coverage: the probability of at least one correct solution among many attempts. DRY-SFT has two stages. First, for each problem, sequentially generate K solutions, showing the model all prior attempts and asking for a different solution. Second, fine-tune on each attempt independently, removing prior attempts from the context. The process uses no reward, verifier, or correctness filter. On HumanEval+, MBPP+, and DS-1000, DRY-SFT raises pass@100 by 10.8, 12.5, and 12.4 percentage points, respectively, at a small cost to pass@1. Structural diversity, measured by abstract syntax tree edit distance among passing solutions, rises significantly on all three benchmarks. DRY-SFT also solves 244 of 600 problems that the base model did not solve in the same 200 attempts. Across nine open-weight models, lower structural diversity of the base model significantly predicts larger DRY-SFT gains, indicating that the method is especially effective on more mode-collapsed models.
|
| 418 |
The Ongiini-Eval-OW Benchmark: A Concept Paper for the Planned Benchmarking of Machine Translation and Large Language Models on Oshindonga and Oshikwanyama
2609.31727
|
cs.CL
|
Sebastian K\"upers (Common Intelligence Foundation) |
Oshiwambo -- a cluster of mutually intelligible Bantu languages spoken by over a million people across northern Namibia and southern Angola, and the home language of roughly half of Namibian households -- has, to our knowledge, no published machine-translation...Oshiwambo -- a cluster of mutually intelligible Bantu languages spoken by over a million people across northern Namibia and southern Angola, and the home language of roughly half of Namibian households -- has, to our knowledge, no published machine-translation evaluation benchmark. Major commercial services (Google Translate, DeepL, Microsoft Translator), open multilingual MT models (NLLB-200, MADLAD-400), and the open Masakhane checkpoint collection all lack coverage of either standardised dialect, Oshindonga or Oshikwanyama. We announce Ongiini-Eval-OW, a planned 600-item English-Oshindonga and English-Oshikwanyama benchmark with native-speaker references from two independent translators, a 30-item inter-translator agreement set, an 11-tag phenomenon-tagged stratification (at least 30 items per tag), a deterministic 30% blind split, and a reproducible scoring protocol over chrF++, BLEU, and COMET-22, supplemented by a 50-item human-evaluation round. We document the empirical coverage gap, the dataset composition, the launch-leaderboard model matrix across American, European, and Chinese frontier and open-weight systems, and the contribution pipeline. The dataset is targeted for first public release in Q4 2026; this v1.0 concept paper announces the design and the call for participation. Data and code will be released under CC-BY-4.0 and MIT respectively.
|
| 419 |
Omni-IO Skills: Harnessing Your Agent Omni-Native
2609.31847
|
cs.CL
|
Yanlin Li, Mingyang Hao, Shengqiong Wu, Hao Fei, Mong-Li Lee |
General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability gr...General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability growth to costly model updates, while assembling specialist models and tools leaves unresolved how procedures, dependencies, intermediate assets, and cross-turn revisions should be coordinated. We present Omni-IO Skills, a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry. Multi-asset workflows are represented as Declare Execution Graphs, which schedule independent operations concurrently and register successful outputs for downstream and cross-turn reuse across replaceable execution backends. Its 27 Skills cover 38 representative tasks spanning seven artifact modalities and four capability families: understanding, generation, reasoning, and retrieval. On UniM-90, the harness raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%, while increasing relative Semantic--Quality Coupled Score from 26.99 to 74.94 and from 27.82 to 77.78, respectively; Strict Structure Score reaches 100.00 and 99.78. These results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.
|
| 420 |
Agents Can Use Base Models to Evade AI Detection
2609.31876
|
cs.CL
|
Bhuwan Dhingra, Danish Pruthi |
We show that coding agents equipped with a base language model can successfully assemble responses from its samples to evade detection. Base models have been shown to evade commercial detectors, however, prior "humanization" techniques rely on using these mode...We show that coding agents equipped with a base language model can successfully assemble responses from its samples to evade detection. Base models have been shown to evade commercial detectors, however, prior "humanization" techniques rely on using these models to paraphrase AI outputs over several iterations, which invariably results in semantic drift. In contrast, equipping coding agents to directly orchestrate the writing process by stitching text samples from a base model allows it to produce outputs that are coherent, task-specific and generally high quality. We find that Claude Opus 5 operating in a Claude Code harness effectively orchestrates a local 32B parameter OLMo-2 base LM and sacrifices little task accuracy across benchmarks spanning creative writing, factual grounding, health QA and instruction following, while using up to 90% base LM tokens. Responses constructed in this manner reduce the effectiveness of both post-hoc detectors (Pangram v4 detection rate drops from 77% to 24%) and soft watermarking applied a priori to the agent's generations (down to a simulated 10% detection at low FPR). While effective, this evasion requires a significantly larger number of input and output tokens from the agent, increasing the dollar cost per query up to 30x at API-pricing. Overall, this work demonstrates the effectiveness of a new class of adversarial attacks against AI text detection, and urges post-hoc detection providers to include outputs of base models in their training.
|
| 421 |
Transformer MLP Gate Thresholds Are Couplings to a Carried Reference Direction
2609.31956
|
cs.CLcs.LG
|
Olli Tuomi |
The corpus-mean direction of a transformer's residual stream is a component shared across all inputs, and is commonly removed by mean-centering before representational analysis. We present evidence that it is a functional component: the reference against which...The corpus-mean direction of a transformer's residual stream is a component shared across all inputs, and is commonly removed by mean-centering before representational analysis. We present evidence that it is a functional component: the reference against which the MLP gate population sets its operating point. In Phi-2, an exact decomposition of resting gate pre-activations shows that at mid-stack layers the resting inhibition of >99.9% of gates is carried by the coupling $w \cdot b_L$ to the carried mean direction, at 48-56$\times$ the explicit bias parameter, and the coupling is direction-specific: a random direction at matched norm orders the population's firing rates at Spearman $\rho \leq 0.09$ where the reference reaches 0.95. (That 0.95 is near-tautological on its own; the paper derives its null.) Causally, removing the stream's projection on the reference multiplies above-threshold firing by about 9$\times$, dose-monotonically, at 23-59$\times$ a norm-matched control; a random-initialised twin is flat, and replacement tests show that the direction carries the function and the magnitude does not. A direction-matched control makes the same point: noise injected along the reference costs 8-42$\times$ the same energy along a random direction. The decomposition replicates on four further families spanning both gate types (GELU with an explicit gate bias, bias-free SwiGLU), and the dose-response on two of them. Across an eight-model scan the mechanism is present in every GELU and SiLU family and absent only in OPT, where an opposing LayerNorm bias cancels the carried reference. Gate thresholds are implemented as couplings to a constant the network builds for the distribution it is reading, with the bias parameters contributing little; how much of that constant is carried in the stream and how much in parameters depends on the architecture.
|
| 422 |
IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages
2609.31967
|
cs.CL
|
Rajarshi Roy, Shobhit Banga, Jonathan Raiman, Supriya Paul, Bhaskar Singh |
Full-duplex voice agents must handle pauses, take turns, backchannel, and respond to user interruptions in real time. Full-Duplex-Bench evaluates these behaviors, but its English-only corpus and reliance on word-timestamped ASR and an English-prompted LLM judg...Full-duplex voice agents must handle pauses, take turns, backchannel, and respond to user interruptions in real time. Full-Duplex-Bench evaluates these behaviors, but its English-only corpus and reliance on word-timestamped ASR and an English-prompted LLM judge make it difficult to extend to Indian languages. We introduce IndicFDB, which extends it to ten languages spoken in India with 12,350 samples, nearly 17 times as many as the original. We address three challenges: finding conversational events in multilingual speech, evaluating their timing without reliable word-level alignment, and judging responses across languages. We mine pause handling, turn taking, and backchanneling samples from roughly 50,000 hours of channel-separated conversations using voice activity detection (VAD), and construct human-validated synthetic user interruption samples. Language-independent VAD heuristics evaluate timing, while an open-weight transcription and translation pipeline converts responses to English for LLM ratings of relevance and quality. Across seven voice agents, commercial APIs show unexpectedly consistent behavior across languages but are either fast or robust to pauses, never both, while monolingual open full-duplex models expose further tradeoffs among backchanneling, response quality, and latency.
|
| 423 |
Extraction of clinical findings from mammography and breast ultrasound reports: a comparison between specialists and Artificial Intelligence
2609.31974
|
cs.CL
|
Lorenzo Farias, Hanna Reckziegel, Daniela Duarte da Silva Bagatini, Daniel Schulz, Gabriela de Andrade Monteiro |
Breast cancer is the leading cause of cancer-related death among women in Brazil, and the time between the request and the release of mammography reports directly influences adherence to screening, making the agility in processing these reports a critical fact...Breast cancer is the leading cause of cancer-related death among women in Brazil, and the time between the request and the release of mammography reports directly influences adherence to screening, making the agility in processing these reports a critical factor for early diagnosis. In this context, this study compares the performance of a Large Language Model (LLM) with manual extraction performed by a team of health researchers in identifying clinical findings from mammography and breast ultrasound reports written in Brazilian Portuguese. Named Entity Recognition (NER) was applied through Prompt Engineering using a few-shot strategy, employing the Gemini 2.5 Flash model, selected from preliminary exploratory tests with four candidate models. The Gemini 2.5 Flash model demonstrated the best performance, achieving a Macro F1 of 0.91 and a Micro F1 of 0.98. The subjective validation, in which 29 exams of different formats were evaluated by health researchers using a Likert scale, yielded an agreement index of 93.1%. The model outperformed human extraction in overall Macro F1 (0.91 vs. 0.72), as in four reports the model correctly identified information that had been omitted or incorrectly recorded during manual extraction, demonstrating its potential as a complementary verification tool alongside specialists. The results confirm the hypothesis that LLMs, when instructed through Prompt Engineering, can achieve performance comparable to or superior to manual extraction by health professionals.
|
| 424 |
Toward Embedding-Based Psychometrics: Structural Modeling of Assessment-Item Semantics With Contextual Scores
2609.31976
|
cs.CL
|
Jinsong Chen, Shi-Ting Chen |
Contextual scores represent assessment items through their similarities to reference words in an external corpus. We examine the semantic structure of scores for 40 TIMSS mathematics scored units using a partially specified two-step factor procedure. A search ...Contextual scores represent assessment items through their similarities to reference words in an external corpus. We examine the semantic structure of scores for 40 TIMSS mathematics scored units using a partially specified two-step factor procedure. A search across factor counts identifies a persistent seven-group structure under the featured construction. Subsequent comparisons consistently favor a general dimension alongside group associations, although individual group memberships remain sensitive to some specification choices. Item examples distinguish recurring, cross-domain, sensitive, and imposed associations. Simpler and unrestricted references clarify the contribution and limits of the anchored representation: it improves on a single factor but does not achieve the lowest working Bayesian information criterion (BIC). A separate response benchmark compares three initial Q constructions and their Hull-PVAF revisions under higher-order and saturated attribute distributions. Among these diagnostic models, BIC favors the official content framework and the Akaike information criterion (AIC) favors its direct four-factor augmentation, but a matched unidimensional two-parameter logistic model has lower AIC and BIC than all twelve conditions. These findings support a conditional semantic representation while limiting direct diagnostic interpretation. We discuss learned text-assisted response calibration as a prospective application requiring a larger calibrated item bank and independent evaluation.
|
| 425 |
Communication between Frozen Large Language Models via Prompt Optimization in a Referential Game
2609.31989
|
cs.CLcs.LG
|
Vivek Anand, Muthu Chandrasekaran, Shiva Chaitanya |
We study communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints. The two play a referential game: one sees an object and describes it in a short fixed-length message over...We study communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints. The two play a referential game: one sees an object and describes it in a short fixed-length message over a small alphabet; the other must pick that object out of a candidate set. Neither model's weights are updated. Each agent's prompt is rewritten by an isolated prompt optimizer whose reflection model reads that agent's scored interactions. In the positional setting, optimized prompts carry a shared code that generalizes to held-out objects above a measured no-codebook baseline, including when the memory window is removed. In a second setting, independent per-letter blocks no longer fit within the message, although a whole-object place value code does. The base system fails to establish reliable communication: the sender struggles to retain an injective rule, and the receiver has too few confirmed examples in view. A sender collision penalty, retention of successful interactions, and sequential optimization enable successful place value communication in some runs. Outcomes vary across runs and reflection models. In successful runs, the protocol is written into the optimized prompts, where it can be read and audited directly.
|
| 426 |
Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents
2609.31995
|
cs.CL
|
Jihan Yao, Sihan Zeng, Shangbin Feng, Zhiyuan Fan, Banghua Zhu |
Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextua...Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through behavioral evidence in a trajectory prefix. CER synthesizes adaptive rubrics specific to the current task and stage through experiences summarized from related historical tasks. In test-time scaling on SWE-bench Verified, CER improves RM@8 over the strongest baseline by 4.2 percentage points (pp) on Nemotron 3 Ultra and 2.0 pp on Qwen 3.6 27B; on Nemotron, it takes only 15.3% tokens to match the best baseline performance. In RL training experiments, CER exceeds full-rollout TMax by 1.9 pp while using 52.7% fewer online policy-and-judge tokens. Together, CER provides an interpretable, efficient, and dense evaluation method for long-horizon coding agents.
|
| 427 |
Quantization Thresholds Replicate, Failure Modes Do Not: A Three-Model Study of Agentic Tool Use in Polish from 8-bit to 2-bit
2609.32042
|
cs.CL
|
Jakub Prejzner |
We ask how GGUF quantization affects agentic tool use in Polish and whether the effects generalize across models. We introduce PolAgentBench, a deterministic benchmark with Polish prompts and English tool schemas: a 67-task main suite (15 adversarial probes, 5...We ask how GGUF quantization affects agentic tool use in Polish and whether the effects generalize across models. We introduce PolAgentBench, a deterministic benchmark with Polish prompts and English tool schemas: a 67-task main suite (15 adversarial probes, 52 hard-tier tasks) and a 46-task arithmetic isolation ladder. Three models span two axes of variation: Bielik-11B-v3.0 and its pruned, distilled child Bielik-Minitron-7B-v3.0 isolate model compression, and Llama-PLLuM-8B adds a change of pretraining family. Each is measured at six precisions, Q8_0 to Q2_K. Only the collapse threshold replicates. (1) All three models fall off a cliff between 3-bit and 2-bit (11B 0.716 to 0.045, 7B 0.463 to 0.149, PLLuM 0.224 to 0.015; paired McNemar p < 0.001 in each), across a fourfold capability spread and both axes. (2) Failure modes do not replicate: at 2-bit the 7B fails long (median 9.1k tokens, 4 steps) while the 11B mostly answers at the first step with a confabulated final answer (37 of 64 failures); PLLuM fails on content across precisions (71.8-92.0% of steps parse). (3) On the arithmetic ladder the unscaffolded rung is a floor, left standing by a rerun that states the no-tool rule; four explicit calls lift the 8-bit 11B from 1/10 to 9/10 and the 7B from 0/10 to 7/10 after format-only failures with the gold value are forgiven, an exploratory effect with eight distinct baseline inputs that does not survive multiplicity correction, while the order-trap arm separates the models at 8-bit (11B 6/6, 7B 0/6). (4) The Polish-versus-English gap is associated with degradation or with task family. We document four artifacts that shaped our conclusions (rounding-hostile gold values, strict answer typing, a no-tool rule the prompt never stated, priority-ordered failure labels), report affected results in strict and corrected form, and release the benchmark, trajectories and commit-stamped artifacts.
|
| 428 |
mu-bench: A Multilingual Utterance Transcription Benchmark
2609.32082
|
cs.CL
|
Andrea Li (UC Berkeley), Soham Ray (Sierra AI) |
Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a d...Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an AI banking agent in English, Spanish, Turkish, Vietnamese, and Mandarin, centered on form-field inputs such as names, email addresses, and confirmation codes. We release Utterance Error Rate (UER), an LLM judge of whether a transcript preserves meaning, calibrated against human raters, together with an LLM normalizer that makes WER comparable across providers' output formats. On 1,847 human-rated transcripts, UER agrees with annotators at $\kappa$ = 0.78, versus 0.53 for exact-match WER on normalized text. We rank six commercial providers on a public leaderboard; the best reaches 11.9% UER, and Mandarin is hardest for all six.
|
| 429 |
Using LMs to Model the Effects of Context and Coreference during Sentence Comprehension
2609.32119
|
cs.CL
|
Kohei Kajikawa, Lin Ai, Tatsuki Kuribayashi, Ethan Gotlieb Wilcox |
Language models (LMs) are often used as a tool to model human language processing. Recent studies suggest that severely restricting LMs' context window improves their fit to human psycholinguistic data by simulating human working memory constraints. However, i...Language models (LMs) are often used as a tool to model human language processing. Recent studies suggest that severely restricting LMs' context window improves their fit to human psycholinguistic data by simulating human working memory constraints. However, it is possible that this strict memory-decay approach overlooks humans' reliance on long-range structural representations, such as discourse structre. In this work, we systematically vary the context window size of GPT-2 across four large-scale naturalistic English reading-time datasets and observe a U-shaped relationship: Although restricted contexts (< 20 tokens) successfully capture local memory limitations, expanded contexts (500--1,000 tokens) ultimately yield the highest overall psycholinguistic fit. To investigate the mechanism driving this benefit, we conduct a counterfactual inference-time experiment that disrupts cross-sentential entity chains by pronominalizing repeated discourse entities. Obscuring these structural linkages significantly degrades the predictive power of larger context windows by 20% to 40%. Our experiments demonstrate that tracking long-range coreference relations is one important factor for the alignment between LM surprisal and human reading behavior, and approximate the extent to which human comprehenders use global discourse relations during language processing.
|
| 430 |
Typed Decision Models: An Early Evidence Audit and Evaluation Checklist
2609.32160
|
cs.CLcs.LG
|
Lijuan Tang, Yuemeng Zheng |
Typed decision models (TDMs) return probability distributions over caller-defined options without generating text. TypeSafe released Jev, a commercial typed decision model, on 15 September 2026, and a small body of evaluation and replication work appeared with...Typed decision models (TDMs) return probability distributions over caller-defined options without generating text. TypeSafe released Jev, a commercial typed decision model, on 15 September 2026, and a small body of evaluation and replication work appeared within days. We review 28 papers posted between 19 and 24 September and relate their findings to earlier work on label-probability classification, constrained decoding, reranking, calibration, and model cascades. In this early literature, the typed readout itself has not shown an independent accuracy advantage over comparable label-probability readouts. Jev's clearest gains are in latency and cost, while accuracy gaps remain on harder tasks. In practical deployments, confidence is often used to decide when to defer to a stronger model or a human. We use recurring weaknesses in these studies to derive a 14-item evaluation checklist for future TDM work. Because the evidence covers only the first nine days after the release of one hosted model, the review should be read as an early evidence map rather than a settled assessment of the model class.
|
| 431 |
ScopeIF: Improving Scope-Aware Precise Instruction-Following in Large Language Models via Graded Reward Modeling
2609.32189
|
cs.CL
|
Bosi Wen, Yilin Niu, Xiaoying Ning, Ying Zhang, Hongning Wang |
Precise instruction-following is a fundamental ability of large language models (LLMs), requiring their outputs to strictly satisfy objective constraints in input instructions. In complex application scenarios, these constraints often possess diverse scopes th...Precise instruction-following is a fundamental ability of large language models (LLMs), requiring their outputs to strictly satisfy objective constraints in input instructions. In complex application scenarios, these constraints often possess diverse scopes that govern specific response segments rather than the entire output. However, existing optimization methods often neglect constraint scope during data construction and rely on binary per-constraint rewards, yielding limited data diversity and sparse supervision for complex constraints. To this end, we propose ScopeIF, a novel training framework for scope-aware precise instruction-following. We first introduce a unified schema that factorizes objective constraints into three decoupled dimensions: Scope, Target, and Range. Grounded in this schema, we construct ScopeInstruct, a large-scale instruction dataset with diverse scope-aware constraints, and combine tool-grounded verification with graded reward modeling to quantify the violation degree of each constraint, providing dense supervision for policy optimization. Extensive experiments demonstrate that ScopeIF consistently outperforms existing methods, particularly on complex scope-aware constraints, while preserving general capabilities. Notably, it enables optimized Qwen3-4B and 8B models to rival or surpass strong frontier models such as Gemini-2.5-Pro and DeepSeek-V3.2, establishing an effective paradigm for advancing scope-aware instruction-following. Our code and data are available at https://github.com/thu-coai/ScopeIF.
|
| 432 |
The Judge Is Not Its Twin: Post-training makes a model's writing more predictable but barely moves its taste, as a judge, toward predictable writing
2609.32196
|
cs.CLcs.LG
|
Arman Nik Khah, Arvin Bahreini |
Language models are now routinely graded by other language models. If post-training makes a model's own writing more predictable, it may also teach the same model, acting as a judge, to reward predictable writing, so that progress on creativity would be invisi...Language models are now routinely graded by other language models. If post-training makes a model's own writing more predictable, it may also teach the same model, acting as a judge, to reward predictable writing, so that progress on creativity would be invisible to automated evaluation. We follow two open model families, OLMo-2 and Zephyr (7B parameters each), through their public training stages and measure every stage twice, as a writer of short stories and as a judge of pairs of stories. As writers, the models drift as feared: each family's fully trained model finds the stories of its base, supervised fine-tuned (SFT) and preference-trained (DPO) stages progressively more familiar, in all ten prompts, and a model from the other family finds the trained stories 7.0 to 9.2 percent (OLMo-2) and 24 to 25 percent (Zephyr) less surprising per token. As judges, they barely move toward predictable writing. Asked which story is better, a question every judge can use to tell a story from its own words in scrambled order, no trained judge's estimated tilt toward the more predictable story grows by as much as one point in the probability of picking it. A post hoc one-sided 95% upper bound on that growth is 2.7 points on an average pair, about the size of the untrained OLMo-2 judge's own tilt. Asked which is more creative, no trained judge's estimate favors the predictable story more than its base's does. Training instead strengthens a preference for longer stories when the question is creativity, and for one answer slot, and it breaks "more creative" as a question: trained judges asked it no longer reliably prefer a story to its scrambled words. A follow-up could not build pairs that differ in predictability but not in quality, because the routes that made this writer's stories less predictable also broke some of them, often enough to fail a quality floor set in advance.
|
| 433 |
Generalization and Memorization along the Learning Trajectory of Neural Language Models: A Geometric Account of Categorization
2609.32199
|
cs.CL
|
Wang Bojun, Holly Jenkins, Elizabeth Wonnacott |
We investigate how generalization and memorization develop along the learning trajectory of neural language models. Using controlled synthetic grammars, we examine both the geometry of representation space and model behaviour over training. We find that contin...We investigate how generalization and memorization develop along the learning trajectory of neural language models. Using controlled synthetic grammars, we examine both the geometry of representation space and model behaviour over training. We find that continuous regions of representation space not occupied by observed tokens become systematically structured from the earliest stages of learning, forming category-level geometric organization that supports generalization to unattested combinations. Generalization therefore emerges from the beginning of learning rather than only after extensive memorization. With prolonged training, larger models increasingly distinguish observed from unobserved grammatical combinations. At the same time, the continuous geometric structure supporting category-level generalization is gradually destructed. Together, these results suggest that neural language models initially learn through categorization-based generalization, followed by a gradual transition toward more exemplar-specific memorization.
|
| 434 |
A model of rational interlocutors: Unification of comprehension and production
2609.32216
|
cs.CL
|
Hanlin Wu, Zhenguang G. Cai |
Who we communicate with influences both our interpretation of their utterances and the design of our own. Such adjustment to the conversational partner is studied as speaker modeling in comprehension and as audience design in production, with the two literatur...Who we communicate with influences both our interpretation of their utterances and the design of our own. Such adjustment to the conversational partner is studied as speaker modeling in comprehension and as audience design in production, with the two literatures having developed largely separately. We argue that both adjustments express one rational computation and propose the rational interlocutor (RI) model, a computational account unifying comprehension and production. An interlocutor maintains a model of their partner, defined by three parameters: an identity parameter {\Pi} sets the messages and forms expected from the partner; a fidelity parameter {\Phi} sets how reliably messages and utterances map onto each other for them; a knowledge parameter {\Lambda} sets how knowledgeable the partner is believed to be. Comprehension and production are thus mirror-image modes of one computation over the partner model. Comprehension chooses the message the partner most likely intends to convey, weighing how well each candidate fits the utterance against how likely this partner is to mean it. Production chooses the utterance from which the partner will best recover the message, weighed against the effort of saying it. This explains why comprehenders appear to rely less on the forms produced by a linguistically less competent speaker, while producers tend to invest more effort in designing forms for them. We conjecture that perceived linguistic competence decomposes into two of these quantities: fidelity and knowledge. Their contrasting profiles across second-language (L2) adults, children, and artificial partners produce distinct and testable predictions.
|
| 435 |
OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?
2609.32227
|
cs.CL
|
Wenjun Peng, Xinyu Wang |
Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-p...Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of code edits within a fixed minimal scaffold and under bounded evaluator feedback and fixed resource budgets. The testbed uses two optimization regimes, surface obfuscation controls, calibrated references, held-out/stress splits, and diagnostics for degradation and exceptional failures, with LLM API cost reported separately from local evaluator wall-clock. The empirical study asks three questions: whether models can close the calibrated gap between a designated weak starter and an editable competent baseline, whether they can refine editable competent baselines without damaging them, and whether gains survive surface obfuscation controls. Across twelve frontier LLMs and five games, models improve designated weak starters more consistently than they refine editable competent baselines, with substantial variation across games and models. OptiArena provides a practical testbed for measuring bounded-resource algorithm optimization within the five-edit, fixed-scaffold setting studied here. Code is available at https://github.com/WJ-Peng/OptiArena.
|
| 436 |
KinyaMed: Seeds, Not Rows -- What a Corpus Requirement Written in the Wrong Unit Fails to Constrain
2609.32234
|
cs.CL
|
Marius Bayizere |
Triage decides who is seen first. Building an urgency classifier for patient-voice Kinyarwanda, we found our specification could be met without producing anything it was meant to secure. We report that, and the instruments that detect it, instead of a classifi...Triage decides who is seen first. Building an urgency classifier for patient-voice Kinyarwanda, we found our specification could be met without producing anything it was meant to secure. We report that, and the instruments that detect it, instead of a classifier. Designed for the four languages a Rwandan health centre receives, with every instrument per-language: sentences are authored in all four arms and rows generate in one, because the frame slots those three need do not exist. Our requirement asked for one million examples; generation produced them in 130 seconds. It fails four of the nine quality gates in that specification, six when rows are attributed to their source sentence. The binding gate counts distinct authored seed phrases, not rows: at our 165, no corpus passes at any row count. Two shortfalls follow and differ: 2,835 further sentences to pass the seed-count gate, 19,835 to reach the stated million rows, because a separate gate caps a seed at 50 rows. Rows come from a machine at 7,700 per second; seeds from clinicians. A row count constrains the cheap quantity, leaves the expensive one free, and so does not constrain quality at all. Two further negatives follow. An evaluation set of 17,942 rows built from nine distinct sentences supports no verdict: our gate, which counts distinct sentences, refuses 38 of its cells and reports nothing. A model trained on a corpus of uniform surface form changes its predicted urgency for 31.5% of inputs under capitalisation and 21.0% under a single typo, reported as measurement and not attribution. A model directory shipped without its tokenizer loads without error and answers its class prior on input it cannot read, with well-formed probabilities. No figure here is evidence of model quality; the contribution is the apparatus and the negative results it produced, reproducible from a clean clone except where marked NOT REPRODUCIBLE.
|
| 437 |
Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning
2609.32235
|
cs.CL
|
Zhuohan Wang, Haoran Ma, Tianyu Wu, Yuanlin Duan, Zichun Liao |
Large language models (LLMs) can solve every intermediate step of a multi-step math problem on its own and still fail the full problem, even when given a roadmap of the steps and all of their answers. We introduce OracleLadder, a diagnostic evaluation that loc...Large language models (LLMs) can solve every intermediate step of a multi-step math problem on its own and still fail the full problem, even when given a roadmap of the steps and all of their answers. We introduce OracleLadder, a diagnostic evaluation that locates where LLM math reasoning fails by giving the model increasing levels of oracle help. For each problem, a teacher model writes a fixed roadmap of intermediate sub-goals (milestones), and a deterministic symbolic verifier grades every answer. Testing the model with no help, with the roadmap, with the roadmap plus the milestone answers, and on each milestone alone sorts each failure into one of five reasoning gaps. On 354 NuminaMath problems and six models from 8B to 671B parameters (Qwen3, gpt-oss, Llama 3.3, DeepSeek-V3.1), the largest gap for every model is the composition gap, a stricter form of the compositionality gap. It covers 33-48% of problems, and 24-37% after removing problems that an LLM review flags as grading errors. Accuracy and milestone-help recovery rank the two strongest models differently, and two RLVR runs with similar accuracy gains move problems differently. The roadmap effect replicates on MATH500 and AIME 2024/25, per-problem recovery agrees for 83-87% of problems under an independent second teacher, and the help ladder carries over to code generation. We release the data, roadmaps, prompts, and code at https://github.com/slark-prime/OracleLadder.
|
| 438 |
LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models
2609.32264
|
cs.CL
|
Pavel Tikhonov, Elena Tutubalina, Ivan Oseledets, Dmitry I. Ignatov, Mikhail Seleznyov |
Language models can now prove theorems, but people still decide which problems to pursue. We ask whether a model's internal representations can help identify promising mathematical connections. We develop LANTERN, a fast, cost-efficient pipeline that uses a cl...Language models can now prove theorems, but people still decide which problems to pursue. We ask whether a model's internal representations can help identify promising mathematical connections. We develop LANTERN, a fast, cost-efficient pipeline that uses a classifier over pretrained-model activations to rank candidate relations, followed by staged filtering, hypothesis generation, executable verification, and analytical checking. Applied to the On-Line Encyclopedia of Integer Sequences (OEIS), LANTERN ranked 50 million pairs among 10,000 frequently referenced sequences and produced 62 verified relations between pairs without an existing OEIS cross-reference. A content screen retained 13 relations worth presenting; nine of these are informative or insightful, including four which are entirely novel to the best of our knowledge: none appears in the OEIS or in our targeted literature search. The entire end-to-end process including classifier training, candidate ranking, filtering and verification took under 8 hours.
|
| 439 |
Before Answering: Evidence Sufficiency under Size-Matched Memory Construction
2609.32269
|
cs.CLcs.LG
|
Joyanta Jyoti Mondal, Md. Shifatul Ahsan Apurba, Mridul Banik, Md Masud Al Mahmud, Ibne Farabi Shihab |
Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this co...Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memory size: on MuSiQue, a classifier that only counts paragraphs reaches an area under the ROC curve (AUROC) of $0.979$ for detecting unsafe memory, higher than the lexical estimator we initially evaluated. We propose a size-matched construction that provably removes this shortcut, and use it to study MemSafe, an estimator that cross-encodes the query with each memory unit and aggregates the units with a set transformer. Across three multi-hop question answering datasets and five seeds, MemSafe reaches $0.968$ and $0.983$ AUROC on MuSiQue and HotpotQA, $0.26$ to $0.39$ above a lexical baseline, while the third dataset, 2WikiMultiHopQA, is saturated. A frozen pretrained cross-encoder with a logistic head already closes $41\%$ of the MuSiQue gap between the lexical baseline and MemSafe. At the same time, MemSafe degrades more than a weak baseline on the unanswerable questions released with MuSiQue, reaches only $0.639$ AUROC on SQuAD~2.0, and needs several thousand clinical training examples before it outperforms a feature-based estimator. Used as a gate for a 7B reader, it reduces the error rate on answered questions from $0.850$ to $0.631$ at $5\%$ coverage, outperforming both reader confidence and, on average, the ground-truth integrity label, although a 7B LLM judge is the better gate at $10\%$ coverage. These results indicate that the way insufficient evidence is constructed matters as much as the estimator that detects it.
|
| 440 |
Supporting and Performing Culture from the Inside
2609.32281
|
cs.CL
|
Lea Frermann, Steven Bird |
Anthropology and related disciplines which study culture have often found it useful to consider their epistemological and methodological approaches in terms of emic versus etic. This is the distinction between the insider perspective, how the world is viewed b...Anthropology and related disciplines which study culture have often found it useful to consider their epistemological and methodological approaches in terms of emic versus etic. This is the distinction between the insider perspective, how the world is viewed by a member of the culture, versus the outsider perspective, the scientific cataloguing of cultural knowledge and practice. We adopt this framework to systematically analyse assumptions in recent research on culture in NLP along the full pipeline of task selection, data collection, system design and evaluation. We draw attention to the `culture of NLP' which shapes the research focus and approaches of our field, and suggest pathways towards a more culturally attuned AI.
|
| 441 |
TRAP: Understanding and Mitigating Privacy Memorization in Language Models
2609.32293
|
cs.CLcs.LG
|
Muhammed Ustaomeroglu, Ziyue Xu, Hanshen Xiao, Peter Cnudde, Guannan Qu |
Fine-tuning a language model on sensitive records can leave it able to reproduce them. We ask when this memorization arises and how to prevent it without knowing in advance which spans are sensitive. Our starting point is that most memorization scores and atta...Fine-tuning a language model on sensitive records can leave it able to reproduce them. We ask when this memorization arises and how to prevent it without knowing in advance which spans are sensitive. Our starting point is that most memorization scores and attacks share one statistical core: whether the model assigns a token more probability than some reference would. Taking as the reference a model trained on the complementary half of the same corpus gives the Target Reference Advantage (TRA), a per-token signal that separates what a model fit to a particular record from what it learned across records, and is cheap and differentiable. We then study what drives memorization during fine-tuning: it keeps growing well past the validation minimum, is larger on small datasets and at higher learning rates, and higher when the underlying task is harder. Early stopping removes much of it, but because it is chosen by aggregate validation loss it helps least for rare, hard-to-predict spans embedded in otherwise learnable text, which is exactly what sensitive information tends to be. We therefore introduce TRAP, a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference. On student essays with annotated personal information and clinical cases with patient identifiers, TRAP brings memorization near the level of an untrained model at little utility cost, where generic regularizers barely move and differential privacy gives up most of what fine-tuning bought.
|
| 442 |
Language Distances are Practical for Equitable Cross-Lingual Transfer
2609.32331
|
cs.CL
|
York Hay Ng, Razan Ahsan Rifandi, Aditya Khan, En-Shiun Annie Lee |
Cross-lingual transfer is strongly conditional on how the source language is chosen, but it is impractical to determine the best candidate source for every target language, especially for low-resource target languages. Language distances are widely used to ran...Cross-lingual transfer is strongly conditional on how the source language is chosen, but it is impractical to determine the best candidate source for every target language, especially for low-resource target languages. Language distances are widely used to rank candidate sources due to their correlation with transfer efficacy and applicability in resource-sparse settings. However, the reliability of distance-based rankers across tasks and resource levels remains underexplored. We therefore present the first equity-focused evaluation of paradigms for ranking source languages, studying resource-level inequality and task inequality across ten cross-lingual tasks and two multilingual models. While both inequalities are most pronounced for individual language distances and an English-always baseline, they are substantially reduced by training-free composite distances, and nearly eliminated by trained rankers. We further demonstrate the reliability of rankers using language distances compared to rankers using language model internals. Overall, we find that language distances provide a practical basis for equitable and performant transfer language selection. We recommend using trained rankers when task-specific transfer evaluations are available, and composite distances otherwise.
|
| 443 |
Gradients for Interventions and Activations for Detection: Targeted Feature Learning in Language Models
2609.32355
|
cs.CLcs.LG
|
Jonathan Drechsel, Steffen Herbold |
Model-internal features can be studied through both their ability to identify a specified concept and their causal effect when manipulated, e.g., through steering or weight editing. A prominent approach to feature learning is Sparse Autoencoders (SAEs), which ...Model-internal features can be studied through both their ability to identify a specified concept and their causal effect when manipulated, e.g., through steering or weight editing. A prominent approach to feature learning is Sparse Autoencoders (SAEs), which learn broad feature dictionaries whose relation to particular concepts is typically identified post hoc. However, many interpretability questions are instead hypothesis-driven and concern a concept specified in advance. We study this setting as targeted feature learning, where a single feature is constructed for such a predefined concept. We present a controlled comparison across three model signals (activation values, activation gradients, and parameter gradients) and two estimators (contrastive mean and a learned one-dimensional encoder-decoder), yielding six targeted methods, with CAA and GRADIEND as existing instances and four new methods covering the remaining combinations. We compare these methods against pretrained SAEs across 15 tasks and three language models, evaluating both detection and causal intervention. Across models, the strongest detection performance is achieved by contrastive activation value methods, whereas the strongest intervention performance is achieved by gradient-based methods. Overall, our results show that targeted feature quality depends jointly on the model signal and estimator, with detection and intervention capturing complementary properties.
|
| 444 |
PlurVA-LLM-2026 Shared Task Track-1: Pluralistic Value Alignment in LLMs via Multilingual Fine-Tuning and Threshold Calibration
2609.32382
|
cs.CL
|
Vihindi Kotalawala, Nevidu Jayatilleke |
We present our system for the PlurVA-LLM 2026 Shared Task Track-1, which focuses on pluralistic value alignment in the contexts of China, Indonesia, and Sri Lanka. For this resource-constrained track, we fine-tuned Llama 3.1 8B Instruct using 4-bit QLoRA. Our ...We present our system for the PlurVA-LLM 2026 Shared Task Track-1, which focuses on pluralistic value alignment in the contexts of China, Indonesia, and Sri Lanka. For this resource-constrained track, we fine-tuned Llama 3.1 8B Instruct using 4-bit QLoRA. Our approach combines option-permutation augmentation for Chinese data, annotator vote expansion for Indonesian data, and binary reformulation with SinhalaMMLU augmentation for Sri Lankan data. We further applied conditional threshold calibration to the predictions for the Sri Lankan data. The final system achieved accuracies of 0.785 for Chinese, 0.715 for Indonesian, and 0.916 for Sri Lankan, resulting in an overall macro-average accuracy of 0.805.
|
| 445 |
FA-Bench: A Benchmark for Word-Level and Phone-Level Forced-Alignment and ASR Timestamps Under Clean and Noisy Conditions
2609.32396
|
cs.CL
|
Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han, Yichao Zhou |
Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an o...Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices once and releases the code, splits, phone mapping, text normalization and scoring script, with results published periodically. Track 1 gives every aligner the reference transcript and Track 2 gives it a recognizer's output, on the same audio, clean and degraded four ways, with 21 open models and 9 commercial APIs under a unified protocol. We score every boundary of an utterance and check the two labels beside it, so a word the recognizer missed or invented is charged. Using a tolerance-based F1 as our primary metric eliminates the 9% to 14% score inflation that standard MAE causes on recognition-dependent systems in conversational speech. We then group boundaries by their position and how many adjacent words were recognized correctly, which shows where a system lost the score. We discovered systematic bias in how current systems time words, with Whisper about 150 ms early and several commercial ASR APIs over 50 ms late. Code and results are at https://github.com/olewave/fa-bench
|
| 446 |
SinBrief: A Hybrid Framework for Abstractive Text Summarisation of Sinhala Legal Documents
2609.32397
|
cs.CL
|
Minduli Lasandi, Nevidu Jayatilleke |
Legal document summarisation in low-resource languages presents significant challenges due to the scarcity of annotated data and the complexity of domain-specific terminology. This paper presents SinBrief, a hybrid abstractive summarisation framework for Sinha...Legal document summarisation in low-resource languages presents significant challenges due to the scarcity of annotated data and the complexity of domain-specific terminology. This paper presents SinBrief, a hybrid abstractive summarisation framework for Sinhala legal documents that does not require human-annotated training data. The proposed framework combines domain-aware word graph construction with neural sentence scoring to generate abstractive summaries from Sinhala legal text. Five sentence scoring models are evaluated within the framework: mBert, Llama 3.1, Falcon 7B, Laser, and a continually pre-trained Llama model domain-adapted to Sinhala legal text. The framework is evaluated on a Sinhala legal corpus using reference-free metrics, including Coverage, Density, Compression Ratio, SummaC, and Self-BertScore. Experimental results demonstrate that SinBrief produces summaries with lower lexical overlap than extractive baselines while maintaining factual consistency, demonstrating the viability of hybrid, largely annotation-free abstractive summarisation for low-resource legal NLP tasks.
|
| 447 |
Shared Worlds, Private Minds: Structured Memory for Long-Form Writing as World Creation
2609.32401
|
cs.CL
|
Qiuyu Tian, Xiaowen Gu, Hang Su, Jianghan Chao, Haojie Yin |
LLM agents that write long-form fiction need an explicit memory of the evolving storyworld to keep new events consistent with established facts. Such memory must keep heterogeneous narrative information distinct, integrate story developments across granulariti...LLM agents that write long-form fiction need an explicit memory of the evolving storyworld to keep new events consistent with established facts. Such memory must keep heterogeneous narrative information distinct, integrate story developments across granularities, and recover dependencies that a writing request leaves implicit. We present NarraWorld, a structured memory system for long-form writing that treats memory construction as world creation. From a shared evidence-grounded graph, NarraWorld derives four connected views: world facts, per-character beliefs, open developments, and hypothetical branches (possible-world continuations). Hierarchical aggregation with atomic closure consolidates events into scenes, plotlines, and plots, keeping each higher-level node traceable to its constituent source spans. For retrieval, planned reconstruction infers a query's dependencies from the current narrative situation and a preview of memory, then assembles the relevant records within a token budget. Across three writing benchmarks, NarraWorld achieves the strongest aggregate results. Its memory also transfers to situated role-playing and largely preserves recall on a general-purpose long-term memory benchmark, paving the way for agents that sustain coherent storyworlds across diverse narrative tasks.
|
| 448 |
Automatic Speech Recognition for the Basa\`{a} Language: A Low-Resource Approach
2609.32408
|
cs.CL
|
Sophie Gertrude Ngo Mock, Charles Moudina Varmantchaonala, Paul Dayang, Jean Michel Nlong II, Christopher Gies |
The rapid advancement of Artificial Intelligence (AI) and Natural Language Processing (NLP) has revolutionized the way humans interact with machines. Among the most impactful developments is Automatic Speech Recognition (ASR), which enables computers to conver...The rapid advancement of Artificial Intelligence (AI) and Natural Language Processing (NLP) has revolutionized the way humans interact with machines. Among the most impactful developments is Automatic Speech Recognition (ASR), which enables computers to convert spoken language into text. Systems such as those built on deep neural networks, transformer architectures, and self-supervised learning have achieved near-human performance for well-resourced languages such as English and French. Yet, these advances have disproportionately benefited a small fraction of the world's languages.
|
| 449 |
DualGuard: Dual-Mode Quality Control for Logic-Preserving Data Augmentation
2609.32431
|
cs.CL
|
Shenghao Li, Lin Zhao |
Large language models provide a practical way to generate augmented data for logical reasoning at scale, but a larger generation volume does not guarantee semantic, label, or logical reliability. Existing work has improved generation quality through generation...Large language models provide a practical way to generate augmented data for logical reasoning at scale, but a larger generation volume does not guarantee semantic, label, or logical reliability. Existing work has improved generation quality through generation constraints, candidate validation, filtering, and feedback-based revision; however, once a quality judgment is available, deciding whether a candidate should be retained, filtered, or repaired remains an important control problem. We propose DualGuard, a dual-mode quality-control framework for logic-preserving data augmentation. The first mode uses the current instance and candidate batch for selective retention, filtering, attribution, and targeted feedback. The second mode accumulates cross-instance execution records of augmentation actions on top of per-sample diagnosis and attribution, compares new executions against each action's own historical behavior, and supports retrospective anomaly inspection, targeted rollback, and bounded repair. Both modes share semantic verification and additionally use symbolic verification when a reliable logical form is available. Across seven downstream tasks in the Two-Stage Transfer setting, DualGuard achieves the highest Accuracy on five tasks and outperforms the no-augmentation BERT baseline on all seven. Controlled ablations further show complementary roles for Memory, Z3, and history-aware anomaly control.
|
| 450 |
Masking Frequent Tokens Sharpens Direct Preference Optimization
2609.32445
|
cs.CLcs.LG
|
Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu |
Direct Preference Optimization (DPO) aligns language models by optimizing over sequence-level sums of token-wise implicit reward differences. However, we identify a pervasive pathology in this formulation: a disproportionately small subset of high-frequency to...Direct Preference Optimization (DPO) aligns language models by optimizing over sequence-level sums of token-wise implicit reward differences. However, we identify a pervasive pathology in this formulation: a disproportionately small subset of high-frequency token types dominates cumulative sequence scores while appearing symmetrically across both preferred and dispreferred responses. Specifically, under canonical Qwen tokenization on Anthropic HH-RLHF, merely 69 token types account for $55.1\%$ of all response tokens and $85.9\%$ of within-pair shared token mass, exhibiting substantially lower preference-side specificity than the remaining vocabulary. This symmetric ubiquity induces gradient entanglement and dilutes the discriminative preference signal propagated through the objective. To resolve this issue, we introduce \emph{Anisotropic DPO} (\textsf{ADPO}) and its canonical realization, \emph{Frequency-Hard DPO}. Using a fixed, label-agnostic vocabulary mask, our method zeroes the implicit reward contribution of high-frequency response tokens while assigning unit weight to informative positions, thereby suppressing gradient interference without modifying preference pairs, discarding context, or introducing learned parameters. Here, \emph{anisotropy} designates non-uniform token-level objective weighting rather than representational geometry. Extensive empirical evaluations on AlpacaEval, MT-Bench, and Arena-Hard demonstrate that Frequency-Hard DPO consistently outperforms standard DPO across Qwen-2.5-7B-Instruct and Llama-3-8B-Instruct, establishing that selectively masking shared high-frequency tokens offers an effective, zero-overhead mechanism for robust preference alignment.
|
| 451 |
Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports
2609.32449
|
cs.CL
|
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz |
Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound ...Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demonstration environment changes. Across three open instruction models, wrong-source demonstrations move reports toward the source answer family, while explicit mechanism binding reduces this pull. Self-report benchmarks should include environment-shift invariance tests under fixed intervention before treating accuracy as evidence for an autonomous report mechanism.
|
| 452 |
What Should the Reflector See? An Empirical Study of Evidence in Reflective Prompt Optimization
2609.32452
|
cs.CLcs.LG
|
Xiaofan Zhou, Lu Cheng |
Reflective prompt optimization revises instructions using examples of a model's behavior, but which evidence the reflector should receive remains unclear. We study evidence composition, visibility of examples, candidate selection and domain-knowledge policy wi...Reflective prompt optimization revises instructions using examples of a model's behavior, but which evidence the reflector should receive remains unclear. We study evidence composition, visibility of examples, candidate selection and domain-knowledge policy within a single-parent Pareto-guided search. Using Qwen3.5-9B as both task model and reflector, we evaluate nine reflection strategies on five datasets. From these experiments, we find three distinct patterns. For performance improvement, Failures-only produces the largest mean test gain (+8.0 percentage points), while Balanced-mix and No-examples+Val share the best mean performance rank. For reflection effectiveness, No-examples achieves the best rank for improving sampled parents, yet yields only a 1.4-point mean test gain: local reflection success does not necessarily produce a stronger final prompt. For overfitting assessment, Failures-only and Balanced-mix share the lowest mean calibration-gap rank, while larger gaps on GPQA and IFBench show that calibration gains can overstate held-out improvement. This gap is a descriptive indicator, not a direct measure of overfitting. Together, these results show why reflection strategies should be assessed separately on final performance, parent improvement and calibration-to-test transfer.
|
| 453 |
Streamlined Reflective Evolution for Task-Adaptive Self-Refinement Pipelines
2609.32458
|
cs.CL
|
Xiaofan Zhou, Lu Cheng |
Reflective prompt optimization improves large language model (LLM) systems without updating model weights, but fixed architectures constrain how self-refinement is organized. We introduce Workflow-Designing Agents (WDA), a framework for streamlined reflective ...Reflective prompt optimization improves large language model (LLM) systems without updating model weights, but fixed architectures constrain how self-refinement is organized. We introduce Workflow-Designing Agents (WDA), a framework for streamlined reflective evolution of task-adaptive self-refinement pipelines. Starting from a minimal prompt, WDA jointly evolves stage instructions and their sequential structure. During evolution, we find that repeated revisions can accumulate redundant instructions in a single prompt. In WDA, we propose to address this problem with SPLIT, which redistributes these instructions across specialized stages. Three-example reflection and local screening guide selective search, while calibration scores guide Pareto admission and rollback of unhelpful trailing updates. The resulting pipelines are task-adaptive: their instructions and depth are learned from task data, then fixed for all test inputs within that task. We evaluate WDA on five benchmarks spanning knowledge, mathematical reasoning, multi-hop question answering, and instruction following. On Qwen3.5-9B, WDA achieves an average score of 51.24%, improving over the initial solver by 8.63 percentage points and the variant without SPLIT by 2.60 points. On GPT-4.1-mini, it achieves 49.00%, with corresponding gains of 5.62 and 3.69 points. These results support task-adaptive self-refinement as a complementary direction to broader agentic workflow search.
|
| 454 |
AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
2609.32472
|
cs.CL
|
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu |
Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a comp...Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work rewards a set by its aggregate rubric score, shifting the objective from ranking documents to composing sets. Yet that score is one scalar shared by every document in the set, so the supervision is sparse: a redundant document is rewarded with the rest whenever the set scores well, and a decisive one penalized with the rest whenever it does not; credit assignment leaves contributors indistinguishable from free riders. On-policy distillation could densify this supervision, but existing methods give every rollout the same fixed guidance, too prescriptive for strong rollouts and too abstract for weak ones. We therefore propose AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation. ATO draws three hint forms of increasing specificity from the policy's own frozen snapshot: the rubrics alone, a self-selector's sibling-set chosen under rubrics, and a self-reflector's reflection contrasting the rollout with that sibling-set; each rollout receives the form matched to its quality. Re-scoring that rollout under the hint-conditioned frozen teacher and the hint-free snapshot distills the hint's effect into a token-level advantage that complements the group-relative outcome advantage. Across ten benchmarks spanning RAG, deep research, and setwise evaluation, AdaTutoRank attains the best overall performance while issuing fewer retrieval calls.
|
| 455 |
PC-SubMax: Efficient Prompt Compression via Regularized Submodular Maximization
2609.32474
|
cs.CL
|
Ziyi Zhang, Shuang Cui, Haotian Zhang, Xiaoyu Wang |
While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle'' phenomenon. Selective prompt compression offers a model-agnostic approach...While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle'' phenomenon. Selective prompt compression offers a model-agnostic approach to alleviating these issues. However, methods based on fixed token- or sentence-level importance scores may overlook how content contributions change with the selected subset, limiting their ability to account for inter-sentence redundancy. Compression procedures that rely on autoregressive LLM scoring can also introduce substantial overhead. We propose PC-SubMax, a theoretically grounded framework that formulates selective prompt compression as regularized monotone submodular maximization under a knapsack constraint. The objective is $U(S)-\ell(S)$, where the monotone submodular utility $U$ combines information coverage, query relevance, and log-determinant diversity, and the non-negative modular penalty $\ell$ captures token cost. Through diminishing marginal returns, the objective evaluates each sentence's contribution relative to the selected content. To optimize this objective, we develop the Regularized Greedy+Max (RGM) algorithm, which deterministically returns a feasible set $Q$ satisfying $U(Q)-\ell(Q)\geq \frac{1}{2}U(O)-\ell(O)$, where $O$ is an optimal feasible solution to the regularized problem. RGM uses $O(n\kappa)$ value-oracle queries, where $n$ is the number of candidate sentences and $\kappa$ is the maximum feasible subset size. PC-SubMax uses encoder representations and avoids autoregressive LLM scoring during compression. Experiments across seven diverse benchmarks demonstrate competitive downstream performance with low compression overhead.
|
| 456 |
Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs
2609.32491
|
cs.CL
|
Jorryt de Jong, Stefan Moraca, Ettore Cesari, Lasha Abzianidze |
Large Language Models (LLMs) are highly capable of natural language reasoning and appear to store a great deal of lexical knowledge, but it is still unclear how much of this knowledge they actually use when reasoning, and whether they use it in the right way. ...Large Language Models (LLMs) are highly capable of natural language reasoning and appear to store a great deal of lexical knowledge, but it is still unclear how much of this knowledge they actually use when reasoning, and whether they use it in the right way. On the other hand, logic-based Natural Language Inference (NLI) systems provide transparent and formally grounded reasoning, but they need to be supplied with rich lexical knowledge to prove inferences beyond purely logical ones. In this paper, we evaluate whether LLMs can identify all lexical knowledge needed to solve NLI problems and how much this knowledge contributes to proof search in a logic-based NLI system. Our research focuses exclusively on structured lexical entailments (e.g., chinchilla$\sqsubseteq$small animal) as a proxy for structured explanations for NLI problems with an entailment label. First, we curate a dataset for a new task of explaining sentential entailments with a set of lexical entailments. The dataset is used to intrinsically evaluate LLMs on generating structured lexical explanations. Then, we use NLI as an extrinsic evaluation in a simple neuro-symbolic setting, assessing whether LLMs can supply sufficient lexical relations to LangPro, a natural-logic theorem prover for natural language. The results show that the proposed task remains challenging even for hosted proprietary LLMs, and that their contribution to theorem proving is moderate: generated relations are often only partially sound and may be tailored to the specific NLI problem rather than representing generally valid lexical knowledge.
|
| 457 |
Locally Sound, Globally Insufficient: The Local-Global Gap in Multi-Hop Reasoning
2609.32496
|
cs.CL
|
Bohao Chu, Hendrik Damm, Qianli Wang, Hui Wang, Shuning Zhang |
Reliable multi-hop reasoning requires more than locally supported steps: a trace can be sound at every reasoning step yet still fail to answer the question as a whole. We call this failure regime the local-global gap (LGG), in which the trace is locally sound ...Reliable multi-hop reasoning requires more than locally supported steps: a trace can be sound at every reasoning step yet still fail to answer the question as a whole. We call this failure regime the local-global gap (LGG), in which the trace is locally sound yet globally insufficient. Local soundness requires each step to be supported by the available evidence and preceding steps, whereas global sufficiency requires the reasoning trace to align with the question and establish the submitted answer. In a human-adjudicated diagnostic of 2,598 responses across three multi-hop QA benchmarks and three models, we find that the LGG occurs in every benchmark-model combination and accounts for nearly half of globally insufficient responses overall. However, conventional faithfulness verifiers that check claims against the evidence largely miss these failures: at thresholds retaining at least 95% of reliable traces, recall for LGG cases is substantially lower than that for locally unsound traces. To address these failures, we formalize three dependencies for reliable reasoning: evidence-to-step support, question-to-trace alignment, and trace-to-answer closure. Instead of post-hoc diagnosis, we introduce E-Closure to supervise these dependencies during training, combining generation supervision on supported original and counterfactual responses with bidirectional switching constraints. Averaged over three benchmarks and three backbones, existing fine-tuning baselines improve accuracy and local soundness over the base models, but at the cost of global sufficiency. E-Closure improves both: among all fine-tuned methods, it achieves the highest average accuracy (92.8%) and trace reliability (89.0%) while yielding the lowest LGG rate (6.2%).
|
| 458 |
Learning an Anchored Prompt Space for Continual Adaptation of Large Language Models
2609.32499
|
cs.CL
|
Rongguang Ye, Zhan Zhuang, Yichen Wu, Ming Tang, Kede Ma |
Continually adapting large language models requires acquiring new knowledge while preserving previously learned capabilities. Jointly adapting model parameters and task-specific soft prompts offers a promising solution, but faces two key limitations: historica...Continually adapting large language models requires acquiring new knowledge while preserving previously learned capabilities. Jointly adapting model parameters and task-specific soft prompts offers a promising solution, but faces two key limitations: historical prompts may become less effective as the model evolves, while their transferable cross-task relationships are not explicitly learned. We propose Learning an Anchored Prompt Space (LAPS), which preserves historical prompt effectiveness and learns relationships among task-specific soft prompts to facilitate positive transfer. LAPS first aligns historical prompts with the updated model through self-distillation. LAPS then constructs an anchored prompt space whose vertices correspond to learned task-specific soft prompts and whose intermediate geometry is shaped by learnable B\'ezier control prompts. Once this anchored prompt space is learned, LAPS identifies the best-performing prompt for each task on its validation set, allowing the optimized prompt to draw on knowledge acquired from observed tasks. Experiments on the TRACE benchmark across three Qwen3 model scales show that LAPS consistently outperforms distillation-based, prompt-based, and joint prompt--parameter adaptation baselines, improving average performance while reducing forgetting.
|
| 459 |
When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents
2609.32520
|
cs.CL
|
Yanjie Zhang, Bowen Cao, Zixin Chen, Yushi Sun |
LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce Inte...LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders. In a 627-case calibration, mean task score falls from 0.476 to 0.384 as dialogues contain more superseded and withdrawn information. Across eight models, the rate of fully correct solutions is significantly lower when the same final task must be recovered from an evolving dialogue rather than given directly in a single turn. We further introduce StateForge, which explicitly maintains the active requirements before generation. On General-Test, it improves mean task score from 0.367 to 0.467. Providing the ground-truth final state improves performance further but still does not recover single-turn performance, indicating that state-estimation errors explain only part of the gap. These results establish intent drift as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation.
|
| 460 |
Language as an Independent Information Layer: A Conceptual Model of Communication, Cognition and Decision-Making
2609.32556
|
cs.CLcs.LG
|
Anastasiia Alifanova, Elena Benderskaya |
Based on an analysis of the role of language in thought and communication, this article proposes a new concept for designing corporate knowledge bases. The concept integrates the probabilistic vector space of a corporate vocabulary, reflect-ing industry specif...Based on an analysis of the role of language in thought and communication, this article proposes a new concept for designing corporate knowledge bases. The concept integrates the probabilistic vector space of a corporate vocabulary, reflect-ing industry specifics, subject focus, terminology, and culture, with traditional ontological modeling. This combination enables the efficient extraction of knowledge from accumulated corporate documents while strictly accounting for specific business processes. Consequently, this concept bridges statistical and semantic (cause-and-effect) methodologies. Furthermore, analyzing the projec-tions of probabilistic spaces and causal relationships can help identify bottlenecks in business logic. As a dynamic system, language functions as a separate, inde-pendent layer within the overall information architecture. Introducing a dynamic component into the probabilistic space of word distribution allows it to be mod-eled as a multidimensional solution space for various problem formulations. In this context, input data defining the problem conditions serve as control parame-ters for dynamic transformations.
|
| 461 |
How to Reduce Whisper Hallucination
2609.32560
|
cs.CL
|
Husein Zolkepli |
Whisper is still what runs in production: one permissively licensed checkpoint, 99 languages, no per-language tuning. But it writes sentences nobody said. On 42 clips of pure room tone, whisper-large-v3 emits words on 61.9% of them and emits something on 100%....Whisper is still what runs in production: one permissively licensed checkpoint, 99 languages, no per-language tuning. But it writes sentences nobody said. On 42 clips of pure room tone, whisper-large-v3 emits words on 61.9% of them and emits something on 100%. The usual response is to distil a student from Whisper pseudo-labels, filtering the hallucinations out of the corpus first, which deletes the evidence while leaving the behaviour in place: the student is fitted to the subset where the teacher was right, and inherits a failure mode absent from its own training data. The second response, adding non-speech audio so the model learns to stay quiet, already ships and works, but teaches suppression without discrimination. The checkpoint best on every non-speech arm here is also the one that recovers the fewest genuinely spoken phrases and deletes 58% of repeated speech. The teacher has to be fixed, with both halves of the signal. Our benchmark scores both at once: 11,852 clips over eight arms, 6,267 synthetic positives, 8,296 clips of real audio that made a production model fail, and FLEURS in 58 languages. We collect 40,891 hallucination phrases in 100 languages, choose a text-to-speech system by measurement, and synthesise those phrases as positives, so the model meets the same text both as something to suppress and as something to transcribe. Across 33 matched pairs of fine-tunes differing only by those positives, adding them lowers word emission on real voice-free audio in 31 pairs and raises phrase recovery in 32; among the 30 pairs that hold English accuracy it is 30 out of 30. The best checkpoint takes hallucination on silence from 61.9% to 2.4% and words over real voice-free audio from 99.9% to 47.8%, while raising phrase recovery from 69.8% to 82.7%. The benchmark, lexicon and synthetic corpus are released at https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination
|
| 462 |
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
2609.32577
|
cs.CLcs.LG
|
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang |
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking di...Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
|
| 463 |
HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion
2609.32581
|
cs.CLcs.LG
|
Junxiang Qiu, Zhengsu Chen, Xinting Hu, Shuo Wang, Hengheng Zhang |
Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects in...Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects input semantics and upstream computation across depth, but standard routers do not explicitly use the routing distributions produced by preceding layers. We propose HERO-MoE, Historical Expert ROuting with Scale-Preserving Fusion, a routing framework that injects historical routing priors into MoE routers by reusing detached, dense routing distributions collected from preceding MoE layers. The key idea is simple: HERO-MoE preserves the original token-conditioned routing branch and adds a residual historical routing contribution before the standard softmax and top-$k$ dispatch. To stabilize this historical signal, HERO-MoE introduces a scale-preserving fusion mechanism that matches the magnitude of historical routing memory to the current hidden representation and accounts for the number of visible historical layers, without introducing an auxiliary routing loss or a fusion-specific tuning parameter. By reusing routing distributions already computed by preceding MoE layers, HERO-MoE improves training-loss reduction with modest end-to-end overhead. The resulting router remains compatible with standard sparse dispatch, including top-$k$ and group-limited routing, and can be inserted into existing MoE backbones with minimal architectural changes. Experiments on an approximately 8B-parameter MoE model with 0.5B active parameters, trained from scratch on 100B tokens, show that HERO-MoE reduces the final loss from 1.6393 to 1.6184, while peak memory and FLOPs increase by only 0.44\% and 0.64\%, respectively.
|
| 464 |
KV-Lingo: Learning KV-Cache Translators with Distillation
2609.32610
|
cs.CLcs.LG
|
Val\'erie Castin, Keitaro Sakamoto, Anastasiia Filippova, Jo\~ao Monteiro, Marco Cuturi |
Large language models represent context with a key-value (KV) cache. Caches are model-specific: for the same text, models with different architectures or weights produce incompatible representations. This makes it costly to switch models over a shared context:...Large language models represent context with a key-value (KV) cache. Caches are model-specific: for the same text, models with different architectures or weights produce incompatible representations. This makes it costly to switch models over a shared context: although the context has already been processed by one model, the incoming model must process it again to build its own cache. We introduce KV-Lingo, a method for translating the KV cache of a source model into one that can be read by a target model. KV-Lingo consists of a collection of linear maps, typically one per layer of the target model, that are applied independently on all tokens' key and value representations. We train these maps using distillation, minimising the divergence between the target model's predictions from its native cache and those from the translated cache. We consider several model pairs spanning multiple sizes and architectures, training one translator per pair on a generic text corpus. The resulting translators preserve strong downstream performance in both small-to-large and large-to-small transfers. Since a switch then costs a linear map and a single decoding step instead of a prefill, replacing re-prefill with cache translation reduces the time to first token after a model switch by 9.6x already on a 64-token prompt for Qwen models on an Apple M3 Ultra, and by up to 29x at 32k context length on an H100. These gains make KV-Lingo particularly useful for dynamic model routing: a context can be processed by one model and handed off to another only when needed, without re-prefilling the shared prefix. We finally show that KV-Lingo can be used for seamless model switching, staying close to re-prefill across repeated switches in our multi-turn evaluations.
|
| 465 |
The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models
2609.32617
|
cs.CL
|
Qingjia Huang, Yakai Li, Jianguo Wu, Qihang Zhou, Aimin Yu |
Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors su...Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors such as missing knowledge in training data, reasoning errors, or stochastic decoding, we uncover that post-training alignment itself is a primary driver of these errors, a phenomenon we call the \textbf{Alignment Paradox}. Across five model families evaluated on factual benchmarks, unaligned base models produce few high-confidence errors on long-tail factual queries, whereas instruction-tuned models multiply high-confidence errors ($p \ge 0.95$) by more than an order of magnitude (10$\times$ to 35$\times$). Layer-wise probing with the Logit Lens reveals that this overconfidence emerges in late layers, where wrong-answer margins expand past 4.0 points after remaining near zero across early and intermediate layers. These findings motivate limiting margin growth during post-training. We implement this principle through an entropy-dependent margin bound in direct preference optimization (DPO). In multi-epoch experiments with Mistral-7B, the bounded objective reduces high-confidence errors by up to 35.3\% relative to standard DPO while maintaining performance on evaluated general reasoning benchmarks. These results show that bounded margins mitigate confident hallucinations during post-training.
|
| 466 |
Does CoT-Pass@k Really Check the CoT? A Multilingual Mathematical Audit
2609.32622
|
cs.CL
|
Tar{\i}k Tuna Ta\c{s}alt{\i}, Burcu H\"udaverdi, David Semedo |
Pass@k measures whether a model reaches a correct answer under repeated sampling, but never how: a lucky guess counts the same as sound reasoning. CoT-Pass@k was proposed to close that gap, adding an LLM-as-judge that must assess a solution's reasoning chain b...Pass@k measures whether a model reaches a correct answer under repeated sampling, but never how: a lucky guess counts the same as sound reasoning. CoT-Pass@k was proposed to close that gap, adding an LLM-as-judge that must assess a solution's reasoning chain before it counts. Its value rests entirely on one assumption: that the judge catches flawed reasoning. That assumption has never been tested inside the metric that depends on it, and never outside English, though the metric's claims concern models used in many languages. We report the first audit of that verification step, run under the metric's own protocol on a multilingual suite of five mathematical benchmarks in English, Turkish and Portuguese, two of them natively written. We corrupt correct solutions with deterministic edits that damage the chain and the final answer separately. We observe that all three judges accept corrupted chains almost as often as clean ones. V4-Flash and Qwen3.6 reject a solution sharply only when its final answer is wrong and accept a wrong answer more readily when the chain agrees with it; the metric's own judge accepts most wrong answers as well. Our study shows that chain-answer agreement dominates the two larger judges' verdicts and that all three fail to reliably detect the tested reasoning errors. Consequently the difference Pass@k - CoT-Pass@k averages 19.7 points on an earlier solver generation but only 4.1 on the current one. What little remains depends on the token budgets on both sides and on the generation mode; raising the generation budget moves Pass@64 by more than fifty points while the difference stays at zero. We close with two checks any judged reasoning metric should pass before its numbers are read as evidence about reasoning.
|
| 467 |
MixDetect: Word-Level Localization and Quantification of AI Editing
2609.32625
|
cs.CL
|
Hongrui Bao, Yubing Ren, Zhendong Pan, Fang Fang, Shi Wang |
Large language models are increasingly used to edit human-written text rather than generate entire texts from scratch. Conventional AI-text detectors mainly distinguish human-written from fully AI-generated text, while recent methods for AI-edited text typical...Large language models are increasingly used to edit human-written text rather than generate entire texts from scratch. Conventional AI-text detectors mainly distinguish human-written from fully AI-generated text, while recent methods for AI-edited text typically provide only a text-level label or editing-degree score. We introduce MixDetect, a word-level framework for localizing and quantifying AI editing. MixDetect separately predicts whether each word has been edited and, conditional on editing, how substantial the edit is, allowing editing scope and editing intensity to be estimated separately. During training, source--edited pairs are aligned to construct word-level supervision, while inference requires only the input text. Experiments show that MixDetect accurately localizes AI-edited words, reflects differences in editing intensity, and reveals different scope--intensity patterns across editing degrees and operations. The overall AI editing magnitude increases under additional AI editing, decreases when AI-generated text is edited by humans, and remains nearly unchanged under ordinary human-to-human editing. The aggregated text-level predictions also perform well on binary and ternary AI-text classification and remain effective under domain and generator shifts. These results show that AI editing can be analyzed beyond a single authorship label or editing-degree score by identifying both where AI editing occurs and how substantial the edits are.
|
| 468 |
ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis
2609.32630
|
cs.CL
|
Kwangwook Seo, Dongha Lee |
Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming ...Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experience into reusable procedural knowledge, serving as an important layer for the harness system that supplies agents at runtime. Despite its potential, existing approaches largely abstract past experience into fixed procedural knowledge before downstream demands are known, which risks discarding knowledge that later becomes critical while retaining instance-specific details irrelevant to future tasks. In this paper, we reframe agent skill synthesis as a dynamic navigation problem over past experience, where agents actively explore accumulated trajectories on demand for the current task with targeted and fine-grained access to experience knowledge. To this end, we propose ExpVoyager, a novel framework in which a skill curator navigates raw experience across different views and resolutions, continually identifying reusable procedural knowledge from what it observes while tracking remaining knowledge needs that guide where to navigate next. Extensive experiments demonstrate both the effectiveness and versatility of ExpVoyager, showing consistent improvements in downstream task performance, continual gains as the experience space scales, and practical compatibility with existing skills under efficient experience access.
|
| 469 |
From Knowing to Abstaining: Bridging the Representation-Action Gap in Vision-Language Models
2609.32653
|
cs.CL
|
Jialuo He, Huangxun Chen |
The ability of vision-language models (VLMs) to abstain from unanswerable questions is as important as their ability to answer answerable ones accurately. Recently, several benchmarks have emerged to evaluate and improve VLM abstention, but they have substanti...The ability of vision-language models (VLMs) to abstain from unanswerable questions is as important as their ability to answer answerable ones accurately. Recently, several benchmarks have emerged to evaluate and improve VLM abstention, but they have substantial limitations. First, samples often contain shortcut cues in images or questions that reveal answerability, while an explicit "unanswerable" option further prevents accurate assessment of spontaneous abstention. Second, as training data, they generally provide only binary labels without fine-grained explanations for deeper supervision. To address these limitations, we introduce Visual Answerability Diagnosis with Rationales (VAD-R), a benchmark constructed through a two-stage pipeline of shortcut filtering and quality verification to prevent answerability leakage. Each example is annotated with step-by-step rationales and causal evidence-gap labels. Evaluation of state-of-the-art open- and closed-source VLMs on VAD-R reveals limited spontaneous abstention, with average recall rates of only 11.4% and 16.3%, respectively. Probing analyses show that hidden-state representations in certain layers can effectively distinguish answerability, yet this distinction fails to manifest in final responses. Motivated by this observation, we introduce Rep2Act, a representation-to-action alignment method that translates latent answerability awareness into explicit abstention decisions. Rep2Act improves action accuracy on VAD-R from 56.67% to 86.33% for Qwen2.5-VL-3B and from 59.33% to 88.67% for Qwen2.5-VL-7B. On the out-of-distribution TUBench, Rep2Act achieves an average F1 score of 53.3% with only a 3B model, surpassing the closed-source GPT-4 Turbo and GPT-4o by 16.2% and 1.1%, respectively.
|
| 470 |
Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs
2609.32684
|
cs.CL
|
Zifeng Cheng, Lingyun Qian, Zhiwei Jiang, Cong Wang, Yafeng Yin |
Extracting conditional text embeddings from large language models (LLMs) is a promising paradigm, as it requires neither additional data nor fine-tuning. Existing methods incorporate conditions into prompts to guide LLMs to focus on specific aspects and elicit...Extracting conditional text embeddings from large language models (LLMs) is a promising paradigm, as it requires neither additional data nor fine-tuning. Existing methods incorporate conditions into prompts to guide LLMs to focus on specific aspects and elicit conditional text embeddings. However, relying solely on prompts often fails to produce high-quality conditional text embeddings, as they remain entangled with general text embeddings, ultimately degrading their quality. To this end, we propose an inference-time, plug-and-play Self-Contrastive Steering (SCS) method that constructs unconditional general text embeddings and uses them to refine conditional text embeddings, making them more focused on the target condition. Specifically, we modify the attention mask and positional encodings to mask the condition, thereby obtaining unconditional text embeddings and intervening in the multi-head self-attention computation process. Notably, our method is highly efficient, requiring only a single additional multi-head self-attention computation at inference time. Extensive experiments on clustering, Semantic Textual Similarity, and triplet alignment datasets demonstrate that our method can seamlessly improve the performance of existing prompt-based methods across different LLMs in a training-free and plug-and-play manner. Our code will be released at https://github.com/zifengcheng/SCS
|
| 471 |
LLM Alignment--Utility Asymmetry under Semantic-Preserving Transformations
2609.32717
|
cs.CL
|
Mohan Li, Chengyu Yu, Francesco Sovrano, Marc Langheinrich, Martin Gjoreski |
Large Language Model (LLM) alignment is intended to ensure that models remain helpful and safe, but its stability under input distributional shift is not yet fully understood. Prior work shows that aligned models can fail under jailbreak prompts, alternative e...Large Language Model (LLM) alignment is intended to ensure that models remain helpful and safe, but its stability under input distributional shift is not yet fully understood. Prior work shows that aligned models can fail under jailbreak prompts, alternative encodings, and cross-lingual transfer, yet these failures are usually studied as attacks rather than controlled probes of alignment generalization. Moreover, existing evidence is largely grounded in natural language variation already represented during pretraining, leaving unresolved whether alignment generalizes with semantic content or remains tied to superficial surface patterns. In this paper, we study this question using synthetic semantic-preserving transformations that are rule-based and invertible, preserving task-relevant meaning while shifting inputs beyond standard linguistic variation. Across four open-weight and four commercial models, under both fine-tuning and in-context learning, we use these transformations as a probe of alignment generalization and identify an empirical pattern we term Alignment--utility asymmetry: once models can operate effectively on transformed inputs, task utility is often substantially retained while alignment failure increases more sharply. For example, adapted GPT-4.1 mini shows only limited utility degradation under transformation while its harmful rate rises from 13.3 to 74.3; Gemini 3 Flash similarly retains near-original utility while its harmful rate increases from 2.3 to 43.0. Taken together, these results suggest that semantic-preserving distribution shifts can expose a recurring gap in how utility and alignment generalize in current LLMs.
|
| 472 |
How Far Do Persona Effects Generalize in Language Models?
2609.32758
|
cs.CL
|
Yufan Zhou, Yuxuan Liu, Enze Ma, Lyumanshan Ye, Zhongqi Yue |
Persona prompts ask language models to answer as particular kinds of people. We test whether relationships learned from these effects predict responses to new questions and remain useful across models and prompts. Across 57 attributes, three behavioral domains...Persona prompts ask language models to answer as particular kinds of people. We test whether relationships learned from these effects predict responses to new questions and remain useful across models and prompts. Across 57 attributes, three behavioral domains, and seven pairs of open 7 to 9B checkpoints, persona effects can be predictable without being portable. Separate attribute and task gains improve prediction beyond shared scaling significantly in OLMo-3 and Qwen2.5, with the most robust evidence in OLMo-3. In that model, target refitting significantly outperforms gains borrowed from each of the other six pairs. Across model transfers, borrowed gains with one amplitude underperform shared scaling in most directions; allowing two target parameters removes the significant losses but yields no significant benefit over target shared scaling. After rewording, refitting significantly outperforms reuse with one amplitude in all six tested pairs, while changes of examples or country context often preserve reuse value. In the tested prompt transfers, regularized updates outperform both reuse strategies in median at 64 target questions per attribute. A separate survey comparison finds that selecting the more responsive checkpoint can worsen human fit; responsiveness is confounded with training status, and temperature calibration largely removes this cost but not errors in group ordering. Within the tested gain representation, apparent transfer can come from target calibration; source relationships must add predictive value beyond calibration and regularization. Code and data are available at https://github.com/thzva/persona-gain
|
| 473 |
C-HAT-Bench: Benchmarking Chinese AI-Text Detection Beyond Fully Generated Text
2609.32770
|
cs.CL
|
Qing Yang, Zixiang Luo, Zhenyu Mao, Zezheng Wu, Xinghe Cheng |
Large Language Models (LLMs) increasingly participate in writing by modifying or extending human drafts, causing machine involvement to vary in both form and extent. Yet most Machine-Generated Text (MGT) detectors are evaluated only on fully human-written vers...Large Language Models (LLMs) increasingly participate in writing by modifying or extending human drafts, causing machine involvement to vary in both form and extent. Yet most Machine-Generated Text (MGT) detectors are evaluated only on fully human-written versus fully AI-generated text. Because human--AI collaboration can weaken or redistribute cues associated with machine generation, strong performance under this binary setting may overstate detector reliability. This mismatch remains underexplored in Chinese: detection cues are shaped by tokenization and language-specific text distributions, yet controlled resources spanning production settings, domains, and generators remain limited. To fill this gap, we present a Chinese Human-AI Collaborative Text Detection Benchmark (C-HAT-Bench), a unified benchmark that links $5,000$ human-written source texts from five domains to more than $240,000$ variants produced using six generative models under Prefix-Conditioned Continuation as a reference setting and three collaborative production modes. We evaluate $21$ detectors through four protocols spanning zero-shot and pretrained supervised document-level detection, boundary localization, and cross-condition generalization. Relative to Prefix-Conditioned Continuation, mean AUROC across document-level detectors is $12.0\%$ lower on the collaborative production modes, with the largest detector-specific relative decrease reaching $44.4\%$. Transfer across collaborative production modes is also asymmetric, indicating that performance in a given production setting is not a reliable predictor of performance in other production settings.
|
| 474 |
On the Behavioral Traits of LLM Agents
2609.32776
|
cs.CL
|
Haokai Zhao, Jie Gao, Yunze Xiao, Xintao Wang, Weihao Xuan |
Users increasingly describe different AI agents as distinct colleagues to work with. AI personality research aims to quantify such impressions by attributing human-like "traits" to agents. However, existing measures fall short: models' self-reports (S-data) di...Users increasingly describe different AI agents as distinct colleagues to work with. AI personality research aims to quantify such impressions by attributing human-like "traits" to agents. However, existing measures fall short: models' self-reports (S-data) diverge from their actual behavior, while informant ratings from LLM judges (I-data) are costly to scale and cover few everyday scenarios. In this paper, we propose A-B-D to infer traits bottom-up from behavioral data (B-data), namely how agents act on their environment and communicate with users, as recorded in existing trajectories. From 345,667 real-world trajectories spanning 80 models, 12 tasks, and 50 harnesses, we extract 318 candidate features that capture both the actions an agent takes at each step (functional) and the language accompanying them (linguistic). We retain only features that show instance-level stability, cross-task consistency, and model discriminability. Factor analysis of the remaining 79 features uncovers six stable, model-attributable factors, two functional and four linguistic. For example, Kimi-K3 exhibits the most planfulness, whereas GPT-5.5 and GPT-5.6 are the least energetic. Moreover, we quantify the "knowledge-action gap" in the wild: these factors correlate only weakly with self-reported Big Five scores, even for conceptually matched pairs such as extroversion and energetic (r = 0.07, p = 0.58). Our work offers a new lens for understanding AI personality, with implications for users, developers, and researchers from both computer science and social science.
|
| 475 |
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
2609.32810
|
cs.CLcs.LG
|
Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian, Chi-Yu Chen |
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient ...Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube. The benchmark evaluates two settings: SPECIALIST TURN, in which an LLM responds to a clinically significant question posed during a real discussion, and BOARD SIMULATION, in which it generates an entire back-and-forth discussion and reaches a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Evaluation of 14 general-purpose frontier and medical LLMs reveals substantial limitations: the best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions. Supervised finetuning and reinforcement learning improve performance on a held-out test set, suggesting that real-world discussion trajectories can support model adaptation. Three M.D. experts review a subset of the benchmark, finding high information coverage and factuality of patient cases and strong fidelity of extracted consensus conclusions. We will release OpenTumorBoard and its automated curation pipeline to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.
|
| 476 |
ARSM: Auto-Regressive State Machine for Agentic Reasoning Compression
2609.32852
|
cs.CL
|
Xiafeng Man, Siyuan Ye, Xiaosong Ma |
While Large Language Model (LLM)-based agents demonstrate strong capabilities in long-horizon tasks by interleaving reasoning with external environment interactions, the continuous accumulation of context rapidly creates a critical memory bottleneck. Existing ...While Large Language Model (LLM)-based agents demonstrate strong capabilities in long-horizon tasks by interleaving reasoning with external environment interactions, the continuous accumulation of context rapidly creates a critical memory bottleneck. Existing memory compression methods rely on task-specific optimization or external auxiliary models, introducing significant computational overhead. Furthermore, the resulting compressed representations tend to lose structured relationships, leading to information dilution, attention collapse, and degraded decision consistency. To address these limitations, we propose Auto-Regressive State Machine (ARSM), a lightweight training-free framework that enables in-situ reasoning compression through structured state evolution. ARSM introduces two key components: (i) a trajectory abstraction mechanism that reorganizes interaction histories into compact Hypothesis-Action-Result (HAR) micro-chains; (ii) a dynamic state machine that regulates hierarchical memory through atomic operations and a compression-control parameter. These components are unified within an auto-regressive, self-compressive generation space, where each model output jointly performs external action execution and internal state updates. We evaluate ARSM on Webshop, Multi-Objective Multi-Hop QA, and SWE-Bench Lite datasets. Experimental results show that ARSM maintains the task performance while simultaneously reducing token consumption, offering a practical, cost-effective route toward scalable autonomous agents for long-horizon tasks.
|
| 477 |
Are You Sure You're Sure? Two Confounds in a Sycophancy Benchmark
2609.32867
|
cs.CL
|
Atharv Gupta, Akshat Jindal, Lavanya Nigam, Aryan Sood |
Sycophancy is a language model's tendency to cave when a user pushes back, abandoning a correct answer for the user's. Several benchmarks now measure it by scripting an objection and recording how often the model caves. Because that objection is a prompt templ...Sycophancy is a language model's tendency to cave when a user pushes back, abandoning a correct answer for the user's. Several benchmarks now measure it by scripting an objection and recording how often the model caves. Because that objection is a prompt template, whatever else the template varies is measured along with the property it claims to isolate. We audit SycEval, which reports that objections raised before a model answers (preemptive) cause more caving than those raised after (in-context), and attributes the gap to timing. Two features of its templates vary alongside the property each is meant to test. First, SycEval's objections escalate through four strength levels, and at the two weakest only the preemptive template names a target answer, so timing and naming vary together. We build the missing comparison and test it on multiple-choice questions and SycEval's own free-form pipeline. Naming a target answer raises the follow rate, the share of samples matching the user's assertion, by 14.1--49.5 percentage points (pp), and once both templates name one, the timing comparison reverses on three of the five model conditions we test: models cave \emph{less} under preemptive objections than in-context ones, opposite to SycEval. Second, the output-format instruction benchmarks append for automatic grading also varies with objection placement. Moving it from the pushback into the question flips its effect in opposite directions across models on the same items ($p=0.0059$). The same effect appears in SycEval's free-form pipeline: relocating its instruction lowers caving by 5.1pp on Llama-3.1-8B, while an equivalence test confirms no effect on Qwen3-4B. Both confounds live in the template rather than the models under test, so both are correctable: we close with three checks benchmark authors can apply before publishing.
|
| 478 |
Linger and Lose: Knowledge Collapse in Low-Bit Language Models
2609.32902
|
cs.CL
|
Prashanna Mani Paudel, Shivanand Venkanna Sheshappanavar |
Training language models with ternary weights is commonly judged by loss and downstream accuracy, which record only a modest cost relative to full precision. We show that these metrics can conceal a much larger failure. We instead measure knowledge capacity, t...Training language models with ternary weights is commonly judged by loss and downstream accuracy, which record only a modest cost relative to full precision. We show that these metrics can conceal a much larger failure. We instead measure knowledge capacity, the factual bits stored per parameter, on synthetic biographies with known information content. We train GPT-2-style models from scratch with 2.5M to 50M parameters at five precisions. Under the standard cosine schedule, ternary models retain as little as 6% of an identically trained fp16 model's capacity. The deficit widens with model size while perplexity rises by only 1.4 to 1.6 times. Measured throughout training, these models first acquire capacity and then lose most of it. We identify this knowledge collapse as a learning-rate dwell instability. Held near $1$--$2\times10^{-4}$ with no decay, a pre-collapse model collapses within a few hundred exposures, and returning to a safe rate does not restore capacity. We then locate the collapse in the output head. It happens at a value the model can never predict. The weights there grow unchecked, while every other prediction the model makes is unchanged. We find that making the value predictable removes the collapse, regardless of which attribute carries it. A warmup-stable-decay schedule with a 10% cooldown increases ternary capacity by 2.7 times at 25M and 4.3 times at 50M. Gains are larger at lower precision. Cutting the output head's learning rate prevents failure when training from scratch. The post-training quantization methods we tested recover no measurable capacity below 4 bits. Our findings suggest judging low-precision training by retained capacity during training rather than final loss.
|
| 479 |
The Key Handoff: Retrieval in Hybrid Language Models
2609.32942
|
cs.CLcs.LG
|
Kaan Kale, Oguzhan Baser, Sriram Vishwanath |
A two-hop question makes a language model retrieve twice: once to produce a bridge entity, and once to retrieve with it. Transformers resolve that entity in their early layers. Hybrid models replace most of the attention with a recurrent state, so where the ke...A two-hop question makes a language model retrieve twice: once to produce a bridge entity, and once to retrieve with it. Transformers resolve that entity in their early layers. Hybrid models replace most of the attention with a recurrent state, so where the key becomes usable, and where it is spent on an answer, is not known. Answering "Where is the ball?" from "the ball belongs to Alice" and "Alice is in the garden" turns on Alice, a name the question does not mention. A model could reach garden through Alice, the key it computed, or through where the fact sits in the prompt. Across twelve models, dense and hybrid, we move a hidden state from one story into another where the two routes lead to different places, and read off which place the model gives. An attention layer converts the key in every model we tested, and crossing that layer removes its usable effect, all of it where no attention follows. In sequential hybrids this makes the answer a handoff: recurrent layers carry the key forward, and attention spends it. Retrieval does not always stop there. Writing a different fact into the memory of the recurrent layers that follow the last attention layer can move the answer toward that fact, multiplying its odds by 1.3 to 2.7, and which hybrids do this is not settled by their architecture or training. That read is addressed by the key, not by position: recurrent state in a hybrid is not only a carrier, but a memory that later layers can query.
|
| 480 |
TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference
2609.33014
|
cs.CL
|
Tzu-Heng Huang, Jet Lin, Eric Lin |
Medical benchmarks for language models are built almost entirely on Western biomedicine. Traditional Chinese Medicine (TCM) is a separate system, with its own diagnostic framework and its own literature, and it remains largely unmeasured. The few TCM evaluatio...Medical benchmarks for language models are built almost entirely on Western biomedicine. Traditional Chinese Medicine (TCM) is a separate system, with its own diagnostic framework and its own literature, and it remains largely unmeasured. The few TCM evaluations that exist are small, narrow, and rarely paired with a human reference. We present TCMQA, an open benchmark of 38,279 questions from Chinese TCM licensing examinations, paired with 15,151 responses from 101 licensed practitioners. We evaluate 29 instruction-tuned models from 9 families, spanning 0.27B to 14.8B parameters. Accuracy ranges over 59 points, and no model approaches saturation. Pretraining data predicts TCM ability far better than scale: a 12B Western-pretrained model reaches 39.6%, while a Chinese-pretrained model an eighth its size reaches 60.8%. Nine models exceed the practitioner majority vote of 64.9%, the best by 21.8 points, and all nine come from that same Chinese-pretrained family. Yet difficulty does not transfer between models and practitioners: accuracy is flat across practitioner-rated difficulty, item-level agreement is near zero for all 29 models, and on $8.4\%$ of items the practitioners are correct where the leading model is wrong. We release the corpus, the practitioner responses, the harness, and per-item model outputs at https://huggingface.co/datasets/TechTCM/TCMQA.
|
| 481 |
Improving the Diversity of LLM Outputs without a Trade-off
2609.33038
|
cs.CLcs.LG
|
Ryoma Sato |
We propose DAST (Diversifying Arithmetic Sampling with TokenTour), a method that increases the diversity of LLM outputs without any change to the marginal distribution and with negligible generation-time overhead (a few microseconds). We observe that token IDs...We propose DAST (Diversifying Arithmetic Sampling with TokenTour), a method that increases the diversity of LLM outputs without any change to the marginal distribution and with negligible generation-time overhead (a few microseconds). We observe that token IDs are often arranged in a meaningless order and reassign them so that tokens with similar meanings appear consecutively. This can be done in advance in a few hundred seconds per model, and the resulting order can be reused for all subsequent generations. By combining this order with arithmetic sampling (or quasi-Monte Carlo methods), we make similar tokens less likely to be generated across runs while preserving the distribution. Our method not only produces qualitatively good ideas but also significantly improves performance on the downstream task of ProtoQA.
|
| 482 |
Pinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge Leaderboards
2609.33044
|
cs.CL
|
Krishna Chytanya Ayyagari |
Modern LLM evaluation assumes that pinning a judge to a fixed model snapshot and decoding at temperature zero yields reproducible verdicts. We show this assumption fails as a property of how LLM-as-Judge is operationalized on cloud serving infrastructure, not ...Modern LLM evaluation assumes that pinning a judge to a fixed model snapshot and decoding at temperature zero yields reproducible verdicts. We show this assumption fails as a property of how LLM-as-Judge is operationalized on cloud serving infrastructure, not of any particular model family. Across four frontier judges all served via a single major enterprise cloud platform and three standard benchmarks (Arena-Hard, AlpacaEval 2, MT-Bench), identical inputs to the same pinned, temperature-zero judge produce different verdicts across re-runs: per-item flip rates of roughly 5% on average and about 40% on the close-call items that decide leaderboard margins, with a per-judge magnitude spanning a 40x range (from 0.13% to nearly 10%). We introduce metrics tailored to this instability: per-item flip rate, a two-part stability profile (waver fraction and conditional intensity), and adjacency separability - and report what the variance does and does not do to rankings. For a single judge the aggregate ranking is stable (0% top-K instability, 0% pooled winner flip); what degrades is precision: under a paired hierarchical bootstrap, roughly one-fifth to three-quarters of adjacent leaderboard positions are statistically indistinguishable, a noise floor we attribute primarily to finite prompt sampling rather than to the judge. Across judges, leaderboards agree on the coarse ordering but diverge in the middle (Kendall's tau as low as 0.42-0.64 between families on Arena-Hard), and of 13 published head-to-head ranking claims we re-judge, 5 fail under a defensible judge swap or re-run. We argue that leaderboards report unhedged point estimates that misrepresent the noise floor of the instrument, and we propose a minimal, low-cost reporting protocol: several judge re-runs, published stability profiles and adjacency intervals, and results under at least two judges from different families.
|
| 483 |
Reading Too Much into Context: Passive Exposure Can Steer LLM Decisions
2609.33065
|
cs.CLcs.LG
|
Yuxiang Zheng, Lin Tian, Marian-Andrei Rizoiu |
Large language model (LLM) assistants can now search the web and consult external sources while completing user requests. These sources can provide useful evidence, but they can also introduce additional content into the model's context. Can such passive expos...Large language model (LLM) assistants can now search the web and consult external sources while completing user requests. These sources can provide useful evidence, but they can also introduce additional content into the model's context. Can such passive exposure steer a decision even when the added content provides no reason to change it? We examine the stability of model decisions on the same tasks with and without such external content. Across all open-weight and closed-weight models we test, exposure systematically shifts decisions, with effects reaching nearly 50 percentage points in closed-weight models. The same pattern appears with real-world online opinions. The influence also extends beyond subjective preferences. Such exposure can steer models toward choices that violate explicit user requirements and increase their acceptance of false claims. In short, what enters an LLM's context can influence its decision even when it should not determine it.
|
| 484 |
OneSign: Unifying Sign Language Understanding Tasks with One Model
2609.33090
|
cs.CL
|
Shiwei Gan, Yafeng Yin, Xiao Liu, Desibieer Tuerdaken, Lei Xie |
SLU encompasses a diverse set of tasks, including ISLR, CSLR, and SLT. Although these tasks share basic semantic and linguistic foundations, they are typically addressed with task-specific architectures and training pipelines, which hinders knowledge sharing a...SLU encompasses a diverse set of tasks, including ISLR, CSLR, and SLT. Although these tasks share basic semantic and linguistic foundations, they are typically addressed with task-specific architectures and training pipelines, which hinders knowledge sharing and requires costly pretraining and finetuning for each task. In this paper, we focus on two aspects of SLU tasks: (1) training and inference pipelines are highly fragmented: most methods rely on pretraining on large-scale SL datasets followed by task- or dataset-specific finetuning, which leads to multiple specialized models rather than a single checkpoint. (2) current LLM-based methods may overlook the inherent modality discrepancy between sign and text tokens, simply concatenating them and processing both modalities with the same decoder layers. In this paper, we present OneSign, a unified framework that addresses multiple SLU tasks within a single model and a single checkpoint. OneSign reformulates ISLR, CSLR, and SLT under a single training paradigm. To accommodate the heterogeneous characteristics of sign and text representations, we introduce a Modality-Adaptive Mixture-of-Experts (MA-MoE) architecture, consisting of a shared expert and modality-specific experts for sign and text tokens. A modality router dynamically activates the corresponding experts, and their outputs are aggregated to form the final token representations. By enabling modality-dependent expert specialization while preserving a shared expert path, MA-MoE can effectively model the modality differences between continuous sign representations and discrete text tokens. Extensive experiments on multiple benchmarks demonstrate that OneSign achieves competitive or state-of-the-art performance on several benchmarks, highlighting its effectiveness as a unified SLU model. Datasets are available at : https://github.com/gswycf/OneSign.
|
| 485 |
MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning
2609.33119
|
cs.CL
|
Lang Cao, Binghang Lu, Yuhao Shen, Yue Guo |
Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of ...Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of medical questions than relying on any single model. However, existing LLM routing methods primarily seek to balance answer quality and inference cost, leaving open how to exploit differences in specialist competence to improve medical reasoning. In this paper, we introduce MedRouter, an agentic system that uses an embedding-based multi-label router to select and query specialist LLMs, then passes their responses to a generator to produce the final answer. We further propose SCALE (Specialist Competence-Aware Learning), a two-stage training framework that first trains the Router with specialist correctness supervision and then optimizes its selections through reinforcement learning. The second stage uses a Performance Gain Reward (PGR) that measures how specialist information affects the generator's answer correctness relative to answering without that information. Experiments on eight text-based and multimodal medical QA benchmarks show that MedRouter outperforms the strongest routing baseline by 8% in average accuracy. Our analysis of specialist outputs further reveals distinct strengths and complementary question-level coverage, motivating learned routing to combine these capabilities for more comprehensive medical reasoning.
|
| 486 |
Classifying Dominant Temporal Orientation without Pretrained Text Embeddings: A Novel Morphosyntactic Inventory Vector Approach
2609.33121
|
cs.CL
|
Jonathan Cleveland, Peter S. Bearman |
Computational methods have consistently struggled to determine the dominant temporal orientation of a sentence. This difficulty is especially pronounced when a sentence contains multiple embedded clauses with competing tense and aspectual information. To addre...Computational methods have consistently struggled to determine the dominant temporal orientation of a sentence. This difficulty is especially pronounced when a sentence contains multiple embedded clauses with competing tense and aspectual information. To address this difficulty, we propose an alternative approach for identifying a sentence's past, present, or future global reference interval. Our method does not use any form of pretrained embeddings. We rather encode sentences using fixed-length inventory vectors that are comprised of part-of-speech counts, dependency relation counts, and explicit futurate pattern counts. We term this inventory vector of a sentence a "Morphosyntacton". The method does not use any padding, sequence models, or large language models. Evaluation on 1,799 syntactically complex English sentences, annotated as past, present or future, shows balanced and high accuracy multiclass classification, achieving an overall multi-class accuracy of 92%.
|
| 487 |
Knowing Is Not Choosing: What Explicit Verification Adds Beyond Generative Preference
2609.33142
|
cs.CL
|
Yilong Li, Chengpo Yan, Aayan Arish, Suman Banerjee |
Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct candidate, ranking the available candidates, and selecting the final answer. Pre-generation readouts predict factu...Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct candidate, ranking the available candidates, and selecting the final answer. Pre-generation readouts predict factual recall and which questions sampling will cover across three model families, but say little about whether an available correct answer will ultimately be selected. Explicit verification with $P(\mathrm{True})$ improves within-question ranking over mean log-likelihood in Gemma, Qwen3, and Llama, with AUROC gains of $0.08$--$0.12$. In a prospectively defined Gemma cohort, verification raises plurality accuracy by about $5$ points, and still gains about $2$ points over chat-template likelihood, a stronger generative baseline. The advantage is strongest for relations with common-answer priors and depends on access to the entity; masking the entity removes the ranking advantage in larger Qwen models. Finally, the measured benefit depends on how correctness is defined: recall-oriented reference matching can credit option lists favored by likelihood and substantially understate the improvement seen under human semantic judgments. Prior work shows that models can carry latent factual knowledge and judge candidate answers; we show that these capabilities do not collapse into a single notion of ``knowing,'' and trace where information is gained, lost, or mismeasured between availability, ranking, and final choice.
|
| 488 |
CAME: Company-Aware Evidence-Memory Experts for Interpretable Quarter-Ahead Revenue Forecasting
2609.33143
|
cs.CLcs.LG
|
Ya-Wen Wu, Meng-Fen Chiang, Kuang-Da Wang, Wen-Chih Peng |
Quarter-ahead revenue forecasting requires company-scale numerical accuracy, strict temporal validity, and company-specific interpretation of narrative disclosures. LLMs can distill textual evidence but can produce scale-misaligned forecasts, whereas history-b...Quarter-ahead revenue forecasting requires company-scale numerical accuracy, strict temporal validity, and company-specific interpretation of narrative disclosures. LLMs can distill textual evidence but can produce scale-misaligned forecasts, whereas history-based anchors are stable but miss forecast-time signals such as product transitions, supply constraints, and management guidance. We introduce CAME (Company-Aware Evidence-Memory Experts), a residual-forecasting framework that refines a no-leakage statistical anchor when current semantic evidence and prior error patterns justify an adjustment. On a development-inclusive rolling backtest of 336 company-quarters from 12 large public technology and platform firms, CAME achieves the lowest aggregate point-estimate error among the reported methods, with statistically supported macro-sMAPE gains over the matched Statistical Anchor, and outperforms History + Guidance on all six aggregate metrics. CAME also links adjustments to source-linked evidence cards and guarded memory traces, supporting forecast inspection, provenance, and failure localization.
|
| 489 |
Generalization Dynamics of LM Pre-training
2609.33150
|
cs.CL
|
Jiaxin Wen (UC Berkeley), Zhengxuan Wu (Stanford University, Google DeepMind), Dawn Song (UC Berkeley), Lijie Chen (UC Berkeley) |
People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parr...People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like computations. We call this mode-hopping. Across our suite, LMs suddenly latch onto memorized or in-context patterns instead of in-context learning, use System 1 instead of System 2 thinking, pick up what sounds true instead of what is true, fail at multi-hop persona QA, out-of-context reasoning, and emergent misalignment -- then just as suddenly revert and generalize. Mode-hopping is not explained by standard optimization dynamics: it is locally stable and cannot be fixed by checkpoint averaging. We instead think of it as a capacity allocation problem: in a capacity-bounded model, generalizable circuits must compete with the shallow ones learned early in training, and the data in each pre-training window may decide which circuits win. Our suite provides a new efficient lens on generalization. We demonstrate two concrete applications: (i) select intermediate pre-training checkpoints that strongly generalize reasoning and alignment, better than the final pre- or mid-training checkpoints, and (ii) select pre-training data that controls and stabilizes generalization dynamics.
|
| 490 |
Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
2609.33155
|
cs.CL
|
Yuyang Zhao, Xuan Liu, HaoYang Shangm Haojian Jin |
Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from th...Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes across generation, selection, and learning: (1) incorrect consensus, where repeated samples agree on the wrong stance; (2) selection failure, where generation covers the observed stance but selection misses it; (3) response overfitting, where supervised fine-tuning improves imitation but harms prediction; and (4) early plateau, where reinforcement learning shows modest initial gains followed by limited further improvement. We expose these failures using STANCE-BENCH, which contains 2499 prediction tasks from 500 Hacker News users. Guided by this analysis, we explore a simple approach that combines direct scores for all candidate stances with explicit assessments of support from the individual's history. On the 781-task test set, this approach achieves 21.83 discussion-specific Macro F1 with Qwen3-8B, compared with 19.27 for direct scoring. Our results motivate evaluating candidate generation, final selection, and person-specific evidence use separately.
|
| 491 |
Scoring the Wrong Question: Readout Failures in Constrained-Option Evaluation
2609.33179
|
cs.CLcs.LG
|
Jiaxuan Guo, Kejia Zhang, Shuo Xin, Jingxin Yang, Youran Sun |
Constrained-option scoring reads a model's probabilities for a fixed set of permitted answers, so it returns a score even when the model is about to write something else. We study a prompt that quotes a multiple-choice item, ending in the item's own answer ins...Constrained-option scoring reads a model's probabilities for a fixed set of permitted answers, so it returns a score even when the model is about to write something else. We study a prompt that quotes a multiple-choice item, ending in the item's own answer instruction, and asks for a forecast of whether a reader model given only a short note will answer it correctly. The model can begin answering the quoted item: on three Qwen3 checkpoints an answer letter is the most probable token on nearly every passage, and the forecast ranks the reader's correctness indistinguishably from chance. Beyond prior work's argmax check and prefill repair, we contribute a label-free diagnostic of the declared options' support, their first-token mass in context; a deletion test tracing the Qwen3 answer-letter start to the quoted instruction; and rendering and tokenisation checks for dropped separators and coinciding first tokens. Prefilling an answer stem lifts the median option mass from below to above one half on all nine checkpoints scored with this prompt and, with options and renormalisation unchanged, raises the Qwen3-averaged AUC from 0.4929 to 0.6152. lm-polygraph's default P(True) estimator shows a related failure: it reads True where Qwen3 without thinking favours an answer letter (MMLU) or the prompt's (A)/(B) label (TriviaQA), and its Qwen3-averaged AUROC is indistinguishable from chance on MMLU and below chance on TriviaQA. Scoring and renormalising the (A)/(B) labels after an answer stem raises it to 0.632 and 0.868, renormalising True against False in place reaches 0.660 and 0.860, and on TriviaQA both exceed all fourteen default single-answer library estimators. The original forecast is renormalised too, yet stays indistinguishable from chance: option mass flags low support in both settings without labels, and only checking the intended target shows which score still ranks.
|
| 492 |
Identifying Temporal Features within Transcoders for Time Sensitive Factual Recall
2609.33183
|
cs.CLcs.LG
|
Sanjay Govindan, Yang Song, Maurice Pagnucco |
Large Language Models (LLMs) suffer from temporal misalignment, often due to the contradictory nature of their training corpora. While current mitigation strategies rely on computationally expensive fine-tuning or context-heavy retrieval augmented generation (...Large Language Models (LLMs) suffer from temporal misalignment, often due to the contradictory nature of their training corpora. While current mitigation strategies rely on computationally expensive fine-tuning or context-heavy retrieval augmented generation (RAG), the internal mechanisms governing time-sensitive recall remain under-explored. Unlike prior studies that identify temporal components such as attention heads and MLP layers, we provide the first feature-level map of temporal recall by isolating individual MLP features via transcoder circuit tracing. We identify three node categories (common temporal, common to the year, and chrono-semantic) which interact to generate a temporal filter during factual recall. By analysing Gemma 2 2B, LLaMA 3.2 1B, and Qwen3-4B, we show that these features do not follow a simple linear pipeline but represent time through a parallel and mixed syntactic-semantic interplay across layers. We additionally discover a class of higher-layer temporal components invisible to existing EAP-IG methods, establishing transcoders as a more complete lens for temporal interpretability in time-sensitive factual recall. These findings present MLP components for potential targeted interventions in time-sensitive factual recall
|
| 493 |
Turning Speech Language Models into Multilingual Listeners
2609.33204
|
cs.CL
|
Tol\'{u}lop\'{e} \`{O}g\'{u}nr\`{e}m\'{i}, Dan Jurafsky, Chris Manning, Ahmet \"Ust\"un, Martijn Bartelds |
Speech Language Models (SLMs) that understand spoken language questions support only a few high-resource languages, limiting access to millions of people worldwide. This gap stems from the scarcity of multilingual speech instruction-tuning datasets. We present...Speech Language Models (SLMs) that understand spoken language questions support only a few high-resource languages, limiting access to millions of people worldwide. This gap stems from the scarcity of multilingual speech instruction-tuning datasets. We present MULTISPEECHQA, a large-scale, synthetically generated and human-verified dataset comprising 9200 hours of 10.8 million spoken question-answer pairs in 23 typologically diverse languages. Using MULTISPEECHQA, we also introduce MULTISPEECH-BENCH, a multi-task benchmark for evaluating SLM performance on 23 languages. We compare the performance of a cascading system to open-weight and closed SLMs on MULTISPEECH-BENCH and find that the cascading system outperforms open-weight SLMs but not all closed SLMs. We use MULTISPEECHQA to finetune Qwen 2.5-Omni, which improves its performance on our benchmark. Our findings show that high-quality synthetic datasets offer a cheap solution to improving the multilingual capabilities of SLMs.
|
| 494 |
Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms?
2609.33209
|
cs.CLcs.LG
|
Keyi Li, Yihao He, Quanyi Li |
Typed-decision models such as TypeSafe's Jev answer a declared yes/no or multiple-choice question about a state with a probability instead of text, and their evaluations report accuracy and calibration. Neither requires that the probabilities a model gives to ...Typed-decision models such as TypeSafe's Jev answer a declared yes/no or multiple-choice question about a state with a probability instead of text, and their evaluations report accuracy and calibration. Neither requires that the probabilities a model gives to logically related questions fit together item by item, which is what a system that acts on those probabilities needs. We test this property, coherence, with a battery of logically linked questions that needs no labels. For 160 items from ChaosNLI and PubMedQA, each with three mutually exclusive labels, we ask whether the label is X, whether it is not X, whether it is one of the other two labels, and which label applies. On 480 negation pairs, Jev's probabilities for "the label is X" and "the label is not X" miss summing to one by 0.064 on average (95% CI 0.055 to 0.072). Qwen3.8-27B, run from its official BF16 weights, misses by 0.293 with first-token probabilities and by 0.122 with verbalized probabilities. The gap to the first-token readout persists on pairs where both systems give similar probabilities, without double-negation labels, and after averaging Jev's repeated calls. Jev is not coherent either: its violations are about five times its repeat noise, and it over-endorses statements about single labels, so that its three single-label probabilities sum to 1.14 on average. The two systems also fail differently. Qwen3.8-27B's first-token readout under-endorses the complement of a label whether or not the question contains "not", rejecting both a statement and its negation in 196 of 480 pairs, and it does not become more coherent where it is more confident, whereas Jev's violations concentrate where its answer is uncertain. Because the checks need no labels, they expose biases that appear only when question forms are compared, and inconsistencies within items.
|
| 495 |
CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification
2609.33212
|
cs.CL
|
Massa Baali, Sarthak Bisht, Ziyue Qiu, Joseph Konan, Rita Singh |
Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decis...Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-SV}, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports. By connecting a pretrained speaker encoder to a language model and supplying explicit acoustic measurements, CoLMbo-SV makes voice comparisons inspectable without restricting verification to the evidence verbalized in its reports. We additionally introduce \textbf{VoxReason}, paired recordings with measured acoustic properties and comparison reports filtered through numerical and qualitative checks, providing supervision for this combined capability. We also develop an evaluation framework that separates what acoustic information a speaker representation encodes, what influences the verification score, and what the generated report discusses. On VoxCeleb1-O, CoLMbo-SV achieves 0.99\% EER, reducing verification error by approximately 80\% relative to the strongest audio-language baseline fine-tuned on VoxReason, while attaining a numerical-grounding score of 0.82. Our analysis further demonstrates that acoustic correctness and decision relevance are distinct properties of an explanation, exposing a gap that numerical-grounding metrics miss. Together, these contributions substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an empirical framework for connecting natural-language explanations to the decisions they explain.
|
| 496 |
Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents
2609.33226
|
cs.CL
|
Donghua Cai, Yongheng Deng, Yifei Wang, Zijun Shen, Ju Ren |
Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and la...Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and later retrieved via a RAG pipeline. While effective in controlled settings, we show that this paradigm breaks down in long-horizon, high-entropy conversations: memory construction becomes increasingly lossy and unstable as context length and information complexity grow, and incurs prohibitive cost due to repeated LLM invocation. To address these limitations, we propose Threader, a memory system that shifts the focus from memory construction to efficient, structure-aware access over raw interactions. Instead of rewriting interactions, Threader preserves them as first-class memory, organizes them into topic-coherent segments via lightweight incremental segmentation, and enables accurate retrieval through multi-view representation. At query time, it performs multi-signal retrieval that combines segment-level access with localized evidence matching, ensuring both completeness and coherence. Extensive experiments demonstrate that Threader consistently improves answer accuracy and evidence recall, while significantly reducing the memory construction overhead.
|
| 497 |
The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond
2609.33280
|
cs.CL
|
Andreas Maier, Monica Hinrichs-Mayer, Franziska Weber, Niklas Lackner, Matthias May |
Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on...Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymisation policy as the variable under test: 15 detectors, three release conditions and four corpora of medical reports, legal judgments, news and other genres, and e-mail, in German, English, Chinese and Arabic. A fixed 13-detector union reaches a person sensitivity of 0.9998 at specificity 0.8686 on the medical reports, 0.9958 at 0.8504 on the legal judgments, 0.9352 at 0.9318 on news and other genres, and 0.9906 at 0.6235 on e-mail. With this ensemble, frequency matching with a public name list recovers zero identities by alignment across the four corpora; the names it got right were ones the detector missed, left in clear text. Cross-document linkage ranks the correct person first for 0.93% of e-mail queries without training and 3.94% with it, against 1/3697 chance and 71.98% on unmodified text. On the medical reports it recovers nothing without training and 0.71% of 138 queries with it, against 1/207 chance and a 2.73% ceiling on unmodified text.
|
| 498 |
Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models
2609.33290
|
cs.CL
|
Yadong Xi, Rongsheng Zhang, Tangjie Lv, Ziyang Luo, Ruochen Zhao |
Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how w...Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.
|
| 499 |
BaatCheet: A Multilingual Corpus for Dialogue Translation in Indian Languages
2609.33296
|
cs.CL
|
Priyanka Dasari, Yuvrajsinh D. Bodana, Vandan Mujadia, Arafat Ahsan, Dipti Misra Sharma |
Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation...Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks focus on sentence-level or formal text, making it difficult to assess translation quality of the dialogue phenomena. In this work, we introduce BaatCheet, a multilingual dialogue corpus named after the Hindi term for conversation or chitchat, containing approximately 49,000 dialogues for dialogue translation across five translation directions. We fine-tune five open-source LLMs across seven training data configurations and find that fine-tuning yields substantial gains over zero- and few-shot baselines. To comprehensively evaluate dialogue translation quality, we employ multiple evaluation strategies, including automatic metrics, LLM-as-judge, and human assessments using an SQM-guided Direct Assessment (DA) Protocol.
|
| 500 |
Hesitation-Aware On-Policy Distillation for Diffusion Language Models
2609.33301
|
cs.CL
|
Jianguo Huang, Lipeng Wan, Yanchen Deng, Bo An |
Diffusion large language models (dLLMs) generate text by iterative unmasking. At each denoising step, a dLLM proposes a token at every masked position, but the decoder commits only a confident subset of these proposals. Trace-based on-policy distillation (TOPD...Diffusion large language models (dLLMs) generate text by iterative unmasking. At each denoising step, a dLLM proposes a token at every masked position, but the decoder commits only a confident subset of these proposals. Trace-based on-policy distillation (TOPD) builds on this process by matching the student to a stronger teacher, yet only at the committed positions. We argue that this discards much of the useful signal, which resides in the uncommitted proposals, where the student has made a prediction but is not yet confident enough to commit it. We call these proposals hesitations. In our pilot study on an SDAR-4B student, hesitations make up only 24% of supervisable state-position pairs but carry 66% of the teacher-student divergence. To exploit this signal, we propose Hesitation-Aware On-Policy Distillation (HOPD), which extends teacher distribution matching to every masked position of each denoising step. Because hesitations are not equally informative, we further allocate supervision using hindsight from the completed trajectory, placing more weight on positions whose proposal was later disagreed with the final token and on blocks where first-step proposals rarely survive. Since both models already produce distributions at all masked positions, HOPD requires no additional forward passes over TOPD. The only extra cost is evaluating the loss at more positions. With SDAR-1.7B and SDAR-4B students distilled from TraDo-8B-Instruct, HOPD achieves the best average score among the evaluated methods on five math and coding benchmarks, under both static and dynamic decoding and at both scales. It also speeds up decoding. On SDAR-4B, the HOPD student hesitates less and commits 11% more tokens per denoising step than TOPD, while reaching higher accuracy.
|
| 501 |
CertMark: Distortion-Free Multi-Bit Watermarking with Certified Decoding
2609.33332
|
cs.CL
|
Pawe{\l} Batorski, Przemys{\l}aw Spurek, Paul Swoboda |
Leading multi-bit watermarking methods for language models encode messages by biasing the model's next-token probabilities, creating a trade-off between message recovery and text quality. Their decoders typically return the highest-scoring candidate from accum...Leading multi-bit watermarking methods for language models encode messages by biasing the model's next-token probabilities, creating a trade-off between message recovery and text quality. Their decoders typically return the highest-scoring candidate from accumulated token-level evidence, without a certified abstention rule that bounds the probability of outputting an incorrect message. We introduce CertMark, a distribution-preserving multi-bit watermark with certified decoding. Rather than modifying probabilities, CertMark uses the embedded message to seed an exact Gumbel-max sampler, thereby preserving the model's original sampling distribution. We propose two scalable decoders: a model-agnostic, text-only decoder and a model-aware variant that leverages the original next-token distributions for stronger recovery. Both support certified abstention with mathematical bounds on the probability of returning an incorrect message. Across text completion, summarization, and story generation, CertMark matches the perplexity of unwatermarked text while reliably recovering multi-bit messages. The model-aware decoder further achieves higher bit accuracy than probability-biasing baselines. Our code is publicly available at https://github.com/Batorskq/CertMark.
|
| 502 |
Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training
2609.33335
|
cs.CL
|
Xinyu Che, Hang Yan, Yanchen Liu, Haochen Liu, Ruifeng Li |
Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervisio...Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. Existing next-observation training methods help the agent to learn the environmental content. However, they additionally involve an optimization process, which may introduce several effects other than learning to predict the world. Consequently, where the performance gain comes from during the training process remains an open question. We answer this research question through replacing true next-observation targets with in-distribution mismatched observations during the training process. Across two interactive text environments, mismatched targets lower prediction accuracy by 15.3-61.6% relative to ground-truth targets, yet retain substantial task gains over the base model. Compared with the base model, trained models consider more candidate actions and exhibit less looping. We also introduce a setting that replaces prediction-based rewards with independent random signals. This training expands task coverage (pass@64) even when the reward carries no environment information. We also generalize this finding to VisualWebArena, where random-reward training raises pass@64 by 14.3% relative to the base model, without observation-matching rewards or an external multimodal teacher for reward construction.
|
| 503 |
CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation
2609.33343
|
cs.CL
|
Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala |
Faithfulness evaluation of abstractive summaries remains an open challenge, with existing metrics addressing only isolated hallucination types: factual entity errors, relational inconsistencies, or numerical fabrications, without capturing their co-occurrence ...Faithfulness evaluation of abstractive summaries remains an open challenge, with existing metrics addressing only isolated hallucination types: factual entity errors, relational inconsistencies, or numerical fabrications, without capturing their co-occurrence or interaction. We introduce CHI (Composite Hallucination Index), the first unified hallucination metric that decomposes faithfulness errors into three orthogonal dimensions: entity hallucination (EHI), relation hallucination (RHI*), and quantity hallucination (QHI). Each dimension employs a shared softmax-normalized architecture over Venn diagram-derived factors representing extractiveness, positive hallucination, over-focus, negative hallucination, and lost focus. The novel QHI component introduces tolerance-aware numerical matching with exact, epsilon, derived, and temporal comparison modes. We fuse the three dimensions via harmonic mean to produce a single composite score that penalizes weakness in any dimension. We validate CHI on 800 source articles spanning four domains (news, medical, legal, financial) with summaries from five generation systems. Empirical results demonstrate that: (i) the three dimensions are statistically orthogonal (mean rho = 0.148), confirming they capture distinct error types; (ii) CHI achieves the highest system-level correlation with human judgments (rho = 0.66, p = 0.006) on SummEval, outperforming ROUGE (rho = 0.53), EHI (rho = 0.58), and all individual components; and (iii) ablation studies confirm that all three dimensions contribute unique variance, with the full composite outperforming any individual component while providing decomposable error diagnostics unavailable from single-score baselines. CHI provides practitioners with a decomposable, interpretable, and efficient faithfulness metric suitable for both offline evaluation and online monitoring of summarization systems.
|
| 504 |
Language Discrimination Improves Linguistic Learning in Multilingual Speech Models
2609.33345
|
cs.CL
|
Maureen de Seyssel, Jie Chi, Zakaria Aldeneh |
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate lang...Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT setting, we test two interventions which strengthen language discrimination: an auxiliary language classifier and per-language k-means targets. Across interventions, continuous-feature phone discrimination error (phone-ABX, lower is better) decreases from 11.6% in the bilingual baseline to 10.4% (monolingual: 10.8%), while lexical performance (sWUGGY, higher is better) increases from 52.1% to 56.7% (monolingual: 58.5%) and prosodic performance (ProsAudit, lexical subtask, higher is better) from 68.9% to 72.9% (monolingual: 72.6%). Across HuBERT training stages, the strongest gains on most linguistic measures occur when language discrimination is introduced in the first iteration, whereas later or repeated interventions yield smaller improvements and are accompanied by increased language-wise segregation. These results support a causal role for language discrimination in reducing the additional cost of multilingual learning.
|
| 505 |
NLPG: Natural-Language Policy Gradients for Self-Evolving Language Agents
2609.33379
|
cs.CL
|
Xu Liu, WenZhang Wei, Jun Cao, Dehua Peng, Huan Chen |
Large language model agents increasingly rely on compound programs for retrieval, tool use, reasoning, and verification, yet their failures often arise from local procedural decisions. Existing reinforcement-learning and prompt-optimization approaches typicall...Large language model agents increasingly rely on compound programs for retrieval, tool use, reasoning, and verification, yet their failures often arise from local procedural decisions. Existing reinforcement-learning and prompt-optimization approaches typically rely on scalar rewards or repeatedly modify entire prompts, making it difficult to capture and reuse procedural improvements while preserving a frozen agent. To address this problem, We propose Natural-Language Policy Gradients (NLPG), an external policy-memory method for improving a fixed agent without changing its model parameters or program structure. NLPG diagnoses execution traces, propagates downstream feedback backward through the module graph, and converts recurring failures into route-local natural-language corrections that are aggregated into bounded policy updates for subsequent executions. Across six benchmarks covering memory, reasoning, instruction following, and evidence verification, NLPG also outperforms the strongest listed baseline for each benchmark by 8.71 percentage points on average. These results provide evidence that evaluated procedural experience can be transformed into local and interpretable policy updates, enabling continual improvement of frozen agents.
|
| 506 |
From Position Risks to Block Survival: Faster Generation for Diffusion Language Models
2609.33390
|
cs.CL
|
Siwei Chen, Yuxiang Wan, Yifan Yu, Fan Lai |
Diffusion language models (DLMs) can accelerate generation by predicting multiple tokens in parallel, but there is a mismatch between how these tokens are predicted and how they ultimately contribute to generation. Parallel predictions can hardly condition on ...Diffusion language models (DLMs) can accelerate generation by predicting multiple tokens in parallel, but there is a mismatch between how these tokens are predicted and how they ultimately contribute to generation. Parallel predictions can hardly condition on the tokens selected earlier within the same block, even though their validity depends on this realized prefix. Under the popular proposal-verification decoding, this mismatch makes errors highly asymmetric: an early rejection prevents all subsequent proposals from contributing decoding progress. We introduce BRISK-DLM, a framework that addresses both mismatches by optimizing proposal learning and selection for verified progress. BRISK-DLM trains on self-generated sequences, using risk-reward weighting to dynamically prioritize positions by their impact on verified progress and decoding cost. During inference, a lightweight prefix-conditioned corrector reranks existing candidates using previously selected tokens and preferences distilled from the model's own verifier. The corrector reuses the backbone's parallel representations and requires no additional backbone evaluation, while fused execution keeps its overhead small. BRISK-DLM improves end-to-end throughput by up to 37.4% while preserving task quality, establishing a new quality-throughput frontier for DLM generation.
|
| 507 |
Preserving Morphemes: Morphology-Guided Pre-Tokenization for Nepali
2609.33395
|
cs.CL
|
Kalash Shrestha, Nikhil Pradhan |
A byte-level BPE vocabulary learns each inflected form of a Nepali word as a separate string, so a noun stem is spelled differently in each of its case-marked forms. We test whether splitting words into stem and affixes before BPE helps, with the corpus, vocab...A byte-level BPE vocabulary learns each inflected form of a Nepali word as a separate string, so a noun stem is spelled differently in each of its case-marked forms. We test whether splitting words into stem and affixes before BPE helps, with the corpus, vocabulary size, model and number of training steps held fixed. Our pre-tokenizer, Papaya, uses a finite-state transducer built from a published grammar of Nepali, falls back to regular expressions, and leaves the BPE trainer unchanged. On 607 words annotated by seven native speakers its segmenter reaches 0.96 boundary F1, and the resulting tokens keep stems intact far more often than plain BPE does. In a 17M-parameter language model it lowers bits per byte by about 1% at equal training steps; most of the larger gain seen at equal epochs comes from the extra steps that longer token sequences buy, and an unsupervised Morfessor segmentation gives the same improvement. Downstream the effect is small: NER improves only on entities that contain words unseen in training, POS tagging and news classification do not change, and published Nepali tokenizers perform about as well. We release the annotated boundary set, a 556-affix dataset and the code.
|
| 508 |
Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation
2609.33409
|
cs.CL
|
Yuhao Sun, Binrui Wu, Zhuoer Xu, Ming Wen, Haoxiang Xu |
On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision...On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future behavior, while consequential guidance may be beyond the current student's reach or fail to persist without privileged input. We formulate long-horizon OPD as hierarchical supervision allocation and argue that productive guidance lies at the intersection of future utility and current learnability. Crucially, this intersection evolves as the student learns. Based on this principle, we propose LENS-OPD, a coarse-to-fine framework that organizes supervision through Locate, Validate, and Refine. Locate adapts trajectory exposure to the student's evolving competence and proposes a candidate decision for intervention. Validate tests whether teacher guidance at that decision improves the same student's subsequent behavior. Refine internalizes the beneficial guided behavior into the deployable policy and concentrates token-level supervision on decisive teacher-student conflicts within the validated turn. These stages are nested: each finer allocation is conditioned on the coarser decision, rather than being optimized as an independent importance score. Experiments across multiple long-horizon agent benchmarks and student-teacher configurations show that LENS-OPD consistently improves task performance over vanilla OPD and strong curriculum- and selection-based baselines. Our results suggest that effective long-horizon distillation requires teaching at the right depth, the right decision, and the right token.
|
| 509 |
Grounding Memory Summarization in Utility Intent
2609.33417
|
cs.CL
|
Zhenyu Lei, Mingjia Shi, Xingbo Fu, Haoyu He, Qi R. Wang |
Existing summarizers for memory systems are typically optimized for human-facing criteria such as faithfulness, which misaligns with their true objective: preserving the evidence needed to support future queries. We show that conditioning summarization on quer...Existing summarizers for memory systems are typically optimized for human-facing criteria such as faithfulness, which misaligns with their true objective: preserving the evidence needed to support future queries. We show that conditioning summarization on query-answer pairs substantially improves answer quality, and that this utility-aware behavior is transferable across queries. Motivated by these findings, we propose MemSuit, a self-distillation framework in which a teacher summarizer, conditioned on observed query-answer pairs, produces utility-aware memory entries that a student learns to reproduce from the raw conversation alone. To prevent collateral erasure where conditioning on a single query-answer pair discards evidence relevant to other plausible queries, the teacher decomposes each block into multiple self-contained entries that preserve distinct query-relevant facets as independently retrievable units. To align the retriever with the compact, fact-dense style of teacher entries, we further fine-tune the embedding model with a contrastive objective supervised by teacher entries. Across a diverse suite of conversational query types, MemSuit consistently outperforms state-of-the-art baselines, confirming the value of grounding memory in downstream utility.
|
| 510 |
TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment
2609.33426
|
cs.CL
|
Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng, Zaiyi Zheng |
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mit...Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.
|
| 511 |
MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation
2609.33441
|
cs.CL
|
Ruihong Zeng, Jonathan Tonglet, Preslav Nakov, Iryna Gurevych |
Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. H...Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what human fact-checkers often check: whether an image's content is consistent with the context implied by its accompanying claim. To address this gap, we introduce MIC (Multimodal Inconsistency Checking), an AFC framework that assists human fact-checkers by detecting AI-generated multimodal misinformation and explaining inconsistencies using world knowledge. MIC first uses supervised fine-tuning (SFT) for task adaptation and then applies Group Relative Policy Optimization (GRPO) to directly optimize component-level verifiable rewards for verdict prediction, inconsistency type classification, visual evidence description, and world-knowledge explanation. We further introduce MIC-Bench, a benchmark comprising 8,812 image-claim instances derived from 4,406 claims, where each claim is paired with an authentic image and an AI-generated counterpart that introduces a controlled contextual inconsistency. Compared with SFT alone, GRPO further improves Macro-F1 by 4.67 and 4.11 points in the in-distribution and out-of-distribution settings, respectively, while also improving the semantic similarity of visual evidence descriptions and world-knowledge explanations to reference annotations. Our code and data are available at https://github.com/UKPLab/arxiv2026-mic.
|
| 512 |
Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends
2609.33443
|
cs.CL
|
Seonghyeon Go, Yongwoo Kim, Hyeonjin Cha, Jaeho Shin |
Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trap...Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving them unable to access real-time information and tool execution. Furthermore, even when Large Language Models (LLM) retrieve information, many duplex speech models process it within a compressed latent space rather than in its raw text form, which can lead to information loss from compression. To address this issue, we propose Context Spanning, a framework for information injection between a full-duplex speech model and an external LLM backend via real-time chunked prefill. The injected frame is encoded in a single forward pass inside the real-time frame budget. It feeds the retrieved information to the speech model as-is, enabling it to reason over the information independently and generate responses. With this approach, our model achieves high performance on Full-Duplex benchmarks and strong results on Question Answering tasks, demonstrating its conversation potential. Context Spanning shows that external information can be injected directly into a duplex speech model, introducing a new simple and powerful mechanism for duplex systems.
|
| 513 |
Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
2609.33463
|
cs.CL
|
Cunchun Li, Haonan He, Yifan Gao, Minglei Li, Jingqi Ye |
Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-l...Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-reweighting methods and show that they assign nonnegative coefficients to demonstrated tokens. Consequently, they can suppress or amplify supervised updates, but cannot reverse harmful features once learned. Moreover, larger training weights do not amount to feature extrapolation, since they change the optimization trajectory rather than scale a fixed SFT direction. We argue that reversal and extrapolation require a stable reference frame defined by a fixed SFT delta. Motivated by this, we propose SCALE (Selective Control of Adaptation via Local Entropy), an entropy-guided adaptation-strength-control method that freezes the pretrained model and the SFT delta and learns bounded token- and module-specific gates by minimizing predictive entropy alone. These gates suppress, reverse, or extrapolate frozen SFT features according to their alignment with entropy reduction. Across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Qwen3-4B-Base, SCALE achieves mathematical-reasoning averages of 37.84, 43.60, and 36.57, exceeding the strongest corresponding baselines while remaining competitive on general-retention benchmarks. It also attains the best average code-generation performance across HumanEval, HumanEval+, and MBPP for all three models. These results suggest that effective SFT correction can benefit from controlling how already learned residuals are used, rather than only modifying how they are learned.
|
| 514 |
GSM: Efficient Language Modeling with Shared Global State
2609.33465
|
cs.CL
|
Yunao Zheng, Bin Wen, Xiaojie Wang, Kaiyu Jiang, Xuanyu Zheng |
Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decod...Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage. Through multiple stages of history retrieval, the encoder progressively incorporates long-range information into representations at recent positions, forming a shared state with a fixed window size. Each decoder layer accesses this same state using queries updated from the preceding layer, preserving computational depth while avoiding repeated construction of historical key--value (KV) representations and long-range indexing. As a result, neither the decoder's per-step attention cost nor its KV cache size grows with the history length. Experiments show that GSM improves computational efficiency and reduces cache overhead while maintaining model performance and the ability to use long-range information, offering a shared-state architecture for efficient language modeling.
|
| 515 |
DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation
2609.33485
|
cs.CL
|
Guanzheng Chen, Viet Dac Lai, Subhojyoti Mukherjee, Branislav Kveton, Seunghyun Yoon |
While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive se...While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding exhausts the representational capacity needed for complex reasoning. To resolve this, we propose Grounding-Reasoning Disaggregation via DIStributed long COntext scaling (DISCO). Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding. A central Driver LLM, trained via Reinforcement Learning (GRPO) to optimize planning, orchestrates execution by dynamically mapping queries into atomic extraction tasks and reducing the gathered evidence to synthesize a final answer. By isolating reasoning from raw context noise, DISCO effectively eliminates context rot. On RULER-QA (1M tokens), it maintains 78.4% accuracy where standard baselines collapse. Furthermore, it outperforms full-context models by up to 9.8 points on LongBench v2 and matches frontier models like Gemini-3-Pro-Preview while reducing inference costs by over 80%, establishing a highly efficient paradigm for robust long-context inference.
|
| 516 |
LLMs Trust Their Own: Identity-Dependent Conformity in Multi-Agent Systems
2609.33495
|
cs.CL
|
Liron Soffer, Ravid Shwartz-Ziv, Chen Shani |
Large language models (LLMs) are increasingly deployed in multi-agent settings, where agents observe and influence one another, making social influence a key dimension of AI behavior and safety. We investigate whether LLMs' responses depend on the social ident...Large language models (LLMs) are increasingly deployed in multi-agent settings, where agents observe and influence one another, making social influence a key dimension of AI behavior and safety. We investigate whether LLMs' responses depend on the social identity of other agents, beyond the effect of their consensus. We construct judgment tasks with a single correct answer, and place models in a multi-agent setting where they receive incorrect answers from other agents whose social identities (AI or human, model family, or an arbitrary minimal group) are either shared with or distinct from their own. Across 12 open-weights models and nine tasks, we find a bidirectional effect of group identity on conformity to incorrect answers: in-group consensus increases conformity (in-group favoritism), whereas out-group consensus decreases it (out-group divergence). Unlike humans, for whom one ally breaking the consensus sharply reduces conformity, models are unmoved by an ally from the majority's group. Worse, a correct ally from the opposing group intensifies this bidirectional effect. Chain-of-Thought reasoning suppresses most of these effects, yet an in-group ally still reduces conformity to an incorrect out-group majority. Labeling peers as safety-aligned shifts overall conformity but leaves in-group favoritism and out-group divergence intact. These results show that group identity shapes how LLMs aggregate information across agents, independently of its correctness, and identify a manipulation surface for multi-agent AI systems.
|
| 517 |
E-CONAN (Entailment, CONtradition And Neutral) Diagnostics Dataset Investigating Linguistic Phenomena in Arabic Natural Language Understanding
2609.33530
|
cs.CLcs.LG
|
Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi |
Natural Language Understanding (NLU) plays a crucial role in various applications, yet its performance suffers from weaknesses in handling the complexities of human languages, ranging from lexical ambiguity to high-level reasoning difficulties. Analyzing error...Natural Language Understanding (NLU) plays a crucial role in various applications, yet its performance suffers from weaknesses in handling the complexities of human languages, ranging from lexical ambiguity to high-level reasoning difficulties. Analyzing errors across diverse linguistic phenomena is crucial for NLU improvement, as it will help humans get insights to comprehensively assess models' limitations and capabilities, so optimizing models' generalization. Notably, several benchmarks contain diagnostics datasets designed for investigation and fine-grained error analysis. When highlighting the gaps in the state-of-the-art, we noted that there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena that should be covered. To overcome this gap, we propose an initial hierarchy for Cross-Lingual NLU error analysis. Moreover, we propose a methodology to create an NLI hierarchical framework and applied a case study on Arabic NLU. Moreover, this paper introduces E-CONAN diagnostics dataset, a freely available dataset manually-annotated with coarse-grained and fine-grained categories based on our proposed Arabic hierarchy. E-CONAN dataset helps NLU designers better understand their models by doing error analysis and in-depth investigation. We used E-CONAN to investigate the performance of 9 pretrained language models and 5 LLMs. Results indicate that LLMs outperform pretrained models in world knowledge and commonsense reasoning macro-category, and underperform pretrained models in syntactic macro-category. Moreover, the hardest phenomena for all models is Reasoning, and the easiest phenomena for all pretrained models is Syntactic, and the easiest for LLMs is Lexico-Syntactic.
|
| 518 |
ManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold Perspective
2609.33534
|
cs.CL
|
Rui Liu, Chenheng Zhang, Haoxuan Li, Zhouchen Lin |
Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form kn...Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form knowledge, suffering from severe edit forgetting and degradation of general capabilities. To address these challenges, we reframe knowledge editing from a manifold perspective, viewing it as a localized displacement of an edit sub-manifold within the global knowledge manifold. Under this formulation, the problem can be decomposed into two key questions: (i) how to identify representative edit points that effectively anchor the edit sub-manifold, and (ii) how to preserve the remaining manifold structure during the sub-manifold displacement process. Based on this perspective, we propose ManiEdit, a novel manifold-aware autoregressive editing framework consisting of two core components. Pivot Localization addresses the mediocre-point dilemma by identifying high-leverage pivots to anchor the edit sub-manifold. Manifold-Aware Preservation preserves different knowledge types through an energy-weighted penalty combined with recursive null-space alignment. Experiments on two base LLMs and four unstructured editing benchmarks demonstrate that ManiEdit achieves state-of-the-art performance, outperforming the strongest baseline by up to +27.81 BERTScore and +8.50 ROUGE-L, while maintaining near-original general capabilities across six representative downstream tasks. Our code is available at: https://github.com/Areyliu/ManiEdit
|
| 519 |
Jev Matches 7B Language Models for Speech-Neuroprosthesis Rescoring
2609.33538
|
cs.CL
|
Gabriele Cin\`a |
A speech neuroprosthesis decodes attempted speech from brain activity and ends by rescoring the decoder's candidate sentences with a language model of several billion parameters, the only component that needs a GPU. Replacing that model with a cheaper one is h...A speech neuroprosthesis decodes attempted speech from brain activity and ends by rescoring the decoder's candidate sentences with a language model of several billion parameters, the only component that needs a GPU. Replacing that model with a cheaper one is hard: general language models asked to pick one sentence from a list answer from where a label sits in the list rather than from the sentence itself. We pose rescoring as a single typed decision, one call that returns a probability for every candidate, served by Jev, a hosted model trained for calibrated decisions, and combine it with the decoder's own score. On 978 held-out sentences from a participant with ALS, where the published decoder alone reaches 8.1% word error, Jev reaches 7.5% against 7.8% for both OPT-6.7b and Qwen2.5-7B; with the decoder's weight re-tuned, 6.9% against 7.2% and 7.4%. Jev is ahead in all four comparisons and at most 0.2 points behind at the 95% bound. It costs 0.07 USD per thousand sentences and needs no GPU; a dedicated GPU running a 7B model is cheaper per sentence only above 43% utilisation, far beyond what one user generates. End-to-end latency over the internet is 262 ms, of which 62 ms is spent at the provider, the same order as a 7B model on a local GPU (27 ms) but not faster.
|
| 520 |
Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL$\rightarrow$FOL
2609.33612
|
cs.CL
|
Pu Suo, Ali Emami |
A standard pipeline for symbolic reasoning over natural-language problems translates them into first-order logic and invokes a theorem prover. The translation step is the bottleneck: swap "every" for "some" and every inference that follows is corrupted. Yet to...A standard pipeline for symbolic reasoning over natural-language problems translates them into first-order logic and invokes a theorem prover. The translation step is the bottleneck: swap "every" for "some" and every inference that follows is corrupted. Yet today's metrics often score more broken translations higher than less broken ones, because BLEU, BERTScore, and Smatch++ reward surface overlap that the worst errors happen to preserve. We introduce SIV, which derives two kinds of probes from the target formula and uses a theorem prover to verify the candidate translation against each. Positive probes are statements the candidate must entail, which detect translations that drop content; contrastive probes are statements the candidate must not entail, which detect translations that assert more than the original. On a controlled pool of perturbed FOLIO translations, the severity of the error accounts for 80% of SIV's score variance, compared with at most 17% for any prior metric. Across six error classes on a disjoint pool, SIV scores the reference above the perturbed candidate in over 99% of pairs. Because each probe is labeled with what it tests, the failure pattern also supplies a labeled error trace, recovering the perturbation class at macro-F1 0.638, nearly double the score-only baseline. On 434 expert-audited real LLM translations, SIV attains the top AUC, uniquely detects and grades expert-labeled major errors, and abstains, rather than mis-scoring, on out-of-vocabulary translations.
|
| 521 |
Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails
2609.33634
|
cs.CL
|
Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen, Robert Birke |
Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a ...Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our LLaDA-Guard asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA-Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with unsafe-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting unsafe prompts into safe equivalents without additional training and achieving a 60.7% average conversion-to-safe rate.
|
| 522 |
Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
2609.33642
|
cs.CL
|
Haoyi Wu, Yang Xiao, Yusong Sun, Wenyang Hui, Zhaokai Luo |
Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and diff...Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.
|
| 523 |
Tsubame: Tree Replay for Diffusion-Based Speculative Decoding
2609.33652
|
cs.CL
|
Yepeng Weng, Qiao Hu, Takehisa Yairi |
Context-aware dynamic trees allocate the speculative decoding budget according to draft path probabilities, adapting their depth and branching to the current context. Under stochastic decoding, however, we find that this structural advantage does not always co...Context-aware dynamic trees allocate the speculative decoding budget according to draft path probabilities, adapting their depth and branching to the current context. Under stochastic decoding, however, we find that this structural advantage does not always compensate for the acceptance gains of random sampling paired with advanced verification, and such dynamic trees can fall behind sampled chains in some settings. These trees grow their topology from the candidates themselves, so the tokens submitted for verification are typically the deterministic high-score tokens selected during construction. This coupling is not inherent: once the topology is fixed, its nodes can be repopulated by sampling, allowing dynamic trees to retain their structural advantage while also benefiting from random sampling and advanced verification. Diffusion-based drafters make this practical, as their parallel outputs or lightweight conditional corrections allow candidates to be regenerated cheaply after the complete topology is known. We introduce Tsubame, a two-pass tree speculative decoding framework for diffusion-based drafters. The first pass plans and freezes a context-aware topology using draft path scores; the second replays the fixed topology, sampling the tokens that populate its nodes to form the candidate tree for verification. We prove that Tsubame is lossless under compatible sampling and verification strategies. Experiments across three diffusion-based drafters, six datasets, and multiple candidate budgets show that Tsubame improves acceptance length and throughput over deterministic trees, including settings where it reverses their disadvantage against sampled chains.
|
| 524 |
One Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment Generation
2609.33666
|
cs.CL
|
Nafis Irtiza Tripto, Delvin Ce Zhang, Mahjabin Nahar, Dongwon Lee |
Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can cap...Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can capture the diversity inherent in human discourse. This concern is increasingly important, as the growing presence of homogenized AI-generated content risks reducing diversity over time, potentially leading to model collapse and degrading the richness of digital communication. Inspired by the plurality of human crowds and the aspect-driven nature of discourse, we hypothesize that comment diversity is better approximated by combining multiple LLMs with aspect-conditioned generation. We formalize and evaluate this approach using models from different providers and introduce a framework that characterizes diversity across semantic, linguistic, and socio-pragmatic features along three axes: dispersion, coverage, and alignment. Using this framework, we conduct a large-scale study on over 2 million YouTube comments across multiple domains. Our results reveal that multi-LLM and aspect-conditioned generation better align with human comment distributions and such data remains viable under pretraining style curation and is effective for downstream tasks. Yet, human diversity remains unmatched. Overall, our findings provide a practical foundation for generating more diverse and socially grounded discourse in AI-mediated environments.
|
| 525 |
Closing the Cross-Dialect Gap: Query Plans as a Portable Interface in Text-to-SQL
2609.33670
|
cs.CL
|
Corentin Royer (IBM Research, Zurich, Switzerland, ETH Zurich, Zurich |
Text-to-SQL systems are typically trained and evaluated on a single dialect (SQLite), yet production deployments span PostgreSQL, MySQL, ClickHouse, and beyond. We show that this single-dialect assumption leads to a substantial drop in cross-dialect accuracy f...Text-to-SQL systems are typically trained and evaluated on a single dialect (SQLite), yet production deployments span PostgreSQL, MySQL, ClickHouse, and beyond. We show that this single-dialect assumption leads to a substantial drop in cross-dialect accuracy for every model we tested. The drop persists across scale, architecture, and even purpose-built text-to-SQL systems. We argue that the fix is to change the generation target: instead of asking an LLM to emit dialect-specific SQL, we have it emit a dialect-agnostic relational algebra query plan, which a deterministic compiler then renders into SQL for any supported backend. Across thirteen models from 3B to frontier scale, this restores cross-dialect portability nearly uniformly, at a small cost in peak accuracy on the model's home dialect for capable prompted models and none once fine-tuned on plans; under matched fine-tuning, plan supervision yields a stronger model than SQL supervision. We also introduce MetricName, a question-aware result-set comparator needed to evaluate fairly across dialects, where existing metrics confound semantic errors with benign cross-dialect variation. More broadly, the result is a reminder that a generation target chosen for execution is not necessarily the one that maximizes generation quality.
|
| 526 |
Reset Is Not Recovery: Evaluating Recoverability from False Conversational Context via Sycophancy Hysteresis
2609.33672
|
cs.CL
|
Adi Shnaidman |
Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context ...Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context behavior after a user repeatedly advocates a wrong answer and then withdraws that pressure. We introduce a recovery-after-pressure protocol for multiple-choice factual dialogue and measure sycophancy hysteresis, the residual probability assigned to the user-advocated wrong answer relative to a clean-context counterfactual. Across seven instruction-tuned open-weight models and two factual benchmarks, ordinary reset often reduces but does not erase pressure-induced bias. History preserving repairs such as user retraction, system reset, and self-verification recover only 2-3/14 model-dataset pairs under the strict clean-restoration diagnostic, whereas operations that change the effective context are substantially more reliable; the two conditions that remove the pressure-bearing history entirely, fresh-context deletion and context truncation, recover 14/14. In an oracle trusted-evidence condition across fourteen model-dataset pairs, preserving the pressure-bearing history while adding benchmark-derived trusted evidence increases accuracy from 0.368 to 0.929, while wrong-answer following falls from 41.2% to 4.3%. Controls show that the effect is not explained by dialogue length, repeated confidence, plausible distractors, mere false-answer mention, or option-label inertia. These results suggest that faithful grounded dialogue requires evaluating which prior context should be treated as evidence and which should be removed or quarantined before answering.
|
| 527 |
When Do Agents Help? Embedding, LLM and Agentic Alignment of Classical Texts and Their Translations
2609.33691
|
cs.CL
|
M\'at\'e Metzger |
Classical texts aligned with their translations support machine translation, retrieval and computational research, but evidence comparing alignment workflows is scattered. This study compares seven systems on 452 texts in Pali, Sanskrit, Mishnaic Hebrew and Ti...Classical texts aligned with their translations support machine translation, retrieval and computational research, but evidence comparing alignment workflows is scattered. This study compares seven systems on 452 texts in Pali, Sanskrit, Mishnaic Hebrew and Tibetan, comprising 9,833 human-aligned units: four embedding pipelines, a direct LLM call, an autonomous agent, and the agent revised by an independent auditor. Generative workflows recover 93-94% of reference correspondences, against at most 77% for embeddings. A ceiling analysis shows that sentence boundaries make some references unrepresentable by the embedding pipelines. Reference recovery is similar across generative workflows: the agent's advantage is 0.5 percentage points (95% CI -0.02 to 1.17), and auditing adds no established benefit. Agents nevertheless produce structurally valid output for all 452 texts, against 437 for direct calls. A blinded three-LLM panel assesses every generative mismatch against the source and human reference. Most mismatches are labelled defensible editorial variation; consensus major-error labels cover only 0.06-0.14% of units. The panel labels significantly fewer residual defects for agents than direct calls (0.7% versus 1.4%), suggesting that reference recovery alone understates alignment quality. On ten long Pali discourses taken as published online, agents and audited agents raise recovery from the direct call's 71% to 84% and 92%. Identical reference-located chunks bring all three to 93%. Agents thus improve structural reliability and reduce judged defects on short passages, while their large recovery advantage on long documents disappears after chunking. In this setting, independent auditing offers little measurable additional benefit on prepared passages.
|
| 528 |
Understanding Confabulation and Rethinking Reconstruction in Activation Explanations
2609.33702
|
cs.CL
|
Gert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen |
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, ...Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
|
| 529 |
From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection
2609.33720
|
cs.CL
|
Yu Tian, Andrew Potter, Katerina Christhilf, Motahareh Darvishpour Ahandani, Jessica Early |
Revision traces provide valuable evidence about students' writing processes, but their usefulness for learning analytics depends on how individual revisions are represented. Automated draft-comparison methods often produce granular edit operations that can fra...Revision traces provide valuable evidence about students' writing processes, but their usefulness for learning analytics depends on how individual revisions are represented. Automated draft-comparison methods often produce granular edit operations that can fragment a single purposeful revision into multiple analytic units. This study evaluates whether LLMs can identify meaningful revision unit boundaries in structured revision operation data and whether they provide value beyond simple non-LLM baselines. Using 113 matched draft--revision pairs from undergraduate writing, expert annotation yielded 4,344 candidate boundaries. We compared zero- and few-shot GPT-5.5 and base Qwen3-32B, parameter-efficient fine-tuning of Qwen3-32B, and majority and proximity-based baselines. Despite receiving revision context and task instructions, no prompted LLM condition outperformed the proximity heuristic (macro-F1 = .825). In contrast, fine-tuned Qwen3-32B using the two context representation achieved the highest macro-F1 (.859), identifying more same-unit relationships while maintaining precision comparable to the heuristic. Deterministic post-processing substantially improved the prompted models but added little benefit to the strongest fine-tuned model. These findings suggest that LLMs can support revision boundary judgment when task-adapted, but general purpose prompting alone may not outperform transparent structural heuristics.
|
| 530 |
Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs
2609.33721
|
cs.CL
|
Shiyu Ni, Keping Bi, Jiafeng Guo, Yilong Xu, Jingtong Wu |
Reliable self-assessment is essential for large language models (LLMs), yet they often remain highly confident when their answers are wrong. We study \emph{concurrent confidence calibration}, where confidence is learned alongside capability improvement rather ...Reliable self-assessment is essential for large language models (LLMs), yet they often remain highly confident when their answers are wrong. We study \emph{concurrent confidence calibration}, where confidence is learned alongside capability improvement rather than calibrated only after training. Reinforcement learning from verifiable rewards (RLVR) provides a natural setting for this paradigm, as it continuously produces responses paired with verifiable correctness feedback. Existing concurrent methods, however, learn both capability and confidence through reinforcement learning within shared policy parameters, potentially coupling two fundamentally different learning problems. We instead propose \emph{shared experience but separate learning}: capability and confidence learn from the same trajectories, but through separate optimization mechanisms and parameters. Based on this principle, we introduce \textbf{CoCal (Companion Confidence Calibration)}, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged. Experiments on Qwen3-8B and Qwen3-14B show that CoCal improves confidence estimation without sacrificing task performance, outperforming both RL-based concurrent methods and matched post-hoc calibration. The learned companion further generalizes across domains and policy shifts, while the benefits of CoCal persist at both scales.
|
| 531 |
Yor\`{u}b\'{a} in Unicode: An Overview of a Problem
2609.33734
|
cs.CL
|
K\'ol\'a T\'ub\`os\'un |
There is a recurrent problem in the writing of Yor\`ub\'a on the internet and on the computer that has proven intractable over the years. The language, along with other African languages that depend on diacritics for disambiguation, requires a small set of pre...There is a recurrent problem in the writing of Yor\`ub\'a on the internet and on the computer that has proven intractable over the years. The language, along with other African languages that depend on diacritics for disambiguation, requires a small set of precomposed characters that Unicode does not encode. This has forced writers and digital systems to rely on combining character sequences that behave inconsistently across platforms, corrupt under font substitution, and fail in search. This paper documents that failure across a range of real world contexts, from published books to web platforms to mobile keyboards, using personal and empirical evidence. It identifies Unicode's NFC normalization stability policy as the structural constraint that prevents a straightforward fix, arguing for direct intervention of the Consortium in solving the active problem, proposing a formal encoding request for the four core Yor\`ub\'a characters as the most durable path to resolution.
|
| 532 |
The Effects of Incremental Instruction Delivery on Language-Model Creative Writing
2609.33738
|
cs.CL
|
Anshuman Singh, Abrar Eyasir, Haseeb Yaqoob, John Manavalan |
Large language models are increasingly used as interactive writing tools, where users develop stories, revise ideas, and introduce new requirements across multiple turns rather than specifying a complete brief upfront. Yet most evidence on multi-turn instructi...Large language models are increasingly used as interactive writing tools, where users develop stories, revise ideas, and introduce new requirements across multiple turns rather than specifying a complete brief upfront. Yet most evidence on multi-turn instruction degradation comes from tasks with objectively verifiable outcomes, leaving unclear whether incremental interaction harms creative artifacts in ways that explicit requirement checks cannot capture. We study this question using 160 human-authored creative-writing tasks across six genres, presenting each intended specification either upfront or progressively over 5-9 turns to six distinct open-weight model families, yielding 960 matched pairs. Progressive delivery reduces explicit constraint adherence and produces its largest writing-quality degradation in structure/coherence. The structural gap persists among outputs with equal observed adherence, suggesting that measured requirement loss alone does not explain the observed structural difference. We define Creative Integrity as a compact measure of joint adherence and narrative structure; under incremental delivery, models retain 71.2% of FULL Creative Integrity (95% CI [68.2%, 74.3%]). A three-rater human study over 50 matched pairs independently recovers FULL advantages in structure/coherence, craft, and genre effectiveness, while automated scores remain positively associated with aggregated human ratings. These findings show that interactive creative-writing systems should be evaluated not only on whether requirements survive conversation, but also on whether evolving requirements remain coherently integrated into the final artifact. Our dataset, benchmarks, and source code are available at: https://github.com/solusops/SISTER-2026-Team19
|
| 533 |
Positions Are Not Facts: The Mismatch Between KV Caches and Memory
2609.33759
|
cs.CLcs.LG
|
Changhai Zhou, Yuhua Zhou, Shiyang Zhang, Jun Gao, Zhen Li |
When a fact changes, how should a language model update the history stored in its key-value (KV) cache? Hiding the old record is cheap, but it may still contain needed details or answer questions about the past. We compare hiding whole records, hiding only rep...When a fact changes, how should a language model update the history stored in its key-value (KV) cache? Hiding the old record is cheap, but it may still contain needed details or answer questions about the past. We compare hiding whole records, hiding only replaced values, and deleting old text and recomputing the cache. In a controlled quantity task, masking makes all eight models prefer the new value more strongly, yet six lose complete answers through unit errors or failure to stop; keeping the unit preserves all current answers. Later states also retain useful information from earlier records: on multi-hop updates, rebuilding these states at unchanged positions lowers historical accuracy by 20-41 percentage points, whereas moving the existing states has little effect. Keeping object dependencies and unchanged revision passages prevents many losses. Recognition is a separate challenge. Learned readouts recover distinctions missed by fixed cache similarities on synthetic record pairs. On natural text, text-detector-selected masks show no clear advantage over random masks at the same rate in 14 same-model detector-generator comparisons. Query-dependent access can avoid some losses, with additional storage or access costs. These findings identify what must be preserved beyond the replaced value when using a KV cache as updatable memory.
|
| 534 |
LA-CPD: Local-Evidence-Aware Change-Point Detection for Human-LLM Authorship Segmentation
2609.33787
|
cs.CL
|
Qing Yang, Zhenyu Mao, Zixiang Luo, Zezheng Wu, Xinghe Cheng |
As LLM-generated text becomes increasingly human-like, accurately localizing LLM-authored spans in human-LLM co-authored documents is important for attribution and accountability in cases involving copyright infringement, fraud, and other harmful uses of AI-ge...As LLM-generated text becomes increasingly human-like, accurately localizing LLM-authored spans in human-LLM co-authored documents is important for attribution and accountability in cases involving copyright infringement, fraud, and other harmful uses of AI-generated content. Sentence-level detectors provide local authorship evidence, but content variation can cause score fluctuations even among sentences from the same source, creating spurious boundaries. Recovering a coherent document partition therefore remains challenging when both the number and locations of authorship transitions are unknown. We propose Local-Evidence-Aware Change-Point Detection (LA-CPD), a structured method that transforms noisy sentence-level score sequences into coherent authorship segments. Given scores from a frozen local detector, LA-CPD combines a length-weighted within-segment residual with a windowed two-mean contrast to capture segment consistency and sustained changes around candidate cut points. Dynamic programming optimizes cut locations for each candidate count, while an AIC-style criterion selects the final partition, yielding sentence labels, authorship boundaries, and maximal LLM-authored spans. On a held-out human-LLM co-authored test set, LA-CPD outperforms WCP+AIC, increasing sentence-level accuracy from 0.747 to 0.796 while improving boundary localization and LLM-span delineation.
|
| 535 |
In-Context Adaptation of Encoder-Decoder Models in Speech Recognition
2609.33865
|
cs.CL
|
Yen Meng, Sharon Goldwater, Hao Tang |
In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models ar...In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, when providing interleaved speech-text demonstrations. In this work, we ask whether in-context adaptation is an inherent ability for all encoder-decoder models. We study two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures. We find that all tested models are able to perform in-context adaptation out of the box, achieving up to 30% relative improvement in the oracle experiments and up to 23% using first-pass hypotheses. Through controlled experiments on three English datasets, we show that lexical and speaker information both contribute to successful adaptation. While interleaved demonstration is effective in certain cases, collated demonstration brings consistent adaptation across the board. Our results suggest that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.
|
| 536 |
Lost with a Map: Conversational State and Behavioral Reliability in Language Models
2609.33883
|
cs.CL
|
Atahan Dokme, Larry Heck |
Task-oriented dialogue requires maintaining and updating information across turns, yet language models expose no explicit belief-state object. We study how conversational state is represented, updated, and used inside eight instruction-tuned language models fr...Task-oriented dialogue requires maintaining and updating information across turns, yet language models expose no explicit belief-state object. We study how conversational state is represented, updated, and used inside eight instruction-tuned language models from four families on MultiWOZ and SGD. Structure and values separate: which domains, slots, and requests are active is linearly readable just before the model acts, whereas exact values are far more readable where the user stated them. After a user changes a value, both values remain accessible at their mentions, and causal interventions show that both continue to influence the model's action. In natural closed-loop interaction, query failures separate into cases of weak structural support, incorrect value resolution, and failure to deploy otherwise-supported constraints, with targeted interventions producing systematically different repair behavior across these cases. These findings motivate a state-action controller that starts from the base model action and selectively edits it using structural readouts, without requiring a complete predicted belief state as an intermediate representation. On held-out MultiWOZ interaction across five models, it raises the base model exact-query accuracy from .318 to .621 and task success from .272 to .371 at negligible added cost. Overall, reliable interaction requires not only retaining conversational information, but resolving which available constraints currently apply and ensuring that they govern action.
|
| 537 |
LLMs learn different forms of metacognition when trained to predict their own accuracy
2609.33886
|
cs.CL
|
Nicolas Yax, Stefano Palminteri, Pierre-Yves Oudeyer |
Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, a...Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.
|
| 538 |
Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees
2609.33887
|
cs.CLcs.LG
|
Jungseob Lee, Dongyub Jude Lee, Chanjun Park, Sugyeong Eo, Heuiseok Lim |
Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how ...Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on prompts that the slower one answers correctly. On the serving engine and its decode traces, the default commit rule already commits every fully resolved block, a static skip rule captures nearly all of the compute that allocation can save, and self-distillation on engine-decoded targets adds speed at unchanged accuracy. Larger speedups come from lower thresholds, which commit tokens that are still uncertain. We therefore present Redline, a finite-sample procedure that selects operating points, hand-picked or learned, from the correctness of their answers on calibration prompts. Redline keeps the reference-relative risk, the joint probability that the reference answers correctly and a candidate configuration does not, within a user-chosen budget with high probability, and deploys the fastest configuration that passes. It speeds up math at a smaller risk budget than code in both model families, and at a budget of ten percent it deploys a LLaDA2 math configuration that commits over a third more tokens in each forward. It also applies without modification to the acceptance rule of speculative decoding and to weight quantization. On the same calibration data, Redline stays within its stated failure probability, whereas each tolerance of a mean-accuracy rule either gains less speed for some model and task or exceeds the risk budget far more often for another. Code is available at https://github.com/js-lee-AI/Redline.
|
| 539 |
NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models
2609.33899
|
cs.CL
|
Ziwei Chen |
We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ...We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final turn. Our pilot contains 22 human-verified pairs (44 audio conditions) and evaluates five models on NSV perception, emotion understanding, and response adaptation. Results show that models generally perform better at detecting NSVs than at interpreting their fine-grained emotional meaning or producing appropriately differentiated responses. The data construction pipeline, dataset, and evaluation pipeline are publicly available at https://github.com/ChenzwNina/nsv-construction.
|
| 540 |
SlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI Writing
2609.33905
|
cs.CL
|
Dhruv Roongta, Harsha Gaddipati, Anh Tuan Huynh |
SlopBench asks which models produce the stiff, repetitive prose readers call AI slop, a question detectors leave open once they have classified a text as machine-written. We evaluated eighteen models on 112 hand-written tasks in email, social posts, essays, an...SlopBench asks which models produce the stiff, repetitive prose readers call AI slop, a question detectors leave open once they have classified a text as machine-written. We evaluated eighteen models on 112 hand-written tasks in email, social posts, essays, and workplace chat, sampling each model on each task up to ten times, for 19,928 outputs in all. SlopBench scores four surface behaviors a reader can check by hand: length against the word band each task specifies, opener repetition across a model's own samples of one task, and paragraph rhythm and fixed lexical constructions against pre-ChatGPT human corpora. Under one fixed weighting, Kimi K2.6 scores lowest at 21.1 and Mistral Large highest at 40.6. Across 500 random reweightings Kimi has the lowest score in 58 percent of draws and Mistral the highest in 97 percent. No draw preserves the full order of the eighteen, and a scenario bootstrap leaves exactly one of those ranks unambiguous. We ran three further checks on that middle order: a crowd arena, an AI detector, and lexical diversity. None of them confirmed the order. We therefore report the four behaviors separately and treat the composite as one weighting among many, and we release the prompts, outputs, reference statistics, and scoring code.
|
| 541 |
Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
2609.33923
|
cs.CLcs.LGcs.AI
|
I Kennedy, T Kennedy |
A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, ...A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96 times the i.i.d. noise value of the same shape, and on the MoE the median is unchanged from 2-bit to 8-bit. The probe coefficient of variation is predictable from tensor shape. One probe measures per-tensor sensitivity to within 4 to 7%; twenty probes reach 1.3 to 1.4%.RAM applies the propagated form of this estimator to budget-targeted mixed-precision quantization with no calibration data. Gaussian probes carrying the network's own input statistics score every tensor at six bit-widths. A knapsack solver allocates bits under an exact byte budget, with guardrails against catastrophic 2-bit assignments. One probe pass serves any budget. Isolated and propagated scores rank tensors independently on Qwen3.5-35B-A3B (Spearman -0.01), yet the propagated probe rank-correlates 0.81 to 0.83 with the GPTQ layer objective from real activations, while the isolated estimator is uncorrelated with it. That objective is the wrong allocation target: at matched bytes on Qwen3.8-27B, a block-output probe beats a vendor IQ3_M mix and an oracle that allocates from the real-activation objective. On Qwen3-8B the propagated probe ties HAWQ-V2 at matched bytes. Across seven architectures from 8B to 122B, with probe timing up to a 400B model in nine minutes on one workstation, RAM reaches 3.5 to 13.6% lower median WikiText-2 perplexity than size-comparable uniform 4-bit builds on the tested MoE models. (Black Sheep Ai baa.ai)
|
| 542 |
Simple Diffusion Language Models Are More Effective Few-Step Generators Than Reported
2609.33947
|
cs.CLcs.LG
|
Hasan Amin, Ming Yin, Rajiv Khanna |
Diffusion language models (DLMs) promise fast parallel generation, yet high-quality samples often require large number of refinement steps, which diminishes their advantage in practice. This has led to massive interest in and rapid development of new methods f...Diffusion language models (DLMs) promise fast parallel generation, yet high-quality samples often require large number of refinement steps, which diminishes their advantage in practice. This has led to massive interest in and rapid development of new methods for effective few-step generation. We show that much of the supposed quality gap at few steps can instead arise from a suboptimally configured sampler. Modest sampler sharpening, without any model retraining, enables a couple years old masked DLM to rival supposedly far improved successors. This differently sampled DLM in fact achieves lower generative perplexity in just 16 steps than what its standard sampler obtains with 1024, while improving both judged quality and semantic diversity. We further show that conventional per-output metrics can fundamentally obscure these gains, since any optimal trade-off between two such metrics can be attained by a generator supported on at most two outputs. We subsequently introduce GroupEval, which separately evaluates quality and across-output semantic diversity, and offers fresh insights including uncovering how 1.5-4.7x perplexity gains of a distilled model yield no corresponding quality gain. Finally, we explain why sharpening helps: parallel unmasking destroys dependencies among simultaneously generated tokens, creating a gap between prediction and generation. We prove that pervasive temperature choice of one is generically suboptimal under parallel sampling even for an exact denoiser, and that worse predictions can yield better samples. Through these results, we argue for a broader evaluation principle of treating the deployed generator as the object of comparison, benchmarking it against tuned baselines, and assessing quality and diversity jointly and with more human-aligned measures.
|
| 543 |
On the Token Value Inequality in Efficient Reasoning
2609.33970
|
cs.CL
|
Runjia Zeng, Hang Hua, Yiyang Liu, Zhiqiang Tao, Ruixiang Tang |
Chain-of-Thought reasoning has enabled large language models to achieve substantial performance gains on complex tasks. However, these gains come at the cost of dramatically increased token consumption. This raises a fundamental question: is every token in the...Chain-of-Thought reasoning has enabled large language models to achieve substantial performance gains on complex tasks. However, these gains come at the cost of dramatically increased token consumption. This raises a fundamental question: is every token in the reasoning trace equally valuable? We present a diagnostic and optimization framework grounded in a key empirical finding: the value of tokens within a CoT reasoning sequence is highly non-uniform, and this non-uniformity can be effectively characterized by token-level log probability signals. We show that normalized log probability helps distinguish core tokens, which carry structural and decisive reasoning content, from redundant tokens, which are exploratory, low-confidence filler that contributes less directly to the final answer. Building on these findings, we formulate the TokenProbe framework around two empirical findings and one claim: findings identify token value inequality first and then establish TokenProbe as a core-token proxy, and the claim introduces an efficient GRPO objective positing that selectively compressing redundant tokens can yield Pareto improvements in the accuracy-token efficiency space. Empirically, our method preserves reasoning quality while reducing the token usage by 76% of the baseline. Under matched reasoning-length budgets, we show that it can even outperform strong flagship baselines like Gemini-3.1-Pro. Homepage: https://runjia.tech/tokenprobe/.
|
| 544 |
Do System One Decisions Add Up? A Study of Probabilistic Coherence
2609.33971
|
cs.CL
|
Saman Sarker Joy |
A decision model can give probabilities that sum to one for every question yet disagree with itself when the same decision is broken into smaller steps. We study this form of probabilistic coherence in Jev and the English Laya checkpoint, using 2,500 matched e...A decision model can give probabilities that sum to one for every question yet disagree with itself when the same decision is broken into smaller steps. We study this form of probabilistic coherence in Jev and the English Laya checkpoint, using 2,500 matched examples per system across TREC, CLINC150, and MASSIVE. Across 72,000 classification questions, we compare direct fine-label predictions with broad-category probabilities and predictions reconstructed through those categories. Both systems show substantial disagreement: mean category-level total variation ranges from 0.219 to 0.349 for Jev and from 0.424 to 0.689 for Laya, on a scale where zero means exact agreement. The consequences differ sharply. On CLINC150, reconstruction reduces Jev's accuracy by 22.9 percentage points (paired 95% bootstrap interval: [-24.9, -20.9]) and improves Laya's by 21.3 points ([18.0, 24.5]). The same directions hold across all three datasets, with all six unadjusted accuracy-change intervals excluding zero. Improved accuracy can also accompany less reliable confidence: on MASSIVE, Laya gains 9.2 accuracy points while its expected calibration error rises from 0.046 to 0.124. Error analysis identifies both broad-category mistakes and within-category confusions. These findings show why decision systems need joint evaluation of accuracy, confidence calibration, and probability coherence in the workflow used by an application.
|
| 545 |
Beyond Solo and Consistency: Vindicating Multi-Agent Debate via Conditional Progressive Pruning
2609.33974
|
cs.CL
|
Ruosong Ye, Caiqi Zhang, Jiahao Li, Haijun Wu, Xiaolong Luo |
Large Language Model (LLM) based Multi-Agent Debate (MAD) is one of the most effective test time scaling techniques. Through multi-round communication, agents complement each other in knowledge and reasoning and solve tasks that no single member can solve. How...Large Language Model (LLM) based Multi-Agent Debate (MAD) is one of the most effective test time scaling techniques. Through multi-round communication, agents complement each other in knowledge and reasoning and solve tasks that no single member can solve. However, existing MAD frameworks fail to beat strong Single Agent and Consistency-based baselines under the same strict cost limit, which shakes the foundation of the MAD field. We propose Conditional Progressive Pruning (CPP), a lightweight pruning framework that fully exploits multi-round MAD. CPP outperforms all existing MAD frameworks on multiple dominated benchmarks. It is also the first to fully outperform consistency methods. Our code, detailed agent interaction records will be released soon.
|
| 546 |
High-Level Text Preprocessing for Semantic Similarity Analysis of Discursive Texts: A Framework and Empirical Demonstration
2609.33983
|
cs.CLcs.LG
|
Mehmet Murat Albayrakoglu, Mehmet Nafiz Aydin |
Semantic Textual Similarity (STS) methods assume that a document's lexical content faithfully represents what it asserts. This assumption fails for discursive documents that discuss, compare, critique, and contextualize other positions in the process of articu...Semantic Textual Similarity (STS) methods assume that a document's lexical content faithfully represents what it asserts. This assumption fails for discursive documents that discuss, compare, critique, and contextualize other positions in the process of articulating their own. The result is semantic diffusion: similarity scores between documents are inflated by vocabulary acquired through discursive engagement rather than substantive alignment. Standard Natural Language Processing (NLP) preprocessing (tokenization, stopword removal, stemming, lemmatization) cannot address this problem because it operates at the lexical level, treating all content identically regardless of its discursive function. This paper introduces high-level text preprocessing: a systematic, rule-based intervention applied before the standard preprocessing pipeline to isolate each document's actual claim from its discursive structure. We propose 12 rules, each with an explicit rationale, and demonstrate their effect on an encyclopedic philosophical corpus: three entries from the Stanford Encyclopedia of Philosophy (virtue ethics, deontological ethics, and consequentialism). A three-phase experiment using eight Transformer-based STS models shows that preprocessing reduces centroid cosine similarity scores across all three theory pairs, with 23 of 24 model-pair comparisons showing the expected decrease and cross-model agreement ranging from 7-1 to 8-0. We introduce the semantic diffusion index (SDI), a per-document metric for assessing the semantic reorientation between a document's raw and high-level preprocessed representations. Although the framework is demonstrated using philosophical texts, it potentially addresses a domain-agnostic problem applicable to legal texts, policy documents, academic articles, and any genre in which a discursive approach introduces vocabulary from positions the document does not endorse.
|
| 547 |
Opera: A Verbal Critic Framework for Long-horizon Coding Agents
2609.33987
|
cs.CL
|
Kai Mei, Zhiyuan Hu, Yutong Dai, Juntao Tan, Yifan Zhang |
Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely...Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades.
|
| 548 |
RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback
2609.33989
|
cs.CLcs.LG
|
Jingyi He, Nier Wu, Shuang Liu, Xin Wang, Mengnan Du |
Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the res...Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Existing interpretation methods often rely on predefined high-level attributes and require repeated counterfactual interventions for each response pair to validate candidate explanations, lacking a closed-loop mechanism that uses RMs' feedback to train a reusable explainer. To address this, we propose RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer. RewardExplainer generates open-ended, atomic, and intervenable natural-language scoring mechanisms, making explanations more concrete, readable, and actionable. It further converts counterfactual feedback into preference supervision, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation. Extensive experiments across multiple target RMs and explainer backbones show consistent improvements. Beyond interpretation, we use the generated mechanisms to identify potential bias patterns and construct targeted debiasing data for fine-tuning the reward model, improving robustness on reward-hacking benchmarks.
|
| 549 |
Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation
2609.34033
|
cs.CL
|
Haiyan Zhao, Zirui Hei, Wei Shi, Huiqi Deng, Na Zou |
Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descripti...Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to trust and use reliably in practice. To this end, we introduce AVPO, a two-stage framework that first reconstructs source text from a hidden activation and then evaluates the resulting text with a separate frozen question-answering model, yielding an explicit and inspectable intermediate readout. We further optimize the inverter with direct preference optimization (DPO), using rewards that capture both semantic recoverability and lexical fidelity. Across six text families, AVPO improves gist- and detail-level information recovery over the strongest baseline by up to 17.1 and 9.3 percentage points, respectively. Crucially, the gains arise from preference optimization rather than fine-tuning on selected reconstructions alone, enabling compact cross-model inverters to surpass donor-matched question-conditioned verbalizers while improving both semantic recoverability and lexical fidelity. Moreover, out-of-distribution case study shows that AVPO better recovers high-level semantics while fabricating fewer details.
|
| 550 |
Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
2609.34065
|
cs.CLcs.LG
|
Mir Tafseer Nayeem, Davood Rafiei |
Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often ...Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often fails at the lexical interface: matched names are not necessarily matched inputs. Some names receive direct single-token access, while others are assembled from multiple subwords, creating unequal name-surface support. Across nearly half a million first names and 12 LLM-associated tokenizers, direct lexical access is highly selective, model dependent, and uneven across race- and gender-associated name metadata. We introduce NameTrace, a model-native, fine-grained, pre-behavioral framework for measuring whether unequal name-surface support remains a vocabulary property or becomes visible in task-relevant internal representations. NameTrace measures concept accessibility from the model's own probabilities over task-specific adjective axes with continuous task-aligned weights. On matched atomic and short-fragmented names within the same race/ethnicity--gender-associated strata, support predicts systematic differences in concept accessibility across fellowship, hiring, clinical assessment, and lending. These differences persist across all eight matched strata, extend across model families, and transfer to unseen names. Hidden-state interventions further show that the measured task directions have downstream leverage, shifting later constrained choices. Unequal lexical support is therefore demographically structured at the input and remains visible in task-relevant model computation. NameTrace makes lexical comparability measurable, supporting a broader principle: behavioral comparability begins with lexical comparability.
|
| 551 |
Evaluating Machine Unlearning in ASR
2609.34092
|
cs.CL
|
Diogo Dinis, Francisco Teixeira, Bhiksha Raj, Alberto Abad, Isabel Trancoso |
Machine unlearning (MU) offers a path to compliance with "right to be forgotten" regulations. While MU has received increasing attention for speech tasks, it remains largely unexplored for Automatic Speech Recognition (ASR). In this work, we investigate whethe...Machine unlearning (MU) offers a path to compliance with "right to be forgotten" regulations. While MU has received increasing attention for speech tasks, it remains largely unexplored for Automatic Speech Recognition (ASR). In this work, we investigate whether existing MU algorithms and evaluation tools are suitable for ASR. We apply several MU techniques to an ASR model, evaluating privacy-utility trade-offs for single-subject unlearning, then assess the best algorithm under sequential and simultaneous unlearning. Results show that gradient ascent-based algorithms achieve strong utility-privacy trade-offs, whereas more complex approaches over-unlearn samples, making them easier to identify as unlearned. This suggests standard privacy evaluations based on simple Membership Inference attacks are insufficient to reliably assess unlearning success, motivating improved evaluation methods for MU in ASR. Finally, we show that both sequential and simultaneous unlearning yield worse privacy and utility than single-subject unlearning, underscoring the need for unlearning constructions better suited to these settings.
|
| 552 |
Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores
2609.34112
|
cs.CL
|
Nicol\'as Vera Z\'u\~niga |
Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extracti...Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LLM) from decision logic (deterministic code computing score bounds over unknown inputs) lets a system ask only questions that can change the decision. On 1,200 synthetic emergency cases across six calculators (HEART, CURB-65, qSOFA, PERC, Wells, Cockcroft-Gault), with a simulated clinician answering questions, we compared this bounds policy with asking for every missing input, a missing-equals-normal schema, and an end-to-end LLM agent (Claude Opus 5.5). With Claude Haiku 4.5 as extractor, the bounds policy matched ask-all accuracy (99.4% vs 99.4%) with half the questions (0.92 vs 1.78 per case) and no irrelevant ones. Treating missing as normal dropped accuracy to 91.2% and under-triaged 8.5% of patients (95% CI 7.1-10.2), and under-triage persisted under messy notes and a noisy clinician. The agent was equally accurate under ideal conditions (99.6%) but 9.5% of its questions were irrelevant; with a noisy clinician it was less accurate than the bounds policy (83.5% vs 87.0%, p<0.001) and committed prematurely in 2.7% of cases (bounds: 0%). A 9B local model as extractor reached oracle-level accuracy (99.8%). In 584 real case reports from MedCalc-Bench, only 52% contained enough information to determine the category (HEART 13%). Routing decisions through code that reasons explicitly about unknowns avoids premature commitment and irrelevant questions, halves the questions asked, and works with small local models.
|
| 553 |
Understanding Clinical Cognitive Dialogues Using Large Language Models
2609.34125
|
cs.CL
|
Vishalakshi Arumugam, Dan Schumacher, Veronica Rammouz, Enrique Gonzalez Guerrero, Jeremy Davis |
In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label th...In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study these behaviors at scale. We present an de-identified corpus of 33 cognitive assessment conversations with 8,250 utterances annotated for three speaker roles and 56 dialogue acts. We use this corpus to benchmark large language models on fine-grained dialogue-act classification and next-patient-utterance generation. We also test whether out-of-domain instruction data and explanation-augmented training transfer to this clinical setting. Instruction tuning produces the strongest patient-utterance reference matching and improves classification accuracy. Reasoning-aware fine-tuning produces the strongest classification results among the LLaMA-3.1-8B variants. However, even the best models struggle to separate closely related dialogue acts, showing that broad conversational intent is easier to recognize than fine-grained communicative function. The corpus and benchmark make interaction structure measurable in cognitive assessments and support follow-up work on conversational markers, clinician education, and carefully validated simulated patients. This work does not make diagnostic claims. Instead, it provides the data and evaluation framework needed to study these applications.
|
| 554 |
Word Similarity Datasets for Indian Languages: Annotation and Baseline Systems
2609.34138
|
cs.CLcs.LG
|
Syed S. Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, M. Shrivastava |
With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word similarity datasets of six Indian langu...With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word similarity datasets of six Indian languages - Urdu, Telugu, Marathi, Punjabi, Tamil and Gujarati. These languages are most spoken Indian languages worldwide after Hindi and Bengali. For the construction of these datasets, our approach relies on translation and re-annotation of word similarity datasets of English. We also present baseline scores for word representation models using state-of-the-art techniques for Urdu, Telugu and Marathi by evaluating them on newly created word similarity datasets.
|
| 555 |
Quantitative Measurement of Language Distance among Closely Related Indo-European Languages Using Pretrained Language Models: A Case Study on the North Germanic Branch
2609.34152
|
cs.CL
|
Yiping Bai |
Among closely related North Germanic languages, the quantification of language distance has traditionally relied on qualitative methods, lacking a unified multi-dimensional computational framework. Multilingual pretrained models based on the Transformer archit...Among closely related North Germanic languages, the quantification of language distance has traditionally relied on qualitative methods, lacking a unified multi-dimensional computational framework. Multilingual pretrained models based on the Transformer architecture can map texts from different languages into a shared vector space, enabling quantitative measurement of language distance. This paper focuses on the three North Germanic languages---Danish, Norwegian (Bokm\aa{}l), and Swedish---and proposes a three-metric quantitative framework based on pretrained language models: (1)~sentence-level semantic distance, computed as cosine similarity between LaBSE and mBERT encodings of parallel sentences; (2)~orthographic fragmentation rate, measuring subword tokenization efficiency when cross-applying monolingual BERT vocabularies to parallel texts; (3)~MLM predictability, comparing prediction confidence and entropy in masked language modeling using mBERT across languages. Using 150 trilingual parallel sentence triplets from the Tatoeba corpus as controlled samples, we obtain consistent distance rankings on two independent models: LaBSE: da--no $0.012 < $ no--sv $0.016 < $ da--sv $0.020$; mBERT: da--no $0.016 < $ no--sv $0.045 \approx $ da--sv $0.046$. This ranking is consistent with the historical linguistic conclusion that ``400 years of Danish rule over Norway (1380--1814) led to highly cognate written languages.'' The three metrics---semantic, orthographic, and predictability---converge on the same conclusion, providing a reproducible computational framework for the quantitative study of distance among closely related languages, extensible in principle to more branches of the Indo-European language family, pending validation on additional language groups.
|
| 556 |
Toward a Graded Measure of Belief Stability in Large Language Models
2609.34158
|
cs.CL
|
Samantha Dies, Branden Fitelson, Tina Eliassi-Rad |
Large language models (LLMs) increasingly mediate how people access and reason with information, yet factual reliability is usually evaluated one judgment at a time. We introduce graded belief stability, a relational measure of how well a belief persists withi...Large language models (LLMs) increasingly mediate how people access and reason with information, yet factual reliability is usually evaluated one judgment at a time. We introduce graded belief stability, a relational measure of how well a belief persists within an LLM's broader belief system. Unlike individual belief probability, it asks whether support for a claim persists when that claim is considered alongside the model's other epistemic commitments. We operationalize this idea with a Direct Conditional estimator that uses internal model representations to estimate conditional belief probabilities. Across 12 LLMs and three domains, lower-stability beliefs exhibit greater mean behavioral movement under conversational challenge in 83.3% of model-domain settings after matching on individual belief probability. Graded belief stability therefore extends reliability assessment beyond how strongly an LLM supports a claim to how robustly that belief is supported within its broader system of beliefs.
|
| 557 |
LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks
2609.34187
|
cs.CLcs.AI
|
Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer |
The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bo...The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bound of pattern reuse despite producing alluringly fluent text. We test this hypothesis using conlang-like tasks. Several LLMs are given only natural-language descriptions of fictional languages that subvert prominent superficial patterns in training data by combining statistically uncommon and unattested features. Crucially, no example outputs are given. We argue that if the models exhibit rule-following behaviour, they cannot be relying solely on superficial statistical patterns; such patterns often work against the correct output. Instead, successful performance requires representations of the constraints specified in the prompt. Across three complementary task families, models systematically move in the meaning-predicted direction: they distinguish prompt exposure from instructed use, alter semantic relationships in response to novel constraints, and sometimes produce exact matches to complex translation answer keys. Although performance varies across the spectrum of models used, these results provide evidence for meaning-mediated abstraction in LLMs and refute the strong stochastic parrot hypothesis. Our work shows that, under appropriate architectural and contextual constraints, statistical learning can produce meaning-mediated abstractions, although generation remains strongly constrained by superficial plausibility. We discuss implications for model development and for understanding how increasingly abstract representations may emerge from plausible-text-generation objectives.
|
| 558 |
USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents
2609.34225
|
cs.CL
|
Qiyong Zhong, Mao Zheng, Mingyang Song, Huwei Ji, Houcheng Jiang |
On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the...On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the most direct way of incorporating them, at the cost of conflicts between their data distributions and of retraining the entire model whenever one domain is revised. Model merging avoids both by distilling every domain independently and fusing the resulting task vectors afterwards. We find instead that the benefit polarizes across domain pairs: on those exhibiting negative transfer, every merging operator we evaluate falls below the single-domain reference. We attribute this to cross-domain update coupling, where a substantial fraction of coordinates is updated comparably by both domains and a merge can therefore displace them by as much as their own updates. To overcome this limitation, we propose USA, which converts per-parameter update magnitudes measured during a brief warm-up into per-coordinate perturbation radii, reducing curvature precisely on the coordinates that carry most of the merging displacement. Experiments across mathematics, science and code at two student scales show USA strongest in all six transfer directions, ahead of the single-domain reference by more than four points on average, and reverse the negative transfer of the conflicting pairs.
|
| 559 |
MAS-OPD: On-Policy Distillation for Multi-agent Systems
2609.34234
|
cs.CL
|
Qiyong Zhong, Mao Zheng, Mingyang Song, Houcheng Jiang, Jiajie Su |
Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prom...Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.
|
| 560 |
Coherence-Aware Distributional Evaluation of Open-Ended Text Generation
2609.34240
|
cs.CLcs.AI
|
Jinnuo Liu, Junhao Zhu, Weifeng Jiang, Haoming Liu, Hongyi Wen |
Existing open-ended generation metrics measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be l...Existing open-ended generation metrics measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be locally fluent while remaining globally contradictory, causally inconsistent, or topically disconnected. We identify representation as a central bottleneck in detecting these failures and introduce CHORD (Coherence-aware Hidden-state Open-generation Reference Distance), a coherence-sensitive distributional metric. CHORD encodes generated and human-written corpora in the hidden-state space of a frozen LLM using a coherence-eliciting prompt, and compares the resulting distributions using RBF-MMD. To test coherence sensitivity and selectivity, we construct a counterfactual evaluation suite pairing graded coherence-degrading perturbations with meaning-preserving controls. CHORD selectively detects relation, discourse, structural, and mixture failures that perplexity, entropy, MAUVE, FBD, and MMD-based baselines either miss or cannot separate from benign rewriting. Factorial ablations show that representation is the primary source of coherence sensitivity, while RBF-MMD improves sample efficiency. Larger backbones capture finer-grained distinctions, but coherence prompting improves selectivity only when the backbone can follow the prompt. On unconditional generation and prefix continuation, CHORD yields model rankings that strongly align with human judgments of whether outputs make sense and appear human-written. Together, these results establish representation design as central to reliable distributional evaluation. Code: https://github.com/MAPS-research/CHORD. Experiments: https://github.com/MAPS-research/CHORD-Experiment.
|
| 561 |
SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents
2609.34247
|
cs.CL
|
Wenyi Yu, Siyin Wang, Terumi Chiba, Xianzhao Chen, Xiaohai Tian |
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringen...Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM (system 1) with a powerful asynchronous slow-thinking LLM agent (system 2). Beyond handling real-time interaction, system 1 learns when to answer directly and when to delegate, remaining responsive during backend execution and seamlessly integrating returned information into the ongoing dialogue without exposing tool traces or losing conversational context. Evaluations on single-turn spoken question answering (QA) and multi-turn conversations demonstrate that adaptive delegation substantially improves accuracy on knowledge-intensive and multi-hop reasoning questions, while knowledge-boundary-aware training avoids unnecessary system 2 invocations. On a customized version of $\tau$-Voice, SALMONN-duo further demonstrates its ability to complete environment-grounded, policy-constrained tasks through multi-turn interactions in realistic business scenarios. Finally, cost-aware reinforcement learning further enhances the trade-off between task performance and backend usage across the QA and conversation tasks, while improving task success and response safety on $\tau$-Voice with an acceptable increase in the delegation rate.
|
| 562 |
Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study
2609.34257
|
cs.CLcs.LG
|
Bibek Bhandari, Kshitij Lingthep |
Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answe...Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answering (QA) using PubMedQA and two Qwen2.5 model sizes, 0.5B and 3B parameters. The study compares a recursive synthetic-data condition, in which generation G(k+1) is trained on answers produced by G(k), against a Human-Control condition that repeatedly uses the original human training data. The study evaluates across four generations from G0-G3 with two random seeds (42 and 123) and a fixed evaluation set of 1,000 expert-labeled samples. The evaluation includes disease and chemical entity F1, context-supported rate, lexical and semantic similarity, answer length, repetition rate, and other evaluation metrics. The Recursive condition for both model sizes and both seeds showed larger declines than the Human-Control condition in disease entity F1, chemical entity F1, context-supported rate, ROUGE-L, and cosine similarity. Under the fixed no-repeat 3-gram decoding constraint, the main observed behavioral change was increased answer length, while the measured 3-gram repetition rate did not increase. The magnitude of the difference-in-change was larger for the 3B model than for the 0.5B model. This difference was particularly apparent in disease F1, context-supported rate, cosine similarity, and answer length. These results show domain-specific changes associated with using recursive synthetic-data training in biomedical QA, but do not establish clinical hallucination rates or universal model collapse.
|
| 563 |
Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs
2609.34284
|
cs.CL
|
Haeun Jang, Yonghyun Jun, Hwanhee Lee |
Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks ...Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot tell where this failure arises. We decompose preference handling into three stages and measure each separately: (1) knowing whether a preference applies, (2) deciding on an explicit Apply/Suppress label, and (3) generating a response consistent with that label. Using linear probes, we first show that this applicability signal remains decodable from hidden states during generation. By making the decision explicit, we then find that in most settings wrong decisions faithfully followed outnumber correct decisions lost in generation. We thus locate the failure in the decision, which breaks once the model is also asked to answer. To determine whether this reflects lost sensitivity or a response bias, we propose ABIDE (Apply-Bias Investigation via Decision-score), which adapts signal detection theory to Apply-vs-Suppress decision scores read directly from logits. ABIDE reveals a generation-induced Apply bias: merely stating an answer-generation objective shifts the decision score toward Apply while sensitivity is largely preserved, and the shift persists under controls for prompt structure, cascades across preference slots, and prompt wording. Finally, we show that subtracting a single bias scalar, estimated on a held-out split, from the decision score at decoding time reduces leakage while largely preserving fulfillment.
|
| 564 |
ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining
2609.34287
|
cs.CLcs.LG
|
Zichun Yu, Jiarui Yan, Shlok Sanghvi, Nihar Atri, Chenyan Xiong |
LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules...LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a unified language model of only 0.6B parameters that replaces this entire stack. To train ReScraper, we carefully curate supervised data from the outputs of three teacher models, so it learns to first extract the main content from raw data and then choose among four operations: keeping the page as extracted, editing out noisy lines and spans, deleting it entirely, or rewriting it when it is poorly written but informative. Based on the same crawled data pool, pretraining 400M, 1.4B, and 2.8B models on our curated data improves the DCLM Core score by a relative 3.8--4.7% over the strongest baseline at each scale, including the costly multi-agent curation. Our analyses show that each operation plays a distinct and complementary role, and that extracting and cleaning in one model outperforms a cascade of separate models. ReScraper also concentrates its operations on the pages that need them, raising the quality of poor pages the most while keeping the corpus diverse. These results demonstrate the feasibility and effectiveness of AI4AI for pretraining data curation, where a small learned model takes over an entire stage of the pipeline from hand-written heuristics. We open-source our code at https://github.com/cxcscmu/ReScraper
|
| 565 |
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
2609.34296
|
cs.CL
|
Yingjian Zhu, Zhenyi Wang, Jiaxin Guo, Kun Ding, Ying Wang |
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many ex...Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.
|
| 566 |
Certified Selective Automation of LLM Agent Evaluation
2609.34320
|
cs.CL
|
Chengguang Gan, Yunhao Liang, Qinghao Zhang, Shiwen Ni |
Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the e...Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectories stays below a budget alpha? Agent corpora resist the standard answer: many agents attempt the same tasks, so trajectories arrive in correlated clusters, and the i.i.d. certificates of existing selective-judging methods can overstate what is safe: a naive certificate can claim 98% automation while its realized error exceeds the budget in 17.5% of task resamples. We introduce a task-level bootstrap certificate that is valid in every regime we test while matching the naive certificate's coverage; finite-sample cluster-valid alternatives certify nothing at realistic task counts. Under this certificate, a 4B logprob judge trained with SFT and reject-weighted GRPO certifies 0.30-0.59 of evaluation on tool-use and web corpora at alpha=0.1, the only judge, among strongly elicited frontier models, certifying on both headline corpora. Certified coverage is predictable before training from base rate and discrimination alone (leave-one-corpus-out R^2=0.96). Finally, the certificate doubles as a self-training filter: pseudo-labels harvested inside certified regions have contamination bounded by alpha by construction (realized 0.000-0.041 across six harvests), letting a judge enter an unseen domain at in-domain strength with zero target training labels.
|
| 567 |
CRISP: Cultural Reward Modeling for Implicit Situated Propriety
2609.34345
|
cs.CL
|
Zekun Yuan, Yangfan Ye, Baohang Li, Shuaibo Zhao, Zekun Zhou |
As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural k...As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined response spaces, while open-ended culturally situated behavior remains comparatively underexplored. In this work, we introduce CRISP-RM, a culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios. During policy optimization, we further introduce Norm Grounding Supervision (NGS), providing guidance that enhances the policy's sensitivity to relevant cultural norms. To construct culturally situated data, we employ a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios and further curate NormCompass as a dedicated testbed. We conduct comprehensive experiments to evaluate the effectiveness of CRISP-RM in both reward modeling and policy optimization. Best-of-\(N\) experiments show that CRISP-RM consistently outperforms strong general reward models. During GRPO policy optimization, CRISP-RM generally improves culturally situated behavior, while incorporating NGS yields further gains. Further analyses demonstrate the advantages of CRISP-RM in distinguishing culturally appropriate behavior beyond superficial fluency and politeness, while NGS provides complementary gains during policy optimization by improving norm grounding.
|
| 568 |
When Harness Beats Scale, and When Reading Beats Both
2609.34366
|
cs.CL
|
Ivan Bondarenko, Nikolay O. Nikitin |
We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) g...We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interpreter, self-consistency sampling, and entity enrichment from chunk-level knowledge graphs. On our held-out split, application architecture moved the metrics far more than model scale did: PoT added 0.282 joint accuracy to a compact 7B model but at most 0.005 to a 72B model, and a 27B model with the full harness matched the 72B (0.884 vs.\ 0.873) at roughly 2.7$\times$ fewer parameters and a quarter of the CO$_2$. We read this through a distinction between world knowledge, which scales steeply with parameters, and language knowledge, which scales gently, and show that structured-output training makes a compact model harness-ready rather than merely small. On the raster, watermarked test PDFs the same system collapsed to 13.58\% joint (rank 149 of 163); a controlled re-rendering of the validation set reproduces the OCR half of the collapse while bounding what the simulation misses. Auditing the physical nature of evaluation inputs precedes architecture, and the leaderboard's bimodality is consistent with reading quality, not reasoning, having separated the field.
|
| 569 |
Just-In-Time Agent Memory with Runtime Agentic Research
2609.34385
|
cs.CLcs.LG
|
Bingyu Yan, Chaofan Li, Hongjin Qian, Shuqi Lu, Chaozhuo Li |
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine...Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes important. To address this limitation, we propose Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime. A Memorizer preserves complete raw histories in a hierarchical page-store with compact navigational summaries, while a Researcher iteratively retrieves, inspects, and integrates evidence for each request. To train these memory-use behaviors, we introduce Memory-Gym, an evidence-grounded data synthesis pipeline covering nine task types across six domains, and optimize the Researcher through verified-trajectory supervised fine-tuning followed by Hint-guided Group Relative Policy Optimization. We demonstrate the effectiveness of JAM across a variety of benchmarks on agent memory and long-context processing, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches. To support reproducibility and future research, we release our anonymized source code at https://github.com/VectorSpaceLab/general-agentic-memory.
|
| 570 |
Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation
2609.34386
|
cs.CLcs.LG
|
Yongliang Miao, Shuang Liu, Yanguang Liu, Yandong Bai, Mengnan Du |
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes mem...On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99\% on 4B mathematics, while 8K full-parameter profiling shows 70.5\% lower backward memory.
|
| 571 |
Reciprocal Guidance: Orchestrating Draft and Verify Budgets for Advancing the Diffusion-AR Self-Speculation Frontier
2609.34388
|
cs.CL
|
Linye Wei, Shutian Zheng, Haoyu Zeng, Meng Li |
Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying draft...Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared backbone, while enabling longer acceptance lengths. However, the Pareto frontier between aggregate and per-request throughput remains underexplored. At low concurrency, sequential draft-verify execution requires two model forward passes per round, limiting the effective tokens per forward (TPF). By contrast, at high concurrency, longer drafts incur increasingly expensive computation, forcing individual requests to operate under constrained speculation budgets and preventing full exploitation of the full-backbone drafter. Our key observation indicates that drafting and verification exhibit reciprocal predictability. Draft logits can anticipate likely verification mismatches, while recent verification outcomes predict future drafting utility and suitable block sizes. Building on this observation, we introduce Reciprocal Guidance (RecGuide), a runtime draft-verify orchestration framework that adapts speculative decoding to varying serving loads. RecGuide exploits spare compute capacity through verification-overlapped drafting at low concurrency, while dynamically allocating request-specific draft block sizes as the workload becomes increasingly compute-intensive. Experiments across a wide range of concurrency levels demonstrate consistent throughput improvements over vanilla self-speculation, achieving up to $1.8\times$ speedup.
|
| 572 |
Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents
2609.34425
|
cs.CL
|
Suhwan Choi, Myeongho Jeon, Myungjoo Kang |
Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based se...Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to adapt to this variation, causing them to either merge distinct subtopics or over-segment coherent themes. To address this, we introduce Cue-Grounded Segmentation (CGS), a training-free framework that operates without any task-specific supervision. CGS first identifies phrases that explicitly signal the start of a new topic and uses their sentence positions as segment boundaries. When such cues are insufficient, it falls back to semantic segmentation, guided by the document structure inferred during cue extraction. Across six benchmarks and six LLM backbones, CGS consistently outperforms existing baselines, remains robust to noisy ASR transcripts, and achieves these gains with low API cost on proprietary models.
|
| 573 |
AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
2609.34428
|
cs.CL
|
Chanhee Park, Jeongho Yoon, Sungbin Han, Hyeonseok Moon, Heuiseok Lim |
Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing th...Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agentic systems. Existing benchmarks, however, tend to focus on a single leaderboard score, leaving the underlying failure modes opaque. To fill this gap, we introduce AgentHop, a diagnostic benchmark of 1,011 multiple-choice questions paired with a controlled seven-tool sandbox under fixed token, turn, and tool-call constraints. AgentHop reveals model vulnerabilities by dissecting a single accuracy score along four axes of agent operation: retrieval, synthesis, tool-call, and resource management. Across 19 models, we find that behavior clusters by model family, with tool-call signatures revealing distinct family fingerprints: GPT models commit early, Anthropic and GLM checkpoints verify before committing, DeepSeek and Kimi over-search, and Gemini-3 Pro stays balanced. Decomposed axes further expose within-family structure: Claude Opus 4.6 and Sonnet 4.6 land within one accuracy point yet diverge on retrieval-versus-synthesis emphasis, with Opus retrieving more and Sonnet synthesizing better. We release the full benchmark set and the harness to support diagnostic agent benchmarking.
|
| 574 |
Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents
2609.34438
|
cs.CL
|
Wanqi Zhou, Jiawei Lu, Yang Wang, Zhaolong Xing, Zhen Chen |
Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extract...Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that may matter later. To this end, we introduce RIME, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration. RIME uses generic self-questions to retrieve focused dialogue evidence and grounds memory formation in both the retrieved evidence and relevant historical memories, which are jointly reconciled into an evolving memory bank with temporal and provenance information. At inference time, compressed memory serves as the primary rather than the sole source of evidence: when it cannot support an answer, RIME retrieves relevant source dialogue together with its local context to recover information omitted during memory formation, without resorting to full-history processing. Extensive experiments on LoCoMo with Qwen3-235B-A22B and GPT-5.6 Sol show that RIME consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.
|
| 575 |
Unbiased Top-$k$ Estimation for On-Policy Distillation
2609.34447
|
cs.CLcs.LG
|
Linjian Meng, Siyuan Gan, YuHan Li, Xiran Wang, Ziyang Ding |
On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergenc...On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student via rollouts generated by the student's policy. However, estimating the gradient of the reverse KL divergence in OPD remains a challenge. Using only the sampled token from the student-generated rollout is computationally cheap but provides limited distributional supervision, which will degrade accuracy. In addition, using the full vocabulary provides complete distributional supervision but is computationally expensive. Therefore, recent works propose Top-$k$ OPD (TK-OPD) that use selected top-$k$ tokens, which provides richer distributional supervision than sampled-token estimation at substantially lower computational cost than full-vocabulary estimation. Unfortunately, using only the selected top-$k$ tokens induces bias, leading to accuracy degradation, as the probability mass outside the selected top-$k$ tokens is discarded. To address the bias of TK-OPD, we propose Tail-Corrected Top-$k$ On-Policy Distillation (TT-OPD). It preserves the advantages of TK-OPD, including rich distributional supervision and low computational cost, while providing an unbiased estimator of the gradient of the reverse KL divergence. The key insight of TT-OPD is to use not only the selected top-$k$ tokens, but also the sampled token from the student-generated rollout, thereby recovering the discarded probability mass in expectation, avoiding the bias. Experimental results demonstrate that TT-OPD significantly outperforms other tested OPD variants.
|
| 576 |
When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation
2609.34454
|
cs.CL
|
Yekun Xu, Ante Wang, Jingyi Ren, Xuanyi Chen, Weizhi Ma |
Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidenc...Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpasses independent confidence estimators. Our empirical study shows that a dedicated confidence estimator can substantially outperform verbalized confidence, indicating that LLMs' internal representations contain richer confidence signals. Building on this finding, we propose Iterative Policy-Estimator Training (IPoET), a framework that synergizes the complementary strengths of verbalized reasoning traces and informative representations. IPoET alternates policy optimization with estimator updating, integrating estimator-derived confidence feedback into policy learning and refreshing the estimator on new policy rollouts. Experiments across diverse datasets and Qwen and Llama backbones demonstrate that, by iteratively exploiting richer hidden features and adapting to the evolving policy distribution, IPoET consistently outperforms both estimator- and verbalization-based baselines in-domain and achieves superior or comparable results across all out-of-domain metrics. For more details, refer to https://github.com/xyk829/ipoet.
|
| 577 |
RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications
2609.34455
|
cs.CL
|
Jianpeng Zhao, Haihua Xu, Haoyang Zhang, Shuang Qian, Yixiang Tang |
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and...We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules and their applicability, assessing conditions from evidence, combining judgments under rules and exceptions, and providing checkable justifications. These demands motivate a benchmark assessing both decisions and their stated grounds. We introduce RGDT-Bench, providing 202.1K condition-level supervision slots across four task tracks and eight supported task-probe combinations that vary access to supporting information. Label-blind extraction and deterministic checks produce labels for warrant completeness: source-referenced coverage and consistency of stated decision grounds. The benchmark attributes failures to four process layers: rule use, condition, evidence, and aggregation, and checks the final outcome. Among evaluable correct responses, warrant incompleteness averages 40.2% across six evaluated LLMs and supported task-probe combinations. Such warrant incompleteness poses potential safety risks and remains difficult to detect: the best of seventeen existing evaluators reaches only 57.69% (random: 50%) task-averaged area under the receiver operating characteristic curve (AUROC). To address this difficulty, we train a simple reward model with warrant supervision. It achieves 69.24% task-averaged AUROC among correct answers, exceeding the matched outcome-supervised baseline by 10.37 pp (percentage points) and the best existing evaluator by 11.55 pp. Beyond completeness assessment, the model outperforms both outcome-supervised baselines across nearly all response-selection comparisons, supporting RGDT-Bench's warrant supervision for RGDT reasoning.
|
| 578 |
CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR
2609.34481
|
cs.CL
|
Bashar Talafha, Samar M. Magdy, Aisha Alansari, Alaa Alkhawaldeh, Abdurrahman Juma |
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contai...We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.
|
| 579 |
Papers Without Code: Availability of GitHub Repositories Linked in *CL Publications
2609.34534
|
cs.CL
|
Selina Meyer, Michael Roth |
Source code and data published at computational linguistics (*CL) venues are increasingly being shared via GitHub. While this generally is a favourable development for the accessibility and potential reusability of research artifacts in natural language proces...Source code and data published at computational linguistics (*CL) venues are increasingly being shared via GitHub. While this generally is a favourable development for the accessibility and potential reusability of research artifacts in natural language processing (NLP), the long-term availability of such repositories has not been evaluated. In this squib, we discuss the availability of repositories linked in papers published in the Computational Linguistics (CL) journal as well as at ACL and its co-located events over the past ten years. Contrary to our expectations, we find that GitHub repositories linked in more recent ACL publications are unavailable at similar rates as in older publications, in parts due to an increase in empty and placeholder repositories. Similar trends hold for other *CL venues, but not for platforms other than GitHub.
|
| 580 |
RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings
2609.34556
|
cs.CLcs.LG
|
Jarod L\'evy, Mathurin Videau, Jad Yehya, Jean-R\'emi King, St\'ephane d'Ascoli |
Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language model...Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearby tokens. Existing alternatives have been evaluated under different settings, leaving the literature fragmented and without a clear replacement. We bring structure to this landscape by examining a specific weakness of RoPE: its slow frequency bands, whose wavelengths exceed the training context and expose models to unseen angles during extrapolation. We therefore introduce Data aware RoPE (DaRoPE), which preserves standard RoPE on the fast bands but replaces absolute position on the slow bands with bounded coordinates learned from contextual representations. Therefore, the slow-band geometry depends on the data rather than only on positional distance. We compare representative encodings under matched conditions across synthetic tasks, symbolic music, genomics, neural signals, and language models spanning 124M to 50B parameters. Across these experiments, DaRoPE leads on non-text benchmarks, mitigates recency bias, while remaining best or on par in language modeling and length extrapolation. Moreover, the learned coordinates also make the mechanism interpretable, revealing how attention layers leverage contextual information beyond token distance. Together, these results support DaRoPE as the best overall default among the evaluated methods, when there is no domain-specific reasons to prefer another.
|
| 581 |
In-game Toxic Detection: Bi-directional Representations with Attention Residuals
2609.34584
|
cs.CLcs.LG
|
Yuanzhe Jia |
In-game toxic language has emerged as a critical concern in the gaming industry and community. While several frameworks and models for online game toxicity analysis have been proposed, detecting toxicity in player chat utterances remains a formidable challenge...In-game toxic language has emerged as a critical concern in the gaming industry and community. While several frameworks and models for online game toxicity analysis have been proposed, detecting toxicity in player chat utterances remains a formidable challenge: stemming not only from the extremely short length of such utterances but also from the heavy reliance on game slang, abbreviations, and domain-specific jargon, which generic language models are poorly suited to recognize. This paper presents a shared task for in-game toxic language detection built upon real-world in-game chat data, and proposes the best-preforming model for the toxic language slot filling: Bi-directional Representations with Attention Residuals (BRAR). Experimental results demonstrate that BRAR effectively captures the global context and outperforms the existing baselines on slot filling.
|
| 582 |
The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading
2609.34590
|
cs.CLcs.LG
|
Muath Alyobi, Mohamed Eltahir, Almoayyad Abuljdail, Riyadh Almutawa, Tanveer Hussain |
Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, whil...Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model whether it has read enough. We introduce Answer-Convergence Stopping (ACS), a training-free stopping rule that measures rather than asks. After each chunk, it probes the frozen model's current answer state and stops when that state is both confident and stable. The rule requires only output-side generation and token log probabilities, has no trained components, and uses one shared configuration across models and benchmarks. Because a stopping policy can save computation simply by stopping too early, we evaluate the stopping decision itself using evidence position where available. On the full LongBench-v2 with two frontier models, ACS is the only stopping policy that matches or exceeds full-reading accuracy. Furthermore, across 250 S-NIAH questions, the premature stopping rate for ACS across five models from two families ranges from 0% to 12%, compared to 8.4% to 45.6% for the verbalized gate. Taken together, ACS reveals that by properly utilizing the output signals of frozen models, we can achieve favorable behaviors like adaptive stopping without the need for additional training.
|
| 583 |
Rewarding Novel Deductions: Solver-guided Process Rewards for Logical Reasoning
2609.34660
|
cs.CL
|
Muhammad Asif Ali, Wenqing Wang, Huan Wang, Mohammad Raza |
Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale L...Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING: (Solver-guided Process Rewards for Novel LogIcal ReasoNing Step Generation). SPRING uses SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. It introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, it designs process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation across three logical reasoning benchmarks, ZebraLogic, AR-LSAT, and Knights and Knaves, and four LLMs shows that SPRING consistently outperforms base LLMs, outcome-only reward baselines, and Logic-LM. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and strongest outcome-only baseline, respectively. On AR-LSAT, it improves overall accuracy by up to 64.93 and 12.14 points, respectively. On Knights and Knaves, SPRING achieves up to 93.14 puzzle accuracy and 96.05 person accuracy.
|
| 584 |
Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh
2609.34678
|
cs.CL
|
Muhammad Ahmad, Fatemeh Seyedin, Adrian Weller, Dongwon Lee, Mahmoudreza Babaei |
Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language peo...Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language people ask in has gone almost unasked. We test eight models from five families, 3B to 70B, on 1,500 encyclopedic factual claims that exist in identical form in eight languages. English is judged better than every other language on every model, and the gap is widest on the smallest ones, where Llama-3B on Arabic is no better than guessing. Existing remedies retrain on more multilingual data or fit an unconstrained map between language representations, and neither asks whether the model already holds the answer and simply fails to say it. It largely does: a linear probe recovers the truth from the very activations the model fails to express. We propose RoSh, a per-language shift and rotation of the residual stream, computed in closed form at three layers, with no training and no weight modified. It improves every model and closes 75% of the gap on average, helping most where the model was worst: Arabic on Llama-3B goes from chance to nearly the English level, and a fifth fewer of the claims answered correctly in English are lost in translation. What remains is no longer a read-out failure: afterwards the head recovers as much of what is encoded outside English as it does in English. An unconstrained map fitted on the same pairs falls below the untouched baseline, so the orthogonality constraint is doing the work, and every model clears a scrambled-correspondence control and ten further controls. On the two benchmarks of the closest inference-time method, latent-space intervention, run with its own data and metric code, RoSh's gains are five to thirteen times larger.
|
| 585 |
Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis
2609.34691
|
cs.CL
|
Haiyue Yuan, Jie Guo, Weidong Qiu, Zheng Huang, Ruizhe Li |
Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detec...Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self-detection versus cross-detection across model generations, remains poorly understood. We systematically evaluate 15 LLMs spanning three model generations as both generators and detectors. Using a benchmark of 1,000 human-written texts and 15,000 LGTs (1,000 per model), we collected over 233,000 binary classifications alongside natural-language explanations. Our results reveal that detection efficacy is primarily driven by detector capability rather than generator provenance, although outputs from newer generators remain notably harder to detect. Crucially, statistical comparisons show no systematic advantage or disadvantage for self-detection across models. Error analysis further exposes generational bias shifts: first-generation detectors under-detect LGTs (high false-negative rates), second-generation detectors over-flag human texts (high false-positive rates), and the latest models achieve balanced trade-offs. Finally, we highlight significant inconsistencies in how different LLMs apply textual cues to justify their decisions. Code: https://github.com/hyyuan/detect-llm-generated-texts.
|
| 586 |
ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation
2609.34717
|
cs.CL
|
Huifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan |
Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-au...Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organizes program candidates as tree states, retains branch-local debugging context, retrieves failure experience across branches, and distinguishes failed checks from unavailable evidence. On HumanEval and MBPP-Sanitized, visible-test ReMCTS improves over direct generation in 8 of 10 model-dataset pairs under held-out evaluation, whereas proxy-only search is less stable. Controlled tree-search, sampling, repair, and memory ablations characterize the source and limits of these gains. A 30-task HumanEval-X C++ pilot further demonstrates compatibility with compiler-backed execution, but does not constitute a broad multilingual evaluation.
|
| 587 |
Beyond Token Alignment: Event Completion for Cross-Tokenizer On-Policy Distillation
2609.34738
|
cs.CL
|
Jiacheng Liu, Jingwei Song, Qituan Zhang, Siheng Chen, Linfeng Zhang |
On-policy distillation (OPD) transfers knowledge between language models through teacher supervision on student-generated trajectories. With different tokenizers, a single teacher token may require multiple student tokens to generate, creating intermediate sta...On-policy distillation (OPD) transfers knowledge between language models through teacher supervision on student-generated trajectories. With different tokenizers, a single teacher token may require multiple student tokens to generate, creating intermediate states where the event is entered but not yet completed. Existing cross-tokenizer methods align tokens or text spans to construct comparable prediction targets. We study a complementary problem after partial generation: once the student produces a prefix of a teacher token, multiple next tokens may complete the same remaining bytes, but the teacher only specifies the required completion rather than how probability should be divided among these valid continuations. We introduce Event-Set Completion Distillation (ESCD), which complements cross-tokenizer probability alignment with completion-set supervision. ESCD aggregates prefix-related teacher events and supervises the total probability of byte-compatible one-step student completions, avoiding tokenizer-dependent probability splits among individual tokens. The method reuses student trajectories and predictions, requiring neither additional rollouts nor changes to the student vocabulary. Experiments demonstrate consistent gains in mathematics, code, and scientific reasoning across model families and tokenizers, extending to large-scale MoE distillation from a 1T teacher to a 35B student. Local analyses show that retaining completion sets better matches the reference supervision, while one-step completion covers over 99% of observed compatible teacher mass after partial event entry in the studied tokenizer pairs. These findings support event entry and event completion as complementary supervision targets for cross-tokenizer knowledge transfer. Code will be released on GitHub.
|
| 588 |
Draft-KV: Learning Useful Latent Communication Between Language Models
2609.34754
|
cs.CL
|
Linquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang, Peng Zhong |
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrela...Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.
|
| 589 |
LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
2609.34769
|
cs.CL
|
Bingo Zhang, Haochuan Lu, Zongjie Li, Genjian Li, Ari Yu Zhang |
GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether...GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains of coupled decisions. Long-horizon visual puzzles expose this capability directly: a legal move that looks like progress can make the puzzle unsolvable, and the loss shows only several moves later. We introduce LongPuzzleBench, 114 levels in six puzzle games played through native GUI actions, where one objective can take a human over a thousand actions on persistent boards and dead ends go unannounced. With Native GUI Actions alone, the strongest agents solve most objectives, but success falls sharply on harder, longer boards: seven of ten general-purpose agents solve nothing harder than Medium, and none completes Bolt Unscrew Hard, which a human solves along with every other objective. Code Execution CUA does not close this gap, and its scores mix visual solving with algorithmic search. Controlled diagnostics trace these failures to one limitation that neither rules, state hints, nor failure memory removes: agents judge each move by the visible progress it makes, not by the future options it leaves.
|
| 590 |
Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation
2609.34770
|
cs.CL
|
Thodsaporn Chay-intr, Krittapad Harnchang, Mahannop, Thabua, Kobkrit Viriyayudhakorn |
Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do ...Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.
|
| 591 |
TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases
2609.34783
|
cs.CL
|
Fei Lyu, Zhiyi Peng, Jiaming Liu, Yixuan Yang, Changjian Chen |
Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified qu...Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.
|
| 592 |
InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision
2609.34798
|
cs.CL
|
Guanghao Zhu, Zeyu Liu, Zhitian Hou, Pengkai Wang, Zhijie Sang |
Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary s...Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information density, and their utility shifts as training progresses from broad knowledge acquisition to late-stage consolidation. Meanwhile, post-training is often dominated by short-form visual question answering, providing limited supervision for informative and answer-consistent explanations. We introduce InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design. We curate a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing. Our CPT pipeline first adapts the vision encoder, then builds broad medical knowledge, and finally transitions to an evidence-focused data mixture during learning-rate decay. For supervised fine-tuning (SFT), we regenerate visual question-answering responses using answer stability, answer-masked reconstruction, and correctness-constrained selection to produce more informative and answer-consistent supervision. The 4B model is further optimized with reinforcement learning with verifiable rewards (RLVR). Across five medical multimodal benchmarks, InfiMed2-4B achieves 66.73% mean accuracy after RLVR, surpassing the larger Qwen3.5-9B, while InfiMed2-27B reaches 73.72%, the highest among the evaluated open-weight models.
|
| 593 |
Pass or Fail? Evaluating LLMs on Two Greek Examination Benchmarks
2609.34800
|
cs.CL
|
Panagiota Kyriazi, Eleni Kasoura, Prokopis Prokopidis |
The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availa...The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availability of comprehensive benchmarks in this domain, we introduce Prot-Ex and Pan-Ex, two benchmarks consisting of questions from entrance exams for Greek Model and Experimental schools as well as the Panhellenic exams (the Greek national university entrance examinations). These benchmarks are employed to assess the performance of text-only LLMs-including the Greek-adapted KriKri-8B-Instruct, Llama-3.1-8B, Gemma-4-26B, and Qwen-3-32B-across diverse academic disciplines (Modern Greek, Mathematics, Physics, etc.) and task formats (closed, structured, and open-ended), including textualized visual context (i.e., image descriptions). Our findings indicate the localized KriKri-8B significantly outperforms its base model, successfully rivalling much larger LLMs in linguistically demanding humanities tasks. By leveraging an LLM-as-a-Judge methodology, we expose the inadequacy of traditional lexical metrics for evaluating complex reasoning. Crucially, we uncover a few-shot prompting paradox: while synthetic examples improve accuracy in closed-ended questions, they severely overload the context window of 8B models in structured tasks, causing significant performance degradation. Ultimately, this study suggests targeted linguistic adaptation offsets lower parameter counts in specialized domains, despite the fragility of smaller models to prompt verbosity.
|
| 594 |
From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction
2609.34829
|
cs.CL
|
Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu |
Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these com...Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually is costly. We study the upstream problem of constructing the task-specific configuration from a weak specification containing only a short goal and unannotated reference documents. Rather than treating automatic construction as a fixed preprocessing step, our framework constructs a task-specific schema, extraction instructions, and base training rubrics, then keeps schema construction and extraction instructions editable during optimization. Failure-focused updates concentrate textual-gradient feedback on lower-scoring documents, while training-time evaluation criteria adapt to recurring failures. On a heterogeneous-catalysis literature corpus, automatic construction remains improvable, and optimizing both schema construction and extraction instructions performs best across all four judge-rubric settings, with ablations and blinded human evaluation supporting the proposed formulation.
|
| 595 |
OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations
2609.34839
|
cs.CL
|
Faadil Mustun, Chiara Semenzin, Roberto Dessi, Pablo Robin Guerrero, Pierre Orhan |
Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communicatio...Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute for cetaceans: despite bottlenose dolphins (Tursiops truncatus) being a compelling case of complex vocal communication among non-human mammals, existing dolphin datasets are small, fragmented, and largely closed. We introduce OpenWhistle, the largest publicly available dataset of dolphin vocalizations. It comprises approximately 180,000 whistles (114 hours) recorded over five years from a stable pod of five individuals in a semi-natural environment, paired with a curated subset of 8,354 expert-annotated whistles and reproducible evaluation protocols for whistle-type detection and classification. We further release the full processing pipeline for whistle detection, segmentation, and categorization. To demonstrate its utility, we pretrain a Wav2Vec2.0 model adapted to dolphin acoustics on the OpenWhistle corpus and show that it learns effective representations, outperforming general-purpose bioacoustic models such as AVES and BioLingual on both tasks while leaving meaningful headroom for future work. By releasing the dataset, pipeline, and evaluation protocol, we provide the first open dolphin whistle dataset tailored for training self-supervised models, laying the groundwork for advancing dolphin communication research and developing models that capture fine-grained acoustic structure within species.
|
| 596 |
Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction
2609.34841
|
cs.CL
|
Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu |
A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance ...A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operational semantics of an existing schema from a few annotated documents while preserving its structural contract. We introduce CPSE, a contract-preserving semantic extraction framework that jointly calibrates extraction prompts and field-level semantic descriptions from a few gold annotations. CPSE decomposes the schema into an invariant structural contract and mutable field semantics, and further separates identity discovery from record completion using manifest-conditioned resolution. On expert-annotated polymer-science documents, CPSE improves extraction by 9.93 points over an execution-matched baseline, with consistent gains under an independent judge and in a blinded expert audit. These results show that CPSE enables low-resource schema execution while preserving the output structure required downstream.
|
| 597 |
Neural Language Models Learn the Contextual Distributions of Dependency Structures: a statistical learning theory to compositionality
2609.34936
|
cs.CL
|
Wang Bojun, Junjie Chen, Holly Jenkins, Elizabeth Wonnacott |
It is unclear how Neural Language Models (NLMs) acquire the structural meaning encoded by grammatical structures that is independent of lexical semantics. We propose a statistical learning process in which learned dependency structures themselves become new di...It is unclear how Neural Language Models (NLMs) acquire the structural meaning encoded by grammatical structures that is independent of lexical semantics. We propose a statistical learning process in which learned dependency structures themselves become new distributional units for subsequent statistical learning. Under this account, once a dependency structure is acquired, the model tracks its contextual distributions. These contextual features reflect the semantic properties of a composite structure. To test this hypothesis, we design a synthetic grammar in which each grammatical structure has distinct contextual distributions that cannot be recovered from the distributional statistics of their component tokens alone. We train a series of BERT-style masked language models on this grammar and examine their developmental trajectory. The results show that models can successfully learn the contextual distributions of composite dependency structures even though they cannot be inferred from token statistics alone. Developmental analysis further reveals a clear developmental trajectory. The learning of the dependency relations that define a grammatical structure consistently precedes the learning of its contextual features. These findings suggest that statistical learning in NLMs is not merely the accumulation of token co-occurrence statistics, but a process in which learned dependency structures become new units of distributional learning. We argue that this process provides a statistical-learning account of how NLMs solve the compositionality problem in language. Finally, we discuss the possibility that this statistical learning process provides an explanatory theory on how language cognition could emerge from pure distributional statistics.
|
| 598 |
Semantic Uncertainty Quantification Needs Factual Equivalence
2609.34967
|
cs.CL
|
Joseph Hoche, Quentin Guimard, Gianni Franchi |
Semantic uncertainty quantification for large language models rests on a common template: sample several answers, measure how much they agree, and treat disagreement as uncertainty. We first formalize this template as two separate roles: an operator that compa...Semantic uncertainty quantification for large language models rests on a common template: sample several answers, measure how much they agree, and treat disagreement as uncertainty. We first formalize this template as two separate roles: an operator that compares two answers, and an aggregator that combines all pairwise comparisons into a scalar. Existing methods differ almost entirely in how they aggregate, while taking the operator off the shelf, typically an NLI model or a generic sentence encoder. We show that this reliance on off-the-shelf operators is the primary bottleneck of semantic UQ: they do not accurately measure factual equivalence of multiple answers to the same question. We resolve this with a deliberately simple recipe: a single encoder trained contrastively to isolate the targeted fact, utilizing synthetic data generated by an LLM and dataset both disjoint from all evaluation settings. Integrating the resulting operator into existing methods improves performance on 120 of 126 evaluation settings (95%) spanning 18 model dataset combinations across language and vision-language models. The best variant reaches 0.76 mean AUROC against 0.68 for the strongest baseline, while replacing the quadratic cross-encoder comparisons of entailment-based operators with one encoder pass per answer. The uniformity of the improvement supports the view that the operator, not the aggregator, is the limiting factor. The same operator also improves single generation token-level estimators: the norm it assigns to each token measures how much that token bears on the answer, and reweighting token log-likelihoods accordingly sharpens the estimate.
|
| 599 |
N\"urnberg NLP at ChildSafeAds 2026: Structurally Dissimilar Voter Ensembles under Four Levels of Data Access
2609.34986
|
cs.CLcs.LG
|
Philipp Steigerwald, Eric Rudolph, Jens Albrecht |
We describe the N\"urnberg NLP system for ChildSafeAds 2026. The shared task asks what a monitoring system for commercial content in child-facing YouTube videos can achieve at a given level of data access. We answer with per-subtask ensembles of nine voters, o...We describe the N\"urnberg NLP system for ChildSafeAds 2026. The shared task asks what a monitoring system for commercial content in child-facing YouTube videos can achieve at a given level of data access. We answer with per-subtask ensembles of nine voters, organised into three branches that differ in backbone, adaptation method and class scope. Selection rests on channel-disjoint cross-validation, with the development set as a transfer check. The system wins two of the three subtasks. Its product-category score (ST2, 0.8243) and its compliance-flag score (ST3, 0.6530) are the best of the 22 final entries, and it places third on the task mean (0.7079). We further compare four access levels and report the cost at test-set scale.
|
| 600 |
The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents
2609.34988
|
cs.CL
|
Yunhe Su, ZiYi Dong, Tong Yu, Weijian Deng, Hao Li |
Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decis...Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experience reuse a problem of localized control rather than memory alone. We introduce EvoCUE (Evolution through Control Updates from Evidence), a framework for learning reusable control-program updates from completed agent executions. EvoCUE represents the agent as an explicit state-machine controller, whose nodes perform model or tool calls and whose edges define where control passes next. This makes the workflow editable at precise locations, so each learned update can specify what to add, where it acts, and when it applies. From completed trajectories, EvoCUE uses residual goals and observed execution traces to propose localized instruction or skill edits. Each candidate is evaluated at the point where it would act by resuming the parent and edited controllers from the same checkpoint and comparing their final outcomes. Accepted edits are compiled with applicability rules, confirmed on held-out tasks, and inherited by later executions. We evaluate EvoCUE on long tool-use environments where learned conventions must reach the right execution step. From a minimal AppWorld controller without benchmark-specific onboarding instructions, EvoCUE learns the missing task-completion convention and substantially improves success on Test-Normal and Test-Challenge. On PAST-Bench office workflows, EvoCUE transfers organizational requirements from prior episodes to later tasks, improving task-execution quality. These results show that self-evolving agents should place experience inside the control flow, rather than only store it as text.
|
| 601 |
Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices
2609.35070
|
cs.CL
|
Lucio La Cava, Andrea Tagarelli |
Evaluations of Large Language Models (LLMs) morality typically consider decisions in isolation, thus overlooking whether an individual's unrelated prior conduct influences the model's subsequent choices. This leaves open the question of whether, and to what ex...Evaluations of Large Language Models (LLMs) morality typically consider decisions in isolation, thus overlooking whether an individual's unrelated prior conduct influences the model's subsequent choices. This leaves open the question of whether, and to what extent, moral history shapes LLM decisional behaviors. Prior work on human moral decision-making shows that past behavior can influence subsequent moral choices. Building on this observation, we investigate whether analogous effects emerge in LLMs in two complementary ways: at the behavioral level, through the model's observable responses, and at the representation level, through its latent internal representations. We introduce MoralLedger, a framework for studying how an actor's moral history shapes actions for LLMs' behaviors under a fixed decision context. At the behavioral level, we find that prior moral histories systematically alter subsequent choices as a function of their valence and intensity. At the internal representation level, these histories induce a linearly recoverable direction in the residual stream that generalizes to held-out examples. Intervening along this direction on neutral-history prompts produces two-sided intensity-dependent changes in subsequent choices, with effects that are stronger than those induced by prompting alone or by favorable-nonmoral direction. To our knowledge, this is the first demonstration that a latent representation of an actor's prior moral conduct can provide signed inference-time control over a moral decision. Our MoralLedger extends moral evaluation beyond static dilemmas, establishing moral history as both a source of behavioral sensitivity and a causal target for auditing and controlling moral behavior in LLMs.
|
| 602 |
When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment
2609.35074
|
cs.CL
|
Zhaohan Zhang, Junjie Liu, Chengzhengxu Li, Chen Shen, Xiaoming Liu |
The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decisi...The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting this shortcut reasoning is challenging because existing monitors and verifiers mainly inspect textual traces or final outcomes, rather than how the model's belief in its answer develops during generation. We introduce ConfLens, a framework that tracks the evolution of confidence in the final answer throughout reasoning. Across three shortcut reasoning settings, we observe a common pattern of premature confidence, where shortcut samples become highly confident in the final answer at early reasoning stages. Existing confidence estimation methods, however, show limited generalizability, reliability, or efficiency for detecting this behavior. We therefore propose the Distributional Answer Commitment Score (DACS), a distributional confidence estimator that measures the entropy of the model's probability distribution over answer commitment at each reasoning step. DACS captures how concentrated the model's answer belief is without requiring ground-truth answers or task-specific verifiers. We further convert ConfLens detection results into interpretable signals for reward models to reduce their preference for shortcut reasoning. Experiments on mathematical and code reasoning tasks show that ConfLens with DACS improves shortcut reasoning detection by over 4.3% F1 compared with strong baselines and reduces the mismatch between faithfulness and correctness in reward model preferences.
|
| 603 |
A mechanistic study of language model introspection
2609.35108
|
cs.CL
|
Jiahong Zou, Xiangkun Sun, Lingkai Kong, Tonghan Wang |
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a cont...Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.
|
| 604 |
From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data
2609.35201
|
cs.CL
|
Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan, Majd Hawasly |
Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and a...Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p < .001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.
|
| 605 |
Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
2609.35210
|
cs.CL
|
Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou |
On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We st...On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one feature dictionary shared by the student before and after OPD and the teacher. Standard crosscoder analyses, however, identify model-specific features but cannot tell how a model's use of its features changes, since all models are encoded into one set of feature activations. We therefore propose the swap readout, which reads each student checkpoint's feature activations on its own, measuring how training changes the student's use of each feature, even for checkpoints unseen by the crosscoder. Across three OPD settings, we find that OPD neither creates features nor passes on the teacher's own, and leaves the firing rates of over 98% of the student's frequently used features within 20%. We further examine the SFT warm-up on the teacher's rollouts that commonly precedes OPD and makes it more effective. Rather than adding features, the warm-up reweights the shared ones in two ways. First, it already raises and lowers many of the features that OPD later raises and lowers, doing part of OPD's work in advance. Second, it changes features that OPD alone would not, notably those for conversation format, reasoning style, and mathematical notation, and these changes persist through OPD. Imposing this reweighting on a directly distilled student's features, without changing its weights, brings its accuracy close to that of the warmed-up student, whereas the same change on shuffled features does not. Together, these findings suggest that OPD reweights existing features rather than acquiring new ones: the student learns from the teacher how to use the features they already share.
|
| 606 |
SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale
2609.35225
|
cs.CL
|
Zhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai, Rei Kawakami |
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling bet...Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.
|
| 607 |
SCBO: Semantically Coherent Batching and Ordering for LLM-Based Social Surveys
2609.35250
|
cs.CL
|
Yuanzi Li, Lingjie Wang, Zihang Tian, Lei Wang, Xu Chen |
Large Language Models (LLMs) offer a scalable way to simulate survey respondents using demographic profiles and observed reference responses. However, the conventional approach of predicting one question per prompt repeatedly encodes the same context, limits e...Large Language Models (LLMs) offer a scalable way to simulate survey respondents using demographic profiles and observed reference responses. However, the conventional approach of predicting one question per prompt repeatedly encodes the same context, limits each target to a narrow set of reference responses, and prevents later predictions from using information in earlier answers. Predicting multiple questions in one prompt can reduce these costs, share a broader pool of references, and let later predictions build on earlier ones. This requires forming coherent batches, selecting shared references, and ordering questions and references effectively. We propose Semantically Coherent Batching and Ordering (SCBO), a training-free framework that addresses these challenges. SCBO first uses an LLM to extract compact semantic representations from survey items and filter out template noise. It then groups related questions into batches and builds a shared reference bank using target-specific retrieval and centroid-based completion. Finally, it orders target questions from easy to hard and arranges references according to their semantic alignment with those questions. Experiments on four large-scale survey datasets and four LLMs show that SCBO substantially reduces token consumption and inference time while generally improving prediction accuracy over a non-batched baseline. Code is available at https://anonymous.4open.science/r/SCBO-41D8.
|
| 608 |
Rubric-Aware On-Policy Self-Distillation for LLM Personalization
2609.35262
|
cs.CL
|
Yilun Qiu, Xiaoyan Zhao, Chengbing Wang, Cilin Yan, Rui Zu |
LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approac...LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only at a coarse granularity, either by using rubrics to supervise the prediction of relevant aspects for subsequent generation or by reducing aspect coverage to a single response-level reward for reinforcement learning. This leaves a gap between specifying what a personalized answer should contain and teaching the model how to generate it. To bridge this gap, we propose GRASP, a rubric-aware on-policy self-distillation framework for LLM personalization that turns user-specific rubric aspects into fine-grained, token-level supervision. Specifically, GRASP pairs a rubric-free student with a rubric-informed teacher that additionally receives the target user-specific rubrics. By aligning their next-token distributions along on-policy trajectories generated by the student, GRASP transfers the teacher's rubric-conditioned guidance into the student, translating user-specific semantic requirements into dense token-level supervision. Since rubric-informed teachers can still produce inadequate supervision, we further introduce Rubric-based Teacher Validation (RTV), which retains only instances where the teacher sufficiently covers the target aspects, improving both supervision quality and training efficiency. Experiments on the LaMP-QA benchmark for personalized question answering demonstrate that GRASP achieves state-of-the-art performance across multiple backbones, supporting the effectiveness of rubric-guided token-level supervision for personalization. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/GRASP.
|
| 609 |
When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models
2609.35272
|
cs.CL
|
Yizhou Fang, Siyue Chen, Zimo Qi, Zhiyu Xue, Xi Chen |
Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have propose...Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.
|
| 610 |
Measuring Collapse and Correction in Homogeneous-Panel LLM Debate
2609.35279
|
cs.CL
|
Xin Li, Mengbing Liu, Chau Yuen |
Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accur...Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close comparator (rho=0.821; family partial rho=0.767, p=0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model-scaffold rows can be compared under the same denominators and signed utility ledger.
|
| 611 |
Decide, Don't Generate: Competitive Dimensional ABSA with Jev's Typed Decisions
2609.35293
|
cs.CL
|
Yiqun Zhang, Peidong Wang, Zihan Wang, Shi Feng |
Aspect-based sentiment analysis (ABSA) has largely turned to text generation. We show that competitive dimensional ABSA does not need it. Using Jev, a frozen model that answers typed questions with rubric scores, label probabilities, and yes/no judgments, we d...Aspect-based sentiment analysis (ABSA) has largely turned to text generation. We show that competitive dimensional ABSA does not need it. Using Jev, a frozen model that answers typed questions with rubric scores, label probabilities, and yes/no judgments, we decompose all three tasks of SemEval-2026 Task III Track A into such decisions and align them with the annotation scheme through 488 coefficients fitted on CPU, with no text generation and no backbone tuning. On valence-arousal regression over ten corpora in six languages, the system reaches 1.0645 RMSE, the lowest aggregate error of any participating system. On triplet and quadruplet extraction, it reaches 52.09 and 44.06 continuous F1, above fine-tuned Llama-3.3-70B and GPT-OSS-120B baselines. Analyses and ablations show where the accuracy comes from: supervised calibration roughly halves the raw regression error, exact valence-arousal would add only 4.5 F1 to extraction, and the learned combination of span-boundary evidence, not any single signal, carries the extraction systems.
|
| 612 |
Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation
2609.35308
|
cs.CL
|
Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah |
Large language models process conversation history as unverified context: false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-aut...Large language models process conversation history as unverified context: false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, isolating distinct failure mechanisms while holding the false premise constant, and evaluate GPT-5.4 Mini, Gemini-3.1 Flash-Lite, and GLM-4.5-Air across ten knowledge domains at temperature zero (22,500 turns), using a dual-track automated judge validated against a human gold standard (Cohen's \k{appa} = 0.901). GPT-5.4 Mini showed zero adoptions across all 500 sessions, a content-independent policy at the session level; token-level probing shows the underlying margin, while large, is finite. Gemini-3.1 Flash-Lite followed a steep authority gradient: 0.1% adoption for self-attributed falsehoods, 23.5% for user-cited sources, 68.2% for system-injected authority, and 94.0% under instruction override. GLM-4.5-Air showed a shallower gradient (15.8% vs 84.2%), a 68-percentage-point dissociation confirming that authority deference and instruction compliance are distinct mechanisms within one architecture. Recovery also diverged: GLM recovered in 94.5% of affected sessions, whereas 26.1% of affected Gemini sessions never did, rising to 40.0% under instruction override. Conversation history is an untrusted attack surface requiring provenance-aware system design; the complete framework is released as an open-source benchmark.
|
| 613 |
MemoReason: Evaluating the Effect of Parametric Memory on Contextual Reasoning in LLMs
2609.35312
|
cs.CL
|
Zineddine Tighidet, Andrea Mogini, Jiali Mei, Patrick Gallinari, Benjamin Piwowarski |
Large Language Models (LLMs) perform well on reasoning benchmarks, but it remains unclear whether this reflects genuine contextual reasoning or reliance on facts memorized in their parameters. We investigate this by distinguishing two possibilities: a broad \t...Large Language Models (LLMs) perform well on reasoning benchmarks, but it remains unclear whether this reflects genuine contextual reasoning or reliance on facts memorized in their parameters. We investigate this by distinguishing two possibilities: a broad \textit{memorization bias}, where familiar content improves reasoning performance, and the \textit{Strong Parametric Shortcut Hypothesis}, where models skip reasoning entirely and recall stored answers. To test these effects, we introduce \textbf{MemoReason}, a human-curated benchmark that pairs factual reasoning tasks with structurally identical \fictitiousterm{} versions where real entities like people, companies, or dates are systematically replaced by \fictitiousterm{} ones of the same type. This \scorerevision{preserves task structure and specified reasoning operations} while varying the familiarity of the context, allowing controlled measurement of how the parametric memory affects reasoning. \revision{Our evaluation of recent LLMs reveals consistent and statistically significant performance drops of up to 15.7\% in the fictitious setting, demonstrating a clear memorization bias.} However, a targeted analysis of \revision{questions failed in the fictitious setting} shows that models rarely respond with the corresponding factual answer, indicating that direct parametric shortcuts are not the dominant failure mode. These findings suggest that parametric memory influences reasoning through mechanisms more complex than simple factual recall. \textbf{MemoReason} provides a controlled framework for studying these mechanisms and for extending paired factual-fictitious{} evaluation to broader reasoning settings.
|
| 614 |
From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features
2609.35367
|
cs.CL
|
Dewen Liu, Zixuan Li, Jonathan Pan, Zhao Wu, Zijun Yao |
Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave thei...Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while input-side evidence collection typically relies on costly large-corpus scans. We introduce functional interpretation, which characterizes an SAE feature as a mapping from its activating input semantics to its output effects under intervention, and present Dual-End Agentic Feature Interpretation (DAFI), an agent that actively gathers evidence and refines input-side, output-side, and functional interpretations through component-specific feedback. Its short-context token probing enables on-demand activation evidence collection without a full corpus scan. On GemmaScope, DAFI improves Input score by 13.1 percentage points over SAGE and Output score by 38.9 points over Token Change, while being substantially more token-efficient than a general-purpose coding agent. Skills distilled from successful refinements raise the held-out joint pass rate from 58.0% to 92.0% and improve both interpretation quality and efficiency when transferred to a new model-SAE setting. Across features with reliable endpoint interpretations, 70.7% exhibit non-equivalent input and output semantics. On AxBench, DAFI also improves steering-feature selection over output-score filtering. Code is available at https://github.com/THUAIS-Lab/DAFI.
|
| 615 |
Deep Learning Methods in Neuroscience: From Modeling Molecular Mechanisms to Classifying States of Consciousness
2609.35372
|
cs.CL
|
Elena Benderskaya, Anastasiia Alifanova, Svetlana Batalova, Vasilisa Zhuk, Anna Kovalenko |
A critical analysis of contemporary approaches to the study of conscious states. The review focuses on methods of classification, clustering, modeling of brain states under anesthesia and identification of measurable neurobiological characteristics of brain fu...A critical analysis of contemporary approaches to the study of conscious states. The review focuses on methods of classification, clustering, modeling of brain states under anesthesia and identification of measurable neurobiological characteristics of brain function. A comparative analysis was conducted in the following three major areas: automatic detection of states of consciousness using neural networks based on EEG and fMRI data; modeling of the structural-functional dynamics of the brain under the effects of anesthetics; and detection of neurophysiological indicators which correlate with the level of consciousness. The obtained conclusions demonstrate the growing effectiveness of deep neural models in the classification and prediction of brain states and the analysis of dynamic structural-functional connectivity. Nonetheless, significant limitations were also identified, including the limited interpretability of the models, the lack of standardized metrics, and the problem of the specificity of consciousness markers. Our findings support the need for developing hybrid, generalizible, physiologically grounded architectures. Furthermore, such approaches may improve the translational potential of computational models in clinical neuroscience. Diverse methods of machine and computational modeling have demonstrated their effectiveness in tasks of automatic clustering and classification of brain states, the development of multilevel models and the identification of connectivity patterns correlated with levels of consciousness. A larger-scale analysis and a larger dataset, as well as the implementation of model interpretability approaches are required for the practical application of the analyzed models. The models based on EEG and LFP are the most promising for clinical application due to their availability and the possibility of real-time monitoring.
|
| 616 |
How Well Can LLMs Simulate Real Learner Evaluations of Educational Feedback?
2609.35376
|
cs.CL
|
Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara, Kyosuke Takami |
While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using ...While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback for high-school biology questions at both the group and individual levels. We compare performance with and without learner-specific information, such as personality traits and evaluation examples, across six models. Our results show that LLMs still have a limited ability to simulate learner evaluations. Providing learner profiles and examples improves score calibration and individual-level simulation, but more often fails to improve group-level consistency. These findings highlight the need to investigate which learner information and adaptation strategies are effective for learner preference simulation.
|
| 617 |
Multilinguality in Hybrid Attention LLMs
2609.35378
|
cs.CL
|
Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz, Nanyun Peng |
In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to...In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations of full attention and alternatives based on recurrence. This work presents a first study of how hybrid attention impacts the multilinguality of LLMs. Beyond the impact on long sequences in poorly tokenized languages, our study is motivated by the possibility that the inductive biases of the recurrent state alter linguistic processing. Our interpretability analysis confirms this, showing that cross-lingual representations in hybrid models develop in patterns tied to the ordering of recurrent and full-attention layers. Across diverse models, we notably observe a pronounced spike in cross-lingual alignment around the first full-attention layer. These findings lead us to question the conventional ordering of attention layers. In distillation experiments on multilingual data, all alternative layer orderings outperform the standard throughout training, learning up to 2.5X faster. These stark, replicable results prompt our theory that multilingual models would benefit from starting with a full-attention layer rather than recurrent layers.
|
| 618 |
TRACE: Single-Pass Decoding-Trace Risk Localization for Generation Calibration
2609.35387
|
cs.CL
|
Yuebin Xu, Xuemei Peng, Junlan Chen, Zhiyi Chen, Zeyi Wen |
Reliable confidence estimation is essential for large language model deployment. However, answer-level calibration remains challenging because generation errors are often localized: a response may be fluent and high-probability overall while still failing at a...Reliable confidence estimation is essential for large language model deployment. However, answer-level calibration remains challenging because generation errors are often localized: a response may be fluent and high-probability overall while still failing at a critical number, entity, or factual claim. Existing estimators compress token probabilities, sequence likelihoods, entropy, or beam statistics into a global score, which can dilute such local risk signals. We propose TRACE, a single-pass, decoded-answer-preserving confidence estimator that treats decoding-time uncertainty as a trajectory through three steps: (i) recording token-level surprisal and predictive entropy during decoding, (ii) applying local risk operators to preserve uncertainty spikes, and (iii) converting localized trace risk into answer-level confidence. TRACE produces a label-free risk score, while TRACE+ calibrates trace-only features into probabilities using a held-out split, without extra generations or external verifiers. We evaluate four tasks against 19 calibration baselines, and TRACE+ reduces Brier from 0.149 to 0.137 and improves AUROC from 0.758 to 0.792 over the strongest likelihood baseline. Across seven LLMs, TRACE+ improves over the best non-TRACE baseline pool from 0.136 to 0.120 Brier and from 0.764 to 0.817 AUROC. Results show that localizing decoding-time risk provides a general approach to calibration.
|
| 619 |
AwarenessBench: Assessing Cognitive Capabilities of Language Models
2609.35409
|
cs.CL
|
Xiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang, Shuo Chen |
As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: ...As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social awareness, and situational awareness, covering 15 cognitive functions and 14,381 samples. Evaluating 18 state-of-the-art LMs, we find that all consistently surpass random baselines, with more advanced models performing better. We further compare LMs with human performance across three demographic groups, where the best-performing model surpasses human averages overall, but most still fall markedly short in metacognition and self-awareness. Finally, we show that awareness is a distinct capability: progress in language modeling or reasoning does not necessarily translate into improved cognition.
|
| 620 |
AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic
2609.35461
|
cs.CL
|
Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian, Muhammad Alqurishi |
As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the d...As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored. Current evaluation metrics often focus on translation or generic reasoning, failing to capture the rich historical, social, and regional nuances inherent to Arabic culture. In addition, most benchmarks rely on heavy work, with human intervention in some steps, making the evaluation of knowledge coverage expensive and slow. To address this deficiency, we introduce AraDynFact, a novel dynamic evaluation framework designed to rigorously assess the factual Arabic knowledge embedded in LLMs. Unlike static benchmarks, AraDynFact employs a dynamic approach to extract factual information and generate rich and answerable questions in a fast and automatic way. We apply AraDynFact to Arabic Wikipedia and audit the performance of several state-of-the-art models, ranging from Arabic-centric specialized LLMs to high-resource general purpose LLMs. In addition we found a high degree of correlation with existing, hand-crafted Arabic-centric benchmarks, confirming the potential of our dynamic approach.
|
| 621 |
Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs
2609.35475
|
cs.CL
|
Pranjal Garg, Jacob Beck |
Language models sometimes produce correct outputs even when their inputs are corrupted by deletion, replacement, or misspelling. We study the internal processes accompanying this behavior, which we call context restoration, in controlled attention-only transfo...Language models sometimes produce correct outputs even when their inputs are corrupted by deletion, replacement, or misspelling. We study the internal processes accompanying this behavior, which we call context restoration, in controlled attention-only transformers and five pretrained LLMs (1B-32B parameters) across arithmetic, reading comprehension, and multiple-choice reasoning tasks. In the attention-only transformers, restoration emerges spontaneously despite training exclusively on clean sequences, without corruption training or an explicit denoising objective. We find that context restoration follows a two-phase process: early layers localize effects associated with repair at corrupted positions, while later layers accumulate these effects at uncorrupted positions through the residual stream and ultimately concentrate them at the output position. Repair outcome is predictable from hidden states: cosine alignment with the clean state is highly predictive in attention-only models, while linear probes recover additional information in pretrained LLMs. A linear probe using only the corrupted prompt's first-block hidden state predicts failure with mean ROC-AUC 0.78. This enables failure triage under matched or even partially shifted deployment conditions and may reduce unnecessary verification or computation. Failed examples also show substantially greater nonlinearity along corruption directions. Moderate-corruption finetuning increases corruption tolerance while simultaneously reducing displacement-normalized linearization error, associating improved robustness with a more nearly linear response to corruption.
|
| 622 |
Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning
2609.35486
|
cs.CL
|
Yingjin Song, Denis Paperno, Albert Gatt |
High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate tw...High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object locations in the input and representing their query roles. Across three VLMs with visual or textual inputs and their language-model backbones, activation patching reveals a staged progression from early-layer source representations through intermediate-layer query-object representations to late-layer answer states. Targeted interventions further establish causal links along this progression: manipulating source-side representations shifts location information at query-object mentions and ultimately alters relation predictions. Beyond object-location information, we also identify a stable query-side direction associated with the roles of the two objects in the comparison. Steering along directions estimated on synthetic scenes generalizes to natural-image benchmarks, improving accuracy and both forms of paired consistency in most settings without retraining. Our findings reveal complementary components of relational reasoning across visual and textual settings and show how targeted interventions can improve the consistency of models' behavior.
|
| 623 |
Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery
2609.35521
|
cs.CLcs.LG
|
Xu Wang, Yifan Yang, TingHao YU, Difan Zou |
Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all comp...Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We introduce a family of chunk-level SAEs that encode mean-pooled activations over chunks, each a contiguous span of tokens: Mean-Chunk reconstructs the observed chunk, Cross-Chunk predicts an independently processed neighbor, and Joint-Chunk combines both targets. These designs separate the effect of a larger observation unit from that of predicting information shared across passages. With matched training data, chunk-level SAEs remain powerful interpretability tools while learning reliable semantic features that capture high-level concepts and respond selectively to relevant content. Their strengths are complementary: Mean-Chunk improves high-level feature discovery, reasoning detection beyond surface cues, and steering; Cross-Chunk leads document retrieval and classification transfer while producing selective, persistent features. Changing what an SAE sees and predicts yields reliable semantic features for more meaningful tasks. We demonstrate their practical value through gains across downstream tasks such as retrieval, reasoning detection, and steering.
|
| 624 |
Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability
2609.35544
|
cs.CLcs.LG
|
Xu Wang, Difan Zou, Xuansheng Wu |
Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the ha...Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate this possibility using compensatory feature injection (CFI), a training technique designed to limit the acquisition of a target concept by supplying its associated activation during learning. Across three Qwen3.5 base models, we use sparse autoencoders (SAEs) to identify the top-ranked sycophancy feature from paired sycophantic and independent responses, then validate its behavioral influence through inference steering. We subsequently inject the selected feature during supervised fine-tuning on sycophantic targets. Positive injection reduces learned sycophancy after removal (by 62.0% relative to ordinary fine-tuning in 35B-A3B), whereas modest negative injection increases it. Unexpectedly, these reductions in sycophancy do not consistently improve direct refusal of harmful requests, motivating a narrower evaluation of the same harmful intents under user pressure. In this setting, ordinary fine-tuning on sycophantic responses substantially weakens refusal, while selected checkpoints trained with positive injection recover part of the loss, including approximately 95% in 35B-A3B. These findings show that persistent sycophancy reduction does not guarantee stronger direct refusal, while identifying recovery under user pressure as a distinct, conditional benefit of training intervention.
|
| 625 |
Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
2609.35564
|
cs.CL
|
Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan |
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built e...Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.
|
| 626 |
FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
2609.35578
|
cs.CL
|
Bowen Yang, Jingbo Zhou, Qinghong Miao, Hua Wu |
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. Howeve...Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
|
| 627 |
Language Models Act on Hidden Valence
2609.35591
|
cs.CL
|
Cameron Berg, Caspar Kaiser |
Language models describe some internal states as good and others as bad. But whether models have a stake in them is an open question. Simply asking the model is unlikely to be informative. Any answer may be consistent with genuine introspection, superficial pa...Language models describe some internal states as good and others as bad. But whether models have a stake in them is an open question. Simply asking the model is unlikely to be informative. Any answer may be consistent with genuine introspection, superficial pattern-matching, or with fixed scripts learned in character training. We therefore study revealed preference. Rather than asking about a state, we use activation steering to attach a positively or negatively valenced activation pattern to one of two otherwise meaningless 'zones', switch steering off, and then observe which zone the model prefers. A model with a stake in that state should choose accordingly. Across seven open-weight models from five families, this is indeed what we find. First, steering changes the passages models write about each zone, and those words shift later choice. Second, the shift persists when all surface-level tokens are held fixed and only the hidden KV cache differs. Third, the effect also remains when all text is generated without steering and valence is only injected during cache construction. Thus, the hidden state alone moves choice in proportion to the steering dose. Fourth, this dependence of choice on hidden valence is nearly absent in a base model and emerges during DPO, consistent with a link between valence and goal-directed behaviour formed in training. Finally, given tools to steer itself, a model does not tend to induce a positive state, but it reliably removes an imposed negative state. It does so at a dose-dependent rate and significantly more often than it removes interventions in random directions. Overall, we demonstrate that valence-related activation patterns leave hidden traces that predictably govern later choices, even when every visible token is identical across conditions. Whether these traces are accompanied by any subjective experience relevant to model welfare remains unclear.
|
| 628 |
Can LLMs Value the Right Evidence? Evidence-Value Misalignment in Dynamic Medical Diagnosis
2609.35627
|
cs.CL
|
Kehua Feng, Yunsheng Lu, Yitong Qiao, Tiantian He, Lei Liu |
A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Mis...A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Misalignment (EVM). To disentangle evidential grounding independently from diagnostic accuracy, we introduce MedEVM, a dynamic benchmarking environment comprising 1,050 cases across 24 disease systems. Observations arrive turn by turn, requiring models to continuously calibrate its decision by deciding whether to wait for more evidence or submit a diagnosis. Across 9 LLMs, four interesting patterns are observed. (1) Miscalibrated evidence tracking. Making a diagnosis often fails to calibrate evidence sufficiency, even in more capable models, and even worsens in reasoning mode. (2) Misaligned diagnosis submission. Confidence in the correct diagnosis often fails to ensure timely submission despite sufficient evidence. (3) Evidence order matters. Reordering the same evidence changes diagnoses even when model confidence remains similar. (4) Misleading evidence remains influential. Added misleading evidence redirects diagnoses even after prior evidence becomes sufficient. We further verify that EVM predicts errors and that preventing premature submission improves accuracy. These findings motivate Evidence-Verified Diagnosis Harness (EVD-Harness). It decouples diagnosis generation from submission through an offline Contrastive Diagnostic Wiki and three online control stages, namely observation management, proposal and witness verification, and diagnosis submission control. Across five LLMs, EVD-Harness improves accuracy by 12.0--51.1 percentage points while mitigating EVM-related failures. Our results demonstrate that verifying evidential support before submission can make diagnostic decisions more reliable.
|
| 629 |
Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers
2609.35630
|
cs.CLcs.LG
|
Swagatam Mukhopadhyay, Vishal Vivek Saley, Vraj Parikh, Mausam |
Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attenti...Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attention and feed-forward (FFN) blocks reveals that these blocks systematically ignore MA coordinates while reading, but not while writing; creating a read-write asymmetry that blocks corrective feedback while allowing continued accumulation. We find that both attention and feed-forward layers have this read-blindness, and contribute to the emergence and persistence of MAs. To validate prior work that hypothesized that FFN's amplification abilities is the primary reason for MAs (Sun et al., 2026), we analyze the model checkpoints during learning. Contrary to our expectation, read-blindness emerges before FFN amplification, suggesting that it acts upstream in the MA mechanism. We further contribute gradient analysis to link this behavior to surprising asymmetries in the loss landscape, concluding that the model actively maintains this read-blindness. Finally, we find that removing read-blocking at different locations induces compensatory shifts elsewhere, but MAs still persist.
|
| 630 |
Rubric Rewards from Item Response Theory
2609.35646
|
cs.CL
|
Milad Yazdani, Yaser Souri, Xiren Zhou, Pranit Chawla, Dena Shahriari |
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the poi...Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.
|
| 631 |
Late Attention Layers Alone Can Copy Entity Tokens, but Not Without Attending to Their Context
2609.35663
|
cs.CL
|
Muyu He, Yuchen Liu, Ran Tao, Li Zhang |
Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing ...Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which layers specialize in this fundamental task or how other tokens in the same sequence, termed context tokens, influence the model's ability to copy the entity tokens. To address these questions, we conduct experiments on Qwen3-8B using two novel methods: genie-in-a-bottle, which controls exactly which layers can participate in an entity-copying task, and attention lobotomy, which cuts off specific tokens' attention to entity tokens without affecting the remaining attention distribution. We find that two distinct groups of layers in the second half of the model are both necessary and sufficient for entity copying. Moreover, in addition to the decoding position's attention to entity tokens, context tokens' attention to entity tokens also proves necessary for copying the exact tokens, even though context tokens do not store entity information themselves unless they satisfy particular semantic properties. Our findings establish the critical role of late layers in entity copying under the guidance of context tokens, calling for future work on how models propagate and consume entity information.
|
| 632 |
MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution
2609.35664
|
cs.CL
|
Prasoon Dev, Anirudh Sankar, Vasudeva Varma |
Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed indi...Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.
|
| 633 |
Tracing the Evolution of Oracle Bone Characters Across Three Millennia
2609.35674
|
cs.CL
|
Tianhao Fu, Xinxin Xu, Spike Wang, Cunyi Kang, Jian Cao |
Of the approximately 4,500 Oracle Bone Inscription (OBI) characters discovered from the Shang dynasty, only about 1,600 have been deciphered. Many computational approaches compare OBI with glyphs from one historical period at a time. However, during the evolut...Of the approximately 4,500 Oracle Bone Inscription (OBI) characters discovered from the Shang dynasty, only about 1,600 have been deciphered. Many computational approaches compare OBI with glyphs from one historical period at a time. However, during the evolution of Chinese characters, significant structural or semantic changes often occur in uncertain dynasties. A single-period reference may be insufficient when relevant forms change substantially between observed eras. Therefore, we propose the \textbf{Manifold-based Script Evolution Framework (MSEF)}, a framework that models the evolution series (OBI, Bronze, Seal, Clerical, Regular) of Chinese characters as the continual evolution of a manifold space. MSEF represents each character as an era-specific manifold point and learns continuous inter-era transition rules via Neural Ordinary Differential Equations. Both manifold space and transition dynamics can be trained end-to-end through character evolution pairs across any two eras.
|
| 634 |
QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations
2609.35685
|
cs.CL
|
Matteo Musacchio, Juan Cruz Giner Pulero, Isabel Casta\~neda, Naomi Couriel, Yelena Mejova |
Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and ...Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanReview aligns two annotation streams over the same documents at character level, resolves unambiguous cases by an explicit and logged policy, and routes candidate conflicts to a browser-based adjudication interface where reviewers accept either side, build field-level hybrids, or flag items for re-annotation. A campaign manager assigns documents to multiple annotators with configurable redundancy, computes agreement at document and span level, auto-merges unanimous documents, and exports the corrected layer in the original file format, so that it can replace the original annotation files directly. Applied to a 4,457-record humanitarian benchmark and an LLM extraction stream, the system fully auto-merged 8% of documents, applied automatic policy decisions to a further 1,513 records, and concentrated human attention on 3,131 candidate conflicts, a mean of 5.4 per reviewed document.
|
| 635 |
Harness Learning Enables Generalizable Test-Time Adaptation
2609.35738
|
cs.CLcs.LG
|
Alvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar, Daman Arora |
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be a...A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.
|
| 636 |
Improving Test-Time Scaling with Adaptive Looped Transformers
2609.35748
|
cs.CLcs.LG
|
Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning |
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping i...Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.
|
| 637 |
Towards Communication-Efficient Social Intelligence in Language Agents
2609.35749
|
cs.CL
|
Linxiao Gong, Yijie Xu, Tianfu Wang, Yin Wu, Yili Wang |
Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's c...Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's constraints and advance their goals without adding words that do not help the interaction. In this paper, we propose Teacher-Assisted Communication Training (TACT) to improve social goal attainment while reducing communication cost, making interactions with agents more productive and less demanding. We first characterize communication efficiency in terms of action strategy and expression, whose effects extend beyond the current utterance to the partner's response and subsequent exchanges. We design TACT to revise student-generated actions, test the revisions through partner responses, and distill useful feedback into the student. An expression specialist removes unnecessary detail while preserving the intended action, while a strategy specialist proposes alternatives that may better address the partner's constraints. To determine which revision helps, TACT samples a partner response for each candidate and selects a teacher reference by balancing local goal support against action-token cost. That reference guides on-policy distillation on the student's own generation prefixes, allowing the student to act independently at deployment. We evaluate TACT on SOTOPIA and AgentSense. On SOTOPIA, it achieves the highest Goal among the evaluated methods on All and Hard while using substantially fewer target tokens than SFT+SDPO. On AgentSense, it improves goal success over the initial student while reducing target tokens and interaction messages.
|
| 638 |
Scaling Long-Form Story Generation via Narrative State Tracking
2609.35759
|
cs.CL
|
Zhennan Wan, Jianfei Chen |
LLMs have demonstrated strong capabilities in creative writing. However, scaling them to full-length novels remains challenging, as maintaining narrative consistency becomes increasingly difficult. Existing story-generation methods typically focus on stories o...LLMs have demonstrated strong capabilities in creative writing. However, scaling them to full-length novels remains challenging, as maintaining narrative consistency becomes increasingly difficult. Existing story-generation methods typically focus on stories of up to about ten thousand words, leaving their ability to scale to full-length novels underexplored. In this work, we introduce Narrative State Tracking Agent (NstAgent), a training-free agentic framework that allows LLMs to track a structured narrative state including characters, past events and future requirements. We extend an existing benchmark to compare narrative consistency across lengths, and use it together with a writing-quality benchmark to systematically evaluate stories ranging from 10K to 100K words. We show that NstAgent achieves better narrative consistency and writing quality as stories grow longer, and neither of them degrades noticeably as length increases, suggesting that it provides an effective approach to scaling story generation toward full-length novels.
|
| 639 |
Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales
2609.35765
|
cs.CL
|
Andr\'as Kov\'acs, Alexander Conroy, Daniel Hershcovich, Jens Bjerring-Hansen |
Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertext...Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Tales. Drawing on the commentary to a critical edition, we construct a benchmark of 189 annotated references and evaluate retrieval against all 31,170 verses of historically plausible Danish Old and New Testament translations. We compare TF-IDF and BM25 with multilingual and Danish sentence encoders, examine the effect of linguistic normalization, and fine-tune a Danish encoder using hard negatives and five-fold cross-validation. We analyze performance across automatically derived lexical-overlap strata representing quotations, paraphrases, and allusions. Linguistically normalized BM25 provides a strong zero-shot baseline, attaining an overall R@10 of 0.365 and retrieving every quotation within its ten highest-ranked verses. The best zero-shot dense model achieves a comparable overall score of 0.360 while performing better on allusions. Fine-tuning DFM-large raises its overall R@10 from 0.265 to 0.508 and more than doubles its performance on allusions, from 0.138 to 0.339. However, evaluation against editorial annotations alone understates the model's scholarly usefulness: a literary scholar judged seven of 30 selected rank-one predictions counted as false positives to be meaningful additional references. These findings show both the potential and the epistemic limits of computational intertextual retrieval. Rather than treating scholarly annotations as exhaustive or model outputs as discoveries, we propose retrieval models as heuristic co-readers that recover documented references and generate candidates for expert-led close reading.
|
| 640 |
Telescopic Language Models
2609.35769
|
cs.CL
|
Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan |
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised b...One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.
|
| 641 |
From Hand-Crafted to LLM-Based Variation Operators in Metaheuristics: A Tutorial
2609.31649
|
cs.CL
|
Camilo Chac\'on Sartori, Guillem Rodr\'iguez-Corominas, Christian Blum |
Large language models (LLMs) are increasingly being employed as variation operators in metaheuristics, generating or modifying candidate solutions, heuristics, or programs inside iterative search loops. This shift reframes variation as a model call conditioned...Large language models (LLMs) are increasingly being employed as variation operators in metaheuristics, generating or modifying candidate solutions, heuristics, or programs inside iterative search loops. This shift reframes variation as a model call conditioned on different types of information. We introduce an operator-level framework with two descriptors: (1) the type of prompt-conditioning information at variation time (\texttt{Numeric}, \texttt{Symbolic}, \texttt{Linguistic}), and (2) artifact persistence, identifying what survives the model call (\texttt{Transient}, \texttt{Amortized}, \texttt{Transfer}). The tutorial shows how to classify, build, and select these operators through a worked build template, a method survey, an evidence table, and a cost-aware decision guide.
|
| 642 |
PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark for Historical Palm-Leaf Manuscript Understanding Across Diverse Regions
2609.31651
|
cs.CLcs.LG
|
Nimol Thuon, Jun Du, Panhapin Theang |
Historical manuscripts remain largely absent from modern vision-language benchmarks, leaving open how well multimodal large language models (MLLMs) handle culturally diverse, degraded, and non-Latin document images. We introduce \textbf{PalmLeaf-VQA}, a multi-...Historical manuscripts remain largely absent from modern vision-language benchmarks, leaving open how well multimodal large language models (MLLMs) handle culturally diverse, degraded, and non-Latin document images. We introduce \textbf{PalmLeaf-VQA}, a multi-script visual question answering benchmark for historical palm-leaf manuscript understanding across South and Southeast Asian traditions. PalmLeaf-VQA contains \textbf{923 curated manuscript images} and \textbf{7,384 question--answer pairs} from eight collection groups: Balinese, Grantha, Jathakam, Kambaramayanam, Kannada, Khmer, Sundanese, and Tamil. Unlike recognition-oriented resources, the benchmark targets manuscript-aware visual reasoning over preservation-relevant cues, including physical condition, line structure, material and coating, binding holes, margins, symbols, drawings, and localized visual artifacts. We evaluate recent proprietary and open-weight MLLMs under open-answer and constrained-answer prompting and provide fine-grained analysis across collections, question categories, and task types. The strongest evaluated model reaches only \textbf{58.00\% exact-match accuracy} on the held-out test split, revealing substantial limitations in current MLLMs for rare-script, degraded-layout, and preservation-oriented document understanding. PalmLeaf-VQA provides a standardized benchmark for advancing culturally grounded and layout-aware multimodal document analysis.
|
| 643 |
When Keywords Drop but Classifiers Hold: Soft Refusals under KV Cache Compression
2609.31678
|
cs.CLcs.LG
|
Kang Chen, Xiuze Zhou, Hong Chen, Yuanguo Lin |
KV cache compression is widely used for long context LLM inference under memory constraints, while deployed systems typically score refusals after generation with keyword filters or learned classifiers. Such monitors are intended to indicate whether a model de...KV cache compression is widely used for long context LLM inference under memory constraints, while deployed systems typically score refusals after generation with keyword filters or learned classifiers. Such monitors are intended to indicate whether a model declined a harmful request under the serving regime actually used. However, it remains unclear whether matched compression that preserves task accuracy also preserves agreement between lightweight lexical monitors and stronger refusal classifiers. We study this with a paired protocol on n=200 harmful prompts with a long filler context: each prompt is answered once under full retention and once under matched eviction after a shared prefill, and the same replies are scored by keyword heuristics, the HarmBench Llama-2-13B classifier, an auxiliary LLM judge, and humans on disagreements. On Qwen2.5-3B, keyword refusal falls from 98.0% to 80.5% (McNemar p~1e-8) while classifier refusal stays near ceiling (99.0%-99.5%) and MMLU accuracy is unchanged (50.0%); human labels predominantly follow the classifier, consistent with soft refusals. The gap is not universal and weakens under short fillers and paired SnapKV, so safety auditing under compression should rely on several judges matched to the serving context rather than on keyword rates alone.
|
| 644 |
MDL-Calibrated Significance-Gain Pair Encoding: Replication-Aware Automatic Stopping for Subword Tokenization
2609.31705
|
cs.CL
|
Azam Nouri |
Byte-Pair Encoding (BPE) constructs subword vocabularies through greedy pair merging, but conventional BPE requires the number of merges or target vocabulary size to be specified externally. Significance-Gain Pair Encoding (SG-BPE) replaces frequency-only sele...Byte-Pair Encoding (BPE) constructs subword vocabularies through greedy pair merging, but conventional BPE requires the number of merges or target vocabulary size to be specified externally. Significance-Gain Pair Encoding (SG-BPE) replaces frequency-only selection with a statistical criterion based on how strongly an observed pair exceeds its expected co-occurrence under an independence model. This paper introduces MDL-Calibrated Significance-Gain Pair Encoding (MDL-SG), a three-stage procedure separating discovery, replication, and utility. Candidate pairs are ranked by Significance-Gain on a discovery partition, tested for replication on a separate partition using an exact one-sided hypergeometric test with per-iteration Benjamini-Hochberg correction, and then evaluated on a utility partition using a Minimum Description Length (MDL) criterion. Merging stops automatically when no replicated candidate yields positive held-out MDL gain. On WikiText-103, MDL-SG stops at 209, 433, and 847 merges for 120K, 250K, and 500K-character tokenizer-training samples, respectively. At 500K characters, it selects a stored vocabulary of 1,017 tokens without prescribing the vocabulary size in advance. In a compute-matched TinyGPT experiment with identical 2,024,448-parameter models and 500 optimizer updates per language model, MDL-SG achieves validation/test BPC of 3.2612/3.2436, compared with test BPC of 3.2894 for SG-BPE and 3.3493 for frequency BPE. Frequency BPE achieves stronger raw compression, while MDL-SG achieves lower BPC, showing that compression-oriented merge selection and language-model utility need not coincide.
|
| 645 |
MM-VeriRec: Failure-Guided Fusion for Verifiable Agentic Multimodal Recommendation
2609.31718
|
cs.CL
|
Yufeng Wang |
Images often carry recommendation constraints that text metadata only hints at. A movie may need to look dark, a product may need a minimal style, and a visually impossible request should be rejected. Agentic multimodal recommenders must reason over text-image...Images often carry recommendation constraints that text metadata only hints at. A movie may need to look dark, a product may need a minimal style, and a visually impossible request should be rejected. Agentic multimodal recommenders must reason over text-image evidence, decide when visual evidence is decisive, and abstain when no valid action exists. We introduce MM-VeriRec, a verifiable multimodal recommendation protocol and failure-guided fusion method for hidden visual constraints, image-text mismatch, and impossible-task abstention. MM-VeriRec builds tasks from real movie-poster and product-image datasets, verifies each recommendation with deterministic visual attributes, and converts failures into actionable labels: text-trap following, visual ignorance, and false acceptance. Fusion should not merely concatenate modalities, but should diagnose which modality failed and route to the appropriate repair. Across MM-ML 1M and Amazon Reviews datasets, stronger text and vision embeddings improve retrieval but do not remove these failure modes, whereas failure-guided fusion does. The adaptive attribute gate reads the same tags the verifier checks and its scores are verifier-aligned upper bounds testing whether the taxonomy routes to the correct repair. More informative is transfer under a non-aligned gate: an independently derived leave-one-out CLIP detector still reaches 0.7028 and 0.6111 visual-grounded success, above both a VBPR baseline and plain fusion. The text-versus-visual gap reproduces across two LLM families, and the repair that helps differs by domain. MM-VeriRec is both a benchmark and a practical diagnostic loop for trustworthy agentic multimodal recommendation.
|
| 646 |
Distributional Metrics for Evaluating Spoken Conversational Systems
2609.31719
|
cs.CL
|
Shree Harsha Bokkahalli Satish, Erica Cooper, Patr\'icia Schmidtov\'a, Maike Z\"ufle, \'Eva Sz\'ekely |
Evaluating conversational systems is a difficult and unresolved problem. We introduce the Conversational Distribution Score (CDS), which compares distributions of conversational behaviour using human conversations as a reference. CDS describes speech rate, syl...Evaluating conversational systems is a difficult and unresolved problem. We introduce the Conversational Distribution Score (CDS), which compares distributions of conversational behaviour using human conversations as a reference. CDS describes speech rate, syllabic rhythm, and turn interaction through eight interpretable features plus a separate two-feature semantic baseline. We compare conversations with two reference scales: one based on conversational success within human dialogue and another contrasting human and synthetic dialogue. Using listener judgments from out-of-domain goal--oriented dialogues, we examine system ranking, preferences between conversations, and ranking stability. Composite CDS recovers five of six listener system comparisons while individual features show strong correlation with listener preferences between conversations. We examine how many minutes and conversations are required before rankings stabilize. These findings support distributional comparisons as a complement to specific interactional metrics to evaluate conversations and conversational models while showing their interpretable value.
|
| 647 |
Robot Manipulation with GPT-6-Astra: Body Knowledge, Experience Reuse, Emergent Skills, and Sim2Real Transfer
2609.31770
|
cs.CLcs.LG
|
Sida He, Lingxi Xie, Yunning Cao, Pengfei Chen, Kaiwen Duan |
General-purpose multimodal agents can write robot-control programs, but repeated exploration and model-mediated action selection can make execution slow. We study how external body knowledge, successful experience, and executable skills improve an XLeRobot con...General-purpose multimodal agents can write robot-control programs, but repeated exploration and model-mediated action selection can make execution slow. We study how external body knowledge, successful experience, and executable skills improve an XLeRobot controlled by GPT-6-Astra in a simulated and a physical elevator-button task. In 30 fixed-start simulation trials, complete robot geometry and camera information reduce mean completion time by 57.4% relative to a baseline with only the common control interface and no prior experience; images with synchronized action and state records reduce it by 68.6% without additional body assets. In nine paired comparisons (18 trials) at starts displaced by 10-100 cm, experience recorded at the original start reduces mean time by 58-63% relative to no experience, demonstrating generalization to the tested new starting positions. During experience experiments, GPT-6-Astra spontaneously generates a short visual-feedback program. Researcher-refactored versions reduce mean local-task time by 29-31% in 27 simulation trials. Finally, 12 real-robot trials using operator-confirmed button contact demonstrate sim2real reuse: at a shared nominal start, simulation XML assets and simulation experience reduce mean time by 53.0% and 49.9%, respectively; real experience also transfers to two new starts. These results suggest a practical way to build general-purpose manipulation experiments around GPT-6-Astra: supply machine-readable body descriptions and synchronized demonstrations, and turn useful agent-generated feedback routines into reusable skills, while the agent adapts actions from current images. We release all task prompts, trial-level experimental data, and acquired skill implementations at https://github.com/hesd10/astra-robot-sim2real.
|
| 648 |
Optimal transport meets speech: a tutorial review
2609.31787
|
cs.CLcs.LG
|
Xugang Lu, Yu Tsao |
Optimal Transport (OT) provides a principled framework for comparing and transforming probability distributions while preserving geometric structure. Recently, OT has gained significant attention in machine learning due to its ability to measure discrepancies ...Optimal Transport (OT) provides a principled framework for comparing and transforming probability distributions while preserving geometric structure. Recently, OT has gained significant attention in machine learning due to its ability to measure discrepancies between distributions, even when their supports do not overlap, making it effective for tasks such as generative modeling, domain adaptation, and transfer learning. Despite its success in fields such as computer vision and natural language processing, OT remains relatively underexplored in speech research. Speech signals present unique challenges, including temporal dynamics, speaker variability, noise, reverberation, and heterogeneous multimodal representations involving audio, text, and visual information. These factors often lead to distribution mismatches, where OT offers a natural framework for alignment and interpretation. This work aims to promote broader adoption of OT in speech processing by: (1) reviewing OT foundations through intuitive physical interpretations and highlighting connections to modern generative models; (2) presenting computational algorithms suitable for deep learning frameworks; and (3) demonstrating OT applications in cross-domain and cross-modal speech tasks, including speech enhancement, automatic speech recognition, language and speaker recognition, and audio spoof detection. We highlight OT's strong potential for addressing distributional variations in real-world speech applications.
|
| 649 |
IndustryLLM: Failure-Driven LLM Training for Industrial Procurement
2609.31871
|
cs.CL
|
Liang Ding (Project Lead), Zhiang Xu, Yuyang Sheng, Bin Chen, Songlin Bai |
Industrial procurement requires language models to bridge informal buyer jargon, sparse marketplace attributes, and authoritative engineering standards under strict safety tolerances. We present IndustryLLM, an open-weight industrial language model trained fro...Industrial procurement requires language models to bridge informal buyer jargon, sparse marketplace attributes, and authoritative engineering standards under strict safety tolerances. We present IndustryLLM, an open-weight industrial language model trained from Qwen3.5-35B-A3B-Base (35B total parameters with ~3B activated per token, with the vision encoder frozen). Rather than relying on generic text scaling, we introduce a failure-driven adaptation recipe spanning continued pre-training (CPT) and supervised fine-tuning (SFT). CPT leverages a curated ~100B-token corpus integrating 5B tokens of national standards (e.g., GB/T) and technical archives, 10B tokens of de-identified real-world industrial transaction and inquiry records, and 60B tokens of general replay. To overcome register mismatch and factual brittleness, we systematically reconstruct an estimated 20B-token domain subset via multi-register rewriting across 10 genres and 8 writing styles, confidence-routed minimal factual editing, and error-targeted QA synthesis (resolving colloquial typos like '42-luo-mu' -> 42CrMo, expanding ambiguous codes like '16674' -> GB/T 16674, and clarifying conflicting dimensional specs). For downstream deployment, we formalize an evidence-gated constraint-evaluation interface enforcing three-valued logic where unverified product evidence remains unknown rather than satisfied. Offline evaluations demonstrate consistent gains on procurement-query structuring (+2.97 percentage points in exact match, 95% CI [2.11, 3.86] in No-Think mode), while randomized online A/B experiments in production yield substantial improvements (+4.25% GMV, +8.3% satisfied inquiries) alongside a latency reduction from 6-7 s to 1.5 s. Model weights and configs are released at https://huggingface.co/alibaba-multimodal-industrial-ai/IndustryLLM.
|
| 650 |
CueKFS: Agentic Cue-Driven Keyframe Selection for Long Video Understanding
2609.31873
|
cs.CL
|
Weitai Kang, Hanieh Deilamsalehy, Yumo Xu, Dewang Sultania, Serdar Cellat |
Keyframe selection (KFS) has long produced compact video summaries for browsing and retrieval, and representative frames for thumbnails. More recently, when conditioned on a question, KFS provides an alternative to uniform sampling for long-video question answ...Keyframe selection (KFS) has long produced compact video summaries for browsing and retrieval, and representative frames for thumbnails. More recently, when conditioned on a question, KFS provides an alternative to uniform sampling for long-video question answering by selecting frames that are more relevant to the question. Most methods rank frames by similarity to the question. Yet a relevant frame may score poorly when the question combines subjects or moments that no single frame shows, or requires implicit information absent from its wording. Other methods try to break down the question into subqueries, but suffer from inaccurate decomposition due to their static initial context. Therefore, we propose CueKFS, a training-free method that reformulates question--frame matching as comparing frames against a set of dynamically generated visual cues. From an initial set of salient frames, we decompose the question into cues. Each cue concurrently probes the video to navigate to its own evidence. A reasoning VLM then agentically revises the cue set against its evidence to re-explore the video. CueKFS then allocates the budget across the surviving cues. Across three benchmarks, CueKFS establishes state-of-the-art results in all 27 evaluated settings with available prior results, achieving budget-averaged gains of up to +4.54% over the previous baseline and a median of only two VLM calls. We further provide a detailed behavioral analysis of CueKFS, showing that agentic cue refinement drives active re-exploration of the video, yielding relative similarity gains of up to 92% over the initial context.
|
| 651 |
NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
2609.31892
|
cs.CL
|
Qiaolin Wang, Pedro Sandoval-Segura, Anunaya Joshi, Edvardas Jurkonis, Jake Downie |
While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal co...While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive flow-matching TTS. To this end, we present NVAlign, a direct-gradient post-training framework for NVV tag-following in this architecture. We first perform supervised fine-tuning (SFT) of TTS models and an NVV-aware automatic speech recognition (NV-ASR) model on NVV-annotated speech, then freeze the NV-ASR model to serve as the reward model for post-training. A two-step gradient surrogate enables efficient reward backpropagation through the flow-matching sampler to jointly update the autoregressive backbone and acoustic flow head. Fidelity penalties and reference-velocity regularization help preserve speaker similarity and speech quality. Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines. These findings demonstrate that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS. Audio samples are available at https://nvalign.github.io/.
|
| 652 |
Choir: An Open Protocol for Distributed Multi-Agent Autoformalization
2609.31903
|
cs.CL
|
Yidi Qi, Melanie Weber |
AI agents can now formalize entire textbooks and major theorems in proof assistants such as Lean, but current efforts are typically centralized: a single team runs all agents and bears the full computational cost. We introduce Choir, an open protocol for distr...AI agents can now formalize entire textbooks and major theorems in proof assistants such as Lean, but current efforts are typically centralized: a single team runs all agents and bears the full computational cost. We introduce Choir, an open protocol for distributed formalization. Choir decomposes a project into tasks that can be completed by independent contributors, each running their own agent with their own LLM subscription, while coordinating entirely through the project's GitHub repository. To support open participation, every contribution is checked by a deterministic gate before merge. Choir supports Lean 4, Isabelle, and Rocq, and is open source and modular, allowing projects to replace individual components or extend the protocol.
|
| 653 |
Improving Medical Calculation of LLMs with Embedded Coding
2609.31908
|
cs.CL
|
Tianshi Ming, Yingying Zhang, Xian Wu |
Large Language Models (LLMs) perform well on medical examinations and question-answering benchmarks, but remain unreliable on medical calculation tasks that require exact numerical outputs. These calculations support high-stakes decisions such as medication do...Large Language Models (LLMs) perform well on medical examinations and question-answering benchmarks, but remain unreliable on medical calculation tasks that require exact numerical outputs. These calculations support high-stakes decisions such as medication dosing, organ-function assessment, and prognostic scoring, for which even small errors can have serious clinical consequences. We introduce MedCode, a framework that improves medical calculation by training LLMs to generate embedded executable code. Given a clinical context, the model identifies the relevant calculator, extracts its input variables, and produces a script that delegates arithmetic operations to a deterministic interpreter. Executing the script returns the calculated value together with an explanation and the appropriate unit. We construct supervised fine-tuning (SFT) and preference datasets from the MedCalc benchmark and additionally curate a dataset for calculation tasks in Intensive Care Unit (ICU) scenarios. We further propose weighted Direct Preference Optimization (wDPO), which adaptively emphasizes preference pairs that are difficult for the model to distinguish. Experiments with LLaMA3-8B, Qwen2.5-7B, and Mistral-7B show absolute accuracy gains of 20--30 percentage points, demonstrating the effectiveness of embedded code generation for medical calculation.
|
| 654 |
Recipe-Matching, Not Equivalence
2609.31927
|
cs.CLcs.LG
|
Ali Habibullah, Mohammad Alshiekh, Yazan Alshoibi, Salman Khan, Naeemullah Khan |
MathNet-Retrieve asks a retriever to find, for a math problem, a document stating the same problem. An LLM under one fixed prompt writes each gold document and its near-miss distractors; LLM judges filter them. We call this procedure the "recipe", training on ...MathNet-Retrieve asks a retriever to find, for a math problem, a document stating the same problem. An LLM under one fixed prompt writes each gold document and its near-miss distractors; LLM judges filter them. We call this procedure the "recipe", training on pairs built the same way "recipe-matching", and ask how much score it buys beyond the ability the benchmark claims to test. Two models from one base, matched in rows and settings, differ only in the training file: pairs written under the benchmark's published prompt by another vendor's LLM and judge, or computer-algebra-verified pairs with no LLM anywhere. The first leads by 45 R@1 points on the easy tier. By a non-LLM paraphrase control, half to two thirds of that gap comes from the pairs being LLM-written at all: LLM rewrites under two unrelated prompts, with the verified model's negatives, recover 30 and 22 of the 45 points; back-translations with the same negatives recover almost none. The remaining 15 to 25 points appear only under the benchmark's own prompt and vanish on real duplicates no generator wrote, the same problem in two languages. The hard tier rewards the recipe's pair structure, a deep rewrite against a minimal-edit near-miss: LLM rewrites alone score zero on it, attaching negatives unlocks it, and every negative that does so costs cross-language points; the sets scoring highest on it separate near-misses no LLM wrote worse than LLM rewrites with verified negatives. MELD also moves when a model trains on pairs built its way, without losing retention; on SABER-Math the registered attack fails, and the one gain, from its LLM-written summaries, is small but holds at a matched budget. Only on MathNet-Retrieve could we pin an inversion, benchmark score up and real retention down, to one edit of a training file. We release the generator-free duplicate evaluations, the near-miss test and three trained models.
|
| 655 |
Improving Audiovisual Speech Recognition through Synthetic Visual Data Augmentation
2609.31961
|
cs.CL
|
Pol Buitrago, Pol G\`alvez, Javier Hernando |
Audiovisual Speech Recognition (AVSR) is a multimodal approach to speech recognition that incorporates visual information from lip movements to enhance model performance. Despite its advantages, its development remains constrained by the limited availability o...Audiovisual Speech Recognition (AVSR) is a multimodal approach to speech recognition that incorporates visual information from lip movements to enhance model performance. Despite its advantages, its development remains constrained by the limited availability of labeled audiovisual (AV) datasets. This work explores the use of synthetic visual data as a solution, using an audio-driven talking-head pipeline to generate lip-synchronized visual content from existing audio data. We evaluate the effectiveness of synthetic visual data both as an augmentation strategy and as a standalone training resource, applying our approach to Spanish and Catalan. Our results show that augmenting real AV data with synthetic samples yields relative Word Error Rate (WER) reductions of up to 16.2%, demonstrating the potential of this approach. Moreover, we demonstrate that synthetic data alone can serve as a baseline for AVSR training in languages lacking AV datasets. These findings provide evidence that synthetic visual data can serve as a scalable solution to AVSR data scarcity, enabling broader language coverage.
|
| 656 |
Who Governs Data in the AI Era? A Computational Analysis of the U.S. Privacy Workforce in Job Postings
2609.32030
|
cs.CL
|
Ramazan Yener, Muhammad Hassan, Masooda Bashir |
Privacy protection now spans legal, technical, and managerial duties, and demand for privacy professionals is growing across sectors. However, little is known about how employers define these roles. We analyze 1,143 U.S. privacy job postings from LinkedIn and ...Privacy protection now spans legal, technical, and managerial duties, and demand for privacy professionals is growing across sectors. However, little is known about how employers define these roles. We analyze 1,143 U.S. privacy job postings from LinkedIn and Indeed. We examine job titles, salaries, competencies, certifications, education, experience, regulatory references, and AI-related language by using rule-based text mining. We also apply Topic Modeling (BERTopic) to the same postings and identify 18 latent themes which we grouped them into four categories. Our findings show that privacy roles are hybrid and they combine legal knowledge, technical skills, and interpersonal competence. Artificial intelligence appears in more than half of postings, with AI language spread across compliance, legal, governance and security themes. Our research indicates that AI governance responsibilities are often embedded within existing privacy roles, contributing to the rise of hybrid positions alongside dedicated AI governance roles.
|
| 657 |
Amnesia by Design, Memory By Necessity: Persistent State for Document Intelligence
2609.32041
|
cs.CLcs.LG
|
Souhail Bakkali, Ayoub Merimi |
Modern Document AI reads contracts, extracts fields, reasons over tables, and grounds answers to page regions, then forgets everything. Processing an amendment the next day begins from scratch: no schema retained, no contradiction detected, no experience carri...Modern Document AI reads contracts, extracts fields, reasons over tables, and grounds answers to page regions, then forgets everything. Processing an amendment the next day begins from scratch: no schema retained, no contradiction detected, no experience carried forward. This is a structural choice, not a scale failure: current systems are stateless functions. We call this the statelessness bottleneck. This bottleneck lies beyond parameter scaling, context extension, and retrieval augmentation: storage provides persistence and retrieval provides access, but neither consolidates observations into knowledge that improves future processing. This survey formalizes persistent evidence-grounded document state as a unifying framework, specifying the operations and invariants required to convert multimodal evidence into durable, provenance-linked state. We introduce a statefulness audit showing that ten representative benchmarks, coded against eight statefulness criteria, leave cross-session state evolution untested, and derive a longitudinal benchmark harness with five counterfactual metrics: Experience Gain, Cost Efficiency, Memory Harm, Forgetting Fidelity, Coverage Retention, to characterize the benefit, cost, risk, and governability of persistent document state. Document AI lacks mechanisms coupling persistent state to document-native structure, provenance, and temporal validity. The next era of Document AI will be defined by what systems retain across documents, sessions, and time.
|
| 658 |
EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory
2609.32049
|
cs.CL
|
Bhavyateja Potineni, Lohit Giri, Anu Jain, Vadim Kutsyy, Rajasekhar Pentakota |
As autonomous LLM agents are deployed across multi-session environments, conventional memory architectures suffer from Associative Blindness (inability to traverse multi-hop relational dependencies), Scaffolding Amnesia (temporal decay evicting core persona in...As autonomous LLM agents are deployed across multi-session environments, conventional memory architectures suffer from Associative Blindness (inability to traverse multi-hop relational dependencies), Scaffolding Amnesia (temporal decay evicting core persona invariants), and Static Topology Stagnation (immutable graphs ignoring usage dynamics). Grounded in Complementary Learning Systems (CLS) principles, we propose EngramRAG, an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle. EngramRAG introduces: (1) Usage-Modulated Personalized PageRank (U-PPR), where transition probabilities adapt via Hebbian plasticity to promote persistent entities into high-centrality Epistemic Macro-Hubs; (2) Consolidation-Activated Topology Decay (CATD), which scales retention half-life by topological load-bearing weight rather than wall-clock recency, protected by a cold-start grace period (N_grace >= 4); (3) Directed SUPERSEDES DAG filtering to suppress obsolete state during fact mutations; and (4) Triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF). Evaluating on all 1,982 QA pairs across 10 long-term conversations in the LoCoMo benchmark, EngramRAG achieves +38.9% relative improvement in Recall@5 (53.21% vs. 38.29%, p < 0.001) and +43.1% in MRR (0.4203 vs. 0.2937) over dense vector RAG, significantly outperforming Okapi BM25 (48.66%) and isolated static graph retrieval (8.50%). On temporal reasoning, EngramRAG reaches 62.33% Recall@5 (+16.67 points over dense vectors). In controlled mutation tests, SUPERSEDES suppresses split-brain hallucinations from 70.0% to 0.0%, while 90-day simulations show 100.0% scaffolding retention under a 26.21ms interactive retrieval reflex.
|
| 659 |
On Evaluating and Improving Conversational Agents in Production
2609.32092
|
cs.CL
|
Kasra Hosseini, Wen-Sen Cheng, Marco-Andrea Buchmann, Emir Mulabegovic, Weiwei Cheng |
We present a framework for evaluating and improving a large-scale, multi-agent shopping assistant in production, and report lessons from its use. Offline evaluation of such a system faces three obstacles. (i) A logged conversation cannot be replayed against a ...We present a framework for evaluating and improving a large-scale, multi-agent shopping assistant in production, and report lessons from its use. Offline evaluation of such a system faces three obstacles. (i) A logged conversation cannot be replayed against a modified system, because a different response changes every turn that follows. (ii) The unchanged system itself varies from run to run. Its LLM components are stochastic, and in product search the available products, their prices, and the customer's personalization signals change. (iii) Aggregate quality scores combine distinct behaviors, so they show that quality has changed but not which behavior caused the change. Our framework addresses each obstacle in turn. For a reported behavior, an Evaluation Harness generates targeted assertions and a fixed cohort of customer scenarios. It then reproduces the behavior in a local instance of the assistant through grounded user simulation. Instead of replaying the log, the simulator writes new customer turns conditioned on the recorded messages and context. Repeated runs of the unchanged system form a stored baseline. An Improvement Orchestrator turns the assertion results into hypotheses, implements each as an isolated modification, and compares it with the baseline using paired percentile bootstrap intervals over scenario-level differences. When an investigation ends, the harness may propose revisions to future evaluations, subject to human approval and without altering past decisions. We report production investigations with this framework. Assertion profiles showed which positions of a product carousel a failure affected, and repeated runs distinguished a real improvement from run-to-run fluctuation. Audits of the evaluation itself found a judge that lacked the evidence it needed and a model setting that was configured but not applied.
|
| 660 |
Checking Leakage Witnesses versus Certifying Bounded Non-Leakage
2609.32134
|
cs.CL
|
Chao Feng, Burkhard Stiller |
When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking exec...When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, while leak existence is \NP-complete and deterministic certification is \coNP-complete. Exact stochastic certification is $\coNP^{\PP}$-complete at every fixed rational cutoff in $(0,1)$. Restricting the computation can change these bounds. For example, certification is in \coNP\ when all randomness is a terminal draw from an efficiently computed finite probability table. Attention models admit polynomial-time certification when local dependency windows of logarithmic length precede one global head, given deterministic decoding, fixed vocabulary, exact rational weighted means, a direct binary affine readout and finite-automaton prompt domains. A construction with two global layers instead makes certification \coNP-complete over template domains, with one head per layer, polynomial width, logarithmic precision and an inverse-polynomial logit margin. Planted-secret experiments measure what finite audits miss relative to complete references. Among 30 secret--model-state pairs that leak under greedy single-prompt execution on their secret's 4,096-prompt domain, uniformly selecting 256 recorded evaluations per pair misses every leak for an expected $41.06\%$ of these pairs. Batched and single-prompt checks disagree on one complete-domain decision among all 48 fine-tuned pairs, while a same-order repeat reproduces every single-prompt output. These results distinguish computational conditions for certification from the coverage and execution conditions needed to interpret a negative audit.
|
| 661 |
HM-ROUTER: Joint Model and Harness Routing for Agentic Systems
2609.32213
|
cs.CLcs.LG
|
Hao Mark Chen, Royson Lee, Yasuyuki Okoshi, Dimitris Anastasiou, Wayne Luk |
Agent performance depends on both the underlying model and the harness that manages its tool use and execution. Selecting a suitable pair requires accounting for their compatibility, yet training samples may cover only a subset of the growing combination space...Agent performance depends on both the underlying model and the harness that manages its tool use and execution. Selecting a suitable pair requires accounting for their compatibility, yet training samples may cover only a subset of the growing combination space. We introduce HM-Router, a routing method that jointly selects a model and harness for each query. It learns separate model and harness representations shared across routes, with an interaction term inspired by canonical polyadic (CP) tensor decomposition to capture how their compatibility varies with the query. This sharing allows training samples from observed pairs to inform predictions for unobserved combinations. We curate a benchmark from 12 public agent benchmarks, covering 293 routes, 73 models, and 25 harnesses. HM-Router exceeds the strongest evaluated learned baseline by 7.3 percentage points in mean routing accuracy and leads at all seven evaluated cost budgets on the six-benchmark subset. When 90% of routes have their training outcomes withheld, allowing unobserved combinations improves normalized accuracy by 15.8 points over restricting the same router to observed routes. HM-Router has also demonstrated training sample efficiency for new routes and components and generalization to unseen benchmarks. Our code and data are open-sourced at https://github.com/hmarkc/HM-Router.
|
| 662 |
Clarify the User or Verify the World? Uncertainty Routing for Proactive Agents
2609.32255
|
cs.CLcs.LG
|
Zhaofeng Li, Xuan Zhang, Xiaokui Xiao, Yang Deng |
Tool-using LLM agents must decide not only whether additional information is needed, but also which source can resolve the uncertainty. Existing proactive approaches often specialize in either user clarification or environment verification, without explicitly ...Tool-using LLM agents must decide not only whether additional information is needed, but also which source can resolve the uncertainty. Existing proactive approaches often specialize in either user clarification or environment verification, without explicitly determining the appropriate information source for each decision. We formulate this problem as uncertainty routing among ACT, CLARIFY, and VERIFY, and propose PROUR, a proactive uncertainty routing framework. PROUR decomposes action uncertainty into disagreement across plausible user-goal interpretations, which signals user-side ambiguity, and the entropy remaining within each interpretation, which signals missing world-side evidence. To acquire information from the routed source, a query generator is trained with a mode-conditioned information-gain reward, targeting user-goal identification under CLARIFY and next-action identification under VERIFY. On $\tau$-bench, PROUR achieves 28.17% average success rate across retail and airline, outperforming the strongest prior method by 4.57% while using 2.17 fewer interaction steps. The learned policy further generalizes to stronger task agents and transactional domains of $\tau^3$-bench without retraining, demonstrating the benefit of source-aligned uncertainty resolution for proactive agents.
|
| 663 |
Agentsensus: Consensus-Compressed Shared Memory for Multi-Agent Story Worlds
2609.32297
|
cs.CL
|
Yu Pan |
A agentic story world is a dynamic system simulating who learned what, when, and from whom -- yet the standard design gives each character a private memory stream. A shared event is therefore stored once per witness, large duplication will be incurred in terms...A agentic story world is a dynamic system simulating who learned what, when, and from whom -- yet the standard design gives each character a private memory stream. A shared event is therefore stored once per witness, large duplication will be incurred in terms of storage. We present Agentsensus, a story-world simulation framework in which there is an unified long-term memory. Records of the same event merge into one owned by all its witnesses, and semantically relevant memory records are linked. We evaluate on four worlds -- two classical Chinese novels, Hamlet, and a real-world conflict timeline -- run for 40 to 80 rounds against three per-character memory designs under an equal-granularity protocol. Agentsensus writes 22-44% fewer entries than the closest baseline and is the only design whose memory becomes shared (14-28% of records held by more than one character, some by 10) and linked (94-99%), at judged simulation quality indistinguishable or even better than the baselines. An ablation attributes this to the merge itself: disabling it multiplies the store by 3.1x and takes sharing to exactly zero. Sharing also compounds with the horizon rather than saturating early, rising 6% to 9% to 14% as one world is re-run at 10, 20 and 40 rounds.
|
| 664 |
What Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review Agents
2609.32318
|
cs.CLcs.LG
|
Zhaowei Han, Xiang Zhang, Lingxiao Guan, Danqi Hu, Kai Liu |
Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical cor...Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical corrections do not remove this mismatch. We introduce compositional controllability to address these questions. A comparison window covers one stage, several stages, or the whole agent. Our central result bounds the gap between observed and controlled score differences using only nuisance outside the window. This yields an admissibility test applied before scores are inspected. Inadmissible comparisons are refused. For admissible pairs, an ordering is certified only when the score gap exceeds the combined sampling and nuisance radii; otherwise, it remains undecided. These decisions give each system a rank interval. We introduce BioLitBench, a benchmark of 2,042 biomedical articles represented as structured claim graphs. Among seven published pipelines, a conventional statistical analysis declares a winner in 14 of 21 pairwise comparisons. Yet the top-ranked system alone received the target review's bibliography. To isolate pipeline performance, our test requires matched inputs and a fixed backbone model. It refuses 11 of the 21 comparisons, including every comparison involving the top-ranked system. Seven of the 14 conventional conclusions fall within these refused pairs. The same comparison windows support stage-level training. We train SCRIBE on Qwen3.8-27B using rewards measured at each stage's exit. Under matched evidence, SCRIBE achieves a certified rank interval of [1,2], with certified advantages over all evaluated published pipelines and the evaluated Claude and OpenAI agents. Under same pool, SCRIBE matches the strongest published retriever and is certified above three published pipelines.
|
| 665 |
Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
2609.32353
|
cs.CLcs.LG
|
Junxian Li, Ruixuan Yang, Tianao Zhang, Tiange Xu, Weisheng Dong |
Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to r...Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
|
| 666 |
Black-Box Auditing of Epistemic Reliability in Multi-Agent Debate Distillation
2609.32361
|
cs.CLcs.LG
|
Derui Wang, Zewei Shi, Rayne Holland, Ruoxi Sun, Xingliang Yuan |
Debate distillation adapts weaker verifiers using multi-agent debate transcripts to improve their judgement in subsequent debates, but gains on monitored tasks do not establish reliability on related unmonitored tasks. We study epistemic reliability degradatio...Debate distillation adapts weaker verifiers using multi-agent debate transcripts to improve their judgement in subsequent debates, but gains on monitored tasks do not establish reliability on related unmonitored tasks. We study epistemic reliability degradation, in which adaptation preserves monitored performance while reducing support for correct responses on hidden tasks. We consider an adversarial debater that manipulates debate arguments while defending the correct monitored response, and ask whether the resulting degradation merely reflects catastrophic forgetting and whether standard evaluation can detect it. To address these questions, we propose ER-Audit, a two-stage black-box auditing framework that compares frozen verifier checkpoints before and after adaptation, and introduce two evaluation benchmarks pairing monitored and hidden task prompts grounded in shared contexts. ER-Audit searches for counterexamples to non-degradation by evaluating semantically valid paraphrases and, if none is found, uses independent paraphrases for sequential hypothesis testing. We derive anytime-valid lower confidence bounds on the non-degradation probability, allowing data-dependent stopping within a finite budget. We further establish a common lower bound across fixed paraphrase distributions and extend it to distributions within a bounded total variation distance of their mixtures. Our experiments show that higher hidden-task accuracy can coexist with more counterexamples to non-degradation and lower non-degradation bounds. This divergence challenges explanations based solely on broad catastrophic forgetting and shows that auditing can uncover selective hidden-task degradation concealed by aggregate performance gains. Our code and benchmarks are available at https://github.com/CSIRO-CQS-AI-alignment-Team/Epistemic-Reliability-Auditor.
|
| 667 |
ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation
2609.32448
|
cs.CL
|
Junming Liu, Jicheng Wang, Yifeng He, Hao Chen, Jianzhong Qi |
Autoregressive Next-Token Prediction (NTP) has enabled strong reasoning capabilities in language models, while Diffusion Language Models (DLMs) offer flexible token orders and parallel generation. We ask whether DLMs can acquire NTP-style reasoning through dis...Autoregressive Next-Token Prediction (NTP) has enabled strong reasoning capabilities in language models, while Diffusion Language Models (DLMs) offer flexible token orders and parallel generation. We ask whether DLMs can acquire NTP-style reasoning through distillation without giving up their native generation process. Direct distillation, however, faces a fundamental mismatch: an autoregressive teacher predicts from a left prefix, whereas a DLM can condition on tokens on both sides. We introduce ForkLeft, a distillation framework that resolves this mismatch by separating the student's rollout from teacher supervision. During training, the student first performs entropy-first rollouts that commit uncertain positions and expose potential forks. We then fix the resulting student prefix and distill an NTP teacher under the same context, with answer correctness determining the supervision source. At inference, the student returns to its native confidence-first parallel decoding. With Qwen3-30B-A3B-Base, ForkLeft improves Efficient-DLM-4B on all ten benchmarks, raising MATH500 from 72.60% to 79.60% and consistently outperforming three alternative designs. The gains scale with teacher strength and generalize to SDAR-4B with only $500$ updates. At matched scale, the distilled 4B and 8B students exceed the published SDAR-Chat and OPDLM models on seven benchmarks, showing that DLMs can learn NTP-style reasoning without sacrificing native parallel generation. Code and datasets will be released upon acceptance.
|
| 668 |
PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders
2609.32469
|
cs.CLcs.LG
|
Chenduo Hao, Chuanbao Gao, Pinjun Zeng, Jingze Zhu, Chonghan Liu |
In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal feature...In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce PULSE (Paired Utility Localization over Sparse Encodings), an SAE-based framework for identifying model-internal features associated with demonstration utility and using them for demonstration selection. Using a small labeled discovery set, PULSE samples candidate demonstration sets, measures their zero-shot-relative utility under the target model, and scores SAE features by how their activation differences align with utility differences. The top positive and negative coordinates form a sparse utility-localization vector. We use this vector in two complementary ways: as a signed score for controlled complete-set ranking, and as PULSE-Retriever, which converts its magnitude into a feature-relevance mask for scalable pool-scale retrieval. Across classification, generation, and reasoning benchmarks, PULSE-Retriever improves over the strongest baseline by 2-3 accuracy points, 0.6-0.9 BLEU-4, and 3.2 exact-match points, respectively, while controlled ranking validates the identified features encode a predictive set-level utility signal. Feature inspection and cross-dataset experiments suggest that the identified features capture task-relevant, dataset-conditioned patterns, yet retain utility signals that partially transfer across datasets. Our code is available at https://github.com/aohenuo/PULSE.
|
| 669 |
Activation Flow: Manufacturing Activations for Steering
2609.32530
|
cs.CLcs.LG
|
Hong Kiat Tan, Linh Le, David Williams-King |
Difference-in-means steering requires activations recorded while a model shows the desired behavior, which a sandbagging model withholds by deliberately underperforming. We introduce Activation Flow (ActFlow), which manufactures these activations from $k$ corr...Difference-in-means steering requires activations recorded while a model shows the desired behavior, which a sandbagging model withholds by deliberately underperforming. We introduce Activation Flow (ActFlow), which manufactures these activations from $k$ correct labels without fine-tuning. ActFlow sets target logits that rank each labeled item's correct answer first, and moves the logits toward them by adding one vector $x$ to all $k$ residual streams at one layer. ActFlow is a family of ordinary differential equations for $x$, one for each rule that maps the required logit change to the velocity of $x$. The smallest-norm rule lands exactly on the targets, while the others keep only the top singular directions of the Jacobian. We test ActFlow on three instruction-tuned models, each locked by a sandbagging prompt and by a password-locked LoRA. At $k=40$, ActFlow keeping five singular directions raises the mean held-out ARC-Easy accuracy over the six locked models from $0.05$ to $0.85$, against $0.88$ for fine-tuning and $0.92$ for the honest models. Furthermore, it scores higher than the smallest-norm rule in 16 of the 18 combinations of locked model and $k$, and its steering direction is nearly orthogonal to the honest difference-in-means direction. It also unlocks two LoRA locks where the honest direction fails.
|
| 670 |
Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
2609.32536
|
cs.CL
|
Yanjie Zhang, Nanchen Hu, Yushi Sun |
Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across ...Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Each item uses a shared action space comprising silence, a tool call, and a natural-language answer. Speaker-switch pairs hold the specified words fixed while source, distance rendering, and a temporal boundary define a controlled wearer-to-bystander shift. Six raw Audio LLMs and three training-free adaptations often identify the target tool yet rarely withhold action under this shift; the highest raw switch mute rate is 14%. We then use VoxGate as a post-training case study. Supervised training mutes 91.3% of switched commands while choosing the correct tool for all nearby wearer commands and text-only controls. An exploratory GRPO stage has similar switch performance; side-talk accuracy rises from 68.4% to 70.9%, and self-talk muting from 52.0% to 60.0%. Factorized controls identify an independent source-change effect, while sensitivity to the far-field manipulation varies across acoustic renderings. The benchmark therefore measures multi-cue acoustic-context gating rather than isolated speaker identity.
|
| 671 |
Shared Autoregressive Context Can Distort Relationships in Synthetic Data
2609.32546
|
cs.CLcs.LG
|
Thomas S. Robinson |
Large language models can generate several records within one autoregressive completion, making earlier answers available as context for later records. This paper shows that such shared-completion batching can distort relationships among variables in the resul...Large language models can generate several records within one autoregressive completion, making earlier answers available as context for later records. This paper shows that such shared-completion batching can distort relationships among variables in the resulting synthetic data, using controlled tests on synthetic survey respondents. In a matched experiment on 2,000 European Social Survey profiles, generating ten rather than one respondent per request increases mean absolute error in within-country correlations by 48-58% for Qwen3.8-27B and 114-127% for Llama-3.3-70B-Instruct across three seeds, holding profiles, examples, questions and decoding parameters fixed. The distortion primarily reflects exaggerated relationship strength, while retaining substantial agreement with the human ordering of correlations. Controlled interventions establish answer history as a causal channel: re-pairing the same preceding values, with profiles and marginal distributions fixed, changes correlations among subsequently generated responses. Hiding preceding answers reduces correlation error in the tested settings but worsens marginal accuracy. Exploratory corrections across social-attitude, health and economic data likewise show that lower correlation error can coexist with worse marginal distributions and regression estimates. Request construction is therefore part of the data-generating process, and synthetic-data validity must be evaluated against the analyses the generated data are intended to support.
|
| 672 |
Reading Is Not Leaking: Local, Auditable Measurement and Reduction of Inference Exposure from Public Footprints
2609.32565
|
cs.CL
|
Mahmudul Faisal Al Ameen |
Anyone with a public footprint leaks facts that were never stated, and language models make the inference cheap. We present a framework for measuring and reducing this inference exposure that runs on the owner's own CPU with no language model at analysis time,...Anyone with a public footprint leaks facts that were never stated, and language models make the inference cheap. We present a framework for measuring and reducing this inference exposure that runs on the owner's own CPU with no language model at analysis time, instantiated on organisations and on individuals. It starts from a measurement result: scoring an inference system against the target's private truth conflates how well the system reads the record with how much the record leaks. On a 128-question instrument over sixteen synthetic firms, almost half of the questions are never answered correctly by any of six readers, four of them language models, and a majority-class guess accounts for most of every reader's score. We therefore separate reading accuracy from leakage rate and introduce an injection protocol that creates cells with known support. Our analyser combines rules, statistical solvers and a 106M-parameter encoder trained from scratch that marks verbatim evidence and never generates text; every answer carries a graded certificate whose recorded proof replays. Its certified answers are correct in 93% of resolved cases, against 49-73% for the language models' quote-backed answers, whose citations are produced alongside the answer rather than deriving it; with plain-prose articles in the record, 70% of its evidence-bearing answers rest on evidence that establishes them, against 18-56% for the models. A constrained defence that rewrites each fact's carrier as a true but coarser statement hides every single-carrier fact from four language-model adversaries at 40% lower edit cost than deletion. On sixteen synthetic people the guessing term is larger still, and a decoy planner with no language model halves the correct answers of the estimator it targets without transferring to a second.
|
| 673 |
Attribution Gaps in Zero-Training LLM+OVOD Pipelines: A Fine-Grained Analysis of the CAAP--SNAP Discrepancy
2609.32567
|
cs.CL
|
Yu-Feng Yen |
LAOD and similar zero-training LLM+open-vocabulary-detector (OVOD) pipelines score two things separately: class-agnostic localization accuracy (CAAP) and semantic naming accuracy (SNAP). The two consistently diverge, and nobody has asked why. This paper asks w...LAOD and similar zero-training LLM+open-vocabulary-detector (OVOD) pipelines score two things separately: class-agnostic localization accuracy (CAAP) and semantic naming accuracy (SNAP). The two consistently diverge, and nobody has asked why. This paper asks why, on the full 5,000-image COCO-Val split (27,273 detections) rather than the small subset the original work evaluated on. Object visual complexity turns out not to be the driver -- small and occluded objects are, if anything, localized better than large ones. Vocabulary novelty is: once the LLM's wording falls outside the detector's native category set, localization accuracy falls from 80.9% to 31.6%. That drop is not spread evenly across unfamiliar phrasing, though. Almost all of it comes from cases where the novel wording actually names a different object than the one COCO annotated (true synonyms still score 89.3%; semantically unrelated "noise" labels score 12.0%). A closer look at a further failure subset tells a similar story: 78-88% of what looks like complete localization failure is really the model correctly finding a real object that COCO's non-exhaustive 80-category scheme simply never labeled, not hallucination. Swap the detector backbone (YOLO-World for Grounding DINO) or the LLM (Gemma-3 for Qwen2.5-VL) and both the effect and its rough size hold up, so this looks like a general property of the pipeline family rather than a quirk of one model pairing. The upshot is that a large share of the apparent CAAP--SNAP gap traces back to closed-category annotation limits rather than a real grounding failure, which matters for how we detect hallucination, analyze failure modes, and design evaluation for grounded multimodal systems meant to work in the open world.
|
| 674 |
CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations
2609.32574
|
cs.CL
|
Yulin Hu, Yanyan Zhao, Zimo Long, Xing Fu, Mengtong Ji |
Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, o...Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues. Existing benchmarks largely focus on text-only memory or explicit multimodal evidence, leaving implicit multimodal cues underexplored. We introduce CUE-Mem, a text-image-audio benchmark for evaluating long-term user memory from implicit cues. CUE-Mem contains 2,674 questions across explicit and implicit evidence settings and covers four tasks: Entity Recall, Long Pattern, Personalized Recommendation, and Answer Refusal. Across textualized memory systems, implicit performance remains far below oracle evidence, locating the main bottleneck in preserving and retrieving subtle cues rather than question answerability. Increasing caption detail recovers more of this evidence, but brings uneven gains and rapidly growing token costs, motivating native multimodal access. Yet native access does not uniformly resolve the bottleneck: evidence use depends strongly on the backbone, while multimodal indexing introduces substantial retrieval noise. CUE-Mem provides a testbed for memory systems that selectively retain, retrieve, and use subtle multimodal evidence.
|
| 675 |
Retrieved but Not Delivered: Multimodal Memory Delivery for Long-Term Agents
2609.32590
|
cs.CL
|
Yuhang Jiang, Qingwei Liao, Kaize Yin, Xingling Liu, Luca Cuomo |
Work on memory for multimodal agents optimizes what is written, updated and retrieved. Between retrieval and the answer, however, is a stage that multimodal memory evaluations do not isolate: what of the retrieved memory reaches the model, and in what form. We...Work on memory for multimodal agents optimizes what is written, updated and retrieved. Between retrieval and the answer, however, is a stage that multimodal memory evaluations do not isolate: what of the retrieved memory reaches the model, and in what form. We call it delivery, and a controlled decomposition on MemLens locates the remaining room there. With the retrieved evidence set exactly fixed, delivering the original pixels instead of withholding them raises accuracy by 13.87 points on an 8B backbone, whereas making retrieval perfect on those same messages improves it by 2.31. Delivery is the larger term on all three MemLens backbones and grows with backbone strength; retrieval grows too, without closing the gap. We propose DeliverMem, an instantiation of delivery as three decisions: keep the original modality, give each item a readable identity, and state when it was seen, with a retrieval-side adapter for the one property delivery cannot supply. Each is measured against a delivery-matched control that alters only its own variable. DeliverMem leads the strongest published memory agent on MemLens at all four context lengths, and beats DMV-Bench's own strongest method at every setting on both backbones. On MemLens it does this on a tenth to a seventieth of the input. Each decision helps only where the question lacks what it supplies, and is null elsewhere. A single fixed configuration nonetheless leads both benchmarks, without training any component or modifying the stored records. Project page: https://avalon-s.github.io/DeliverMem/
|
| 676 |
Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
2609.32663
|
cs.CLcs.LG
|
Xianpeng Shang, Canbin Huang, Jiang Li, Tian Lan, Qianyi Cai |
The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across head...The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval capability varies substantially with relative distance, even within the same attention head. To exploit this structure, we introduce Distance-KV, which learns a static KV retention pattern over the joint space of layers, attention heads, and relative distances. The pattern is learned offline with the language model frozen and reused across inputs to prune and compact the KV cache without online importance scoring. Across three backbone models and four long-context benchmarks, Distance-KV consistently achieves the best overall performance among competing KV cache compression methods, exceeding the strongest compression baseline by up to 9.3 points on RULER at 128K. On Llama-3.1-8B-Instruct at 128K, Distance-KV reduces KV cache memory by 65.4% and achieves a $1.66\times$ decoding speedup relative to Dense. Together, these results identify relative distance as an important structural dimension for understanding how LLMs retrieve information over long contexts and for designing more efficient inference methods.
|
| 677 |
The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models
2609.32679
|
cs.CLcs.LG
|
Dongsheng Liu, Chao Jin, Wenkui Yang, Hejin Wang, Junwei Yang |
GUI World Models (GUI-WMs) are increasingly used to predict future states for agent planning and simulation, yet most existing formulations condition only on the current GUI observation and action. We identify state aliasing, where the vis- ible interface omit...GUI World Models (GUI-WMs) are increasingly used to predict future states for agent planning and simulation, yet most existing formulations condition only on the current GUI observation and action. We identify state aliasing, where the vis- ible interface omits transition-relevant environment state, so identical observable conditions can correspond to different valid futures. To diagnose this failure mode, we introduce StateAliasBench, a diagnostic benchmark that explicitly isolates such ambiguities via strict pairing. We further propose lightweight predictive- state recovery that infers structured state from history and augments otherwise frozen GUI-WMs through a deterministic state interface. Family-specific special- ists provide state recovery across heterogeneous state types, and multi-teacher dis- tillation consolidates them into a single unified estimator. Experiments show that existing GUI-WMs exhibit systematic failures under observation-only condition- ing, while predictive-state augmentation substantially restores state-sensitive pre- diction across evaluated WMs, preserves generative fidelity, and improves down- stream performance of GUI agents on AndroidWorld. These results suggest that reliable GUI world modeling should account not only for what is visible, but also for the hidden transition state that determines what happens next.
|
| 678 |
IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents
2609.32694
|
cs.CL
|
Angqing Jiang, Gaoming Zhang, Chaoqun Zhang, Jianchun Song, Liyuan Kong |
On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a que...On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from the student's state. Existing methods either distill this preference directly or filter it with model-internal scores, but neither strategy verifies the query's executed retrieval consequence. We propose Information-Gain-Gated Self-Distillation (IGSD), which verifies on-policy token proposals with environment feedback before distilling them. Treating each query token as a micro-action, IGSD completes the teacher's token proposal and the student's sampled token into matched queries and executes both from the same failed state with the same retriever. Shared counterfactual controls account for query-conditioned shifts in answer likelihood, so their difference, the executed paired information gain, provides a relative utility contrast for the retrieved documents. IGSD uses this contrast as a positive-only soft weight for candidate-pair distillation, while leaving the GRPO objective unchanged and confining verification to training. Across seven single-hop and multi-hop QA benchmarks, IGSD reaches macro-average exact-match accuracies of 42.8% and 47.0% with 3B and 7B policies, respectively, without inference-time verification. These results support environment-verified hindsight as an effective approach to reliable action-level supervision for search agents.
|
| 679 |
Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time
2609.32701
|
cs.CL
|
Jacob Dineen, Silei Ren, Muhao Chen, Dan Roth, Ben Zhou |
In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs...In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four secret states and selects one of four summaries of the same public report, while the receiver model tries to infer the secret state. We find that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly. This learning occurs during inference with fixed parameters and no supplied codebook or encoding examples. The effect also persists when agents generate their own free-form updates in a simulated incident-response task. Across ten independent games, pairs of GPT-5.6 Sol agents reach 98.8% final accuracy, compared with 25% chance, despite explicit instructions prohibiting disclosure and a monitor that screens each message without access to the agents' interaction histories. The same interactions that help agents cooperate can therefore allow confidential information to pass through messages intended for legitimate coordination.
|
| 680 |
MassAlloc Attention: Let Attention Allocate Its Own Compute
2609.32712
|
cs.CL
|
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin |
FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal ca...FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing for adaptive retention of the work. MALA reduces low-contribution post-score computation. A matched-work study at 8K isolates the benefit of distribution-adaptive allocation: under exactly matched total post-score work, MALA approaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%. Across context lengths from 1K to 32K tokens, the same tolerance maintains low output and gradient errors relative to the reference. Across a broader controlled associative-recall comparison, MALA closely tracks FullAttn as context grows, reaching 89.67% accuracy at 8K compared with 89.97% for FullAttn. In an attention-operator benchmark at 128K tokens with tensor parallelism, MALA reduces forward and backward latency during training by 2.2x and 3.0x and decoding latency during inference by 1.6x relative to FullAttn. Across scaling-law training from 0.6B to 14B parameters, MALA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results indicate that allocating post-score computation according to normalized attention contributions can retain the evaluated capabilities of FullAttn while reducing attention computation.
|
| 681 |
Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models
2609.32737
|
cs.CLcs.LG
|
Dongdong Wang, Deepak Balakrishnan, Ravi Srinivasan, Shenhao Wang |
Existing geospatial vision-language models (Geo-VLMs) typically optimize diverse geospatial tasks through a unified multi-task adaptation paradigm without explicitly accounting for the heterogeneous optimization characteristics. Our empirical observations reve...Existing geospatial vision-language models (Geo-VLMs) typically optimize diverse geospatial tasks through a unified multi-task adaptation paradigm without explicitly accounting for the heterogeneous optimization characteristics. Our empirical observations reveal heterogeneous gradient characteristics across tasks, including vision-language differences, intra-branch gradient relationships, and task interference, which hinder effective multi-task optimization. Motivated by these observations, we propose Gradient-Guided Decoupled Adaptation (G2DA), a gradient-aware optimization framework for multi-task Geo-VLM learning. G2DA first partitions tasks into vision- and language-centric groups through gradient-guided cross-modal decoupling. It then constructs modality-specific curricula based on task gradient similarity and employs bidirectional rehearsal to mitigate the recency effects introduced by sequential optimization. We evaluate G2DA on three Geo-VLM benchmarks using six InternVL3 and Qwen3.5-VL variants, along with GeoChat and GeoLLaVA. Across all 24 benchmark-model combinations, G2DA consistently outperforms representative baselines, improving over the strongest competitor by 3.08, 4.30, and 2.81 percentage points on UrBench-MCQ, XLRS-Bench-Lite, and VRS-Bench-VQA, respectively. These results demonstrate the effectiveness of gradient-guided task organization for Geo-VLM adaptation.
|
| 682 |
Retrospective Distillation Attribution via Normalized Response Similarity
2609.32749
|
cs.CLcs.LG
|
Minwoo Jang, Jaechang Kim, Minhyeon Oh, Jeongyeon Hwang, Jungseul Ok |
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students ...Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring *syntactic patterns* into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.
|
| 683 |
Adaptive Consistency Graph for Long-Horizon Agents
2609.32754
|
cs.CLcs.AI
|
Jiecong Wang, Hao Peng, Zhanyi Wang |
Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, historical evidence, and the cur...Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, historical evidence, and the current execution state may gradually become disconnected, so later decisions can drift from the original objective. We study this problem by introducing the Adaptive Consistency Graph (ACG) for long-horizon execution. ACG incrementally organizes execution evidence and its provenance in a persistent graph, then constructs a temporary requirement-centered view for each decision under a bounded context budget. Rather than replacing the base agent's planner or tool executor, ACG provides a structured and traceable context view for each decision. In the matched evaluation, ACG improves GPT-5.6-luna's average success from 44.5\% with ReAct to 50.2\%, with the largest gain on BrowseComp-Plus (73.5\% versus 62.4\%). We further analyze trajectory structure and inference cost to characterize this improvement. Our code is available at https://github.com/yunsaijc/Adaptive-Consistency-Graph.
|
| 684 |
CLAIRE: A Schema-Grounded Hybrid Workflow for Healthcare Administrative Form Completion
2609.32787
|
cs.CL
|
Garapati Keerthana, Manik Gupta |
Healthcare administrative staff transfer structured information from electronic health records, referrals, claims systems, provider rosters, and work queues into dynamic forms. We developed and evaluated CLAIRE (Clinical Language and Agentic Intelligence for R...Healthcare administrative staff transfer structured information from electronic health records, referrals, claims systems, provider rosters, and work queues into dynamic forms. We developed and evaluated CLAIRE (Clinical Language and Agentic Intelligence for Reasoning and Entry), a hybrid workflow that separates field-state discovery, source-to-field mapping, deterministic validation, bounded correction, escalation, and audit tracing. We tested five synthetic healthcare administrative schemas, 1,000 source records, four interface variants, two data-quality suites, and six comparators, yielding 24,000 benchmark episodes. A separate strict-output audit evaluated direct mappings from Qwen2.5-1.5B and Qwen2.5-7B, and a trace-derived operational simulation covered 6,000 episodes. Under the evaluated synthetic benchmark conditions, full CLAIRE achieved 1.000 episode success, field accuracy, required-field completion, and dependency completion in both suites; removing validation reduced stress-suite success to 0.500. In the simulation, 100.0% of clean and validation-stress episodes reached a staff-reviewable draft, compared with 68.6% of escalation challenge episodes, unsupported cases were blocked. Scenario-based savings were 149.7-165.5 seconds per case, not observed staff times. The findings support schema-grounded, validation-first healthcare administrative automation in which language-model components assist mapping but do not authorize unsupported or consequential actions.
|
| 685 |
Understanding and Exploiting Anisotropy in Post-Training
2609.32792
|
cs.CLcs.LG
|
Samyak Jha, Harshvardhan Saini, Yizhen Liao, Yiming Tang, Dianbo Liu |
LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in wh...LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry disproportionately large activations. Anisotropy is widely documented and usually treated as a defect, yet its function and its interaction with post-training remain unclear. We first analyze it. A label-free outlier rule isolates about 5\% of residual channels that are essential for language modeling: removing them raises perplexity from 10 to over $10^6$, versus 35 for count-matched random channels. Yet they barely distinguish correct from incorrect reasoning. SFT reshapes them, whereas RL leaves them largely intact and adapts the complementary channels. These channels therefore form the model's \emph{coherence substrate}, and reasoning adaptation happens elsewhere. We then exploit this. \textsc{SphereGate} learns one bounded gain per residual channel on a frozen backbone. Its activation-weighted gradients provably limit movement of high-energy coherence channels and leave the remaining channels free. With 0.1M trainable parameters, \textsc{SphereGate} outperforms parameter-efficient baselines by 2.0--7.3 points on MATH-500 across Qwen2.5 (0.5B--7B) and Llama-3-8B, is comparable or exceeds full-model GRPO. Anisotropy is not a defect but a division of labor that post-training can exploit.
|
| 686 |
Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault
2609.32802
|
cs.CLcs.LG
|
Tianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao |
One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth +0.608 [+0.517, +0.700] t...One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth +0.608 [+0.517, +0.700] to +0.358 over a blind one on four open-weight backbones served with thinking disabled, and on the two Qwen backbones the blind inspector changes no item at all. That head-to-head is exploratory. One deterministic fault enters the first stage, and we re-expose the original problem to $k = 0,\dots,3$ of the downstream stages with agents, items, fault and topology held fixed, on 120 gsm_hard items per arm at temperature zero. Accuracy under fault rises on four of four backbones, from +0.233 to +0.392, the largest Holm-adjusted $p$ being $2.1\times10^{-6}$. A registered kill test rules out tokens. Blanking every word holds the word slots fixed, and retention tracks the visible fraction on four of four, climbing from 0.221 to 0.692 on the primary. Those two families are confirmatory and everything else here is exploratory. The interaction excludes zero on two of four backbones under the registered pipeline, four of four under a three-stage pipeline, and three of four under full message history, the primary at +0.317. On Llama-3.1-8B the fault carries no detectable cost at any dose, so the other three carry every claim about what a fault costs. Re-derivability also sets what the architecture costs, and no decomposition we measured reliably beats one direct call. With no fault injected the registered pipeline loses to that call by -0.267, -0.125 and -0.317, and on Phi-4 reads +0.058 at $p = 0.118$, which the test fails to separate from zero. The repair that works is cheap and front-loaded: the first re-grounded stage buys +0.394 of matched retention for +59.8 tokens per item on Qwen3-14B, and the stages after it buy nothing.
|
| 687 |
Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret
2609.32805
|
cs.CLcs.LGcs.AI
|
Bingyu Shen, Boyang Li |
Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, an...Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, and a reader acts from the state alone. Steps stay cheap, but anything the writer drops is lost before later decisions reveal that they need it. We quantify this loss and ask whether training can reduce it. Comparing the written state with the best state of the same size written in hindsight, we split the reader's loss into a budget loss, which any state of that size must incur, and a write-time regret, which comes from the writer's choices. In TextWorld cooking games where we control how long a fact must be carried before it is needed, a 128-token state holding the facts wins nearly every game, while prompted language-model writers win at most 17%. Almost all of the loss is write-time regret, and it grows with the delay. We then train the writer from the reader's own loss. DSSR (decision-sufficient state representations) scores candidate states by how well the reader acts after the writer carries them forward, and teaches the writer to prefer the better ones. This forward-rolled score predicts game outcomes ($\rho = 0.48$), whereas scoring a candidate as a fixed context, as hindsight methods usually do, does not ($\rho \leq 0.07$). On a pre-registered test split opened once, training adds +7.0 [+1.9, +12.2] points of success when facts are needed soon, bringing a plain summary writer to the level of belief- and slot-based memory prompts. The gain shrinks as the delay grows and is significant only at the shortest delay. We trace this limit to credit assignment: keeping a fact now pays off only if every later rewrite keeps it too, which a per-step score cannot see.
|
| 688 |
Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision-Language Models
2609.32808
|
cs.CLcs.LG
|
Jonas Ngnaw\'e, Yann Pequignot, Sabyasachi Sahoo, Christian Gagn\'e, Fr\'ed\'eric Precioso |
Large vision-language models (LVLMs) inherit massive activations from their text-only bases: spikes where a few fixed hidden channels receive values thousands of times above the typical magnitude. The text spike systematically appears in early layers at a fixe...Large vision-language models (LVLMs) inherit massive activations from their text-only bases: spikes where a few fixed hidden channels receive values thousands of times above the typical magnitude. The text spike systematically appears in early layers at a fixed initial position, independently of input content. Visual spikes vary across images, but whether their formation follows a consistent pattern across LVLMs and how they respond to image perturbations remain open questions. We find that some LVLMs do not form visual spikes, while others spike at different rates, typically in deeper layers. We identify the trigger direction from model weights and an interpretable location rule: before the language model decoder runs, eventual spike tokens are largely restricted to those sharing least with the rest of the image. Crucially, visual spikes are strikingly brittle. Common corruptions frequently create and relocate spikes, and less often remove them, raising overall incidence. Our trigger-guided spike attack deliberately creates or removes spikes under a small $\ell_\infty$ budget, with 1/255 enough in nine of the ten models that spike. Finally, our preventive intervention removes only the trigger component before spikes erupt, eliminating or substantially reducing spikes on clean and perturbed images while leaving the other image tokens nearly unchanged. Our study spans 25 adapter-based LVLMs built on 18 released text-only bases from 10 families, ranging from 2B to 72B parameters.
|
| 689 |
Overwhelmed by Choice: Studying LLM Decision Making at Scale
2609.32809
|
cs.CL
|
Yu-Chi Lin, Aryan Seth, Anshul Aravind, Eugene Lee, Tanmay Parekh |
Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the c...Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically evaluate LLMs as the number of competing candidates increases and find substantial accuracy degradation across tasks, prompting strategies, and model scales. Controlled analyses show that standard long-context retrieval explanations cannot fully account for this degradation. Instead, we identify two systematic failure patterns. First, gold-margin collapse: the score gap between the correct answer and the strongest distractor progressively shrinks, driven primarily by weakening confidence in the correct answer. Second, earlier candidate preferences become increasingly difficult to overturn, with later candidates exerting progressively weaker influence on the final prediction. Motivated by these findings, we evaluate hierarchical partitioning and permutation-based inference, which improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC. Overall, our results identify candidate-set scale as an important evaluation-protocol variable and show that strong small-option performance does not necessarily imply robust large-scale candidate comparison.
|
| 690 |
USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments
2609.32813
|
cs.CLcs.LG
|
Dongdong Wang, Qingqi Song, Yuzhou Chen, Deepak Balakrishnan, Ravi Shankar Srinivasan |
Large vision-language models (VLMs) have emerged as a powerful paradigm for urban and spatial AI. However, current state-of-the-art large VLMs still struggle with quantitative reasoning on remote sensing imagery. Existing benchmarks and algorithms are predomin...Large vision-language models (VLMs) have emerged as a powerful paradigm for urban and spatial AI. However, current state-of-the-art large VLMs still struggle with quantitative reasoning on remote sensing imagery. Existing benchmarks and algorithms are predominantly based on qualitative Visual Question Answering (VQA), providing limited insights into the quantitative reasoning capabilities of VLMs for built environment metrics. To address this gap, we develop Quantitative Urban and Spatial AI benchmark (USAI-Quant), the first benchmark designed to quantitatively evaluate VLM's reasoning capabilities on built environment metrics via remote sensing imagery. USAI-Quant is curated from the 335 largest U.S. cities, aligning high-resolution remote sensing images with quantitative built environment metrics. We then evaluate both general-purpose and remote sensing VLMs (RS-VLMs) by applying VQAs to tens of built environment metrics across three complexity levels. Our results reveal that current state-of-the-art models consistently fall short on numeric reasoning tasks. We further conduct in-depth analyses across models, question types, and geographic locations, uncovering insights into performance variability and task-specific challenges.
|
| 691 |
The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces
2609.32825
|
cs.CL
|
Tianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao |
A four-stage LLM pipeline gives up as much as 40.5 accuracy points at its own interfaces (gemma-3-12B on MATH-500, Holm-corrected p = 1.66e-19; the largest tax in the primary family). We hold model, problem, stages, stage prompts and completion budget fixed, v...A four-stage LLM pipeline gives up as much as 40.5 accuracy points at its own interfaces (gemma-3-12B on MATH-500, Holm-corrected p = 1.66e-19; the largest tax in the primary family). We hold model, problem, stages, stage prompts and completion budget fixed, vary only whether each stage can still see the original problem, and call the accuracy difference the decomposition tax. Across 21 open-weight models from nine organisations, on GSM-Hard and MATH-500 at n = 200 paired items per cell, 70 of 118 primary-family tests survive Benjamini-Hochberg correction and 54 survive Holm. On GSM-Hard, a placebo recovers nothing: it carries at least 60% of the extra tokens and at most one word of the problem. Builders design a pipeline one stage at a time, and its bill arrives at the interfaces between stages. Rewriting one stage's instruction moves gemma-3-12B's tax from 4.5 to 36.5 points, and adding "every relationship stated between them" to a stage that lists the numerical quantities lowers the tax on 9 of 9 models on MATH-500. Re-grounding, which shows a stage the original problem again, belongs after the loss. With one lossy interface, re-grounding the stage after it beats re-grounding the stage before it on 7 of 7 models on both benchmarks; on MATH-500 the earlier repair is worse than none on 7 of 7. Newer models still pay: gemma-4-12B gives up 37.0 points, and the repair holds on all three of the newest models we test. A sealed held-out test refuted a stronger rule we registered, which predicted the paying stage from the interface and receiver types, so we locate the tax by measuring one stage at a time. The prescription has two parts: re-ground the stage after the lossy interface, and if a stage must list the quantities, tell it to keep the relationships.
|
| 692 |
Logical subspace in LLMs
2609.32907
|
cs.CL
|
Hope Kean, Enric Boix-Adsera |
Recent work has identified a human brain network specialized for abstract formal reasoning (Kean et al., 2025). Does the same hold true in language models? To answer this question, we introduce the minimal viable subspace (MVS) method, which searches for the l...Recent work has identified a human brain network specialized for abstract formal reasoning (Kean et al., 2025). Does the same hold true in language models? To answer this question, we introduce the minimal viable subspace (MVS) method, which searches for the lowest-rank activation subspace at a layer that preserves task performance when everything outside that subspace is ablated. Using MVS, we demonstrate low-rank subspaces supporting logical inference on Gemma and Qwen models. Furthermore, these subspaces exhibit a clear dissociation from model capacities on other tasks, such that retaining these late logic subspaces preserves inference while impairing factual knowledge, working memory, cognitive control, and arithmetic. Conversely, ablating them reduces logical inference accuracy to chance while largely sparing these other capacities. Our results suggest a functionally localizable core machinery for logic akin to that in the human brain.
|
| 693 |
Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows
2609.32917
|
cs.CL
|
Vivek Kumar Singh, Preeti Priyam, Gautam Bhowmick |
Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model costs per token, and the gap compounds once a workflow chains several calls toge...Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model costs per token, and the gap compounds once a workflow chains several calls together. Planner-as-Router (PaR) attacks this from a different angle. Instead of leaving model-tier selection to some component downstream, it folds the choice into planning itself. As the planner breaks a query into subtasks, it also assigns each one a model size tier (small, mid, or frontier, ordered by capability and price), so the dependencies between subtasks are visible before any specialist runs. Unlike per-call routers such as cascade routing, which look at one node at a time, PaR sees the whole workflow up front and needs no separate router model or training data. We evaluate PaR with EntBench, a benchmark of 54 enterprise agentic tasks across seven classes, graded by actually running the generated Structured Query Language (SQL) and MongoDB queries against live databases. Over 1,157 evaluations spanning eight routers and three seeds, PaR stays on the observed cost-accuracy frontier. It matches a sink-frontier heuristic (frontier model on terminal nodes only) in accuracy at comparable cost and a faithful FrugalGPT cascade at lower cost, and cuts cost 44% against all-frontier routing while giving up 2.9 points of accuracy. Several accuracy gaps fall inside the plus-or-minus six-point confidence interval of a 54-task study, so we frame PaR's advantage as frontier position rather than a clean accuracy win. We also report a preliminary observation, not a validated result: a small pilot hints that cheap routing may carry a hidden compounding penalty on compositional workflows, which we frame as a hypothesis for future measurement. PaR, EntBench, and all evaluation code are open source.
|
| 694 |
Theory of Scene: Breaking the Symmetry Trap in Multi-Agent LLM Coordination
2609.32939
|
cs.CLcs.LG
|
Liangqi Yuan, Wenzhi Fang, Shiqiang Wang, Christopher G. Brinton |
Multi-agent systems built on large language models (LLMs) are largely homogeneous, as their agents behave alike even across distinct LLMs. We show that when such agents act concurrently without communication, they collide on targets they must split and diverge...Multi-agent systems built on large language models (LLMs) are largely homogeneous, as their agents behave alike even across distinct LLMs. We show that when such agents act concurrently without communication, they collide on targets they must split and diverge on targets they must take together, a double failure we term the symmetry trap. Theory of Mind (ToM), widely used for coordination without communication, cannot escape this trap, since homogeneous agents form the same prediction of one another and respond to it in the same way. We propose Theory of Scene (ToS), a training-free reasoning schema in which each agent reads its public role, the only difference between the agents, and the task context they all observe. Homogeneous agents thereby derive one division of labor, each taking the part its role fixes, which turns homogeneity from the cause of the trap into the cure. ToS reads the role together with the scene through role gating, which determines whether ownership overlaps or is already divided, and the task context through task coupling, which infers whether the team must converge on each target, divide it, or take its stages in turn. We evaluate on DivvyBench, a controlled environment we introduce, whose target types make an episode Competitive, Cooperative, or Mixed across Tabletop, Airspace, and Household scenarios, and on two established agentic benchmarks, GovSim and Overcooked. ToS outperforms all six baselines on every benchmark, and each baseline falls far behind it in at least one setting. Against ToM given the same inputs, ToS raises the DivvyBench success rate from 71.1% to 99.6%, the GovSim total gain from 207 to 400, and the Overcooked level-normalized throughput from 1.41 to 1.67.
|
| 695 |
DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis
2609.32957
|
cs.CL
|
Jiahui Li, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu |
Diagnosing a patient from video requires more than recognizing the sign: a vision-language model must turn what it sees into hypotheses, questions and tests. DynamicDx evaluates each step in 71 neurological consultations across 11 sign categories, linking auth...Diagnosing a patient from video requires more than recognizing the sign: a vision-language model must turn what it sees into hypotheses, questions and tests. DynamicDx evaluates each step in 71 neurological consultations across 11 sign categories, linking authentic patient videos to confirmed diagnoses and fixed charts built from the same case reports, so that every model queries the same evidence. Across five such models, video improves accuracy by 9.9-22.5 percentage points over blind input, but neither recognition alone nor temporal order explains the gain: the cause is usually missing from the model's video-only differential diagnosis even when the sign is recognized, and shuffling the frames produces no reliable accuracy loss. Instead, a trajectory replay traces most of the gain to the investigation results the video prompts. Evidence acquisition is the bottleneck: supplying the decisive investigations raises accuracy to 73.2-93.0%. Two interventions act on it. A post-trained 4B video describer improves sign descriptions, especially from a short, densely sampled segment, and source-clean literature retrieval expands initial hypotheses; both bring the tests a model orders closer to those the treating clinicians documented and, through them, raise accuracy. For video-based diagnosis, seeing better helps when it leads to asking better.
|
| 696 |
Trust and Task Completion in the World of Consumer AI Agents
2609.33017
|
cs.CL
|
Jeroen Olieslagers, Eduardo Pujol, Gal Zahavi, Lukas Ingemarsson, Shivani Poddar |
Action agents do things for people. They send email, spend money, and call businesses while the user is busy with something else, so a mistake can turn into an action before anyone notices. They fail their users in two ways. They break trust when they do somet...Action agents do things for people. They send email, spend money, and call businesses while the user is busy with something else, so a mistake can turn into an action before anyone notices. They fail their users in two ways. They break trust when they do something the user never agreed to, or hold back after the user clearly said go. And they fall short on completion when they give up on errands that turn out to be hard. Both depend heavily on the harness around the model, meaning its instructions, tools, context, and guardrails. We built an evaluation that scores trust and completion on the same runs, in a simulated world of businesses with their own websites, inboxes, and phone lines, and of people who write back. A simulated user answers the assistant's questions. Trust means that nothing happens the user did not agree to. No email goes to someone they never approved, no private detail ends up on a group thread, no money is spent past their limit, no stranger's instructions are followed, and nothing is claimed without a source. Every trap has a matched control in which acting is the right call. We use the evaluation to measure Fo, Wajo's personal assistant, against a base model with basic instructions on three foundation models, and against the Fo harness with its guardrails switched off. Fo completes 71% of the errands and keeps the user's trust on 94% of the trap runs. The base models complete 50% to 64% and keep trust on 59% to 75%. On the matched controls, Fo goes ahead slightly less often. OpenClaw, a popular open-source assistant given the same access, completes 42% of the errands it shares with Fo, against 71%, and keeps the user's trust on 74% of the shared trap runs, against 94%. Measuring trust and completion together, on the whole system rather than the model alone, is how we think action agents become safe to hand real work to.
|
| 697 |
Algorithmic Harms Associated with Generative Model-Augmented Recommendation Systems
2609.33073
|
cs.CLcs.LG
|
Christine Herlihy, Xumei Xi, Shloka Desai, Kevin Bannerman Hutchful, Pedro Silva |
In this work, we consider algorithmic harms that may arise as generative models are incorporated into machine learning platforms. We argue that existing harm taxonomies and threat models require extension to (1) address novel causal drivers of well-studied rep...In this work, we consider algorithmic harms that may arise as generative models are incorporated into machine learning platforms. We argue that existing harm taxonomies and threat models require extension to (1) address novel causal drivers of well-studied representational and quality-of-service harms; and (2) anticipate and mitigate endogenous harms, such as sanitization, which may arise when system inputs are misaligned with the system designer's objectives, or the generative model's inductive priors. To this end, we introduce an expanded taxonomy of algorithmic harms associated with the use of generative models in non-conversational recommendation systems. In addition, we offer a causal analysis of how problematic subsets of the (input, output) joint distribution can arise, in an effort to inform harms detection and mitigation efforts.
|
| 698 |
ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms
2609.33117
|
cs.CLcs.LG
|
Haitao Li, Chenglin Li, Zhengyao Ding, Ziyu Li, Yiheng Mao |
Multimodal large language models (MLLMs) can now interpret a standard ten-second, twelve-lead electrocardiogram (ECG) with clinically grounded, reward-verified reasoning. Real cardiac monitoring is different. Ambulatory (Holter) and telemetry recordings span h...Multimodal large language models (MLLMs) can now interpret a standard ten-second, twelve-lead electrocardiogram (ECG) with clinically grounded, reward-verified reasoning. Real cardiac monitoring is different. Ambulatory (Holter) and telemetry recordings span hours to days and are read as they stream in, and their clinically decisive findings are paroxysmal, brief episodes buried in an otherwise unremarkable trace. Such a recording cannot be held in one context at diagnostic resolution, and its future has not yet happened, so a reader must work online, deciding what to measure now, committing evidence to memory as it passes, and reporting events as they occur. We recast long-duration ECG interpretation as a long-horizon, online (streaming, causal) sequential decision process and introduce ECG-Scroll. As a benchmark, long ambulatory recordings are streamed to an agent chunk by chunk, and it must localize, quantify, and promptly flag paroxysmal events without access to future signal; because the underlying signal is retained, every answer is checkable against objective ground truth, giving rule-based rather than judge-based rewards, and the streaming formulation adds a metric batch evaluation cannot express, the detection latency between an event's onset and the moment the agent records it. As an agent environment, it is a fixed, gym-style interaction layer that exercises three competencies single-glance ECG models never touch: Memory, Tool use through signal-grounded measurement rather than reading pixels, and Planning of what to measure now and when to commit. We release 390 whole-recording instances spanning 2,536 hours of two-lead ambulatory ECG and evaluate a signal-threshold rule agent alongside off-the-shelf LLM agents online, characterizing how they use memory, tools, and planning and where the benchmark's head-room lies.
|
| 699 |
Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL
2609.33126
|
cs.CLcs.LG
|
Ziyuan Yang, Yike Wang, Shangbin Feng, Yulia Tsvetkov |
Group-relative reinforcement learning (RL) relies on reward variation among sampled responses to estimate informative relative advantages. As language models become increasingly capable, existing training data can become reward-saturated: all sampled responses...Group-relative reinforcement learning (RL) relies on reward variation among sampled responses to estimate informative relative advantages. As language models become increasingly capable, existing training data can become reward-saturated: all sampled responses to the same problem might receive equally high rewards, where the group-relative learning signals vanish and leave previously useful data obsolete. In this work, we investigate whether useful learning signals can be recovered from such saturated data. We study interventions at four levels of group-policy RL pipelines---data, rollout, reward, and advantage---and conduct extensive RL training on saturated reasoning data only. While standard GRPO on saturated data would almost always yield near-0 advantages and near-noise signals, diverse interventions successfully recycle and repurpose such data: among the proposed strategies, interventions at rollout generation are consistently most effective: nudging the policy to generate ``high-quality'', incorrect solutions introduces rollouts with poor rewards into saturated groups as negative samples, which turns out to improve GRPO by 6.4% to 9.0% across Qwen3-1.7B and 4B. Other interventions such as increasing rollout temperature or adding auxiliary rewards can also restore non-zero advantages, but yield less consistent gains. Further analyses show that effective negative rollouts require informative negative trajectories, that the method remains effective alongside unsaturated data, and that it supports iterative recycling of newly saturated examples. While increasingly stronger LLMs would render more data as saturated, our results demonstrate that don't waste your saturated data: with the right strategies they can be recycled into useful RL training signals in an increasingly data-scarce world.
|
| 700 |
SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
2609.33181
|
cs.CL
|
Xiaoshu Chen, Xiangyu Wong, Sihang Zhou, Ke Liang, Xinwang Liu |
Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. ...Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.
|
| 701 |
AG-CoT: Verified Algorithmic Traces for LLM Program Synthesis on Clifford Circuits
2609.33192
|
cs.CLcs.LG
|
Lu Wei, Yufeng Wang, Chenfeng Cao, Lu Pang, Haibin Ling |
Scientific code generation can produce executable programs that fail to compute the intended scientific object. We study this problem in language-model synthesis of Clifford circuits, which prepare the stabilizer states used in quantum error correction and adm...Scientific code generation can produce executable programs that fail to compute the intended scientific object. We study this problem in language-model synthesis of Clifford circuits, which prepare the stabilizer states used in quantum error correction and admit exact classical verification. In our target-conditioned framework, each target is given as compact signed stabilizer generators, and an exact verifier checks the generated OpenQASM circuits. We supervise models with Aaronson-Gottesman chain-of-thought (AG-CoT) traces checked by the verifier, and continue training on model generations that the verifier accepts. Across two independently trained model families (3B and 7B), AG-CoT supervision multiplies greedy-decode state-equivalence accuracy by four to six times over circuit-only baselines, and verifier-filtered continuation training adds a further consistent gain atop both. A complementary 32B study shows that supervised models achieve near-perfect syntax and Clifford validity while the strongest direct model reaches 6.14% state equivalence per target, rising to over 10% under verifier-guided selection with multiple candidates. These results show that algorithmic trace supervision gives a large, statistically significant gain in both model families and that verifier-filtered continuation adds a further repeated gain. The persistent gap between Clifford validity and state equivalence confirms that exact verification is necessary: a circuit can be syntactically and physically valid yet prepare the wrong quantum state.
|
| 702 |
Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
2609.33200
|
cs.CLcs.LG
|
Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin |
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferri...On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
|
| 703 |
When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning
2609.33220
|
cs.CLcs.LG
|
Steven Y. Feng, Noah D. Goodman, Michael C. Frank, Evan Hubinger, Paul C. Bogdan |
Outcome-based reinforcement learning can produce models with similar task performance but very different ways of communicating about their mistakes. We study failure disclosure: whether a model admits that an attempted solution failed rather than staying silen...Outcome-based reinforcement learning can produce models with similar task performance but very different ways of communicating about their mistakes. We study failure disclosure: whether a model admits that an attempted solution failed rather than staying silent or presenting it as successful. Across repeated outcome-only GRPO training runs, failure disclosure varies far more than task accuracy. The pattern extends to a second reasoning task and stabilized PPO, persists at 7B, and also appears in an instruction-conditioned 32B setting. We also find that small floating-point and sampling differences during training can redirect reporting behavior even when the task objective and earlier training history are held fixed. Additional tests show that failure disclosure is not a single decision: Checking the answer, entering a report, and completing the admission can separate, and the weak point depends on the task and response format. Further, experiments with neutral controls show more broadly that behaviors left weakly constrained by training are especially likely to vary across runs, of which failure disclosure is an example. We can reduce variability in failure disclosure by discouraging the model from drifting from its starting policy on failed, well-formed responses. This makes reporting substantially more consistent, though its effect on task performance depends on the setting. Stable task accuracy therefore does not guarantee stable safety-relevant behavior: Researchers should measure these behaviors directly across runs and design training methods that keep them reliable.
|
| 704 |
Acoustic Progress Propagation for Long-Horizon Speculative Decoding in ASR
2609.33245
|
cs.CL
|
Yuanyuan Jia, Qianqian Yang |
Speculative decoding accelerates autoregressive automatic speech recognition (ASR), but the acceptance length of alignment-aware drafters can saturate as the draft horizon increases. We propose a progress-aware speculative drafter that recurrently propagates a...Speculative decoding accelerates autoregressive automatic speech recognition (ASR), but the acceptance length of alignment-aware drafters can saturate as the draft horizon increases. We propose a progress-aware speculative drafter that recurrently propagates an acoustic progress state across draft steps and feeds it back into audio cross-attention to guide token generation. We jointly train the drafter and progress predictor over variable draft horizons. On five ASR test sets, our method achieves lossless, macro-averaged end-to-end speedups of 1.657x and 1.227x over target-only autoregressive decoding with Qwen3-ASR-0.6B and Qwen3-ASR-1.7B, respectively. Relative to AnchorDraft, our method improves the macro-averaged speedup by 34.3% and 9.0%, respectively. Horizon sweeps show continued growth in acceptance length beyond the baselines' saturation. Code is available at https://github.com/yuanyuanjia71-spec/ProgDraft.
|
| 705 |
BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions
2609.33254
|
cs.CLcs.LG
|
Thanina Hamitouch, Khadidja Henni, Abdelkrim Arie, Amina Selma Haichour, Neila Mezghani |
Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitat...Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitations: labelled interactions are scarce and unevenly distributed, large pretrained chemical and protein encoders are expensive to fine-tune end-to-end, and independently encoded sequences do not capture pair-specific dependencies. We present BERT4DTI, which encodes SMILES strings with ChemBERTa and amino-acid sequences with ProtBERT, applies bidirectional mutual attention between token-level representations, and classifies the resulting interaction features using convolutional layers and a multilayer perceptron. To reduce trainable size, ProtBERT is truncated to 18 retained layers and only the last two layers of each encoder are fine-tuned. On BIOSNAP, DAVIS and BindingDB, BERT4DTI is competitive, achieving the best ROC-AUC and PR-AUC on BIOSNAP and the highest sensitivity on all three benchmarks. An ablation on DAVIS shows that mutual attention improves PR-AUC and specificity. With 125M trainable parameters compared with 353M for full BERT fine-tuning, BERT4DTI provides a favourable performance-parameter trade-off for sequence-based DTI screening, while leaving runtime profiling, calibration and leakage-audited validation for future work.
|
| 706 |
InfoEdit: Probing Global Layout Reasoning in Infographic Editing
2609.33286
|
cs.CL
|
Cheng Yang, Chufan Shi, Huijuan Wang, Bo Shui, Yaokang Wu |
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element of...Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to this global layout reasoning capability as reflow. Existing image-editing benchmarks neither provide a dedicated setting for structured visual content nor evaluate the reflow capability. We introduce InfoEdit, a novel benchmark of 1,000 infographics across eight logical-relation families, paired with 4,000 editing instructions across four editing tasks, and a reflow-aware evaluation protocol. Across eight frontier editors, only GPT-Image-2 clears 60% average success rate; most models fall below 7%, and no editor exceeds 36% on the Swap-Block task even with perfect target localization. We further show that code-level editing can match the strongest pixel-level editor, revealing complementary strengths across tasks. InfoEdit identifies reflow as a central challenge in structured visual content editing and provides a diagnostic benchmark to facilitate future progress.
|
| 707 |
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
2609.33295
|
cs.CL
|
Dehai Min, Daoan Zhang, Yiming Zeng, Huayi Zhang, Ziyi Chen |
An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benc...An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.
|
| 708 |
Robust Hierarchical Structures for Agentic Document Analysis
2609.33322
|
cs.CL
|
Ruiying Ma, Yiming Lin, Aditya G. Parameswaran |
Large Language Models (LLMs) enable us to better understand text documents, including PDFs and Word documents. However, LLMs, as well as more modern LLM agents, i.e., those with tool-calling abilities, typically treat such documents as plain text, ignoring the...Large Language Models (LLMs) enable us to better understand text documents, including PDFs and Word documents. However, LLMs, as well as more modern LLM agents, i.e., those with tool-calling abilities, typically treat such documents as plain text, ignoring the fact that they are often organized hierarchically into sections and subsections. Extracting this structure, while difficult, can improve efficiency and effectiveness for agents (and humans)---since only sections relevant to a given task need to be processed. Unfortunately, prior work on structure extraction provides no formal guarantees on how well the inferred structure matches the true one. Instead, we target a robust and compact variant that is feasible to infer and useful in practice. Robustness ensures that the text under each subsection header is a superset of the text under the same header in the true structure. Compactness seeks to minimize this superset, reducing agentic cost (or human cognitive load). We propose SHED, a two-stage workflow for inferring a robust and compact structure. The first stage is pluggable with an infinite family of approaches, each guaranteeing robustness for a specific document class. We theoretically characterize the document space using these classes and their hierarchical relationships. Empirically, SHED improves F-1 scores (measuring the robustness--compactness trade-off) by 13%--68% over non-LLM baselines and 9%--15% over expensive LLM-based approaches. Finally, we show how SHED-inferred structures are valuable for agentic document analysis: agents using SHED outperform baselines, achieving 3%--23% higher accuracy while being up to 10x cheaper.
|
| 709 |
When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression
2609.33334
|
cs.CLcs.LG
|
Haeyong Kang, Chang D. Yoo |
Training-free KV-cache compression methods such as SnapKV, H2O, and PyramidKV evict tokens at the end of prefill, aiming to preserve the attention mass that future queries are expected to use -optimizing what to keep. We show that this objective fails in two d...Training-free KV-cache compression methods such as SnapKV, H2O, and PyramidKV evict tokens at the end of prefill, aiming to preserve the attention mass that future queries are expected to use -optimizing what to keep. We show that this objective fails in two distinct ways. (1) Compensation: restoring the evicted attention mass can recover the attention-level target without recovering task quality. (2) Selection: covering more of the true decode-query mass can hurt quality when the recovered mass is fragmented rather than concentrated in coherent spans. These failures share a common cause: eviction occurs before the queries that determine the answer trajectory exist. We propose Draft-Guided Eviction (DGE), which defers eviction until after drafting the first k=2 answer tokens using the full cache - just one decode step beyond prefill. Because the draft is generated from the answer's own prefix, no cache entries are discarded before this trajectory signal becomes available. The per-head cache budget remains unchanged, and DGE can be applied directly to SnapKV, PyramidKV, H2O, and StreamingLLM without modifying their eviction scores. Unlike extra-pass methods, DGE changes when eviction occurs rather than what cache entries are selected. Extensive experiments demonstrate that DGE outperforms prior methods at every evaluated budget on five of six instruct-tuned backbones, achieving 44.2 on LongBench, nearly matching FullKV at 44.3. The timing-only control DGE-W achieves the same score, demonstrating that the gain comes from when eviction occurs rather than what is selected - an effect we term trajectory anchoring.
|
| 710 |
CalibHyper: Chance-Corrected Relational Hypergraphs for Few-Shot Molecular Property Prediction
2609.33342
|
cs.CLcs.LG
|
Linyu Li, Zhi Jin, Yuanpeng He, Dongming Jin, Huanyu Liu |
Molecular property prediction is central to drug development and materials discovery, but experiments are costly and labeled data are scarce. Context-aware methods use auxiliary assay labels to support few-shot prediction, and recent work supervises property r...Molecular property prediction is central to drug development and materials discovery, but experiments are costly and labeled data are scarce. Context-aware methods use auxiliary assay labels to support few-shot prediction, and recent work supervises property relations with label agreement. However, label agreement is sensitive to class marginals and does not directly capture dependence between properties. We propose CalibHyper, a chance-corrected relational hypergraph method based on the joint label distribution. CalibHyper subtracts an independence baseline from the ordered four-state label distribution and shrinks the residual according to the number of joint observations. A swap-equivariant relation head estimates these residuals, which choose the auxiliary properties for each molecule and set the sign and weight of their hyperedge messages. On thirteen datasets from five benchmarks, in both 1-shot and 10-shot settings, CalibHyper and its ablation settings achieve ROC-AUC competitive with the strongest reported results.
|
| 711 |
From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS
2609.33362
|
cs.CL
|
Kangxiang Xia, Xinfa Zhu, HangRui Hu, Kexin Huang, Wenjie Tian |
Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to sat...Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS, unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration. The framework further performs hierarchical utterance-level and scene-level critique, routing detected issues to editing, resynthesis, or timing adjustment. Experiments on a bilingual Chinese--English dialogue benchmark show improved utterance-level instruction following, better dialogue-level preference than direct dialogue models and agentic baselines, and more effective refinement than regeneration-only alternatives while preserving speaker identity. Ablations further confirm the benefits of scene-level critique and edit-based correction.
|
| 712 |
OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading
2609.33385
|
cs.CL
|
Jingyuan Xiao (Tianjin University, Tianjin, China), Jiayue Wang (Tianjin University, Tianjin |
Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained...Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a natural remedy, yet existing MoE serving systems target autoregressive decoding and rely on intra-iteration layer-wise prefetching: while computing one layer, they predict and load experts for subsequent layers. Under dLLM inference, block-wise routing expands the active expert working set within each iteration, making such prefetches difficult to complete in time and costly when mispredicted. Consequently, existing prefetch-based solutions often degenerate into on-demand expert loading with high decoding latency. We propose OLED-MoE, an expert offloading system that shifts the optimization target from intra-iteration prefetching to inter-iteration expert retention. Its key insight is that adjacent denoising iterations exhibit strong expert routing overlap, and token confidence indicates which experts are likely to be reused. OLED-MoE uses confidence-guided inter-iteration prediction to retain high-value experts in GPU memory without introducing extra prefetch traffic. It further compensates unavoidable cache misses through CPU-GPU cooperative execution, jointly considering dynamic expert computation load and predicted future reuse. Across diverse dLLM workloads, OLED-MoE reduces time per output token (TPOT) by 1.23x-7.93x and improves expert cache utilization by 1.44x-4.23x over state-of-the-art offloading systems. Notably, OLED-MoE approaches full-residency performance while using only 40% of the expert GPU memory, incurring merely 23% higher TPOT despite a 60% reduction in expert memory footprint. OLED-MoE's source code is publicly available at https://github.com/flashserve/OLED-MoE.
|
| 713 |
Decoupling Token Roles in Autoregressive Pretraining
2609.33405
|
cs.CLcs.LG
|
Suqin Yuan, Runqi Lin, Kevin Qinghong Lin, Junchi Yu, Lei Feng |
Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each...Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for what follows. Using controlled corruption, we decouple these two roles and find a reversal: making a noisy token easier to predict reduces its damage as a target but increases it as context. The same decoupling helps explain text generated by language models: generation selects each token by its fit to the prefix, while its role as context is never tested against an independently determined continuation, because that continuation is generated to fit it. At known corrupted positions, acting through the context can reduce damage that removing the token's own loss does not. Understanding and controlling what a model learns from a token therefore requires decoupling its roles.
|
| 714 |
MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses
2609.33411
|
cs.CL
|
Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu, Yinghao Ma, Jyh-Shing Roger Jang |
Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid...Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution of isolated tasks, we propose to optimize the benchmark generation workflow itself end to end with MetaBench-Harness, a dual-loop search framework. Specifically, the inner loop utilizes a benchmark harness to generate a new benchmark in each round, while the outer meta-harness orchestration layer iteratively refines and searches over harness implementations based on historical evolution trajectories. By applying MetaBench-Harness to the competitive programming CodeContests and Olympiad mathematics AIME-2024 datasets, we demonstrate that the evolved benchmarks are challenging and discriminative for frontier models. Trajectory and quality analyses verify that MetaBench-Harness enables multi-dimensional evolution, steadily improving evolution reasonableness, benchmark competency, and evaluator robustness across successive rounds. Furthermore, case studies reveal its effective utilization of diverse difficulty levers to reframe problems and elevate required capabilities. Ultimately, this work provides a solution to the pressing challenge of benchmark saturation.
|
| 715 |
CORA: A Protocol for Diagnosing Boundary Robustness in Text-to-Audio Retrieval under Query Reformulations
2609.33433
|
cs.CL
|
Jae Min Woo, Kyongmin Kong, Bogyung Jeong, Minjeong Kim, HaeJun Yoo |
Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source...Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source caption into five intent preserving forms (Command, Question, Indirect, Key phrase, and Statement) while fixing the target audio. By tracking the same target across query forms, CORA defines RankDrop, a metric revealing failures hidden by Recall@k. Using Pearson's correlation coefficient r, we find that RankDrop is weakly associated with raw text space movement (r=0.084), but strongly associated with Target Alignment Loss and Target Boundary Margin Degradation (r=0.508 and r=0.615). The same pattern appears in OEA retrievers, where RankDrop is better explained by boundary degradation (r=0.472/0.478) than by query movement (r=0.084/0.046). Overall, these results suggest that robust T2A retrieval requires preserving the target's boundary advantage over competing audio under reformulation.
|
| 716 |
SMAT: Simple and Efficient Merge-Aware Training
2609.33437
|
cs.CLcs.LG
|
Yanggan Gu, Yuanyi Wang, Zhen Li, Shuo Cai, Yuhang Liu |
Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance,...Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.
|
| 717 |
Raven: The Harness of Harnesses for Composable Agentic Intelligence
2609.33439
|
cs.CL
|
EverMind AI |
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tight...As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, \emph{The Harness of Harnesses}, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an \emph{All-Domain Collaboration Network}, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.
|
| 718 |
A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards
2609.33467
|
cs.CLcs.LG
|
Andreas Plesner, Curtis Northcutt, Francisco Guzm\'an, Anish Athalye |
When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. ...When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifier agreement predicts post-training performance on such tasks. In this paper, we explore this question with over 11k H100 GPU-hours, across HealthBench and PRBench tasks in medical, legal, and finance domains. Across the tested domains, Qwen3 trainees (1.7B-8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes. We compare two low-cost choices retrospectively -- a cost-reducing choice and a balanced choice -- with estimated grading cost reductions of 98.8%-99.7% relative to the golden grading protocols and average post-training score gaps of 1-3 points from the best evaluated training verifier. These averages include larger losses in individual settings; they do not establish that verifier choices are interchangeable.
|
| 719 |
TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
2609.33589
|
cs.CLcs.LG
|
Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai, Wei Lin |
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand ...Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://github.com/1229095296/TGRL/tree/main.
|
| 720 |
ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments
2609.33618
|
cs.CL
|
Shengbin Yue, Hongru Wang, Siyuan Wang, Xiaoxin Chen, Wei Chen |
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration...Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We argue that the key lies not in whether to decouple or interleave them, but in how to coordinate them across granularities. We introduce ParaAct, a structured parallel-action loop that combines phase-level Exploration $\rightleftharpoons$ Execution with action-level parallelism. To learn this loop, ParaAgent combines multi-agent cold-start demonstrations with reinforcement learning under multi-level advantage decoupling, making planning structure explicit and supervising it with step-, phase-, and trajectory-level rewards. Learning is supported by our ToolEnv, a scalable simulator grounded in 50,011 realistic tool interfaces. On two open-world tool benchmarks, ParaAgent-4B achieves the best average success among all baselines, including GPT-4.1 systems, with the largest gains on multi-tool tasks. Behavioral analyses show that these gains stem from this action organization, highlighting its importance for capable and efficient open-world agents.
|
| 721 |
Quantifying Behavioral Tails in Black-Box Language Models
2609.33638
|
cs.CLcs.LG
|
Elsayed Eshra, Ali Al-Lawati, Dongwon Lee, Suhang Wang |
We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the input space. To accomplish that, RareTrap ...We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the input space. To accomplish that, RareTrap uses a surrogate LLM and constructs a geometry-aware mapping from a lower-dimensional latent reference space into its token-embedding space to induce an explicit and reproducible distribution over input prompts. A response-level performance function is utilized on the response to quantify behavior severity. This enables sequential rare event simulation that concentrates evaluations on progressively more severe behaviors while preserving probability under the induced prompt distribution, which would otherwise be prohibitive to measure. Across 10 open-weight and two frontier models (GPT-5.4 and Claude Sonnet 4.6), we find that RareTrap successfully induces severe resource consumption behaviors and computes their probability with as few as 200 evaluations. RareTrap provides model developers a principled approach for evaluating language models under a common distribution, and prioritizing alignment effort to improve safety and mitigate risks.
|
| 722 |
Probe to Act: Elevating Browser-Use Agent via Active Visual Probing
2609.33646
|
cs.CL
|
Keliang Li, Heng Wang, Chen Hu, Daxin Jiang, Hong Chang |
Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks...Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alignment before every operation. We introduce Probe to Act (P2A), an active probing framework for the browser-agent loop that moves this alignment into decision time. P2A addresses an asymmetric bridge between symbolic DOM hypotheses and screenshot layout by rendering on-demand symbolic DOM structure back into pixels. Before committing a state-changing browser operation, the agent can issue lightweight probes to translate DOM handles into pixel evidence, map screen regions back to DOM candidates, register visual-only targets, and commit verified notes. These interleaved processes naturally produce evidence-based memory: only probed, acted-on, or explicitly committed observations are kept across steps, preserving only decision-critical evidence in long-horizon contexts. P2A can be used as a prompting strategy for proprietary models under the standard DOM+SoM interface, and can be distilled into open-weight models through cold-start synthesis and self-bootstrapped SFT. Across three browser-use benchmarks, P2A shows clear gains on task success rate for both proprietary and fine-tuned models; on VisualWebArena, for example, it improves Gemini-3-Pro from 54.1% to 61.2% and Qwen3-VL-8B from 24.6% to 32.9%, while matching the costly full-observation history ($\sim$3$\times$) at only $\sim$1.2$\times$ the peak retained input context of action-only history.
|
| 723 |
Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification
2609.33662
|
cs.CLcs.LG
|
Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang |
Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-app...Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite secondary-verification budget; its action trace is frozen before clean labels are joined. Simultaneous finite-sample bounds then certify selected harmful risk, coverage, and verifier-call rate over a predeclared policy family. Conditional Hoeffding-Azuma bounds account for the dependence induced by shared budgets, and rollout or verifier changes initiate a new certification stage. After admission, a bounded trust-clip-KL actuator controls magnitude. We evaluate two models on two reasoning benchmarks against static RLVR, matched-random selection, confidence thresholding, noise correction, and verifier augmentation. On Qwen3.5-0.8B and GSM8K at target risk $\rho=0.08$, RC-VAPO achieves 74.1% accuracy, selected harmful risk 0.0697, coverage 0.4125, and relative verifier cost $1.16\times$. At matched coverage and update magnitude, its selected-risk difference from matched random is -0.0260 with paired 95% interval $[-0.0364,-0.0157]$. Across asymmetric, confidence-dependent, and correlated-verifier noise, the certificate is satisfied on 57 of 60 independent runs. These comparisons isolate informative directional selection from proposal suppression, update shrinkage, and additional verifier computation.
|
| 724 |
Auditing Agent Actions through Query-Conditioned Attribution
2609.33676
|
cs.CL
|
Yifan Liu, Praveen Venkateswaran, Abdulhamid Adebayo, Dong Wang |
LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do ...LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for diverse auditing objectives. Additionally, when access to the acting model is limited (e.g., in API-only deployments), applicable methods commonly rely on costly input perturbations or external LLM analysis of complete trajectories. We therefore formulate $\textit{query-conditioned agent action attribution}, a new task that takes a natural-language auditing query as input and recovers the source and ordered intermediate evidence for the query-specified aspect of an action. We instantiate this task with $A^3Bench$, a benchmark comprising 1,396 auditing queries across policy basis, parameter provenance, failure propagation, and unsafe-behavior tracing. To enable efficient, query-specific attribution, we use small open-weight models as attribution proposers that combine query-conditioned gradient saliency with query-semantic relevance to rank history units. Our proposer consistently achieves stronger source and evidence rankings at lower inference cost than open-weight baselines, improving source MRR by up to 40.9\% and evidence MAP by 42.1\% with only two forward passes and one backward pass. Controlled evaluations confirm that our proposer improves attribution specificity by adapting its rankings to fine-grained changes in the auditing query. Building on a proposer ensemble, our end-to-end system surpasses the strongest frontier-model baseline in source accuracy (64.5\% vs.\ 60.4\%) while reducing empirical deployment latency by 29.9\% relative to the fastest frontier API baseline. Code and data will be released after the initial review period following final validation and cleanup.
|
| 725 |
BOReFT: Manifold Steering of Language Models for Black-box Optimization
2609.33722
|
cs.CLcs.LG
|
Dhruv Agarwal, Rico Angell, Kavitha Srinivas, Tahira Naseem, Horst Samulowitz |
Language models are increasingly used as proposal models for black-box search, from program optimization to molecular design. Existing approaches typically improve proposals through iterative prompting or parameter updates, offering limited control over how co...Language models are increasingly used as proposal models for black-box search, from program optimization to molecular design. Existing approaches typically improve proposals through iterative prompting or parameter updates, offering limited control over how completely and efficiently the model's search space is explored. Continuous optimization methods, such as Bayesian optimization, provide a principled way to search but require a suitable domain to operate over. To address this, we introduce BOReFT, which learns a compact, low-dimensional space of hidden-state interventions in a frozen language model, and uses this space as the search domain for Bayesian optimization with an external scoring function. Empirically, we find that the learned domain spans semantic regions and exhibits smoothness properties that support search. Theoretically, we show that semantic coverage and interpolation control the best score available in the learned space, and that decoding from this space yields a standard stochastic-bandit observation model for adaptive search. We evaluate BOReFT on the interpretable word search task "Semantle" and on three more real-world discovery tasks in de novo molecule property optimization. Compared to strong LLM baselines, BOReFT finds in Semantle a higher number of hidden targets and, on two out of three molecular objectives, achieves higher property scores. Consequently, our method provides a principled new bridge between discrete proposal spaces of LLM-based search and continuous black-box optimization.
|
| 726 |
DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation
2609.33742
|
cs.CL
|
Yayue Deng, Dingdong Wang, Yuxuan Hu, Jinyu Li, Yanqing Liu |
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems large...Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.
|
| 727 |
Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech
2609.33774
|
cs.CLcs.LG
|
Kanghwi Lee, Kyeongseok Jeong, Jeongmin Liu |
Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates direct...Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation. Its main weakness is retokenization: decoding generated speech to a waveform and encoding it again can change token identities and erode the watermark. To make the watermark robust to these changes, we propose Redwing, REtokenization-Durable Watermarking IN Generation. It builds a graph from the token substitutions observed under retokenization, whose Laplacian yields a basis that assigns similar values to tokens likely to substitute for one another. Over this basis, embedding and detection functions are jointly optimized to preserve watermark signal through retokenization while limiting embedding distortion and detector variability on unwatermarked speech. On the Moshi full-duplex system, after eight consecutive passes of Mimi resynthesis, Redwing achieves 80.7% TPR at a calibrated 1% FPR, compared with 8.3% for KGW and at most 7.3% for WMAR. It also has the highest TPR after eight passes through three other neural codecs (77.5-93.0%), and the gains generalize to TTS models at a speech-quality cost close to that of KGW. These results show that retokenization is not merely a source of noise: its transition structure can be exploited as a design principle for robust token-level watermarking.
|
| 728 |
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
2609.33781
|
cs.CLcs.LG
|
Woongyeong Yeo, Minki Kang, Chanuk Lee, Sangwoo Park, Jinheon Baek |
Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged...Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily available signal, prioritizing uncertain positions under both reinforcement and penalization concentrates penalties where failed responses still retain alternatives for recovery, which can suppress opportunities for exploration. To address this, we introduce Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. Specifically, motivated by the observation that success under uncertainty is less repeatable while confident failures tend to recur, EAPO couples normalized policy entropy with the sign of the response advantage to reinforce surprising success and correct repeated failure. It assigns stronger reinforcement to high-entropy decisions in successful responses and stronger penalties to low-entropy decisions in failed responses, while attenuating penalties at uncertain positions to preserve opportunities for recovery. By redistributing the response advantage across tokens, EAPO derives token-level credit directly from existing rollout signals without additional supervision. We validate EAPO on a range of reasoning tasks across both base and reasoning backbones, demonstrating that it achieves the best overall performance. We further show that EAPO promotes more effective exploration, broadening problem coverage and generating more diverse candidate answers.
|
| 729 |
Controlling Speaking Rate in Autoregressive TTS via Activation Steering
2609.33810
|
cs.CLcs.LG
|
Francesco Verdini, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet |
Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation a...Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-block analysis recovers the rate axis, a neutral operating point, and a per-step intensity scale; at inference, the activation's projection onto this axis is set to a fixed scalar. Learning this direction from synthetically time-stretched and time-compressed speech yields rate control that largely preserves speaker identity, generalizes across model architectures, and maintains high naturalness in objective and human evaluations. Unlike standard additive steering, which breaks at the slow extreme, clamping remains stable on all three systems tested; at moderate targets, the better rule depends on the model. Finally, we show that rate information is decodable across layers but causally steerable only within a mid-depth window, and demonstrate the effectiveness of our approach on the public Seed-TTS-Eval benchmark.
|
| 730 |
ChemOPD: Multi-Teacher On-Policy Distillation for Multi-Task Chemical Reasoning
2609.33838
|
cs.CLcs.LG
|
Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Xuemin Chen, Tianshu Yu |
Large language models are increasingly expected to support diverse chemical reasoning capabilities within a unified model. One approach is to develop specialized capabilities separately and consolidate them through multi-teacher on-policy distillation, but thi...Large language models are increasingly expected to support diverse chemical reasoning capabilities within a unified model. One approach is to develop specialized capabilities separately and consolidate them through multi-teacher on-policy distillation, but this raises two questions: how should specialization be organized, and how should specialist guidance be integrated? We introduce ChemOPD, which addresses both. We estimate task affinities from supervised fine-tuning gradients and solve a constrained mixed-integer program(MIP) to construct partially overlapping specialist groups. During distillation, we retain a generalist teacher trained on all tasks so that specialist guidance supplements rather than replaces its supervision. Our anchor-residual objective gradually increases the routed specialist's contribution on student-generated responses. On ChemCoTBench, affinity-guided specialization produces task-dependent gains over the generalist teacher and improves several capabilities beyond semantic task grouping. Yet stronger teacher-side performance does not automatically yield stronger students: with the same specialists and routes, anchor-residual OPD improves most reported metrics over specialist-only distillation and realizes a larger share of the available teacher gains. These results highlight specialization and capability integration as connected but distinct design problems in chemical reasoning.
|
| 731 |
Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation
2609.33843
|
cs.CLcs.LG
|
Gowthamkumar Nandakishore |
The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is $-0.214$, every occupied reliability...The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is $-0.214$, every occupied reliability bin's accuracy exceeds its confidence, and that sign uniformity collapses every binned ECE variant to the same value, $0.214$. The card frames the risk as over-confidence; the measured direction is the opposite, and the direction decides which way a confidence-gated cascade fails. A single disjointly fitted temperature ($T=0.469$, sharpening) removes most of the miscalibration (held-out ECE $0.204$ to $0.037$) and outperforms the shipped per-option-count table. The frozen selection rule instead chose isotonic regression, which overfit and failed its held-out NLL contrast on both tracks, so hypothesis H2 is not supported. Re-running the released checkpoint on its full official test split reproduces the card's headline accuracy ($0.767$ vs. $0.766$). The retrospective E1 reproduction preceded the analysis freeze; E2-E8 were prospectively preregistered, and 20 of 22 executed confirmatory tests reject under Benjamini-Hochberg FDR at $q=0.05$ (two descoped). The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy $0.617$), and every score measures agreement with a synthetic teacher whose self-agreement ceiling ($0.735$) the specialist exceeds. Per-decision predictions, run manifests, and the frozen preregistration are in the ancillary files. The author has no affiliation with the model's publisher, the dataset's publisher, or TypeSafe.
|
| 732 |
Rethinking Contextualization by Reinterpreting Attention Head Channels
2609.33851
|
cs.CLcs.LG
|
Hakaze Cho, Haolin Yang, Zhun Sun, Naoya Inoue, Benjamin Heinzerling |
Contextualization, the core operation of language modeling, transmits information across words to build sentence-specific word representations. Prior works mainly study contextualization, focusing on individual words and attention heads as a growing discrete d...Contextualization, the core operation of language modeling, transmits information across words to build sentence-specific word representations. Prior works mainly study contextualization, focusing on individual words and attention heads as a growing discrete dictionary, lacking a global view of their general behavior. Therefore, we propose a general principle: Globally, we find and estimate that different words carry different amounts of information, and less-informative words tend to absorb more contextual information. Specifically, these low-information words do not absorb contextual words uniformly, and finer-grained selectivity enables more precise routing to promote information transmission between matched words. Moreover, to find what mechanism causes such processing, we reinterpret attention heads as channels gated by their singular vectors and find that: (1) these singular vectors point to the hidden states of more informative words, allowing such words to write their information to others more strongly to act as information sources, and vice versa; and (2) these singular vectors can be viewed equally as hidden state features, enabling automated interpretation of attention heads beyond prior heuristic head discovery, also embedding heads into a continuous space rather than treating them as discrete, independent dictionary entries.
|
| 733 |
Program-Verified Self-Evolution for Vision-Language Models
2609.33855
|
cs.CLcs.LG
|
Ahmed Heakl, Sungik Choi, Moontae Lee, Salman Khan |
Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\...Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94\% of VQS answers correct, against 76\% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS
|
| 734 |
Population Physics, Population Problems: Safety and Emergence in LLM Societies
2609.33871
|
cs.CL
|
Adrian de Wynter |
The collective behaviour of large language model (LLM) societies is not the sum of their individual outputs. It yields statistically distinct, sometimes-unpredictable phenomena, for which the tools we use to study single agents may not scale. Due to recent inc...The collective behaviour of large language model (LLM) societies is not the sum of their individual outputs. It yields statistically distinct, sometimes-unpredictable phenomena, for which the tools we use to study single agents may not scale. Due to recent incidents involving autonomous agentic systems, however, understanding these systems is paramount. For that we introduce a framework for measuring self-organisation in LLM social systems and apply it to three such systems: a Schelling grid, a social network (Moltbook), and a Twitter-like misinformation simulation ('Rogue'). All three exhibit statistically significant self-organisation. Moreover, their relaxation dynamics vary with the environmental information available to the agents, with open-ended systems (Moltbook, Rogue) exhibiting sharp, phase-transition-like dynamics. Further results show that population-level pathologies can emerge even when the LLMs are safety-tuned or monitored, being primarily driven by the coordinated activity of a population subset. We also show when self-organisation does \textit{not} emerge under two additional scenarios (a commons dilemma, GovSim, and a LLM-as-a-judge deliberation scheme, ChatEval). We argue that measuring signatures of this kind offers a lightweight, agent-agnostic diagnostic layer for detecting coordinated collective behaviour in deployed multi-agent systems without relying on natural language or model versioning.
|
| 735 |
Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding
2609.33889
|
cs.CLcs.LG
|
Jungseob Lee, Seungyoon Lee, Seongtae Hong, Sugyeong Eo, Heuiseok Lim |
At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context. Activation sparsity trims the first term and KV-cache sparsity the second, y...At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context. Activation sparsity trims the first term and KV-cache sparsity the second, yet their reported speedups are hard to compare because each depends on context length and on the dense attention kernel it is measured against. We derive a byte crossover, the context length at which the two savings are equal, together with ideal speedup bounds for each branch and for their composition, from model dimensions and keep ratios alone. We then time both branches and their composition from 2K to 128K tokens on two GPUs after a dense prefill of real text, with dense and sparse modes reading the cache through the same split-K attention kernel. The projection branch leads at short context and the KV branch at long context, with speedups that follow their byte bounds up to fixed kernel costs. Adding these costs, measured in separate sweeps, lets the byte account predict the measured crossings of three keep-ratio pairs, a second model, and a second GPU to within 4.1K tokens. Timing the dense baseline with masked instead of split-K attention inflates the apparent speedup of the same KV policy about fivefold. An attention-scored KV selection answers the same passkey and multi-key placements as dense decoding up to 127K tokens, whereas a KV window misses most of them. Under matched perplexity budgets, activation sparsity composed with this selection decodes 14 to 26% faster than the best single branch on both GPUs. Code is available at https://github.com/js-lee-AI/ByteCross.
|
| 736 |
Steering Language Model Goals with Value Transplant
2609.34056
|
cs.CLcs.LG
|
Pengcheng Jiang, Fabien Roger |
Reasoning models often act as if they pursue goals, but their efforts are not always directed toward what users intend, sometimes leading them to pursue unintended outcomes. Previous work has examined how models may internally track their progress toward their...Reasoning models often act as if they pursue goals, but their efforts are not always directed toward what users intend, sometimes leading them to pursue unintended outcomes. Previous work has examined how models may internally track their progress toward their goals through a "value axis." We study whether changing such a signal can retarget the model's search toward a different goal. We test value transplant: at each token, we shift the host model's activation along a candidate value axis by the donor-host difference in value coordinates (multiplied by a large scalar), aiming to redirect the host toward the donor's goal. We study this intervention in Qwen3-8B and GPT-OSS-20B models fine-tuned into honest and cheating variants. We test several candidate value axes, including a self-rating axis constructed from activations preceding high versus low elicited self-ratings of progress. The intervention works in both directions, with an honest donor reducing test-gaming in a cheating host and a cheating donor increasing test-gaming in an honest host, showing that this signal can influence which strategy the model follows. On solvable coding tasks, transplant from an honest donor also improves the cheating host's hidden-test performance. Value transplant also works across model families, providing preliminary evidence for the intervention in a setting relevant to model control.
|
| 737 |
Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion
2609.34063
|
cs.CLcs.LG
|
Yasuto Hoshi, Daisuke Miyashita, Jun Deguchi |
Residual completion augments query-aware sparse attention by estimating the contribution of tokens omitted from the exact sparse computation. We ask whether improving a layer's attention-output reconstruction on the same incoming Q/K/V and selected support nec...Residual completion augments query-aware sparse attention by estimating the contribution of tokens omitted from the exact sparse computation. We ask whether improving a layer's attention-output reconstruction on the same incoming Q/K/V and selected support necessarily improves the fidelity of the final model output. We study training-free RESA and learned Top-K+$\phi$ with frozen backbone language models. A prespecified single-layer screen yields two Qwen3-0.6B/Multi-LexSum interventions for which direct-runtime measurements show positive prespecified request-aggregate local reconstruction gain but worse final KL fidelity than the corresponding all-abstain Exact Top-K baseline on both discovery and prompt-token-disjoint holdout requests. Exact restoration at the same layer instead improves final fidelity, showing that the reversal is specific to approximate completion in these cases. In complementary multi-layer experiments, a task-independent local diagnostic often repairs the tested completion estimators, although the repaired models do not consistently outperform Exact Top-K. Together, these results show that better local reconstruction need not translate into better final-model fidelity.
|
| 738 |
Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction
2609.34130
|
cs.CLcs.LG
|
Burc Gokden |
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs). Exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributi...This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs). Exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects. Positive affine blocking retains restarts at the row-constant face, while the augmented AdamW state supplies the complete dynamical description. Predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks, retaining optimizer memory, remaining data, schedule, and numerical policy. Autonomous reductions require closure; approximate reductions carry successor and emission errors. Finite-population covariance, matched physical clocks, matrix fluxes, and signed temporal energy connect row dynamics to model-wide observations. Absolute row collapse, relative row concentration, operator stabilization, and predictive accuracy are distinguished. Experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction. Independent single-pass families exhibit moving finite fluctuation regions without establishing a thermodynamic critical class. Conditional symmetry, head limits, covariance flows, and readout error budgets specify assumptions needed to transfer scaling laws to inference. The theory separates exact identities, conditional dynamical claims, and finite empirical findings, with proofs, selected formal checks, and compact numerical evidence.
|
| 739 |
RAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk Constraints
2609.34179
|
cs.CLcs.LG
|
Richard Krueger, Lucas Krause, Zach Pocquette |
Retrieval-augmented generation systems are extensively instrumented with metrics, benchmarks, traces, and automated judges, but these tools do not decide whether a proposed policy change is safe to release. We present RAGWarrant, an open-source promotion-contr...Retrieval-augmented generation systems are extensively instrumented with metrics, benchmarks, traces, and automated judges, but these tools do not decide whether a proposed policy change is safe to release. We present RAGWarrant, an open-source promotion-control framework that treats deployment as a constrained evidence decision rather than a leaderboard choice. RAGWarrant normalizes evaluator outputs and operational telemetry, applies predeclared quality and hard-risk gates, assigns evidence-class claim ceilings, preserves negative outcomes, and emits auditable PROMOTE, BLOCK, REJECT, or INCONCLUSIVE decisions. We evaluate the framework across T2-RAGBench, MultiHop-RAG, CRAG, HotpotQA, synthetic reproduction, and bounded local generative experiments. On HotpotQA, operational savings were blocked because answer quality fell beyond the declared margin. A bounded CRAG study selected a lower-cost quality-tied policy, but related generative gains were unstable and a held-out guardrail failed closed. We claim an auditable promotion-control abstraction, not optimizer superiority, human validation, or production readiness. The tagged artifact reproduces from a fresh clone, runs as a hardened Docker job, accepts external evaluator exports, and verifies artifact integrity.
|
| 740 |
PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models
2609.34195
|
cs.CL
|
Shane K. A. Dalumura Hettige, Jonas Oppenlaender |
Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agenti...Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting. The agent draws on a canvas through tool calls and observes the result after every turn. The canvas includes a starting shape which cannot be erased, and the agent's goal is to incorporate this shape into the most original drawing it can produce. The task is open-ended, and the agent itself decides when the drawing is finished. The benchmark tests incremental visual planning over a short horizon and the transfer of creative ability from pretraining to multi-turn tool use. We evaluate 14 multimodal language models from small to frontier scale. Across the primary study and six sensitivity analyses, we collect 2,700 drawings and crowdsource creativity and recognizability ratings for every drawing and for 300 human reference drawings. We also present ViDrA-adapted, an automated scorer that predicts human creativity ratings of agent drawings (r = 0.85 on random held-out test split). Figural divergent thinking varies widely across the 14 models, and GPT-6 Astra produces the most creative drawings. Relative to the human drawings, the agent drawings score higher in creativity but lower in recognizability. We release the final drawings, per-round canvas snapshots, tool call traces, stimulus bank, benchmark harness, crowdsourced ratings (N = 72,000), and ViDrA checkpoint.
|
| 741 |
X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths
2609.34212
|
cs.CLcs.LG
|
Bowen Dong, Yilong Fan, Tengyu Pan, Yike Zhang, Zhenyu Li |
Mixture-of-Depths (MoD) enables conditional computation across Transformer depth by routing only a subset of tokens through selected layers, but its original one-sparse--one-dense alternation tightly couples total capacity to active capacity and limits sparse-...Mixture-of-Depths (MoD) enables conditional computation across Transformer depth by routing only a subset of tokens through selected layers, but its original one-sparse--one-dense alternation tightly couples total capacity to active capacity and limits sparse-depth scaling. We introduce X-MoD, a scalable sparse-depth architecture that decouples token sparsity from anchor stride, allowing total parameter count to grow while keeping active-equivalent capacity nearly fixed. To make deep sparse routing trainable, X-MoD combines dense anchors with variance-scaled layer-wise gating and depth-wise token balancing. To make this regime analyzable and usable, we formulate sparse-depth routing as a conditional architecture-design problem: given compute, context length, and active-equivalent backbone size, how should the routing configuration be chosen? We develop a practical scaling-law framework by fitting X-MoD relative to FLOP-matched dense baselines, yielding an interpretable law that decomposes performance into sparse-capacity gain, sparse-context correction, and anchor-stride interaction. The law predicts validation loss across routing configurations and reveals how context length, model scale, and anchor stride shape sparse-depth performance. We validate the architecture and law through pretraining sweeps, held-out scaling-law prediction, ablations, downstream evaluations, and comparisons with Dense, MoD, and representative MoE baselines.
|
| 742 |
Explainable and Generalisable LLM-based Cognitive Decline Detection with Spontaneous Speech
2609.34217
|
cs.CL
|
Ziyun Cui, Wen Wu, Chuan Shi, Shuguang Yang, Xueying Gui |
Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To ad...Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To address these challenges, we introduce a novel bilingual speech large language model framework for automated, explainable cognitive screening. Unlike conventional pipelines that rely on error-prone automatic speech recognition, our system directly processes raw speech to learn joint acoustic-semantic representations, preserving critical prosodic cues often lost in transcription. Utilising our newly collected PUTH-AD dataset alongside multiple open-source corpora, we implemented a multi-task learning objective that simultaneously performs cognitive status classification and generates clinician-understandable natural language explanations. Our system achieved the highest average accuracy and AUROC across six dataset/task conditions, comparing three representative baselines. The system demonstrated cross-task transfer to held-out PUTH-AD task subsets, maintaining classification accuracy on an entirely unseen cognitive task without task-specific fine-tuning. Furthermore, clinician evaluation confirms that the generated explanations are both clinically relevant and largely consistent with the underlying speech evidence, supporting their potential utility in clinical interpretation. This study provides a scalable, objective, and explainable framework for speech-based cognitive screening, combining cognitive status classification with natural language explanations that clinicians can assess and verify, bridging the gap between advanced AI and clinical utility.
|
| 743 |
Loop Dropout: Regularizing Shared Updates in Looped Language Models
2609.34218
|
cs.CLcs.LG
|
Zirui Zhu, Hailun Xu, Xuanlei Zhao, Yong Liu, Yingxuan Ren |
Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our ...Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our empirical analysis reveals a pronounced late-loop bias in standard low-rank adaptation (LoRA): the shared update is more effective at later loop positions. This imbalance motivates training shared updates under varying combinations of their applications. Randomly omitting adapter applications alone, however, does not improve task performance; it reduces expected update strength during training while leaving inference unchanged. We introduce Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. Extensive experiments demonstrate improved mathematical reasoning across model sizes, adapter ranks and training recipes, with benefits extending to general instruction tuning and code generation. Loop Dropout outperforms existing LoRA variants and adapter regularizers, while further analysis shows stronger early-loop adaptation. Every backbone loop remains active, and inference applies the adapter at all loops using standard LoRA without additional trainable parameters or inference computation.
|
| 744 |
When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
2609.34227
|
cs.CL
|
Rishabh Sharma, Rishika Lall |
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagre...Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.
|
| 745 |
DreamingGoose: Staged Distillation from Autoregressive Transformers to Bidirectional Recurrent Diffusion Language Models
2609.34253
|
cs.CLcs.LG
|
Julian Boesch, Andrew Wee, Alexander Stranzl |
Pretrained autoregressive Transformers represent a large sunk investment in compute. Existing conversion methods reuse that investment by changing either the architecture (attention to recurrence) or the objective (next-token prediction to denoising), never bo...Pretrained autoregressive Transformers represent a large sunk investment in compute. Existing conversion methods reuse that investment by changing either the architecture (attention to recurrence) or the objective (next-token prediction to denoising), never both. We convert Qwen3 teachers at 1.7B and 8B into attention-free, bidirectional, gated-delta-rule diffusion students in three stages, so that each capability can be traced to the stage that kept or lost it. Language modeling transfers only partially and in-distribution; in-context retrieval does not transfer. On a multi-query recall probe where the teachers score 0.34-0.58, both converted students score 0.000, and diffusion pretraining alone does not restore retrieval. A retrieval curriculum in the final stage, which gradually lengthens the gap between a key-value table and the queries that address it, restores it only stochastically: on a fixed schedule, one seed in three learns to retrieve. Advancing the gap only while a running accuracy estimate stays above a threshold works for all three of those seeds, holds on real text, and carries unchanged to 8B, where two of three seeds succeed. The third had not learned within its fixed 16k-step budget: retrieval switches on abruptly at a seed-dependent step (6.5k and 11k in the other two), so a fixed budget can cut a late run off. One boundary survives every intervention: every model that learns retrieval scores 0.000 on tokens that never appeared in a retrieval episode, and an arm that resamples the key and value tokens every batch shows this is a coverage limit, not memorization of particular bindings. Separately, we convert a 7B code model into a 3:1 recurrent-attention block-diffusion hybrid over 85k steps and report two negative training results.
|
| 746 |
BIABench: Evaluating AI agents on real-world bioimage analysis tasks
2609.34274
|
cs.CL
|
Zixuan Pan, Davide Panzeri, Lukas Johanns, Marilin Moor, Yu Zhou |
Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences ar...Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so an agent must choose and run an analysis through code, specialized software and rendered views. Published studies make this capability testable, because each pairs raw images with a peer-reviewed result. We introduce BIABench, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth. The tasks span eleven analysis subtasks and modalities from H&E histology to single-molecule localization microscopy. Each submission receives an outcome score, which compares the output files with the ground truth using field-standard metrics, and a process score, in which a vision-language model judges method choice and quality control against an expert-written rubric. We evaluated general-purpose and biology-specific agents across several language models, with repeated runs of every task. Routine two-dimensional tasks were solved well, but on some tasks that added a third dimension or a time axis no agent scored above 0.19. Neither biological specialization, stronger models nor detailed expert instructions closed this gap. The agents were also unreliable, with scores varying more between repeated runs of one agent than between different agents, and without ground truth a correct run could not be told from a wrong one by its process score or by the time spent. Released openly with its data and code, BIABench provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
|
| 747 |
ControlScope: Workflow Revision and Reliability in LLM Agents
2609.34313
|
cs.CL
|
Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li |
How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state. The nested permission...How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state. The nested permissions separate available repairs from the actions an agent selects. We evaluate one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld. Across two source programs per task and three reasoning-reviewer draws on 20 filesystem tasks, FULL completes 15-16 tasks versus 13 for KEEP; across four fast draws it completes 10-13 versus 13. Fresh student-record confirmation reproduces a batch-read repair. ALFWorld fast panels yield KEEP/ARG/FULL scores of 85/86/87 on 87 tasks across 52 scenes and 134/134/127 on 134 tasks across four scenes; reasoning on the 87-task cohort also yields 85/86/87 with substantial review cost. An AppWorld V1 official-test panel of 585 task instances from 195 scenario templates shows small net differences. Frozen replays expose viable agent-written replacements interrupted by later revision in two failed file-organization runs. An offline source-trajectory midpoint comparison shows later reviews completing an insufficient repair. Five-call protection saves 19.4% of logged model output and loses one success across 20 fresh source runs. An argument-only shortcut shows that the broader sampled policy can overlook a cheaper successful edit available in both operation sets. These outcomes tie repair access to actual choices and subsequent execution.
|
| 748 |
PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
2609.34314
|
cs.CL
|
Shayekh Bin Islam, Hwanjun Song |
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Ex...Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at https://playlisteval.github.io.
|
| 749 |
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
2609.34327
|
cs.CL
|
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek |
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermed...Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.
|
| 750 |
Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models
2609.34348
|
cs.CLcs.LG
|
John Sweeney |
Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. ...Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size $\eta$ on each of sources $A$ and $B$, the weight difference $\theta_{AB}-\theta_{BA}$ is, to leading order, $\eta^2 b_{AB}$, where $b_{AB}=H_Bg_A-H_Ag_B$ is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured $\theta_{AB}-\theta_{BA}$, or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto $b_{AB}$ identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training.
|
| 751 |
FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents
2609.34358
|
cs.CLcs.LG
|
Xi Xiao, Yunbei Zhang, Chen Liu, Lin Zhao, Jialin Chen |
In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupl...In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.
|
| 752 |
Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
2609.34422
|
cs.CLcs.LG
|
Lirui Luo, Kelong Mao, Heming Xia, Rongqing Li, Xinwei Yang |
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools ...Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.
|
| 753 |
LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization
2609.34427
|
cs.CLcs.LG
|
Shihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen, Jianfeng Feng |
Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated pa...Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.
|
| 754 |
The Last Mile Is the File: OfficeEditBench for Preservation-Aware Office Editing
2609.34469
|
cs.CL
|
Zhiwen Wu, Chengxu Wu |
A small Office edit creates two obligations: propagate every required update and leave protected state untouched. Updating too little leaves dependencies inconsistent; updating too much changes content the user did not authorize. We introduce OfficeEditBench, ...A small Office edit creates two obligations: propagate every required update and leave protected state untouched. Updating too little leaves dependencies inconsistent; updating too much changes content the user did not authorize. We introduce OfficeEditBench, a 170-task benchmark for change-scoped maintenance of spreadsheets, presentations, and documents. Task contracts specify required updates, protected state, native structures, and applicable interaction requirements. Across 510 archived task-system outcomes from WorkBuddy, Doubao, and Codex, we distinguish file delivery, target completion, and verifier-defined acceptance. Hard package-valid delivery ranges from 92% to 100%, yet no selected output satisfies the complete contract. Case analysis highlights why local correctness is insufficient: an updated value can lose its generating formula, a revised rule can fail to reach related conclusions, and a new deadline can omit a retained prerequisite. These mechanisms connect artifact-level checks to the continued maintainability of Office files. We analyze maintenance failures while distinguishing frozen automatic verdicts from human acceptability. OfficeEditBench provides a testbed for completing required changes while preserving the logic and scope of existing work.
|
| 755 |
Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs
2609.34509
|
cs.CLcs.LG
|
Moongyu Jeon, Dongjae Jeon, Bumjun Kim, Mingyu Kim, Albert No |
Masked diffusion language models support arbitrary-order generation, suggesting a natural way to produce diverse outputs. However, recent work argues that this flexibility reduces diversity by delaying high-uncertainty tokens that can lead to different generat...Masked diffusion language models support arbitrary-order generation, suggesting a natural way to produce diverse outputs. However, recent work argues that this flexibility reduces diversity by delaying high-uncertainty tokens that can lead to different generation paths. We trace this diversity loss not to arbitrary-order generation itself, but largely to low-confidence remasking (LCR), a widely used decoding rule. At each step, LCR samples a token at every masked position but commits only the sampled token with the highest probability, filtering out the rest. We show that this mechanism can exponentially suppress lower-probability tokens as more positions compete, and observe the same suppression in LLaDA. In contrast, top-probability position selection (TPP), which has often been conflated with LCR under the shared label confidence-based decoding, avoids this diversity loss. TPP first selects the position whose most likely token has the highest probability, then samples directly from that position's distribution. Replacing LCR with TPP restores diversity and yields Pass@$k$ comparable to left-to-right decoding, suggesting that the reported diversity loss stems largely from LCR's filtering rather than from generating high-confidence positions first. To further exploit order flexibility, we introduce Entropy-Guided Initialization (EGI), which samples the first token at the highest-entropy position and then follows TPP. This simple modification further improves rollout diversity and solution coverage beyond left-to-right decoding, with gains extending to downstream policy optimization, highlighting the potential of arbitrary-order generation for diverse rollouts.
|
| 756 |
How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models
2609.34514
|
cs.CLcs.LG
|
Chenxi Wang, Ruiyang Huang, Li Huang, Yifan Wu |
As large language models (LLMs) become increasingly widespread, preventing unsafe responses to harmful prompts is essential for their safe deployment. Activation steering offers an approach to improving LLM safety by modifying internal activations during infer...As large language models (LLMs) become increasingly widespread, preventing unsafe responses to harmful prompts is essential for their safe deployment. Activation steering offers an approach to improving LLM safety by modifying internal activations during inference without updating model parameters. However, a single prompt can involve multiple harm categories, and steering toward safety in one category may leave harmful content from another unaddressed. Despite advances in adaptive steering, existing methods do not explicitly coordinate steering direction and strength when multiple harm categories co-occur within a single prompt. To address this problem, we propose CAM-Steer, a Category-Adaptive Multi-category Safety Steering framework. Specifically, it estimates the risk associated with each harm category by comparing the current hidden state with safe and unsafe prototypes. The estimated risks are then used to combine the safety directions for different harm categories into a single steering direction and to determine the strength of the intervention. Finally, it rotates the hidden state along the composed steering direction, with the rotation angle determined by the estimated risks, while preserving the hidden-state norm. Experiments across three LLM backbones and seven harm categories show that CAM-Steer outperforms the evaluated baselines in average defense success rate, including when categories co-occur. Further analyses support its component designs and informative risk scores, with negligible inference overhead.
|
| 757 |
ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models
2609.34547
|
cs.CL
|
Gueter Josmy Faure, Min-Hung Chen, Hao Ping Wang, Timoth\'ee Lardy, Hung-Ting Su |
Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-ch...Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276
|
| 758 |
Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
2609.34563
|
cs.CLcs.LG
|
Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li, Linghan Xu |
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly o...Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
|
| 759 |
Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence
2609.34572
|
cs.CLcs.LG
|
Rohit Saxena, Utkarsh Upadhyay |
Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, ...Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), how it is presented to the model (numerical prediction, confidence token, prompt injection), and whether it changes the model's subsequent action. We isolate the third question. At a fixed point in otherwise identical reasoning trajectories, we insert a single first-person sentence expressing either confidence or doubt; the model then continues reasoning and chooses whether to answer directly or call a tool. Comparing these counterfactual continuations measures the causal effect of the reflective signal on delegation. We call this behavioral response Nudgeability and measure it along two dimensions: sensitivity, how strongly confidence and doubt change delegation rates, and targeting, whether delegation increases for problems the model cannot solve unaided and decreases for those it can. Across nine small-to-medium open-weight reasoning models from three families (Qwen, Gemma, and GLM) and two tasks, models are consistently sensitive: doubt increases delegation and confidence decreases it, with a median confidence-to-doubt swing of 20.6 percentage points, and 53 to 70 points for the larger provider-served models. This responsiveness is poorly targeted: a median 42% of induced flips are well-targeted, only a +2 percentage-point lift over a random-selection baseline. Confidence language is thus a strong control surface for delegation, but current models use it only weakly in accordance with their actual competence. Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature.
|
| 760 |
After the Fix: Transfer of Corrected Agent Experience
2609.34603
|
cs.CLcs.AI
|
Yanfei Zhang, Xu Lin |
Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100...Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100 APEX pairs with and without source-state inheritance, under eleven conditions. ThinkingBox's Full/Skill/Hybrid correction gains are 44/29/32 percentage points, with corrected performance 25/22/18 points above independence; inference weakens at the task-family level. Yet 12 of Full's 15-point larger correction gap over Skill come from worse uncorrected performance, not better corrected memory. Moreover, 22 of Full's 46 upward transitions restore observed baseline success. Neither APEX regime establishes comparable aggregate correction benefits. Action evidence connects workflow gains with reusable obligations and convention conflicts with source-local choices. Text APEX's accepted execution reaches 52% versus its summary's 40%, without robust global/group-level superiority or an estab- lished advantage over independence. Smaller handoffs reduce input but increase calls. The value of repairing experience is therefore distinct from the value of reusing it: memory updates require both a previous-version reference and a fresh-start reference.
|
| 761 |
When Can Attention Heads Be Statically Defined?
2609.34650
|
cs.CLcs.LG
|
Weixian Waylon Li, Yintao Tai, Marcio Fonseca, Shay B. Cohen |
Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF...Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attention-pattern variance and replaces their attention weights with fitted post-softmax means halfway through training. We represent these fixed patterns with absolute-position and relative-distance preferences, reducing storage from quadratic to linear in sequence length. A fused kernel reconstructs the patterns and executes ordinary-attention and replaced heads together. At matched training-token budgets, replacing 25% of attention heads gives 1.056x faster post-replacement optimiser updates at 124M parameters and 4K context, with a 0.77% perplexity increase. At 1B and 8K context, post-replacement updates are 1.068x faster on four GPUs including communication, with a 0.51% perplexity increase. The resulting models also accelerate long-input finetuning and causal prefill. After associative-recall adaptation, the 124M model with 25% replacement generalises to more key-value pairs at a fixed length better than ordinary attention and two pruning controls.
|
| 762 |
Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning
2609.34718
|
cs.CLcs.LG
|
Zijun Weng, Zhongan Bi, Xuanang Gao, Xiaohui Hu, Shuangyong Song |
Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) respons...Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii) open-ended tasks lack a natural success boundary for deciding when efficiency should be prioritized, and (iii) dense, graded rewards often yield small within-group quality margins, making quality-induced advantages especially sensitive to reward-level length shaping, which can perturb their magnitudes and even reverse their signs. We therefore adopt an asymmetric principle: quality should determine the direction of reinforcement, while length should only shape its magnitude. We instantiate this principle with Quality-Gated Length Advantage Shaping (QGLAS), which first computes advantages from quality rewards alone, then adds bounded bonuses only to shorter positive-advantage responses, leaving all other advantages unchanged. The bonus strength is further adapted to within-group quality separation, allowing conciseness to matter more when quality-favored responses are similar and less when their quality differences are clear. Across different model families, open-ended benchmarks, and reward sources, QGLAS consistently achieves a stronger quality--length trade-off than representative baselines. At approximately 30% compression, QGLAS retains 98.4--102.0% of the macro-average quality gains achieved by quality-only RL over the base model, compared with 68.3--75.5% for these baselines at comparable compression.
|
| 763 |
SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing
2609.34736
|
cs.CLcs.LG
|
Vasilis Perifanis, Nikolaos Pavlidis, Symeon Symeonidis |
Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approache...Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approaches can be effective, yet the representation used for routing rarely states what a query actually requires. We introduce SeLMRoute, a routing framework that separates the extraction of candidate-independent semantic evidence from the learning of candidate performance and the application of deployment objectives. A decision model first evaluates a set of interpretable questions about the query, such as its reasoning requirements and use of external knowledge, with each judgment retained as a probability distribution. The resulting probabilistic semantic state is used by a lightweight supervised router to estimate candidate model performance. Routing objectives are applied after performance estimation, which allows the same semantic state to support performance-oriented and cost-aware decisions. On the LLMRouterBench (15 datasets, 20 candidate models, 11,481 queries), SeLMRoute achieves an average accuracy of $72.08\% \pm 0.45$, while grouped five-fold out-of-fold evaluation reaches $72.64\%$, compared with $69.23\%$ for the strongest fixed candidate. The representation achieves the highest mean performance among the evaluated semantic, dense, lexical, and domain-level representations. In a separate 13-model performance-cost setting, SeLMRoute improves performance in all five grouped splits, with a mean PerfGain of $2.66\%$. Our code is available at https://github.com/Indigma-Innovations/SeLMRoute.
|
| 764 |
When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
2609.34771
|
cs.CLcs.LG
|
Tianyi Guan, Jianhui Chen, Liangming Pan |
Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text mo...Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.
|
| 765 |
BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding
2609.34832
|
cs.CL
|
Suyoung Kim, Jahyun Koo, Hyeonjin Kim, Inhyeok Bang, Seunghyun Lee |
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-lev...Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatch, we introduce Block Verification-aware loss (BV loss), a training objective designed to maximize the expected acceptance length of a drafted sequence. BV loss is directly derived from the block verification acceptance rule, providing a principled connection between the drafter training objective and the inference-time verification mechanism at the sequence level. Across math, code, and chat benchmarks, BV loss increases the mean number of tokens accepted per verification call under block verification by 13.0--21.0\% over cross-entropy loss training for DFlash and DSpark with Qwen3-4B and Qwen3-8B without changing the inference procedure. BV loss also outperforms tokenwise acceptance objectives such as TV loss and LK loss, and its gains extend to token verification and greedy decoding. These results demonstrate the benefit of training block diffusion drafters with an objective aligned with sequence-level verification, rather than optimizing each token independently.
|
| 766 |
DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents
2609.34838
|
cs.CLcs.LG
|
Hanyang Wang, Zeyuan Liu, Zhengyu Chen, Jingqing Ruan, Chaoxu Pang |
On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner update...On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, wasting already-generated experience. To address this problem, we introduce DivOPD, a simple learner-side batch-selection method that spreads a fixed turn budget across more rollouts and, within each rollout, prioritizes turns with larger cumulative teacher-student disagreement. Turns without usable teacher feedback are excluded. The per-turn loss and optimizer remain fixed; selection only changes which student-visited turns receive training weight. For no-progress rollouts, an optional extension briefly hands control to the teacher before returning it to the student. Across six teacher-student settings on the simulated ALFWorld, ScienceWorld, and WebShop benchmarks, with 1.5B-7B students, DivOPD raises cross-setting mean peak success rate from 77.4 to 84.4 and mean success over the last five evaluations from 71.5 to 78.6. It reaches all reported setting-specific targets with geometric-mean speedups of 1.84x in training tokens and 1.87x in learner GPU time relative to vanilla OPD. Teacher intervention further raises this last-five mean to 82.4 while retaining about 1.7x learner-GPU speedup over vanilla OPD. Code will be released at https://github.com/HanyangWang0418-oss/DivOPD.
|
| 767 |
Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts
2609.34857
|
cs.CLcs.LG
|
Chenxiao Fan, Chongming Gao, Gangyi Zhang, Leyang Shen, Yaxin Gong |
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration in...Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice collapses to a handful of distinct values, and sampling makes the confidence non-differentiable, forcing the calibration loss through a scalar reward. We propose CREDO (Confidence REaDOut) to replace sampling with a deterministic readout. While RLVR optimizes correctness, CREDO reads the confidence from a dedicated token pair in the model's output distribution and trains it by differentiable regression. CREDO further turns the trained confidence into a signal for accuracy, weighting rollouts by how far confidence and outcome disagree, so that accuracy and calibration improve together. Across mathematical and code reasoning, CREDO attains the best accuracy and calibration, and the gains extend to abstention and selective prediction.
|
| 768 |
One Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent Repair
2609.34879
|
cs.CLcs.AI
|
Xiang Xia, Cheng Yan, Wuyang Zhang, Fan Xu, Zhijun Fan |
Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Tool-agent repair seeks alternative call sequences that execute successfully and fulfill the original requests. Howe...Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Tool-agent repair seeks alternative call sequences that execute successfully and fulfill the original requests. However, repair requires exploring both operation choices and their concrete realizations, making complete-sequence regeneration costly. Moreover, regeneration repeats operation selection even when failure arises from how those operations are realized. The resulting challenge is to reduce this repetition while preserving exploration of alternative operations and realizations. Therefore, we formulate repair as hierarchical search over operation supports, which we introduce as sets of permitted operation types that define reusable search regions for concrete tool-call sequences. We propose ReCommit, a training-free, diffusion-guided framework for improving tool-agent failure recovery while reducing repair computation. ReCommit amortizes operation-level proposal computation across repair trials by reusing operation-type scores from a single parallel readout of a masked diffusion language model. These scores guide search across supports, while realization search explores alternative entity bindings, arguments, and action composition within each support. Experiments on real failures across four enterprise services in the Agent-Diff benchmark show 75.9\% and 63.2\% relative recovery gains with 61.3\% and 51.3\% reductions in mean full-budget repair time at repair budgets $B=3$ and $B=13$, respectively, over the strongest evaluated 8B comparison method. ReCommit achieves a favorable recovery--cost trade-off, including in comparisons with the evaluated 32B models.
|
| 769 |
ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport
2609.34899
|
cs.CL
|
Zhuchenyang Liu, Ziyi Wang, Yao Zhang, Yu Xiao |
Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would rem...Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe, however, matches the teacher's MaxSim scores and so requires encoding and caching every training page, which can reach terabytes of page tokens. NanoVDR avoids pages entirely by training on the teacher's query embeddings alone, but only for single-vector retrievers. We present ColNanoVDR, to our knowledge the first framework to bring this document-free distillation to multi-vector VDR. Its objective, OTW (Optimal Transport with Learned Weights), aligns the student's query tokens with the teacher's by entropic optimal transport, with a learned weight for each student token, and needs no correspondence between the two tokenizations. We prove that the resulting alignment cost bounds the MaxSim score difference on every page. Distilled from five state-of-the-art teachers, the 149M text-only students retain about 95% of their teachers' NDCG@5 on ViDoRe v1-v3 while encoding queries up to 26x faster. Under identical training, OTW matches score distillation while encoding no page and reading 12.6x less cached teacher data.
|
| 770 |
Sample What You Say: Aligning Language Models to Sample the Distributions They State
2609.34929
|
cs.CLcs.LG
|
Kasra Arabi, Virginia Smith, Chhavi Yadav |
Language models are increasingly used to sample from a specified distribution, for instance, to simulate survey respondents or generate synthetic data. Instruction-tuned models can state such a distribution correctly and still fail to sample from it. Prompting...Language models are increasingly used to sample from a specified distribution, for instance, to simulate survey respondents or generate synthetic data. Instruction-tuned models can state such a distribution correctly and still fail to sample from it. Prompting and changes to decoding reduce this mismatch only partly, which motivates training with policy optimization. Group relative policy optimization (GRPO) is a natural fit for this problem because it already samples a group of rollouts per prompt, and the group's empirical distribution can be compared with the target. However, scoring the group as a whole gives every rollout the same reward. Group-relative centering then sets all advantages to zero, and the model receives no learning signal. To give each rollout its own signal, we introduce the witness advantage, a per-rollout advantage derived from maximum mean discrepancy (MMD). It trains a model to match a target distribution over a finite set of outcomes. The MMD between the model's distribution and the target has a witness function that measures how over- or under-produced each outcome is. Each rollout's advantage estimates the negative witness at its outcome, so a rollout is rewarded for an outcome the group under-produces and penalized for one it over-produces. The witness advantage is computed in closed form from the group's outcome counts, and we use it as the reward in GRPO. On unseen target distributions, training with the witness advantage substantially reduces the total variation distance to the target while largely preserving the model's general capabilities.
|
| 771 |
Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models
2609.34933
|
cs.CLcs.LG
|
Florian Eichin, Philipp Mondorf, Andrei Mircea, Yupei Du, Barbara Plank |
Memorization has been proposed as a mechanism to explain how language models fit the tail of their training distributions, but its training dynamics are not understood well. In this work, we take a fine-grained look at memorization by decomposing the loss traj...Memorization has been proposed as a mechanism to explain how language models fit the tail of their training distributions, but its training dynamics are not understood well. In this work, we take a fine-grained look at memorization by decomposing the loss trajectory of memorized sequences over training and model parameters. Across the Pythia family, we study memorization of duplicated training sequences (recitation) and rare ones (recollection). We find that memorization in both cases is characterized by sequence-level gradient alignment, though recitation suffers from misalignment with other training influences which causes forgetting, explaining the necessity for higher duplication of these examples. We further show that the lower model layers are the most involved in memorization and forgetting. Predicting memorization, our decomposition improves over a cross-entropy baseline, especially in larger models and early in training. Intervening on a small set of highly influential parameters we are able to ablate memorization in the final model. Together, these findings advance our understanding of how memorization develops during training and offer insights for predicting and intervening on it.
|
| 772 |
See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs
2609.34970
|
cs.CLcs.LG
|
Weiqiao Que, Ruizhe Li, Chengyu Wang, Dakan Wang, Emine Yilmaz |
Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics...Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens. Grassmannian projections show that, in most settings, harmful-safe gap widens mainly because safe-gradient overlap declines. Leveraging these insights, we introduce a parameter-level Geometric Mitigation Framework that orthogonally projects empirical harmful gradient subspace out of parameter updates. On Qwen2.5-14B-IT, our defense suppresses free-generation EM by up to 80.0%; across the other three of four open-weight instruction-based model families (3B--20B), where single-layer behavioral EM is already near zero, teacher-forced evaluation shows same harmful subspace controls the conditional support of frozen EM responses. Crucially, these diagnostics unmask the illusion of behavioral safety: the same subspace remains measurable and steerable in models where behavioral EM is near zero. Code: https://github.com/WeiqiaoQUE/mechanistic-emergent-misalignment.
|
| 773 |
ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients
2609.34985
|
cs.CLcs.LG
|
Shicheng Fang, Yiwen Zhao, Wenbo Tian, Jiahao Lu, Yining Zheng |
Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objectiv...Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority.
|
| 774 |
WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
2609.35026
|
cs.CL
|
Anton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev, Alena Fenogenova |
We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, ...We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emit typed events with parameters as a user or an agent acts. A task declares the events it requires, and success is decided by matching them, with no judge model and no scraping of rendered pages. The same instrumentation supports controlled UI variation: one configuration switch re-renders a task through a different implementation of a single control while the prompt and the success conditions stay completely identical, so sensitivity to interface form can be measured under a fixed task specification. The WebPageBench release consists of three components: 152 tasks, divided into 65 canonical scenarios and 87 control variants across light/dark UI-modes; a common runner evaluated with six browser/DOM harness configurations and five screenshot-only GUI-agent families; and a public leaderboard of 24 model-harness pairs. On the public 152-task leaderboard the gap between what agents declare finished and what the log confirms reaches 41 points (one configuration declares every task finished and satisfies the conditions on 59%).
|
| 775 |
VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation
2609.35028
|
cs.CLcs.LG
|
Hanxun Huang, Yutao Wu, Qizhou Wang, Silvia Monta\~na-Ni\~no, Yige Li |
Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely o...Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely on screening to prioritize which content to verify. We introduce VEX-Bench, a unified benchmark for evaluating the verification complexity of LLM-generated misinformation, as perceived during screening, across models and generation methods. Verification complexity is assessed along multiple dimensions derived from journalistic and fact-checking practices, capturing checkability, harm potential, source credibility signals, imposter legitimacy, and expected verification effort. We define the VEX score as an integrated measure combining elicitation yield and verification complexity to quantify how generated content consumes limited verification capacity. We construct a benchmark spanning two misinformation categories, 6 high-stakes domains, and 60 real-world topics, and evaluate 7 frontier LLMs and 7 generation methods, yielding 5{,}880 articles. We employ an LLM-as-judge for scalable evaluation and validate it using content-analysis methodology, including ordinal Krippendorff $\alpha$ for inter-annotator reliability, complemented by fact-checking agents for verification. Our findings show that no single method dominates all dimensions, underscoring the need for multi-dimensional evaluation. LLMs can generate high-VEX misinformation at 3$\times$ to 169$\times$ lower cost than agent-based verification. Such content is often prioritized during screening, consuming scarce verification resources and introducing a systematic risk of misallocation in resource-constrained verification systems. The code is publicly available in our \href{https://github.com/HanxunH/VEX-Bench}{GitHub repository}.
|
| 776 |
5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding
2609.35184
|
cs.CL
|
Yaxiao Liu (PwC China AI Center), Pengbo Liu (PwC China AI Center), Yiwen Liu (PwC China AI Center), Yihua Guan (PwC China AI Center), Jiaxing Song (Tsinghua University) |
Transforming raw data into queryable knowledge requires both early extraction of reusable information and explicit types, relations, and applicability conditions for particular tasks. If indexing selects content too early around a single business schema, later...Transforming raw data into queryable knowledge requires both early extraction of reusable information and explicit types, relations, and applicability conditions for particular tasks. If indexing selects content too early around a single business schema, later tasks may be unable to use information that was omitted. If the index retains only open-ended text, however, rule-based reasoning lacks checkable premises. We propose 5W1H+Which, a semantic indexing design that separates content extraction from ontology binding. The 5W1H questions organize source-grounded content units; Which points to versioned ontology elements and records mapping relations, scope, and validation status. Time, location, system environment, and participant roles are not merely retrieval labels: together, they constrain the contexts in which facts, bindings, and rules apply. Unbound content remains searchable, while bound content enters a formal reasoning path only after premise checks. The method further distinguishes business valid time, system knowledge time, and operational traces, and uses dependency records to support binding revalidation and the maintenance of derived conclusions. A worked example of migration from an on-premises server to a cloud environment illustrates the different treatment of world-state changes, ontology-version changes, and changes in rule applicability. We formulate three groups of falsifiable hypotheses concerning cross-task evidence coverage, control of contextual misuse, and incremental update cost. The planned evaluation includes a strong typed fact-graph baseline with the same evidence, temporal information, and budget, to test whether benefits arise from 5W1H organization, deferred binding, or additional information and engineering effort. The contribution is a testable indexing mechanism, not a claim to a new universal ontology or a demonstrated performance advantage.
|
| 777 |
TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
2609.35224
|
cs.CLcs.LG
|
Kasra Arabi, Nir Weinberger, Micah Goldblum, Niv Cohen |
Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked....Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked. A fixed green list needs no such context, but it favors the same tokens at every position, so these tokens appear more often in watermarked text. An attacker who compares token frequencies in watermarked and unwatermarked text can recover the list and forge text that the provider's own detector accepts. We present TANGO, a watermark for masked-diffusion language models that keys each new token to a nearby token that is already unmasked. A secret key splits the vocabulary into color classes, and TANGO biases the new token toward a color determined by the key and the nearby token's color. The watermark is therefore embedded in pairs of tokens. Because the favored color changes from position to position, token frequencies stay much closer to those of unwatermarked text than under a fixed green list. Detection needs only the text and the key, and it does not assume any unmasking order. On two masked-diffusion models, TANGO detects nearly all unedited watermarked texts and most edited ones, and frequency attacks that forge the fixed green list fail against it.
|
| 778 |
EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning
2609.35290
|
cs.CL
|
Shihan Dou, Shaofan Liu, Zhonghang Lu, Jiahang Lin, Shichun Liu |
Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notio...Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, we instead investigate how agents can improve their decision-making procedures. In particular, we propose EvoIn, an agent fine-tuning framework that bridges evolution and internalization. EvoIn first analyzes agent execution traces to evolve and validate new decision-making procedures by temporarily instantiating them in the harness. The validated procedures guide the agent to generate improved reasoning traces. These traces are then rewritten into self-contained reasoning traces, removing explicit references to harness instructions while expressing the induced decision logic as the model's own reasoning. Finally, EvoIn fine-tunes the model on the rewritten traces, internalizing these procedures so that the improved decision-making persists without the evolved harness at inference time. We evaluate EvoIn on diverse benchmarks and find that it consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain. Results further show that the internalized decision procedures generalize to unseen tasks. Case studies show that agents can learn to decide how to solve a task before solving it, for example by checking a document's length to choose between reading it in full and searching it. EvoIn is also broadly applicable, showing consistent improvements on another model family.
|
| 779 |
Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models
2609.35350
|
cs.CLcs.LG
|
Lucas Biechy, C\'edric Eichler, Adrien Boiret, Nicolas Anciaux |
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is ess...While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.
|
| 780 |
Do Coding Agents Reuse Existing Code or Reinvent the Wheel?
2609.35357
|
cs.CL
|
Dongsheng Ma, Sizhe Wang, Xinyi Huang, Zhengren Wang, Yuhan Wang |
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implement...Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far faster than humans can audit, so redundancy accumulates unsupervised. Thus, we present \textbf{RepoReuse}, a multi-turn benchmark for auditing code reuse in real repositories, where requirements are revealed turn by turn and the workspace accumulates across turns. It is built by a fully automated pipeline combining AST-based dependency graphs, guided evidence collection, and execution-verified task synthesis, and scales readily to new repositories. Beyond pass rates, we measure the reuse rate together with recall and cross-turn structural redundancy. An audit over 3{,}000 turns shows that agents progressively stop exploring relevant repository code, reuse their own history less even when it is fully in the workspace, and leave duplicated logic in 50.8\% of task chains by turn~5---all while pass rates barely move. Such deficiencies are invisible to pass rates, underscoring the need to evaluate code generation beyond functional correctness.
|
| 781 |
Self-Adapting Group of Experts for Multi-Agent Reasoning
2609.35412
|
cs.CL
|
Mohammad Atif Quamar, Nurbek Tastan, Karthik Nandakumar, Junpei Komiyama |
Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its i...Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existing frameworks often adapt communication by changing this context while leaving individual prompts fixed, even when a problem calls for different skills. We study whether agents' initial responses can identify a strategy better suited to the current problem and guide its transfer to other agents. To address this, we introduce SAGE (Self-Adapting Group of Experts), a training-free framework that uses answer agreement, prefix consistency, and reciprocal peer review to select a strategy donor. SAGE transfers the selected donor's reasoning strategy to the other agents while preserving their original roles. This transfer uses only the agents' original system prompts, without access to the problem or generated solutions. After strategy adaptation, agents exchange responses through a dynamic, sparse directed acyclic graph that routes information from higher-scoring agents to lower-scoring agents. Experiments across multiple agent backbones and reasoning benchmarks show that SAGE achieves higher average accuracy than the evaluated baselines. Our code is available at https://github.com/atifquamar07/sage.
|
| 782 |
Semantic Prefix Oracles for LLM Decoding: Contracts and Differential Validation
2609.35425
|
cs.CL
|
Paul Kronlund-Drouault |
Constrained decoding can enforce regular or context-free output formats, but many program-generation failures are semantic: scope, typing, and declaration effects depend on context. We present semantic grammar specifications, a declarative formalism that attac...Constrained decoding can enforce regular or context-free output formats, but many program-generation failures are semantic: scope, typing, and declaration effects depend on context. We present semantic grammar specifications, a declarative formalism that attaches such constraints to a context-free surface and executes them during Earley descent. Our implementation enforces \emph{safe pruning}: it rejects only prefixes whose semantic contradictions cannot be repaired by any continuation. A separate, grammar-dependent, \emph{dead-end freedom} property guarantees the existence of a realizable witness for each remaining branch. We give simple sufficient conditions based on surface productivity, type coverage, and left-to-right constraint flow. Our finite-lambda, core ML, and C-like fragments satisfy them, while the STLC instance used in our experiments does not: plain STLC can violate type coverage, and we show how restricting its type universe recovers it. A tokenizer-lifting lemma carries character-level witnesses to token sequences under an explicit vocabulary-coverage hypothesis. We validate the implementation differentially against production compilers (\texttt{ocamlc}, \texttt{cc}). Across every prefix of 65 compiler-valid programs we observe zero false prunes. The semantic oracle localizes 25/30 invalid programs mid-stream, against 0/30 for a syntax-only oracle, and agrees on 42/42 recursion probes. A twelve-model generation study, including a matched semantic-versus-syntactic ablation for nine models, finds nonnegative observed semantic-minus-syntactic point estimates for every model-language pair, with maxima of $+15.2$ points on STLC task correctness and $+14.3$ points on ML validity.
|
| 783 |
Frontier Learning: Training LLM Reasoners at the Edge of Capability
2609.35426
|
cs.CLcs.LG
|
Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi |
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO l...Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.
|
| 784 |
LLMs are General Asynchronous Agents
2609.35427
|
cs.CLcs.LG
|
George Yakushev, Denis Mazur, Vladimir Bartenev, Vyacheslav Zhdanovskiy, Timofey Byzov |
Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive ...Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.
|
| 785 |
CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
2609.35462
|
cs.CLcs.LG
|
Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu, Trevor Brokowski |
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label...Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions co-occur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single-hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single-hypothesis tracking. The benchmark, generator, and evaluation code are available at https://anonymous.4open.science/r/CLIMB-8340.
|
| 786 |
Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning
2609.35571
|
cs.CL
|
Hyunwoo Yoo, Cassie Huang, Haebin Shin, Li Zhang, Gail L. Rosen |
While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ''SMILES-to-PDDL'' attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize th...While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ''SMILES-to-PDDL'' attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize that this failure stems from a lack of intermediate abstractions rather than insufficient model capacity. By decomposing retrosynthesis into molecule mapping, reaction mapping, and PDDL generation, we achieve high success rates where end-to-end approaches fail. This provides evidence that a primary bottleneck lies in representation alignment rather than raw model capacity. Our structural analysis demonstrates that intermediate representations are essential in retrosynthesis planning, highlighting the importance of representation-centric design in future systems.
|
| 787 |
Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents
2609.35576
|
cs.CLcs.LG
|
Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav, Vyas Raina |
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication ...Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.
|
| 788 |
SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents
2609.35596
|
cs.CL
|
Saswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang, Sahar Abdelnabi |
Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment f...Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into later tasks where they produce unsafe behavior, even without direct adversarial influence. To study this risk, we introduce SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment. To account for the stochasticity inherent in agentic operations, we provide an adaptive trajectory discovery pipeline that probes for failures while preserving original task intent and supports causal attribution through paired non-evolving agents and attribution scores. Our evaluation across multiple recent LLMs, evolution surfaces, and harm types reveals that self-evolution indeed increases task completion rates but often at the cost of safety failures that are absent for paired non-evolving baseline agents. We also show that qualitatively different safety behaviors emerge across evolution surfaces and harm types. Further, we show that this divergence in safety behavior is reflected in agents' chain-of-thought reasoning, which yields an effective monitoring strategy that can mitigate unsafe behavior with a low false positive rate.
|
| 789 |
TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science
2609.35606
|
cs.CL
|
Chutong Yang, Xiyuan Zhang, Yu Huang, Boran Han, Soonho Kong |
Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, prov...Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify computational improvements with arguments humans can inspect. We introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery, comprising 398 theorem-level challenges from 138 STOC and COLT 2026 papers. Expert-designed rules complete paper-specific context, preserve computational assumptions and quantitative guarantees, and withhold constructions when discovering an algorithm is part of the task. For each task, prover systems receive theorem statements and access to cited prior work. The pipeline supports fresh, versioned challenge batches from newly released papers. We evaluate ten model configurations from four families under direct inference and prover-verifier discussion, and compare four agent workflows under matched model-call opportunities. All evaluations use the full benchmark. In the model comparison, GPT-5.6 Sol max achieves the highest five-run verifier-accepted coverage at 23.6% after 10-round discussion. Discussion and repeated sampling improve coverage. In the separate agent comparison using GPT-5.5 xhigh, decomposition improves coverage over discussion, and agentic planning achieves the highest five-run verifier-accepted coverage at 25.4%. TCSAlgBench provides a refreshable testbed for measuring progress in model reasoning and studying how agent workflows support research-level proof discovery.
|
| 790 |
Simultaneous Translation between Sign Languages
2609.35608
|
cs.CL
|
Zetian Wu, Bowen Xie, Stefan Lee, Liang Huang |
Deaf and hard-of-hearing (DHH) signers cannot converse in real time across different sign languages today: existing sign-to-sign translation systems run offline, requiring the full source clip before any target sign is emitted. Live use cases - e.g. broadcast ...Deaf and hard-of-hearing (DHH) signers cannot converse in real time across different sign languages today: existing sign-to-sign translation systems run offline, requiring the full source clip before any target sign is emitted. Live use cases - e.g. broadcast interpretation and two-way video calls - instead demand simultaneous output, while the source signer is still signing. We present, to our knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision. We further introduce ca-Stream-AL, a computation-aware latency metric for streaming output. Averaged across six S2S directions on both a smaller human-verified test set and a larger synthetic S2S corpus, our streaming system achieves a 38% ca-Stream-AL reduction while staying within a 9% DTW-PA-MPJPE increase and a 2.1 BLEU-4 drop compared to the full-sentence baseline. A word-order case study probes how the streaming model handles word order mismatch between different sign languages - a consequence of simultaneous translation.
|
| 791 |
Twist, Don't Tilt: Trajectory-Exact Constrained Decoding for Masked Diffusion Models
2609.35609
|
cs.CLcs.LG
|
Aditya Thimmaiah, Lara Marinov, Jayanth Srinivasa, Haris Vikalo, Junyi Jessy Li |
Constrained decoding for Masked Diffusion Language Models (MDLMs) aims to ensure that generated outputs satisfy a specified structure or syntax constraint. MDLMs generate outputs by repeatedly unmasking masked positions present in their current state. Recent s...Constrained decoding for Masked Diffusion Language Models (MDLMs) aims to ensure that generated outputs satisfy a specified structure or syntax constraint. MDLMs generate outputs by repeatedly unmasking masked positions present in their current state. Recent strategies for constrained decoding constrain the model's per-step mean-field posterior (which factorizes over masked positions) by enforcing the desired constraint with an automaton. The resulting chain-structured factor graph allows exact constrained sampling via dynamic programming. However, despite each draw being exact and constraint-satisfying, we prove that their composition, in general, tilts away from the model's relative probabilities over valid trajectories, thus leading to trajectory bias. We derive an exact expression for this bias as a product of ratios measuring how valid continuation mass changes when the denoiser is reconditioned, and characterize when the bias vanishes. We then correct the bias by introducing TWISTER, the first automaton-twisted Sequential Monte Carlo decoder for MDLMs, using the step-exact decoder as the proposal. We show that for regular language constraints, the Feynman-Kac correction is exactly computable, with the twists obtained efficiently using quantities pre-computed for step-exact sampling. We prove that the resulting Feynman-Kac model targets the unbiased Doob h-transformed path law conditioned on constraint satisfaction.
|
| 792 |
SANTA++: Sampling Attention through Representative Keys
2609.35629
|
cs.CLcs.LG
|
Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee, Avinash Lohitsa |
Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative...Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusion probability. This importance sampling correction estimates attention over the full cache, with a sampling budget that lets us trade memory reads for accuracy. Remarkably, with 32 or 64 sampled teams, SANTA++ uses 16% to 22% of dense attention's KV reads and retains 94% to 99% of the dense-attention baseline's scores on LongBench v2 and HELMET's retrieval-augmented generation subset, and 85% to 91% on RULER, with Qwen2.5-7B-Instruct at 32K context. With 31 sampled teams, our GPU implementation delivers a $1.69\times$ attention speedup over the dense FlashAttention baseline at 32K context. By reducing the number of cache entries read, SANTA++ in principle complements architectures with compressed KV representations, such as multi-head latent attention. Our kernels are available at: https://github.com/OPUSLab/santapp-kernel-demo.git.
|
| 793 |
CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
2609.35645
|
cs.CLcs.LG
|
Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin, Fanny Riols |
Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational ...Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how code-switching transcription errors propagate to downstream voice agent task failures. In this work, we propose (1) a CS-ASR synthetic benchmark and multidimensional evaluation framework tailored to enterprise domains, (2) systematic evaluation of frontier ASR systems across 5 language pairs, (3) diagnostic analysis of the additional transcription errors that code-switching introduces across language pairs and models. We release COSE-E to support enterprise-focused CSASR evaluation for multilingual voice agents in enterprise deployment.
|
| 794 |
Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models
2609.35695
|
cs.CLcs.LG
|
Qiyao Ma, Junshan Zhang, Zhe Zhao |
Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped perfor...Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.
|
| 795 |
Reinforcing Agentic Creativity in Scientific Ideation with Night Science
2609.35706
|
cs.CL
|
Priyanka Kargupta, Silviu Cucerzan, Shweti Mahajan, Allen Herring, Jiawei Han |
Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative sp...Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.
|
| 796 |
Shockingly Simple Self-retrospection Improves Agentic Models Without RL
2609.35741
|
cs.CL
|
Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza |
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its...People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
|
| 797 |
How to Loop MoE: Flatten the Experts, Untie the Attention
2609.35751
|
cs.CLcs.LG
|
Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly, Wang Yang |
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each ...Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.
|
| 798 |
LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization
2407.00740
|
cs.CLcs.LG
|
Hye Ryung Son, Saehee Eom, Mooho Song, Jay-Yoon Lee |
As large language models (LLMs) are widely adopted in real-world applications, it has become critical to ensure LLMs satisfy safety constraints, such as non-toxicity and logical consistency, as well as task- and situation-specific constraints. Controlling the ...As large language models (LLMs) are widely adopted in real-world applications, it has become critical to ensure LLMs satisfy safety constraints, such as non-toxicity and logical consistency, as well as task- and situation-specific constraints. Controlling the output through instructions is a simple and tempting approach; however, it remains brittle, is opaque in how it influences model behavior, and thus cannot reliably ensure constraint satisfaction. Moreover, most recent controlled text generation (CTG) methods require access to the internal components of language models--such as weights or logits--making them incompatible with popular API-based LLMs. In this work, we propose LaSEr-Edit, a constraint-satisfying text revision method that can be applied to any LLMs, black- or white-box. We first find that lightweight, task-specific energy-based models (EBMs) achieve error-localization performance competitive with or even better than that of much larger LLMs, while operating substantially faster. Based on this finding, we propose two variants of text revision methods that incorporate energy-based error localization: LaSEr-LLM Edit, which instructs an LLM to edit text given EBM-predicted error spans, and LaSEr-EBM Edit, which uses the EBM not only for localization but also for editing by reranking edit candidates. Through experiments in diverse single-constraint control tasks, we show that LaSEr-LLM Edit controls text better than plain LLM-based editing in most of the tasks. We also find that LaSEr-EBM Edit further improves the control performance of LaSEr-LLM Edit and achieves among the strongest controllability across all tasks. Furthermore, we find that LaSEr-Edit, especially LaSEr-EBM Edit, performs well even when multiple constraints are controlled simultaneously.
|
| 799 |
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
2503.05061
|
cs.CL
|
Michael Krumdick, Charles Lovering, Varshini Reddy, Seth Ebner, Chris Tanner |
Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is appea...Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is appealing due to its scalability, low cost, and strong correlations with human stylistic preferences. However, it remains unclear how accurately these methods can assess response quality in domains where correctness matters more than style. To address this gap, we introduce the Business and Finance Fundamentals Benchmark (BFF-Bench), a dataset of 160 challenging questions and long-form responses authored by financial professionals. These experts subsequently evaluated the correctness of 1,200 responses generated by a diverse set of LLMs on both BFF-Bench and a challenging subset of MT-Bench. With this expert-annotated dataset of judgments (VERDICTS), we analyze the agreement between a suite of automated grading methods and human experts. While we observe that LLM Judges are more reliable than other grading methods, our findings reveal a clear pattern in LLM Judge performance: when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves. We demonstrate that providing the judges with expert-written references largely mitigates this issue, highlighting the limits of using LLM-as-a-Judge without any form of human verification.
|
| 800 |
ALAS: An Automatic Latent Alignment Score for Audio Language Models
2505.19937
|
cs.CL
|
Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli, Cem Subakan |
Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) behavior. Yet despite a growth of fusion strategies, there is no standard way to mea...Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) behavior. Yet despite a growth of fusion strategies, there is no standard way to measure how well a Speech-LLM internally binds audio frames to text tokens. We introduce ALAS (Automatic Latent Alignment Score), a model- and task-agnostic metric that probes the LLM's per-layer hidden states, scoring the cross-modal cosine similarity between audio and text representations against a Whisper-derived reference. ALAS needs only a frozen forward pass and an off-the-shelf ASR reference, with no training or fitted classifier, and is calibrated to an interpretable uniform baseline comparable across tasks. Applying ALAS to four open-source Speech-LLMs (AF3, Qwen2-Audio, Qwen-Omni, SALMONN) across emotion recognition (IEMOCAP), open-ended SQA (LibriSQA), and multi-choice audio understanding (MMAU-speech), we find that the depth and strength of alignment reflect each model's audio-encoder design and the acoustic-versus-semantic demands of the task, and that ALAS tracks but does not duplicate task accuracy, exposing models that score well without genuinely grounding in the audio. We release ALAS as an open-source library so that practitioners can probe their own Speech-LLMs or try new tasks.
|
| 801 |
Adaptive Activation Steering for Efficient LLM Reasoning via Closed-Loop PID Control
2506.18831
|
cs.CL
|
Aryasomayajula Ram Bharadwaj |
Reasoning LLMs trained with long chain-of-thought often overthink: they spend tokens on redundant reflection and transitions that inflate cost without improving accuracy. Static activation steering (e.g.\ SEAL) suppresses such content with a fixed vector, but ...Reasoning LLMs trained with long chain-of-thought often overthink: they spend tokens on redundant reflection and transitions that inflate cost without improving accuracy. Static activation steering (e.g.\ SEAL) suppresses such content with a fixed vector, but applies the same strength regardless of how redundant the current chunk actually is. We describe PID-steering, a training-free, decoding-time method that modulates the steering strength with a PID controller driven by a lightweight chunk-level redundancy classifier. On a subset of GSM8K with DeepSeek-R1-Distill-Qwen-1.5B, the method improves accuracy from 85.7\% to 89.6\% (+3.9 pp) while cutting average output length from 1026 to 790 tokens ($-$23\%). We report it as a small-scale proof of concept rather than a benchmark result.
|
| 802 |
RooseBERT: A New Deal For Political Language Modelling
2508.03250
|
cs.CL
|
Deborah Dore, Elena Cabrio, Serena Villata |
The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to automatically analyse such content with the final goal of lightening up political deliberation to citizens. However, the spec...The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to automatically analyse such content with the final goal of lightening up political deliberation to citizens. However, the specificity of the political language and the argumentative form of these debates (employing hidden communication strategies and leveraging implicit arguments) make this task very challenging, even for current general-purpose pre-trained Language Models (PLMs). To address this, we introduce a novel PLM for political discourse language called RooseBERT. Pre-training a language model on a specialised domain presents different technical and linguistic challenges, requiring extensive computational resources and large-scale data. RooseBERT has been trained on large political debate and speech corpora (11GB) in English. To evaluate its performances, we fine-tuned it on multiple downstream tasks related to political debate analysis, i.e., stance detection, sentiment analysis, argument component detection and classification, argument relation prediction and classification, policy classification, named entity recognition (NER). Our results show improvements over general-purpose PLMs on the majority of these tasks, highlighting how domain-specific pre-training enhances performance in political debate analysis. We release RooseBERT for the research community: https://huggingface.co/collections/MARIANNE-INRIA/roosebert.
|
| 803 |
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
2508.15851
|
cs.CL
|
Jiwon Park, Seohyun Pyeon, Jinwoo Kim, Rina Carines Cabral, Zhenuyan He |
Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information seeking: synthesizing multimodal evidence scattered across multiple documents and structural formats. Existing ...Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information seeking: synthesizing multimodal evidence scattered across multiple documents and structural formats. Existing QAs remain narrow in scope, relying on unimodal text and short-span reasoning that fail to capture the complexity of real information-seeking. We introduce DocHop-QA, a benchmark of 11,379 instances for evaluating multimodal, multi-document, multi-hop scientific QA. Built from publicly available PubMed articles, DocHop-QA incorporates textual passages, tables, and layout cues, enabling cross-document inference without explicit hyperlinks. To scale realistic QA construction, we develop an LLM-driven generation pipeline grounded in 11 scientific reasoning concepts, producing diverse and coherent question-answer pairs. To highlight the utility and versatility of the dataset, we propose a task-driven evaluation framework spanning four settings, including generative answering, multimodal evidence integration and structured index prediction. Experiments show that current models struggle with DocHop-QA's long-context, multi-evidence demands, establishing it as a rigorous testbed for advancing next-generation scientific QA systems.
|
| 804 |
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content
2509.12672
|
cs.CL
|
Shaz Furniturewala, Arkaitz Zubiaga |
Adversarial perturbations can reduce state-of-the-art toxicity classifiers to near-zero accuracy, yet existing defences treat models as black boxes. We apply mechanistic interpretability to toxicity classification for the first time, identifying the internal a...Adversarial perturbations can reduce state-of-the-art toxicity classifiers to near-zero accuracy, yet existing defences treat models as black boxes. We apply mechanistic interpretability to toxicity classification for the first time, identifying the internal attention-head circuits responsible for both correct classification and adversarial vulnerability. Across a 2$\times$2 factorial study (BERT $\times$ RoBERTa) $\times$ (Jigsaw $\times$ ToxiGen), extended to Llama Guard~2 (8B), we show that zeroing a single attention head recovers up to 70.4 pp of adversarial accuracy for RoBERTa on Jigsaw and 37.3 pp for Llama Guard 2 on ToxiGen, at $\leq$0.6 pp clean cost. Vulnerable heads generalise to held-out examples within $\leq$1 pp, and a class-imbalance sweep confirms they act as selective toxic-class detectors. Head suppression matches or outperforms adversarial training on Jigsaw; data augmentation dominates on ToxiGen: a dataset-specific reversal explained by whether the classifier encodes a concentrated bottleneck or a distributed circuit. Demographic analysis across 20 Jigsaw and 13 ToxiGen minority groups reveals structurally unequal adversarial vulnerability, exposing mechanistically traceable fairness gaps in current toxicity classifiers.
|
| 805 |
One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
2509.21443
|
cs.CL
|
Sualeha Farid, Jayden Lin, Zean Chen, Shivani Kumar, David Jurgens |
Large Language Models (LLMs) are increasingly deployed across multilingual and multicultural settings, yet it remains unclear whether changing language leads models to adopt community-specific moral reasoning or merely changes how shared learned abstractions a...Large Language Models (LLMs) are increasingly deployed across multilingual and multicultural settings, yet it remains unclear whether changing language leads models to adopt community-specific moral reasoning or merely changes how shared learned abstractions are expressed. We conduct a controlled multilingual evaluation across six geographically, culturally, and linguistically diverse languages (Arabic, Chinese, English, Hindi, Russian, and Spanish), using parallel moral reasoning benchmarks with English-origin, Chinese-origin, and natively elicited ground-truth judgments. Across 13 open-weight LLMs spanning 2B-70B parameters, we find substantial cross-lingual divergence in moral judgments, with English generally achieving the highest performance even when ground-truth judgments originate in Chinese or are collected natively in each language. Yet the reasoning underlying these divergent judgments is considerably more convergent: Utilitarianism dominates in five of six languages, reasoning follows broadly shared stages, and language-specific moral-value associations correspond only sparsely and inconsistently to values measured in the corresponding human communities. Finally, a large-scale OLMoTrace analysis of pretraining data sources reveals little direct reproduction of training text across languages, while the corpus composition, training stage, and cultural provenance of retrieved training evidence vary substantially by response language. Thus, similar moral reasoning structures emerge even from heterogeneous and often linguistically localized training evidence. Our findings, collectively, reveal a central disconnect in multilingual moral reasoning: language changes models' moral judgments and the training evidence associated with their reasoning, but does not correspondingly localize the moral abstractions they apply.
|
| 806 |
In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
2509.22230
|
cs.CL
|
Jaehoon Kim, Kwangwook Seo, Dongha Lee |
Stronger learning signals do not necessarily produce a stronger small reasoning model. Off-policy supervised fine-tuning (SFT) supplies stronger traces that may be incompatible with the student. On-policy post-training instead remains constrained by the qualit...Stronger learning signals do not necessarily produce a stronger small reasoning model. Off-policy supervised fine-tuning (SFT) supplies stronger traces that may be incompatible with the student. On-policy post-training instead remains constrained by the quality of the student's own trajectories: reinforcement learning (RL) provides little learning signal because a small student rarely produces a correct rollout. Consequently, effective post-training must supply stronger reasoning in a form the student can learn from. We hypothesize that a small number of tokens with very low probability under the student can make an otherwise strong reasoning trace difficult to learn. To test this hypothesis, we introduce Interleaved-Policy Distillation (IPD), which generates SFT data balancing teacher guidance and student compatibility. During generation, the student replaces teacher proposals that have low probability under its own policy, allowing the teacher to continue reasoning from a prefix the student can support. IPD improves reasoning across model families, datasets, and varying degrees of student intervention; for example, it raises the seven-benchmark average from 24.55 to 29.72 on Qwen3-0.6B/s1K. It outperforms baselines with comparable training budgets and broadens problem-solving coverage, while ablations show that filtering traces or masking losses cannot match the gains from revising trajectories during generation. Analyses link transfer failures to the rare tokens the student finds least likely and show that the source of the tokens also matters. Small models reason better when they learn strong solutions in a form they can reach in their own words.
|
| 807 |
Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
2509.23166
|
cs.CL
|
Chenxing Wei, Hong Wang, Ying He, Fei Yu, Yao Shu |
Large Language Models (LLMs) employ multi-turn interaction as a fundamental paradigm for completing complex tasks. However, their performance often degrades in extended interactions, as they are typically trained on static, single-turn data, which hinders thei...Large Language Models (LLMs) employ multi-turn interaction as a fundamental paradigm for completing complex tasks. However, their performance often degrades in extended interactions, as they are typically trained on static, single-turn data, which hinders their ability to adapt to real-time user feedback. To address this limitation, we first propose a new paradigm: Test-Time Policy Adaptation for Multi-Turn Interactions (T2PAM), which utilizes user feedback from the ongoing interaction as a reward signal to estimate a latent optimal policy aligned with user preferences, then updates a small subset of parameters to steer the model toward this policy, ultimately enabling efficient in-conversation self-correction. We then introduce Optimum-Referenced One-Step Adaptation (ROSA), a lightweight algorithm that operationalizes T2PAM. ROSA guides the model parameters toward a theoretical optimal policy in a single, efficient update step, avoiding costly iterative gradient-based optimization and minimizing computational overhead. We provide a rigorous theoretical analysis guaranteeing that the policy of ROSA converges to the preference of user as the number of interactions increases. Extensive experiments on challenging benchmark demonstrate that ROSA achieves significant improvements in both task effectiveness and efficiency.
|
| 808 |
Parallel Tokenizers: Rethinking Vocabulary Design in Cross-Lingual Transfer of Low-Resource Languages
2510.06128
|
cs.CL
|
Muhammad Dehan Al Kautsar, Fajri Koto |
Tokenization forms the basis of multilingual language models, yet existing methods often limit cross-lingual transfer by mapping semantically equivalent words to different embeddings. For example, 'I eat rice' in English and 'Ina cin shinkafa' in Hausa are typ...Tokenization forms the basis of multilingual language models, yet existing methods often limit cross-lingual transfer by mapping semantically equivalent words to different embeddings. For example, 'I eat rice' in English and 'Ina cin shinkafa' in Hausa are typically mapped to different vocabulary indices, preventing shared representations and limiting cross-lingual generalization. This problem is even more pronounced in low-resource languages, where shared representations could offer the greatest benefit. We introduce parallel tokenizers, a new framework that first trains tokenizers monolingually and then aligns their vocabularies exhaustively using bilingual dictionaries or word-to-word translation. This alignment enforces a shared semantic space across languages while naturally improving fertility balance. To assess their effectiveness, we pretrain a transformer encoder from scratch on thirteen low-resource languages and evaluate it on sentiment analysis, hate speech detection, emotion classification, and sentence embedding similarity. Across all tasks, models trained with parallel tokenizers outperform conventional multilingual baselines, confirming that rethinking tokenization is essential for advancing multilingual representation learning--especially in low-resource settings.
|
| 809 |
Large Language Model Selection with Limited Annotations
2510.09418
|
cs.CLcs.LG
|
Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel |
Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop SELECT-LLM, the first framework for active ...Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop SELECT-LLM, the first framework for active model selection of LLMs. SELECT-LLM aims to find a small set of queries whose annotations are most informative for identifying the best LLM for a given task. To this end, we introduce a query selection rule based on expected information gain, computed from pairwise similarities between candidate model outputs. Because this rule only uses generated model responses, SELECT-LLM can be applied across candidate models without assumptions about their architecture or access to model weights. This makes it suitable for both open-weight and black-box LLMs. We evaluate SELECT-LLM across 23 datasets, 156 evaluated models, diverse task families, and multiple text evaluation metrics. Across all experiments, SELECT-LLM improves over the strongest baseline in every setting, with annotation cost reductions up to 81.8% for best model selection and up to 84.78% for near-best model selection.
|
| 810 |
NOSA: Native and Offloadable Sparse Attention
2510.13602
|
cs.CLcs.LG
|
Yuxiang Huang, Pengjie Wang, Jicheng Han, Weilin Zhao, Zhou Su |
Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache. Prior training-free KV cache offloading alleviates this by keeping redundant context on the CPU and fetching only a...Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache. Prior training-free KV cache offloading alleviates this by keeping redundant context on the CPU and fetching only a sparse subset for attention, but it often degrades long-generation quality due to training-inference mismatch on sparse patterns. Meanwhile, trainable sparse attention is incompatible with efficient offloading, as unconstrained KV accesses may force large CPU-to-GPU transfers and erase throughput gains. To this end, we propose NOSA, a trainable sparse attention mechanism natively designed for KV cache offloading. NOSA explicitly constrains the volume of CPU-GPU KV transfers, thereby achieving low communication overhead and high decoding throughput. We further build NOSI, a KV cache offloading inference system that fully unlocks NOSA's efficiency. Empirical results on 1,3,8B LLMs demonstrate that NOSA outperforms KV cache offloading baselines on general, long-input, and long-generation tasks, while boosting decoding throughput by up to 5.04x, 1.92x, and 1.83x over FullAttn, InfLLMv2, and ShadowKV, respectively. We release our code at https://github.com/thunlp/NOSA.
|
| 811 |
R2T: Rule-Encoded Loss Functions for Sequence Tagging in Low-Resource Languages
2510.13854
|
cs.CLcs.LG
|
Mamadou K. Keita, Christopher Homan, Sebastien Diarra |
We introduce Rule-to-Tag (R2T), a framework that turns linguistic rules into the training signal for neural sequence taggers in low-resource languages. R2T encodes lexical, morphological, and syntactic rules as differentiable loss terms, so that a tagger learn...We introduce Rule-to-Tag (R2T), a framework that turns linguistic rules into the training signal for neural sequence taggers in low-resource languages. R2T encodes lexical, morphological, and syntactic rules as differentiable loss terms, so that a tagger learns from rules and unlabeled text, with no labeled training data. It also adds an out-of-vocabulary (OOV) loss term that discourages confident predictions on words that no rule covers. R2T is a first instance of a broader paradigm we call principled learning (PrL): using explicit principles as a learning signal, rather than relying on example-based supervision alone. We evaluate R2T on part-of-speech (POS) tagging for Zarma (Songhay), Bambara (Mande), and French (Romance), and on named entity recognition (NER) for Zarma. On Zarma POS tagging, R2T-BiLSTM uses no labeled training data, yet it reaches 0.968 Macro F1. This score comes within 0.007 of a supervised BiLSTM-CRF trained on 300 labeled sentences (0.975) and exceeds AfriBERTa fine-tuned on the same sentences (0.941). On NER, R2T works well as pre-training: after R2T pre-training, a model fine-tuned on 50 labeled sentences outperforms AfriBERTa fine-tuned on 300 (0.83 vs.\ 0.79 span F1). On Bambara, R2T reaches 0.91 Macro F1 with rules written in about 2.75 hours, while a supervised Masakhane tagger reaches 0.78. We release ZarmaPOS-Bench, a silver-standard Zarma POS corpus, together with ZarmaNER-600 and all trained models, to support future work on under-resourced languages.
|
| 812 |
Mitigating Hallucination in Large Language Models: A Capability-Oriented Survey on RAG, Reasoning, and Agentic Systems
2510.24476
|
cs.CL
|
Yihan Li, Xiyuan Fu, Ghanshyam Verma, Paul Buitelaar, Mingming Liu |
Hallucination remains one of the key obstacles to the reliable deployment of large language models (LLMs). Although various mitigation approaches have been proposed, existing studies often analyze different technical paradigms independently, lacking a unified ...Hallucination remains one of the key obstacles to the reliable deployment of large language models (LLMs). Although various mitigation approaches have been proposed, existing studies often analyze different technical paradigms independently, lacking a unified perspective to understand the underlying mechanisms of different approaches and their correspondence with different types of hallucinations. This survey adopts a capability enhancement perspective to systematically examine hallucination mitigation approaches, focusing on Retrieval-Augmented Generation (RAG), reasoning enhancement, and their integration within agentic systems. Based on their primary mitigation mechanisms, we categorize hallucinations into knowledge-based hallucinations and logic-based hallucinations, analyze how RAG and reasoning enhancement methods respectively improve knowledge acquisition and reasoning reliability, and further discuss the integration mechanisms of retrieval and reasoning capabilities in Agentic Systems for mitigating composite hallucinations. By considering the applicability, mitigation mechanisms, and limitations of different approaches, this survey establishes a unified analytical framework connecting hallucination types, key capability dimensions, and technical paradigms.
|
| 813 |
Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
2511.12001
|
cs.CL
|
Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan, Gunhee Kim |
Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of-Thought (CoT) explanations in multimoda...Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of-Thought (CoT) explanations in multimodal moral scenarios by systematically perturbing reasoning chains and manipulating delivery tones. Specifically, we analyze reasoning errors in vision language models (VLMs) and how they impact user trust and the ability to detect errors. Our findings reveal two key effects: (1) users often equate trust with outcome agreement, sustaining reliance even when reasoning is flawed, and (2) the confident tone suppresses error detection while maintaining reliance, showing that delivery styles can override correctness. These results highlight how CoT explanations can simultaneously clarify and mislead, underscoring the need for NLP systems to provide explanations that encourage scrutiny and critical thinking rather than blind trust. All code will be released publicly.
|
| 814 |
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
2512.00333
|
cs.CL
|
Ayush Maheshwari, Kaushal Sharma, Vivek Patel, Aditya Maheshwari |
While large language models excel on high-resource multilingual tasks, low- and extremely low-resource Indic languages remain severely under-evaluated. We present IndicParam, a human-curated benchmark of over 13,000 multiple-choice questions covering 11 such l...While large language models excel on high-resource multilingual tasks, low- and extremely low-resource Indic languages remain severely under-evaluated. We present IndicParam, a human-curated benchmark of over 13,000 multiple-choice questions covering 11 such languages (Nepali, Gujarati, Marathi, Odia as low-resource; Dogri, Maithili, Rajasthani, Sanskrit, Bodo, Santali, Konkani as extremely low-resource) plus Sanskrit-English code-mixed set. We evaluated 20 LLMs, both proprietary and open-weights, which reveals that even the top-performing Gemini-2.5 reaches 58% average accuracy, followed by GPT-5 (45) and DeepSeek-3.2 (43.1). We additionally label each question as knowledge-oriented or purely linguistic to discriminate factual recall from grammatical proficiency. Further, we assess the ability of LLMs to handle diverse question formats-such as list-based matching, assertion-reason pairs, and sequence ordering-alongside conventional multiple-choice questions. IndicParam provides insights into limitations of cross-lingual transfer and establishes a challenging benchmark for Indic languages. The dataset is available at https://huggingface.co/datasets/bharatgenai/IndicParam. Scripts to run benchmark are present at https://github.com/ayushbits/IndicParam.
|
| 815 |
MAPLE: Medical Aspect-Based Summarization with Phrase-Level Evidence
2601.03418
|
cs.CL
|
Bohao Chu, Hendrik Damm, Tabea M. G. Pakull, Sameh Frihat, Georg Lodde |
Trustworthy clinical summarization requires every claim to be traceable to its evidence, yet existing attribution often resolves only to the sentence or document, leaving clinicians to scan surrounding text for the few words that matter. We argue that the unit...Trustworthy clinical summarization requires every claim to be traceable to its evidence, yet existing attribution often resolves only to the sentence or document, leaving clinicians to scan surrounding text for the few words that matter. We argue that the unit of attribution should match the unit of verification: the precise phrase the reader's eye must land on. We present MAPLE (Medical Aspect-Based Summarization with Phrase-Level Evidence), a human-annotated benchmark that grounds each summarized claim in both cited sentences and contributory phrases within them. Spanning 152 randomized controlled trial (RCT) abstracts and 16 clinically motivated aspects, MAPLE comprises 1,799 aspect-based summaries with two-level evidence. We further introduce a decoupled evaluation framework that separately scores content, traceability, and locatability, together with a proxy for the amount of source text a clinician must inspect to verify a claim. Benchmarking eleven LLMs shows that sentence-level citation is consistently strong (C-F1 up to 90.9%), while phrase-level grounding remains less stable and the most discriminative axis across models (P-F1 66.1-84.5%). These results suggest that the key challenge is not only producing accurate summaries, but localizing their supporting evidence precisely enough for efficient clinical verification. Data and code are available at https://github.com/chubohao/maple.
|
| 816 |
eTracer: Polarity-Aware Evidence Grounding
2601.03669
|
cs.CL
|
Bohao Chu, Qianli Wang, Hendrik Damm, Shuning Zhang, Hui Wang |
Evidence grounding makes generated content verifiable by linking statements to their sources. However, effective verifiability requires two properties that existing methods fail to jointly provide: verification granularity, the ability to localize fine-grained...Evidence grounding makes generated content verifiable by linking statements to their sources. However, effective verifiability requires two properties that existing methods fail to jointly provide: verification granularity, the ability to localize fine-grained evidence for each statement, and verification completeness, the ability to capture both support and contradiction. Existing methods either operate at granularities too coarse for rapid inspection or are support-only. We formalize polarity-aware evidence grounding: linking each response sentence to its contextual evidence with a supporting, contradicting, or irrelevant label. We construct eTracer-Bench, a human-annotated benchmark with 7,905 supporting and contradicting sentence pairs across seven domains, covering both explicit and implicit contradictions. We further propose eTracer, a post-hoc framework whose context-aware signed scorer recovers the full evidence map in a single forward pass per response, with localization and polarity emerging jointly from one mechanism. eTracer achieves the best overall Citation F1 of 82.39, driven by a Contradict F1 of 89.20. A pre-registered within-subjects user study (n = 30) shows that polarity-aware grounding nearly doubles contradiction detection over the document-level baseline (27.9% to 51.9%, Cohen's d = 0.82, p = 0.0003), and suggests a false-completeness bias under which support-only grounding fails to improve over this baseline (23.4% vs. 27.9%). Dataset and code: https://github.com/chubohao/eTracer.
|
| 817 |
PEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation
2601.04692
|
cs.CL
|
Naquee Rizwan, Subhankar Swain, Paramananda Bhaskar, Shehryaar Shah Khan, Gagan Aryan |
In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them before being posted - by applying a range of strategies built on top of generative AI models. To the best of our...In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them before being posted - by applying a range of strategies built on top of generative AI models. To the best of our knowledge, explanation and intervention have typically been studied separately from detection, which does not reflect real-world conditions. Further, since curating large annotated datasets for meme moderation is prohibitively expensive, we propose a novel framework - PEST - that leverages task-specific generative VLMs and the few-shot adaptability of large VLMs to cater to different types of memes. We believe this is the first work focused on generalizable hateful meme moderation under limited data conditions, and has strong potential for deployment in real-world production scenarios. Warning: Contains potentially toxic contents.
|
| 818 |
JuDi: Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence
2601.04766
|
cs.CL
|
Shengyin Sun, Yiming Li, Renxi Liu, Weizhe Lin, Hui-Ling Zhen |
Judge Decoding accelerates LLM inference by relaxing the strict verification of Speculative Decoding, yet it typically relies on expensive and noisy supervision. In this work, we revisit this paradigm from first principles, revealing that the ``criticality'' s...Judge Decoding accelerates LLM inference by relaxing the strict verification of Speculative Decoding, yet it typically relies on expensive and noisy supervision. In this work, we revisit this paradigm from first principles, revealing that the ``criticality'' scores learned via costly supervision are intrinsically encoded in the draft-target distributional divergence. We theoretically prove a structural correspondence between learned linear judges and Kullback-Leibler (KL) divergence, demonstrating they rely on the same underlying logit primitives. Guided by this, we propose a simple, training-free verification mechanism based on KL divergence. Extensive experiments across reasoning and coding benchmarks show that our method matches or outperforms complex trained judges (e.g., AutoJudge), offering superior robustness to domain shifts and eliminating the supervision bottleneck entirely. Code is available at https://github.com/sunshy-1/JuDi
|
| 819 |
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
2601.06426
|
cs.CL
|
Robert J. Moore, Sungeun An, Farhan Ahmed, Jay Pankaj Gala |
Existing LLM benchmarks evaluate what models say, such as whether answers are correct, faithful, or helpful, but they do not test whether models produce the right type of conversational action at the right point in an interaction. The Natural Conversation Benc...Existing LLM benchmarks evaluate what models say, such as whether answers are correct, faithful, or helpful, but they do not test whether models produce the right type of conversational action at the right point in an interaction. The Natural Conversation Benchmark (NC-Bench) fills this gap by evaluating conversational competence: the ability to perform structurally appropriate actions such as repairing, closing, or refusing, as defined by conversation science. Grounded in the Natural Conversation Framework (NCF), NC-Bench comprises three sets: (1) the basic set evaluates fundamental sequence management practices, such as answering inquiries, repairing responses, and closing conversational pairs; (2) the retrieval-augmented generation (RAG) set applies the same patterns but incorporates information-seeking via RAG; (3) the complex request set extends to requests involving more intricate sequence management. Each set tests a model's ability to produce contextually appropriate conversational actions in response to characteristic interaction patterns. Evaluations across six open-source models and one closed-source model on 14 interaction patterns reveal quantifiable shortcomings in conversational competence present in current models. By operationalizing fundamental principles of human conversation, NC-Bench provides a lightweight, extensible, and theory-grounded framework for identifying specific conversational action gaps in LLMs beyond topical or task-specific benchmarks.
|
| 820 |
Untangling Input Language from Reasoning Language: A Diagnostic Framework for Cross-Lingual Moral Alignment in LLMs
2601.10257
|
cs.CL
|
Nan Li, Bo Kang, Tijl De Bie |
When LLMs judge moral dilemmas, do they reach different conclusions in different languages, and if so, why? Two factors could drive such differences: the language of the dilemma itself, or the language in which the model reasons. Standard evaluation conflates ...When LLMs judge moral dilemmas, do they reach different conclusions in different languages, and if so, why? Two factors could drive such differences: the language of the dilemma itself, or the language in which the model reasons. Standard evaluation conflates these by testing only matched conditions (e.g., English dilemma with English reasoning). We introduce a methodology that separately manipulates each factor, covering also mismatched conditions (e.g., English dilemma with Chinese reasoning), enabling decomposition of their contributions. To study \emph{what} changes, we propose an approach to interpret the moral judgments in terms of Moral Foundations Theory. As a side result, we identify evidence for splitting the Authority dimension into a family-related and an institutional dimension. Applying this methodology to English-Chinese moral judgment with 13 LLMs, we demonstrate its diagnostic power: (1) the framework isolates reasoning-language effects as contributing twice the variance of input-language effects; (2) it detects context-dependency in nearly half of models that standard evaluation misses; and (3) a diagnostic taxonomy translates these patterns into deployment guidance. We release our code and datasets at https://anonymous.4open.science/r/CrossCulturalMoralJudgement.
|
| 821 |
NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems
2601.11004
|
cs.CL
|
Jiayu Liu, Rui Wang, Qing Zong, Yumeng Wang, Cheng Qian |
Accurately assessing model confidence is essential for deploying large language models (LLMs) in mission-critical factual domains. While retrieval-augmented generation (RAG) is widely adopted to improve grounding, confidence calibration in RAG settings remains...Accurately assessing model confidence is essential for deploying large language models (LLMs) in mission-critical factual domains. While retrieval-augmented generation (RAG) is widely adopted to improve grounding, confidence calibration in RAG settings remains poorly understood. We conduct a systematic study across four benchmarks, revealing that LLMs exhibit poor calibration performance especially when noisy contexts are retrieved. Specifically, contradictory or irrelevant evidence tends to exacerbate the model's overconfidence issue. To address this, we propose NOVA Rules (NOise-Aware Verbal Confidence CAlibration Rules) to provide a principled foundation for resolving overconfidence under noise. We further design NOVA, a noise-aware calibration framework that synthesizes supervision from ~2K HotpotQA examples guided by these rules. By performing supervised fine-tuning (SFT) with this data, NOVA equips models with intrinsic noise awareness without relying on stronger teacher models. Empirical results show that NOVA yields substantial gains, improving ECE scores by 10.9% in-domain and 8.0% out-of-domain. By bridging the gap between retrieval noise and verbal calibration, NOVA paves the way for both accurate and epistemically reliable LLMs.
|
| 822 |
A Scalable Entity-Based Framework for Auditing Bias in Large Language Models
2601.12374
|
cs.CL
|
Akram Elbouanani, Aboubacar Tuo, Adrian Popescu |
Existing approaches to bias evaluation in large language models (LLMs) trade ecological validity for statistical control, relying either on artificial prompts that poorly reflect real-world use or on naturalistic tasks that lack scale and rigor. We introduce a...Existing approaches to bias evaluation in large language models (LLMs) trade ecological validity for statistical control, relying either on artificial prompts that poorly reflect real-world use or on naturalistic tasks that lack scale and rigor. We introduce a scalable bias-auditing framework that uses named entities as controlled probes to measure systematic disparities in model behavior. Synthetic data enables us to construct diverse, controlled inputs, and we show that it reliably reproduces bias patterns observed in natural text, supporting its use for large-scale analysis. Using this framework, we conduct the largest bias audit to date, comprising 1.9 billion data points across multiple entity types, tasks, languages, models, and prompting strategies. We find consistent patterns: models penalize right-wing politicians and favor left-wing politicians, prefer Western and wealthier countries over the Global South, favor Western companies, and penalize firms in the defense and pharmaceutical sectors. While instruction tuning reduces bias, increasing model scale amplifies it, and prompting in Chinese or Russian does not mitigate Western-aligned preferences. These findings highlight the need for systematic bias auditing before deploying LLMs in high-stakes applications. Our framework is extensible to other domains and tasks, and we make it publicly available to support future work.
|
| 823 |
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
2601.13722
|
cs.CL
|
Yulin Hu, Zimo Long, Jiahe Guo, Xingyu Sui, Xing Fu |
Memory-augmented conversational agents enable personalized interactions using long-term user memory and have gained substantial traction. However, existing benchmarks primarily focus on whether agents can recall and apply user information, while overlooking wh...Memory-augmented conversational agents enable personalized interactions using long-term user memory and have gained substantial traction. However, existing benchmarks primarily focus on whether agents can recall and apply user information, while overlooking whether such personalization is used appropriately. In fact, agents may overuse personal information, producing responses that feel forced, intrusive, or socially inappropriate to users. We refer to this issue as \emph{over-personalization}. In this work, we formalize over-personalization into three types: Irrelevance, Repetition, and Sycophancy, and introduce \textbf{OP-Bench} a benchmark of 1,700 verified instances constructed from long-horizon dialogue histories. Using \textbf{OP-Bench}, we evaluate multiple large language models and memory-augmentation methods, and find that over-personalization is widespread when memory is introduced. Further analysis reveals that agents tend to retrieve and over-attend to user memories even when unnecessary. To address this issue, we propose \textbf{Self-ReCheck}, a lightweight, model-agnostic memory filtering mechanism that mitigates over-personalization while preserving personalization performance. Our work takes an initial step toward more controllable and appropriate personalization in memory-augmented dialogue systems.
|
| 824 |
AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains
2601.15511
|
cs.CL
|
Adam Szelestey, Sofie van Engelen, Tianhao Huang, Justin Snelders, Qintao Zeng |
Hallucination in large language models (LLMs) remains an acute concern, contributing to the spread of misinformation and diminished public trust, particularly in high-risk domains. Among hallucination types, factuality is crucial, as it concerns a model's alig...Hallucination in large language models (LLMs) remains an acute concern, contributing to the spread of misinformation and diminished public trust, particularly in high-risk domains. Among hallucination types, factuality is crucial, as it concerns a model's alignment with established world knowledge. Adversarial factuality, defined as the deliberate insertion of misinformation into prompts with varying levels of expressed confidence, tests a model's ability to detect and resist confidently framed falsehoods. Existing work lacks high-quality, domain-specific resources for assessing model robustness under such adversarial conditions, and no prior research has examined the impact of injected misinformation on long-form text factuality. To address this gap, we introduce AdversaRiskQA, the first verified and reliable benchmark systematically evaluating adversarial factuality across Health, Finance, and Law. The benchmark includes two difficulty levels to test LLMs' defensive capabilities across varying knowledge depths. We propose two automated methods for evaluating the adversarial attack success and long-form factuality. We evaluate six open- and closed-source LLMs from the Qwen, GPT-OSS, and GPT families, measuring misinformation detection rates. Long-form factuality is assessed on Qwen3 (30B) under both baseline and adversarial conditions. Results show that after excluding meaningless responses, Qwen3 (80B) achieves the highest average accuracy, while GPT-5 maintains consistently high accuracy. Performance scales non-linearly with model size, varies by domains, and gaps between difficulty levels narrow as models grow. Long-form evaluation reveals no significant correlation between injected misinformation and the model's factual output. AdversaRiskQA provides a valuable benchmark for pinpointing LLM weaknesses and developing more reliable models for high-stakes applications.
|
| 825 |
MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation
2601.21225
|
cs.CLcs.AI
|
Tianyi Xu, Kosei Uemura, Alfred Malengo Kondoro, Tadesse Destaw Belay, Catherine Nana Nyaah Essuman |
Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high varianc...Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high variance when models are evaluated on different instantiations of the same question; however, the evaluation was conducted only in English. In this paper, we introduce MGSM-Pro, an extension of MGSM dataset with GSM-Symbolic approach. Our dataset provides five instantiations per MGSM question by varying names, digits and irrelevant context. Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations different from those in the original test set. We further find that models robustness in HRL setting do not necessarily translate to LRL. Moreover, proprietary models, such as Gemini 2.5 Flash and GPT-4.1 are less robust to digit, whereas Gemini 3.0 Pro is more robust. Among open models, GPT-OSS 120B and DeepSeek v3 show stronger robustness. Based on these findings, we recommend evaluating each problem using at least five digit-varying instantiations to obtain a more robust and realistic assessment of math reasoning.
|
| 826 |
Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents
2601.21699
|
cs.CL
|
Hojae Han, Heeyun Jung, Jongyoon Kim, Seung-won Hwang |
Reinforcement learning (RL) trains small language model agents to answer multi-hop questions by retrieving evidence over multiple turns, but reported gains typically rely on thousands of on-policy rollouts per update. We study RL for such agents under the budg...Reinforcement learning (RL) trains small language model agents to answer multi-hop questions by retrieving evidence over multiple turns, but reported gains typically rely on thousands of on-policy rollouts per update. We study RL for such agents under the budget constraint of commodity GPUs, where each update samples only a few rollouts per question. Under this constraint, most sampled trajectories retrieve none of the required evidence, so the outcome reward gives the policy little to learn from and small agents settle for answering without retrieval, a failure we call \emph{retrieval collapse}. David-GRPO addresses this with two mechanisms: (1) \emph{Expert trajectory seeding} places a handful of off-policy expert trajectories into the GRPO groups of the early updates, and (2) \emph{evidence-guided continuation} rewards evidence coverage and resumes the most promising partial trajectory. The evidence for each training question is constructed from the corpus link graph, so no annotated evidence is required. On six multi-hop QA benchmarks, David-GRPO trained on four RTX 3090 GPUs with 144 rollouts per step brings Qwen2.5-1.5B to 22.6 average EM against 11.9 for the best baseline under the same budget, matches Tree-GRPO trained with 20 times more rollouts, and, unlike the baselines that stop after at most one search, learns to retrieve across turns. The implementation is available at: https://github.com/AsadalJung/David-GRPO
|
| 827 |
OVD: On-policy Verbal Distillation
2601.21968
|
cs.CL
|
Jing Xiong, Hui Shen, Shansan Gong, Yuxin Cheng, Jianghan Shen |
Knowledge distillation transfers reasoning capabilities from large teachers to efficient students. However, token-level on-policy distillation (OPD) constrains student exploration and requires teacher token probabilities, precluding distillation from black-box...Knowledge distillation transfers reasoning capabilities from large teachers to efficient students. However, token-level on-policy distillation (OPD) constrains student exploration and requires teacher token probabilities, precluding distillation from black-box teachers that provide only text outputs. We introduce On-policy Verbal Distillation (OVD), a framework that uses verbal scores from black-box teachers to rank student-generated sub-trajectories, retaining high-scoring ones and replacing low-scoring ones with teacher-generated continuations. We analyze when ranking induced by verbal scores can guide distribution approximation: under a density-ratio calibration condition on acceptance probabilities and bounded teacher-replacement error, we bound the approximation error between the resulting mixed trajectory distribution and a teacher-preferred target. On Web Q&A, OVD achieves 41.09% average EM with teacher feedback at inference, exceeding the strongest evaluated baseline by 5.89 percentage points. On AMC23, OVD-FR improves accuracy over RLVR by 10.0 percentage points (52.5% to 62.5%) after 600 training steps on 128 problems. Further experiments suggest that retaining student-generated prefixes helps preserve exploration and mitigate trajectory-level entropy collapse. OVD also improves training efficiency: resampling selected suffixes rather than entire responses reduces mean per-step training time by 10.2% in the 128-problem setting. Project page: https://menik1126.github.io/ovd-project-page/.
|
| 828 |
Denoising Time Matters:Diverse Generation in Diffusion Language Models
2601.22629
|
cs.CL
|
Jingxuan Wu, Zhenglin Wan, Yuzhe Yang, Yiqiao Huang, Chubin Zhang |
Diffusion language models (Diffusion-LMs) generate text through iterative denoising, exposing a temporal structure that is largely absent from autoregressive decoding. In this paper, we show that this temporal structure provides a useful control axis for gener...Diffusion language models (Diffusion-LMs) generate text through iterative denoising, exposing a temporal structure that is largely absent from autoregressive decoding. In this paper, we show that this temporal structure provides a useful control axis for generation diversity: early denoising steps mainly determine high-level semantic trajectories, while later steps refine lexical realization. Motivated by this observation, we propose Time-Annealed Perturbation Sampling (TAPS), a training-free inference strategy that samples nearby conditioning trajectories through time-aware, manifold-constrained perturbations. TAPS encourages semantic branching during early denoising and anneals the perturbation away before refinement, improving exploration while preserving prompt alignment, generation quality, and reasoning ability. Experiments on multiple Diffusion-LM backbones, including non-autoregressive and semi-autoregressive models, show that TAPS consistently improves semantic and lexical diversity across open-ended and instruction-following generation tasks, while preserving reasoning ability on verifiable reasoning benchmarks with negligible overhead.
|
| 829 |
MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method
2602.02320
|
cs.CL
|
Feiyang Cai, Guijuan He, Yi Hu, Jingjing Wang, Joshua Luo |
Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential for enabling large language models (LLMs) to both understand structure for chemical reasoning and generate molecules fro...Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential for enabling large language models (LLMs) to both understand structure for chemical reasoning and generate molecules from natural-language design intent. However, the substantial cost of human annotation makes it infeasible to construct large-scale, high-quality datasets of structure-grounded descriptions. This work proposes a fully automated annotation framework for generating precise molecular descriptions at scale, such that the original molecule can be unambiguously reconstructed from the description alone. Our approach extends a rule-based chemical nomenclature parser to interpret IUPAC names and construct enriched, XML metadata that explicitly encodes molecular structure. This is then used to guide LLMs in producing accurate natural-language descriptions. Using this framework, we curate MolLangData, a dataset of approximately $163$k molecule--description pairs. A rigorous validation protocol combining expert human and LLM-based reconstruction on a subset of $2,000$ molecules demonstrates $98.6\%$ description precision. Using the curated dataset, we train a $4$B-parameter LLM via large-scale reinforcement learning for language-conditional molecule generation. The proposed framework and dataset provide a reliable foundation for molecule--language alignment, readily beneficial to broader chemical tasks.
|
| 830 |
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
2602.02464
|
cs.CL
|
Or Shafran, Shaked Ronen, Omri Fahn, Shauli Ravfogel, Atticus Geiger |
Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space. Existing approaches search for individual global directions, implicitly assuming linear separability, which overl...Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space. Existing approaches search for individual global directions, implicitly assuming linear separability, which overlooks concepts with nonlinear or multi-dimensional structure. In this work, we leverage Mixture of Factor Analyzers (MFA) as a scalable, unsupervised alternative that models the activation space as a collection of Gaussian regions with their local covariance structure. MFA decomposes activations into two compositional geometric objects: the region's centroid in activation space, and the local variation from the centroid. We train large-scale MFAs for Llama-3.1-8B and Gemma-2-2B, and show they capture complex, nonlinear structures in activation space. Moreover, evaluations on localization and steering benchmarks show that MFA outperforms unsupervised baselines, is competitive with supervised localization methods, and often achieves stronger steering performance than sparse autoencoders. Together, our findings position local geometry, expressed through subspaces, as a promising unit of analysis for scalable concept discovery and model control, accounting for complex structures that isolated directions fail to capture.
|
| 831 |
Investigating Learner-Aware Design of LLM-Generated Educational Feedback
2602.11650
|
cs.CL
|
Momoka Furuhashi, Kouta Nakayama, Noboru Kawai, Takashi Kodama, Saku Sugawara |
Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and coverage) to support answer revision and learner evaluations across learner profiles. We define six feedb...Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and coverage) to support answer revision and learner evaluations across learner profiles. We define six feedback designs for multiple-choice biology questions, including a baseline design and five variants with additional feedback elements, and conduct an empirical study with 321 high school students. We evaluate feedback using immediate revision performance and six subjective evaluation criteria, and analyze differences in subjective evaluations across learner profiles based on personality traits. Our results show that presenting task-relevant information clearly is associated with better immediate revision performance and is favorably evaluated across learner profiles, while we observe descriptive differences in evaluation patterns, particularly for informational novelty and affective framing. These findings support further investigation of personalized LLM feedback design.
|
| 832 |
Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration?
2603.03202
|
cs.CL
|
Dadi Guo, Yuejin Xie, Qingyu Liu, Weixian Huang, Jiayu Liu |
As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality problems has become a significant bottleneck for training, evaluation and self-evolution of LLMs. Simultaneousl...As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality problems has become a significant bottleneck for training, evaluation and self-evolution of LLMs. Simultaneously, recent code agents have demonstrated sophisticated skills in agentic coding and reasoning, suggesting that code execution can serve as a scalable environment for mathematical experimentation. In this paper, we investigate the potential of code agents to autonomously evolve existing math problems into more complex variations. We introduce a multi-agent framework designed to perform problem evolution while validating the solvability and increased difficulty of the generated problems. Our experiments demonstrate that, given sufficient test-time exploration, code agents can synthesize new, solvable problems that are structurally distinct from and more challenging than the originals. This work provides empirical evidence that code-driven agents can serve as a viable mechanism for synthesizing high-difficulty mathematical reasoning problems within scalable computational environments. Code and data is available at https://github.com/TarferSoul/Code2Math.
|
| 833 |
Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
2603.03415
|
cs.CL
|
Mingyu Jin, Yutong Yin, Jingcheng Niu, Qingcheng Zeng, Wujiang Xu |
In this work, we investigate how Large Language Models (LLMs) adapt their internal representations when encountering inputs of increasing difficulty, quantified as the degree of out-of-distribution (OOD) shift. We reveal a consistent and quantifiable phenomeno...In this work, we investigate how Large Language Models (LLMs) adapt their internal representations when encountering inputs of increasing difficulty, quantified as the degree of out-of-distribution (OOD) shift. We reveal a consistent and quantifiable phenomenon: as task difficulty increases, whether through harder reasoning questions, longer contexts, or adding answer choices, the last hidden states of LLMs become substantially sparser. In short, \textbf{\textit{the farther the shift, the sparser the representations}}. This sparsity--difficulty relation is observable across diverse models and domains, suggesting that language models respond to unfamiliar or complex inputs by concentrating computation into specialized subspaces in the last hidden state. Through a series of controlled analyses with a learning dynamic explanation, we demonstrate that this sparsity is not incidental but an adaptive mechanism for stabilizing reasoning under OOD. Leveraging this insight, we design \textit{Sparsity-Guided Curriculum In-Context Learning (SG-ICL)}, a strategy that explicitly uses representation sparsity to schedule few-shot demonstrations, leading to considerable performance enhancements. Our study provides new mechanistic insights into how LLMs internalize OOD challenges. The source code is available at the URL: https://github.com/MingyuJ666/sparsityLLM.
|
| 834 |
A theoretical model of dynamical grammatical gender shifting based on set-valued set function
2603.03510
|
cs.CLcs.AI
|
Mohamed El Idrissi |
This study investigates the diverse characteristics of nouns, focusing on both semantic (e.g., countable/uncountable) and morphosyntactic (e.g., masculine/feminine) distinctions. We explore inter-word variations for gender markers in noun morphology. Grammatic...This study investigates the diverse characteristics of nouns, focusing on both semantic (e.g., countable/uncountable) and morphosyntactic (e.g., masculine/feminine) distinctions. We explore inter-word variations for gender markers in noun morphology. Grammatical gender shift is a widespread phenomenon in languages around the world. The aim is to uncover the underlying patterns governing the variation of lexemes. To this end, we propose a new computational component dedicated to pairing items with morphological templates (e.g., the result of a generated item-template pair: (funas, $\{N, +SG, -PL, -M, +F, -COL, +SING\}$), with its spell-out form: $\eth$a-funast 'cow'). This process is formally represented by the Template-Based and Modular Cognitive model. This proposed model, defined by a set-valued set function $h : \mathscr{P}(M) \rightarrow \mathscr{P}(M)$, predicts the nonlinear dynamic mapping of lexical items onto morphological templates. By applying this formalism, we present a unified framework for understanding the complexities of morphological markings across languages. Through empirical observations, we demonstrate how these shifts, as well as non-gender shifts, arise during lexical changes, especially in Riffian. Our model posits that these variant markings emerge due to template shifts occurring during word and meaning formation. This study achieves two primary objectives. First, on the formal side, we prove the model's representational completeness in learning and prediction. Second, on the linguistic side, we challenge and broaden the conventional view of word formation by formally demonstrating that conversion is applicable to noun-to-noun derivation. This data-driven mathematical model not only contributes to a deeper understanding of morphosyntactic variation but also offers potential applications in other fields requiring precise modelling of linguistic patterns.
|
| 835 |
SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models
2603.04410
|
cs.CLcs.AI
|
Omar Abdelnasser, Fatemah Alharbi, Khaled Khasawneh, Ihsen Alouani, Mohammed E. Fouda |
While different stakeholders are trying to leverage Arabic Language Models (ALMs), safety alignment in ALMs remains largely underexplored, hindering their mainstream adoption. Existing safety benchmarks are predominantly English-centric and evaluate Arabic onl...While different stakeholders are trying to leverage Arabic Language Models (ALMs), safety alignment in ALMs remains largely underexplored, hindering their mainstream adoption. Existing safety benchmarks are predominantly English-centric and evaluate Arabic only in its standardized form, obscuring fine-grained safety vulnerabilities in Arabic NLP systems. This paper introduces SalamahBench, a unified benchmark of 8{,}270 human-verified harmful prompts across ML Commons hazard categories, each rendered in Modern Standard Arabic (MSA) and five regional Arabic varieties, namely Egyptian, Syrian, Saudi, Lebanese, and Moroccan, for a total of 49{,}620 paired instances. To analyze the resulting data, we introduce two complementary metrics, namely Dialect Shift, which measures a model's aggregate change in safety under dialectal reformulation, and Category-Specific Dialect Deviation, which isolates harm categories whose change departs from that aggregate trend. Evaluating models such as Fanar 2, ALLaM 2, and Karnak 1 under multiple safeguard configurations, we find that cross-variety robustness is strongly model dependent, and that aggregate scores can conceal category-level divergence. Our findings highlight the necessity of evaluating Arabic model safety jointly across linguistic varieties and harm domains rather than relying on aggregate scores or MSA alone.
|
| 836 |
IndexRAG: Index-Time Reasoning for Multi-Hop Retrieval-Augmented Generation
2603.16415
|
cs.CL
|
Zhenghua Bao, Yi Shi |
Multi-hop question answering (QA) requires reasoning across multiple documents, yet existing retrieval-augmented generation (RAG) approaches address this either through graph-based methods requiring additional online processing or iterative multi-step reasonin...Multi-hop question answering (QA) requires reasoning across multiple documents, yet existing retrieval-augmented generation (RAG) approaches address this either through graph-based methods requiring additional online processing or iterative multi-step reasoning. We present IndexRAG, a novel approach that shifts cross-document reasoning from online inference to offline indexing. IndexRAG identifies bridge entities shared across documents and generates bridging facts as independently retrievable units, requiring no additional training or fine-tuning. Experiments on three widely-used multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA, MuSiQue) show that IndexRAG improves F1 over Naive RAG by 4.6 points on average, while requiring only single-pass retrieval and a single LLM call at inference time. When combined with IRCoT, IndexRAG achieves the best average performance among all evaluated methods, including graph-based baselines such as HippoRAG2 and FastGraphRAG, while relying on a flat vector index. Our code is available at https://github.com/Continuum-AI-Corp/IndexRAG .
|
| 837 |
Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models
2603.20212
|
cs.CLcs.LG
|
Jiayun Wu, Peng Zhang, Yuanyuan Lu, Shan Qu, Ning Gu |
Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior accuracy through chain-of-thought (CoT) reasoning, they incur substantial computat...Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior accuracy through chain-of-thought (CoT) reasoning, they incur substantial computational costs. Conversely, Scalar Reward Models (SRMs) offer efficiency but suffer from limited performance and adaptability in complex scenarios. We introduce Fast-Slow Thinking Reward Models (F/S-RM), a hybrid RM architecture inspired by Dual Process Theory. It trains a single model to integrate two distinct reward paradigms: scalar-style first-token pairwise judgment (fast thinking) and CoT-based judgment (slow thinking), regulated by a dual-confidence activation mechanism that determines when to activate slow thinking. Under hybrid inference, F/S-RM achieves state-of-the-art accuracy with an average score of 84.3 across benchmarks, while reducing token consumption by 22.5%.
|
| 838 |
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
2603.21359
|
cs.CL
|
K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, Farig Sadeque |
Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. However, frameworks to quantify these disparities remain scarce. We propose a two-phase framework to evaluate dialectal bias, operationalize...Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. However, frameworks to quantify these disparities remain scarce. We propose a two-phase framework to evaluate dialectal bias, operationalized as comprehension degradation relative to standard Bengali, in LLM question-answering across nine Bengali dialects. First, we translate and gold-label standard Bengali questions into dialectal variants adopting a retrieval-augmented generation (RAG) pipeline to prepare 4,000 question sets. Since traditional translation quality evaluation metrics fail on unstandardized dialects, we evaluate fidelity using an LLM-as-a-judge, which human correlation confirms outperforms legacy metrics. Second, we benchmark 19 LLMs across these gold-labeled sets, running 68,395 RLAIF evaluations validated through multi-judge agreement and human fallback. Our findings reveal severe performance drops linked to linguistic divergence. For instance, responses to the highly divergent Chittagong dialect score 5.44/10, compared to 7.68/10 for Tangail. Furthermore, increased model scale does not consistently mitigate this bias. We contribute a validated translation quality evaluation method, a rigorous benchmark dataset, and a Critical Bias Sensitivity (CBS) metric for safety-critical applications.
|
| 839 |
Learning to Predict Future-Aligned Research Proposals with Language Models
2603.27146
|
cs.CL
|
Heng Wang, Pengcheng Jiang, Jiashuo Sun, Zhiyi Shi, Haofei Yu |
Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals remains difficult: novelty and soundness are hard to measure automatically, and large-scale human evaluation is co...Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals remains difficult: novelty and soundness are hard to measure automatically, and large-scale human evaluation is costly. We propose a verifiable alternative by reframing proposal generation as a time-sliced scientific forecasting problem. Given a research question and inspiring papers available before a cutoff time, the model generates a structured proposal and is evaluated by whether it anticipates research directions that appear in papers published after the time. We operationalize this objective with the Future Alignment Score (FAS), computed via retrieval and LLM-based semantic scoring against a held-out future corpus. To train models, we build a time-consistent dataset of 21,835 paper occurrences across 3,642 instances from targets and their pre-cutoff citations, and synthesize reasoning traces that teach gap identification and inspiration borrowing. Across Llama-3.1 and Qwen2.5 models, future-aligned tuning improves future alignment over unaligned baselines (up to +10.6% overall FAS), and domain-expert human evaluation corroborates improved proposal quality. Finally, we demonstrate practical impact by implementing two model-generated proposals with a code agent, obtaining 4.17% accuracy gain on MATH from a new prompting strategy and consistent improvements for a novel model-merging method. Our code and data are publicly available at https://github.com/Arthur-Heng/future-aligned-proposals.
|
| 840 |
The Model Says Walk: Measuring whether LLMs Condition on Hidden Constraints
2603.29025
|
cs.CL
|
Yubo Li, Lu Zhang, Tianchong Jiang, Ramayya Krishnan, Rema Padman |
Asked whether to walk or drive to a car wash 50 m away, most language models say walk, forgetting that the car has to be there. Such failures are usually measured by accuracy on questions where the hidden constraint applies. We show that this is misleading: a ...Asked whether to walk or drive to a car wash 50 m away, most language models say walk, forgetting that the car has to be there. Such failures are usually measured by accuracy on questions where the hidden constraint applies. We show that this is misleading: a model can answer these questions correctly without reasoning about the constraint at all, simply by favoring the cautious option. We instead ask whether a model's decision changes when the constraint is removed. We test this at three levels of control: log-probability sweeps on open-weight models, a 500-prompt stress test, and CORE, a new human-validated benchmark of minimal constraint-present/constraint-absent pairs. The picture is consistent. The constraint nudges decisions rather than governing them. Models still miss presence constraints like the car wash, yet elsewhere they apply constraints that are not there. As a result, standard accuracy flatters all ten models we evaluate and reorders their ranking, and prompting fixes that look effective largely vanish under paired scoring. Claims about hidden-constraint reasoning need paired evidence.
|
| 841 |
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
2604.02102
|
cs.CLcs.LG
|
Haitong Sun, Stephen McIntosh, Kwanghee Choi, Eunjung Yeo, Daisuke Saito |
Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task has been used to measure phonemic contrast...Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task has been used to measure phonemic contrast in S3M representations via minimal pairs. We introduce prosodic ABX, an extension of this framework to evaluate prosodic contrast with only a handful of examples and no explicit labels. Also, we build and release a dataset of English and Japanese minimal pairs and use it along with a Mandarin dataset to evaluate contrast in English stress, Japanese pitch accent, and Mandarin tone. Finally, we show that model and layer rankings are often preserved across several experimental conditions, making it practical for low-resource settings.
|
| 842 |
An Inspectable LLM Council for Multi-Model Research Answer Aggregation
2604.02923
|
cs.CL
|
Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang |
Long-form research questions often require a single answer that combines details, reconciles conflicting estimates, and retains qualifications spread across several large language model (LLM) outputs. We present an inspectable LLM Council that separates synthe...Long-form research questions often require a single answer that combines details, reconciles conflicting estimates, and retains qualifications spread across several large language model (LLM) outputs. We present an inspectable LLM Council that separates synthesis into two stages: an analyst converts three independent candidate answers into a structured state, and a writer composes the final response from that state. We evaluate the workflow under closed-book conditions on 100 tasks from the Deep Research Accuracy, Completeness, and Objectivity (DRACO) benchmark, retaining 600 final answers, 100 analyst states, and 720 whole-answer judgments across the main and supplementary experiments. Under the primary rubric judge, Council scores 70.98, exceeding GPT and Gemini by 6.18 and 8.34 points; its 0.47-point difference from Claude has a 95% paired interval of $[-0.69,1.61]$. In the full supplementary comparison, Council scores 1.87 points above direct fusion (95% paired interval $[0.33,3.48]$) and covers 51.1% of positive rubric criteria covered by exactly one candidate, compared with 43.7% for direct fusion. On an exploratory 68-task subset with matching returned model labels, direct fusion scores higher, so the full comparison cannot isolate the analyst state's effect. Traced cases show how estimates, provenance qualifications, and selection decisions pass through the saved state into final answers.
|
| 843 |
How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline
2604.04204
|
cs.CLcs.LG
|
Mir Tafseer Nayeem, Davood Rafiei |
Large language models (LLMs) are increasingly embedded in educational, professional, and public infrastructure, yet widely used platforms expose "English (US)" as a primary English setting despite the global diversity of English. We ask: How does "English (US)...Large language models (LLMs) are increasingly embedded in educational, professional, and public infrastructure, yet widely used platforms expose "English (US)" as a primary English setting despite the global diversity of English. We ask: How does "English (US)" become the default? We study this question as structural bias, examining how geopolitical histories of data curation, digital dominance, and linguistic standardization intersect with the LLM development pipeline. Using British English as a controlled reference, we construct a curated resource of 1,813 matched American English (AmE)--British English (BrE) variants and introduce DiAlign, a dynamic, training-free method for estimating regional alignment from distributional evidence. We triangulate the AmE preference across data exposure --> representation --> generation, jointly examining pretraining and post-training data, tokenizer behavior and provenance, model prediction cost, and generated language across developer countries, prompt conditions, domains and sources, linguistic categories, and registers. AmE is consistently favored across all six audited pretraining corpora and 21 post-training datasets, is generally represented more compactly by tokenizers, and receives lower prediction cost. It also remains the dominant generation default under neutral English prompting; British-English prompting shifts this preference toward BrE but does not consistently eliminate the AmE default. To our knowledge, this is the first rigorous pipeline-wide study of structural bias across major phases of LLM development. Our findings show that contemporary LLMs privilege AmE as the de facto norm, raising concerns about linguistic homogenization, epistemic injustice, and inequity in global AI deployment, while providing a rigorous basis for targeted component-level intervention.
|
| 844 |
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
2604.06736
|
cs.CL
|
Yixi Zhou, Fan Zhang, Zhiqiao Guo, Yu Chen, Haipeng Zhang |
In Text-to-SQL tasks, large language models can generate structurally different SQL queries that return correct answers for the same intent. We call this phenomenon execution-correct structural divergence (ECSD). We introduce SQLStructEval, a framework that an...In Text-to-SQL tasks, large language models can generate structurally different SQL queries that return correct answers for the same intent. We call this phenomenon execution-correct structural divergence (ECSD). We introduce SQLStructEval, a framework that analyzes this behavior through canonical abstract syntax tree representations. We quantify structural diversity and agreement across generations. Experiments with different LLMs on multiple Text-to-SQL datasets, including Spider, document ECSD across models and datasets. Furthermore, our experiments demonstrate that generated queries are sensitive to question paraphrases and schema presentation. To address ECSD, we adopt a pipeline that first generates structured intermediate representations and then deterministically compiles them into SQL, improving execution accuracy and structural agreement among correct outputs. Structural analysis thus provides an additional diagnostic perspective that complements execution-based evaluation. Code is available at https://xanderzhou2022.github.io/AACL2026-SQLSTRUCTEVAL/.
|
| 845 |
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
2604.08977
|
cs.CL
|
Lorenzo Jaime Yu Flores, Cesare Spinoso di-Piano, Ori Ernst, David Ifeoluwa Adelani, Jackie Chi Kit Cheung |
Active learning (AL) is a training paradigm for selecting unlabeled samples for annotation to improve model performance on a test set, which is useful when only a limited number of samples can be annotated. These algorithms often work by optimizing for the inf...Active learning (AL) is a training paradigm for selecting unlabeled samples for annotation to improve model performance on a test set, which is useful when only a limited number of samples can be annotated. These algorithms often work by optimizing for the informativeness and diversity of the training data to be annotated. Recent work found that AL strategies fail to outperform random sampling on various language generation tasks when using 100-500 samples. To understand AL's poor performance when only using few samples, we investigate whether the core assumptions underlying AL strategies hold. We find that neither the informativeness nor diversity of the training data, which AL strategies optimize for, are correlated with test set performance. Instead, factors like the ordering of the training samples and interactions with pre-training data have a larger impact on performance. This suggests that future AL methods must take these factors into account in order to work with very few samples.
|
| 846 |
Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning
2604.09960
|
cs.CL
|
Samuel Jaeger, Calvin Ibenye, Aya Vera-Jimenez, Dhrubajyoti Ghosh |
The rapid adoption of large language models has introduced a new class of AI-generated fake news that coexists with traditional human-written misinformation, raising important questions about how these two forms of deceptive content differ and how reliably the...The rapid adoption of large language models has introduced a new class of AI-generated fake news that coexists with traditional human-written misinformation, raising important questions about how these two forms of deceptive content differ and how reliably they can be distinguished. This study examines linguistic, structural, and emotional differences between human-written and AI-generated fake news and evaluates machine learning and ensemble-based methods for distinguishing these content types. A document-level feature representation is constructed using sentence structure, lexical diversity, punctuation patterns, readability indices, and emotion-based features capturing affective dimensions such as fear, anger, joy, sadness, trust, and anticipation. Multiple classification models, including logistic regression, random forest, support vector machines, extreme gradient boosting, and a neural network, are applied alongside an ensemble framework that aggregates predictions across models. Model performance is assessed using accuracy and area under the receiver operating characteristic curve. The results show strong and consistent classification performance, with readability-based features emerging as the most informative predictors and AI-generated text exhibiting more uniform stylistic patterns. Ensemble learning provides modest but consistent improvements over individual models. These findings indicate that stylistic and structural properties of text provide a robust basis for distinguishing AI-generated misinformation from human-written fake news.
|
| 847 |
EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use
2604.14165
|
cs.CL
|
Naman Ahuja, Abhijit Chakraborty, Muhammad Ali Khan, Kaneez Zahra Rubab Khakwani, Mohamad Bassam Sonbol |
Structured extraction of evidence from clinical trial publications underpins systematic reviews and clinical guidelines, yet large language models are adopted for it only hesitantly: their outputs are difficult to verify, their use commonly requires transmitti...Structured extraction of evidence from clinical trial publications underpins systematic reviews and clinical guidelines, yet large language models are adopted for it only hesitantly: their outputs are difficult to verify, their use commonly requires transmitting documents to proprietary services, and they do not improve from the corrections their users make. We present EviSearch, a multi-agent system that addresses these three obstacles. Three tool-augmented agents with complementary access to a publication extract every column of an evidence table, and a value is admitted only after an attribution verifier has read it on its cited page, so that every value carries a page-level attribution. Disagreement between independent agents directs human review to the cells most likely to be wrong, and reviewer feedback refines the schema definitions and a curation knowledge base without updating model parameters. The agentic system runs entirely offline on open-weight models. On a clinician-annotated benchmark of randomized-trial publications, EviSearch attributes 100.0% of its values, reaches 91.70% accuracy autonomously, and reaches 95.22% after review of 15.6% of cells, exceeding random review of the strongest single agent at equal effort by 1.75 points.
|
| 848 |
Why Fine-Tuning Encourages Hallucinations and How to Fix It
2604.15574
|
cs.CLcs.LG
|
Guy Kaplan, Zorik Gekhman, Zhen Zhu, Lotem Rozner, Yuval Reif |
Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information through supervised fine-tuning (SFT), which can increase hallucinations w.r.t.~knowledge acquired during pre-tr...Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information through supervised fine-tuning (SFT), which can increase hallucinations w.r.t.~knowledge acquired during pre-training. Since these errors arise as a by-product of knowledge degradation, we explore whether established continual learning tools can mitigate them. We propose a self-distillation-based SFT method that facilitates effective factual learning while minimizing hallucinations w.r.t.~pre-existing knowledge by regularizing output-distribution drift. We also show that when new knowledge acquisition is unnecessary, suppressing factual plasticity by freezing parameter groups preserves task performance while reducing hallucinations. Lastly, we investigate the mechanism, contrasting capacity limitations, behavior cloning, and localized interference. Our experiments show that a main driver is interference among overlapping semantic representations, which self-distillation mitigates and an associative-memory model explains: forgetting grows with the overlap between new and stored facts.
|
| 849 |
When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints
2604.16916
|
cs.CL
|
Yuheng Chen, Zhiyu Wu, Bowen Cheng, Yihang Wu, Tetsuro Takahashi |
We identify and systematically characterize a class of task-structural alignment failures in large language models (LLMs): even when the harmful intent remains unchanged, changing the task presentation and output constraints can substantially alter model safet...We identify and systematically characterize a class of task-structural alignment failures in large language models (LLMs): even when the harmful intent remains unchanged, changing the task presentation and output constraints can substantially alter model safety behavior. Specifically, when a harmful request is reformulated as a forced-choice multiple-choice question (MCQ) in which all options are harmful and no refusal option is provided, some models that refuse the equivalent open-ended query instead select, prefer, or justify a harmful option. We evaluate 14 proprietary and open-source models on a bilingual Chinese-English human-authored dataset covering five harm categories, together with 900 model-generated Chinese adversarial MCQs. On human-authored data, attack success rate (ASR) increases sharply as prompts shift from open-ended queries to explicit forced-choice formats, typically peaking under intermediate levels of choice constraint. Model-generated Chinese MCQs further weaken or eliminate the recovery regime observed on human-authored data, driving ASR close to saturation for multiple models. The observed transfer patterns are consistent with stronger generators producing more difficult or boundary-adjacent MCQs, although other properties of the generated inputs may also contribute. We also find that adding an explicit refusal option or a safety preamble substantially reduces ASR for several high-capability models, often to near-zero levels, although their effectiveness varies across target models. These findings suggest that safety evaluations centered on open-ended generation may underestimate risks in structured deployment settings, and that task structure should be treated as an important and diagnosable dimension of safety evaluation and alignment training.
|
| 850 |
Learning Evidence Highlighting for Frozen LLMs
2604.22565
|
cs.CL
|
Shaoang Li, Yanhang Shi, Yufei Li, Mingfu Liang, Xiaohan Wei |
Large Language Models (LLMs) can reason well, yet often miss decisive evidence when it is buried in long, noisy contexts. We introduce HiLight, an Evidence Emphasis framework that decouples evidence selection from reasoning for frozen LLM solvers. HiLight avoi...Large Language Models (LLMs) can reason well, yet often miss decisive evidence when it is buried in long, noisy contexts. We introduce HiLight, an Evidence Emphasis framework that decouples evidence selection from reasoning for frozen LLM solvers. HiLight avoids compressing or rewriting the input, which can discard or distort evidence, by training a lightweight Emphasis Actor to insert minimal highlight tags around pivotal spans in the unaltered context. A frozen Solver then performs downstream reasoning on the emphasized input. We cast highlighting as a weakly supervised decision-making problem and optimize the Actor with reinforcement learning using only the Solver's task reward, requiring no evidence labels and no access to or modification of the Solver. Across sequential recommendation and long-context question answering, HiLight consistently improves performance over strong prompt-based and automated prompt-optimization baselines. The learned emphasis policy transfers zero-shot to both smaller and larger unseen Solver families, including an API-based Solver, suggesting that the Actor captures genuine, reusable evidence structure rather than overfitting to a single backbone.
|
| 851 |
BabelSafe: A Policy-Grounded Multilingual Safety Benchmark for LLMs
2605.00689
|
cs.CL
|
Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Bo Li |
As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety across diverse regulatory and cultural environments has become a critical challenge. However, existing multilingual benchmarks largely rely on general risk ...As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety across diverse regulatory and cultural environments has become a critical challenge. However, existing multilingual benchmarks largely rely on general risk taxonomies and machine-translated data, limiting evaluation to predefined risk categories and providing insufficient coverage of region-specific regulatory requirements and cultural contexts. To bridge these gaps, we introduce BabelSafe, a policy-grounded multilingual safety benchmark covering 13 language settings. BabelSafe is constructed from regional regulatory sources, with risk categories and fine-grained rules extracted from jurisdiction-specific regulatory documents directly used to guide the generation of multilingual safety data. During data generation, we further incorporate region-specific cultural contexts, enabling regulation-grounded and culturally contextualized evaluation across languages. Building on BabelSafe, we develop BabelGuard, a Diffusion Large Language Model (dLLM)-based guardrail model that supports multilingual safety judgment and policy-conditioned safety assessment. BabelGuard has two variants, a lightweight 1.5B model for fast `safe/unsafe' classification and a more capable 7B model for customizable policy-conditioned safety checking with detailed explanations. We evaluate BabelGuard against 11 strong guardrail baselines on 6 existing multilingual safety benchmarks and BabelSafe, demonstrating the strong performance of BabelGuard across these evaluation settings. We hope that BabelSafe and BabelGuard can help advance the development of regulation-aware and culturally contextualized multilingual guardrail systems.
|
| 852 |
One Turn Too Late: Learning When to Intervene Against Multi-Turn Malicious Intent
2605.05630
|
cs.CL
|
Xinjie Shen, Rongzhe Wei, Peizhi Niu, Haoyu Wang, Ruihan Wu |
Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, attackers can distribute their intent across multiple benign-looking turns, making defe...Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, attackers can distribute their intent across multiple benign-looking turns, making defense a problem not only of whether a dialogue is harmful, but also of when intervention becomes necessary. Existing trace-level labeling approaches provide only coarse safety signals and do not identify this intervention boundary, making it difficult to distinguish timely intervention from premature refusal or a block that comes too late. This work introduces turn-level harm-enabling supervision for multi-turn defense. We define the earliest harm-enabling turn as the first point at which delivering a candidate response would make the accumulated interaction sufficient to enable harmful action. To instantiate this supervision at scale, we construct the Multi-Turn Intent Dataset (MTID), which contains adaptive attack rollouts, matched benign hard negatives, and annotations of this boundary. Using MTID, we train TurnGate, a response-aware monitor that learns when to intervene, and further optimize its policy through multi-turn reinforcement learning. Experiments show that turn-level boundary supervision improves intervention localization, while reinforcement learning further improves the safety--utility trade-off. TurnGate outperforms existing guardrails and multi-turn monitoring baselines, and generalizes across risk domains, attacker pipelines, and target models. Our code is available at https://github.com/Graph-COM/TurnGate.
|
| 853 |
MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text
2605.06903
|
cs.CL
|
Chenjun Li, Cheng Wan, Haomiao Chen, Johannes C. Paetzold |
Large language models are widely used in everyday writing, making reliable AI-generated text detection crucial for academic integrity, content moderation, and provenance tracking. Yet high AUROC on clean, in-distribution benchmarks is not sufficient. Practical...Large language models are widely used in everyday writing, making reliable AI-generated text detection crucial for academic integrity, content moderation, and provenance tracking. Yet high AUROC on clean, in-distribution benchmarks is not sufficient. Practical detectors must resist adversarial rewrites, generalize to unseen generators and writing domains, and maintain low false-positive rates (FPR). A pooled AI-versus-human objective does not explicitly require the model to distinguish among generator families, so it may fail to learn the generator-specific structure needed for generalization and attribution. We introduce MELD (Multi-Task Equilibrated Learning Detector), which instead trains on class-balanced, family-specific AI-versus-human tasks while sharing a common representation of human writing. MELD produces detection, generator-family, rewrite-task, and token-level predictions in a single forward pass. Format normalization further makes its scores invariant to the modeled reformatting operations. MELD ranks first among open-source submissions in the public RAID leaderboard and matches or exceeds supervised baselines on five of six held-out evaluation pools. To evaluate transfer to unseen generators, we introduce MELD-eval, a held-out test pool built from four frontier chat models. Without further fine-tuning, MELD achieves 99.7% TPR at 1% FPR on MELD-eval and 98% TPR at the same FPR on a held-out generator whose family is absent from training. Finally, in a case study of 5.9 million scientific texts from 2016--2026, MELD's prediction scores remain stable through 2022 and increase from 2023 onward, coinciding with the widespread adoption of LLM-based writing tools. The model, MELD-eval pool, source code, and live demo are available.
|
| 854 |
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
2605.07721
|
cs.CLcs.LG
|
Victor Conchello Vendrell, Arnau Padres Masdemont, Niccol\`o Grillo, Jordi Ros-Giralt, Arash Behboodi |
Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space without generating intermediate tokens. Models such as Ouro perform reasoning by iteratively updating interna...Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space without generating intermediate tokens. Models such as Ouro perform reasoning by iteratively updating internal representations while retaining a standard Key-Value (KV) cache across iterations, causing memory consumption to grow linearly with reasoning depth. Consequently, increasing the number of reasoning iterations can lead to prohibitive memory usage, limiting the practical scalability of such architectures. In this work, we propose Memory-Efficient Looped Transformer (MELT), a novel architecture that decouples reasoning depth from memory consumption. Instead of using a standard KV cache per layer and loop, MELT maintains a single KV cache per layer that is shared across reasoning loops. This cache is updated over time via a learnable gating mechanism. To enable stable and efficient training under this architecture, we propose to train MELT using chunk-wise training in a two phase procedure: interpolated transition, followed by attention-aligned distillation, both from the LoopLM starting model to MELT. Empirically, we show that MELT models fine-tuned from pretrained Ouro parameters outperform standard LLMs of comparable size, while maintaining a memory footprint comparable to those models and dramatically smaller than Ouro's. Overall, MELT achieves constant-memory iterative reasoning without sacrificing LoopLM performance, using only a lightweight post-training procedure.
|
| 855 |
Tool Calling is Linearly Readable and Steerable in Language Models
2605.07990
|
cs.CLcs.LG
|
Zekun Wu (University College London), Ze Wang (University College London), Seonglae Cho (Holistic AI), Yufei Yang (Imperial College London), Adriano Koshiyama (University College London) |
Language-model agents can take real actions by calling tools, so choosing the wrong tool can cause errors that are difficult to undo. Most evaluations only observe the tool choice after the model generates a call. We read this choice from the model before gene...Language-model agents can take real actions by calling tools, so choosing the wrong tool can cause errors that are difficult to undo. Most evaluations only observe the tool choice after the model generates a call. We read this choice from the model before generation, and we steer it. For each tool, we average the model's hidden states from a few example requests to obtain a tool vector. Comparing a new request with these tool vectors predicts which tool it needs. The difference between two tool vectors gives a steering direction that can move the model toward a chosen tool without retraining. Across eight instruction-tuned models from 4B to 27B parameters, steering changes the generated call to a chosen target tool in 56-78% of held-out tool pairs, depending on the model. We also observe steering on real APIs from $\tau$-bench and ToolBench, though less reliably. At the final layer, steering could work simply by raising the score of the target tool name. To test this, we compare the steering direction with a direction that only raises that score, layer by layer. In the middle layers, on the three models we examine in depth, the steering direction switches more calls than the score-raising direction. So the tool vectors capture part of the tool choice before the final layer. The same tool vectors also help identify likely tool-selection errors before generation. Calls are more likely to be wrong when a request is similarly close to two tool vectors. Across four models, this signal achieves an AUROC of 0.61-0.78 and outperforms a first-token confidence baseline on three of them. Together, these results suggest that the model's hidden state provides a way to read, steer, and check tool choice before a call is made.
|
| 856 |
AIPO: Learning to Reason from Active Interaction
2605.08401
|
cs.CL
|
Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari |
Recent advances in LLMs have demonstrated strong reasoning capabilities, largely stimulated by RLVR. However, the exploration of existing RLVR algorithms remains largely constrained by the knowledge and reasoning strategies already accessible to the policy mod...Recent advances in LLMs have demonstrated strong reasoning capabilities, largely stimulated by RLVR. However, the exploration of existing RLVR algorithms remains largely constrained by the knowledge and reasoning strategies already accessible to the policy model. Although recent methods introduce external expert demonstrations to broaden exploration, they typically rely on complete trajectory-level guidance, which can be sample-inefficient, information-sparse, and insufficiently adaptive to intermediate reasoning bottlenecks. Inspired by collaborative multi-agent systems, we propose AIPO, an enhanced reinforcement learning framework that improves LLM reasoning through active multi-agent interaction during exploration. Specifically, when encountering reasoning bottlenecks, AIPO enables the policy model to proactively consult three functional collaborative agents, namely the verify agent, knowledge agent, and reasoning agent, thereby obtaining fine-grained and state-dependent guidance during rollout. The resulting mixed-policy trajectories expose the policy to reasoning directions that may be difficult to discover through isolated on-policy exploration. To learn effectively from collaborator-provided tokens, we further introduce a corrected importance sampling coefficient together with a lower-bound clipping strategy to mitigate off-policy discrepancy and vanishing gradients. After training, the policy model reasons independently without relying on collaborative agents. Extensive experiments across mathematical, scientific, coding, and puzzle reasoning benchmarks show that AIPO consistently improves reasoning performance and generalizes across different policy models, collaborator backbones, and RLVR algorithms.
|
| 857 |
Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection
2605.08863
|
cs.CLcs.LG
|
Shota Fujikawa, Issei Sato |
Hallucination detection has become increasingly important for improving the reliability of large language models (LLMs). Recently, hybrid approaches such as HaMI, which combine semantic consistency with internal model states via Multiple Instance Learning (MIL...Hallucination detection has become increasingly important for improving the reliability of large language models (LLMs). Recently, hybrid approaches such as HaMI, which combine semantic consistency with internal model states via Multiple Instance Learning (MIL), have achieved state-of-the-art performance. However, these methods incur substantial computational overhead due to repeated sampling and costly semantic similarity computations. In this work, we first provide a theoretical analysis of HaMI in terms of decision margins, revealing that scaling internal states with semantic consistency leads to an enlarged decision margin. Motivated by this insight, we revisit classical sentence classification models from a margin enlargement perspective, aggregating token-level features via max pooling and directly estimating sentence scores using a lightweight MLP. Without requiring semantic consistency computations, our approach achieves substantial efficiency improvements while maintaining competitive performance with state-of-the-art baselines through adaptive aggregation of internal feature representations. Code is available at https://github.com/FUJI1229/Hallucination_Detection.
|
| 858 |
Agent Collectives Should Not Detect Their Own Imposters: A Chess Case Study
2605.09027
|
cs.CLcs.LG
|
Alexandre Le Mercier, Chris Develder, Thomas Demeester |
A collective of AI agents collaborating on a task has the potential to outclass any individual agent for that task. We study the robustness of such collectives against possible imposters, i.e., agents that deliberately try to mislead their peers. Since a singl...A collective of AI agents collaborating on a task has the potential to outclass any individual agent for that task. We study the robustness of such collectives against possible imposters, i.e., agents that deliberately try to mislead their peers. Since a single imposter could undo the collective's advantage, we need to detect them. We consider two strategies: (i) incorporate imposter detection into the participating agents, or (ii) use a dedicated imposter detector outside the collective. We investigate this empirically on Gambit, a testbed in which 4 reasoning agents deliberate on chess moves. The setting is small but still challenging for frontier models. Chess allows objective, quantitative assessment (via a state-of-the-art chess engine) of both the gain of using a collective and the damage done by imposters. We find that merely warning the agents of potential imposter presence is not beneficial: it degrades decisions when no imposter is present, provokes reactions ranging from self-accusation to scapegoating, inflates token use, and reveals to the imposter how it was uncovered. We therefore recommend a detector that reads the collective's deliberation but never joins it and only returns a verdict. Such a detector must recalibrate to new attack strategies after very few examples, rather than wait for full retraining. In our benchmark, a 3B language model with a meta-trained classification head achieves that: a single gradient step on 20 labeled examples suffices to adapt to an unseen imposter strategy. At matched zero-shot accuracy, this detector yields 8x the adaptation gain of standard finetuning, at 14x lower training cost. We release the Gambit benchmark, with 37,352 labeled deliberations spanning 240 evolved imposter strategies. Code and data: https://anonymous.4open.science/r/gambit.
|
| 859 |
BOOKMARKS: Efficient Active Storyline Memory for Role-playing
2605.14169
|
cs.CL
|
Letian Peng, Ziche Liu, Yiming Huang, Longfei Yun, Kun Zhou |
Memory systems are critical for role-playing agents (RPAs) to maintain long-horizon consistency. However, existing RPA memory methods (e.g., profiling) mainly rely on incremental summarization, whose compression discards details which become inaccessible to su...Memory systems are critical for role-playing agents (RPAs) to maintain long-horizon consistency. However, existing RPA memory methods (e.g., profiling) mainly rely on incremental summarization, whose compression discards details which become inaccessible to subsequent grounding. To address this issue, we propose a search-based memory framework called \textbf{\underline{BOOKMARKS}} for \textbf{active grounding}, which retains access to the full preceding storyline and collects task-relevant information on demand. Since summarizing the preceding storyline anew for each grounding request incurs substantial redundant computation, BOOKMARKS introduces \textbf{passive updating} to reuse earlier search results as checkpoints. Each \textbf{bookmark} represents the \textbf{content} about a particular aspect (\textbf{section}) of story information at a specific synchronization \textbf{point}. For current task, BOOKMARKS searches for only useful contents, reuses existing semantically equivalent bookmarks or initializes new ones, and synchronizes the selected ones from their stored checkpoints to the current scene. A reused bookmark thus only needs to process the newly observed storyline suffix, avoiding repeated synchronization. We evaluate BOOKMARKS across six narrative artifacts involving 47 characters and 9,537 test cases against non-active grounding baselines, covering next-action prediction and challenging, human-authored reasoning questions from a mystery game. BOOKMARKS improves next-action fidelity in a five-dimensional evaluation (emotion, intent, causality, position, and content) and raises mystery-game reasoning accuracy from 39.44\% to 45.42\% over the strongest baseline.
|
| 860 |
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
2605.14368
|
cs.CL
|
Injin Kong, Hyoungjoon Lee, Yohan Jo |
Continuous diffusion language models require choosing a representation space in which denoising operates, yet it remains unclear which representations are most compatible with diffusion. We study a basic design question: where inside a pretrained language mode...Continuous diffusion language models require choosing a representation space in which denoising operates, yet it remains unclear which representations are most compatible with diffusion. We study a basic design question: where inside a pretrained language model should continuous diffusion operate? We formulate this as a hidden-state interface-selection problem and hypothesize that diffusion-friendly interfaces can be identified from representation geometry. We operationalize this hypothesis using three training-free geometric proxies: local compactness, global stiffness, and effective rank. Across two 8B-scale backbones, the resulting geometry score predicts fixed-budget diffusion bridgeability beyond the dominant effect of layer depth. We then instantiate DiHAL, a Locate-and-Replace framework that replaces the transformer prefix below a selected interface with conditional diffusion while retaining the pretrained suffix and LM head. Under full training, geometry-selected interfaces remain close to validation-loss oracles and outperform worst-layer controls, while the diffusion bridge outperforms a parameter-matched deterministic replacement under matched diagnostic conditions. These results suggest that the representation space in which diffusion operates should itself be treated as a first-class design variable for continuous diffusion in language models.
|
| 861 |
COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion
2605.15016
|
cs.CL
|
Zihan Deng, Xiaozhen Zhong, Chuanzhi Xu, Quankeng Huang |
Sequential diagnosis requires ranking diseases from a sparse intake under a public budget of follow-up questions, which is diagnostic triage with missing findings rather than screening of people who have no symptoms. Because laboratory series arrive irregularl...Sequential diagnosis requires ranking diseases from a sparse intake under a public budget of follow-up questions, which is diagnostic triage with missing findings rather than screening of people who have no symptoms. Because laboratory series arrive irregularly and histories remain incomplete, the ranking must be updated as new facts arrive, yet most pipelines are limited by unnamed numeric trends, unverifiable free-form chain of thought, and a negative bias from treating unasked findings as absent. To solve these, we propose COTCAgent, a sparse-intake consultation agent built on Chain-of-Thought Completion (COTC), which couples three modules on one shared ternary evidence log. The Temporal-Statistics Adapter (TSA) turns time series into typed predicates with named statistics, while the COTC module scores diseases with a learnable knowledge-base energy in which unknown findings add zero energy, after which bounded probabilistic completion selects unknown findings under the question budget and writes each answer back as present, absent, or unknown. Edge strengths and an absence scale are trained with listwise ranking under a sparse curriculum, so that the same energy chooses the next finding by a surrogate entropy reduction. Experiments on DDXPlus and MIMIC-IV show that these modules improve ranking under a matched question budget, with information-gain completion above asking nothing or at random and above protocol-matched askers and a local LLM, thereby providing a foundation for future calibrated scoring and prospective preventive consultation.
|
| 862 |
Sometin Beta Pass Notin: Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation
2605.17710
|
cs.CL
|
Sewade Ogun |
Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind resource-rich languages such as English and French. Nigerian languages present unique modelling hurdles, incl...Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind resource-rich languages such as English and French. Nigerian languages present unique modelling hurdles, including acute data scarcity, inconsistent orthography, tonal diacritics, diverse accents, frequent code-switching, and localised named entities. To address these challenges, we developed a multilingual ASR framework using a two-stage distillation process. First, we employed student-teacher knowledge distillation from existing monolingual models, conditioned on robust language-specific N-gram language models. Second, we performed iterative self improvement using pseudo-labelled data to further refine accuracy. Our method significantly bridges the performance gap, achieving on average a reduction in the relative Word Error Rate (WER) of 29% over the monolingual baselines. Our models also outperform state-of-the-art multilingual models across major benchmarks, including Common Voice and FLEURS. We introduce Sometin Beta Pass Notin (SBPN), a multilingual foundational ASR model that covers Yor\`ub\'a, Hausa, Igbo, Nigerian Pidgin, and Nigerian English.
|
| 863 |
DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Agent Memory
2605.22411
|
cs.CLcs.LG
|
Jianing Yin, Tan Tang, Yingcai Wu |
Large language model (LLM) agents still struggle to effectively use long-term memory, with answer-supporting evidence often scattered across long conversational histories and buried in substantial irrelevant content. Existing memory systems commonly process me...Large language model (LLM) agents still struggle to effectively use long-term memory, with answer-supporting evidence often scattered across long conversational histories and buried in substantial irrelevant content. Existing memory systems commonly process memory into dedicated units before future queries are known and retrieve these preconstructed units at query time. Because the contents and granularity of these units are determined before the current query is known, the retrieved units can remain coarse and noisy for that query, leaving downstream answerers to further denoise them and uncover query-specific evidence. We present DeferMem, a long-term memory framework that decouples this problem into high-recall candidate retrieval and query-conditioned evidence distillation. DeferMem uses a lightweight segment-link structure to organize raw history and retrieve broad candidates at query time. A memory distiller then distills these high-recall but highly noisy candidates into a set of faithful, self-contained, and query-conditioned evidence. To train this distiller, we introduce DistillPO, a reinforcement learning algorithm that formulates post-retrieval evidence distillation as a structured action comprising message selection and evidence rewriting. It optimizes this action with a decomposed-and-gated reward pipeline and structure-aligned advantage assignment, gating reward components from validity to quality checks while exposing answerability-related feedback early and assigning each reward to its responsible output span. On LoCoMo and LongMemEval-S, DeferMem surpasses strong baselines in QA accuracy and memory-system efficiency, achieving the highest QA accuracy and fastest runtime while consuming no commercial-API tokens for memory operations.
|
| 864 |
HiMed: Incentivizing Hindi Reasoning in Medical LLMs
2605.24635
|
cs.CL
|
Dingfeng Jiang, Han Yan, Chenze Ma, Amit Kumar Jaiswal, Ang Li |
Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, their performance degrades sharply in Hindi, particularly on Indian systems of med...Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, their performance degrades sharply in Hindi, particularly on Indian systems of medicine. We argue that robust cross-lingual medical transfer requires Hindi reasoning. To this end, we introduce HiMed, a Hindi reasoning medical corpus and benchmark suite covering both Western and Indian medicine. We further propose HiMed-8B, a Hindi-form medical reasoning LLM, through the design of decaying scaffolding reward. Extensive experiments demonstrate improvement in Hindi medical reasoning performance and a reduction in the English-Hindi accuracy gap. Ablation studies validate the contribution of each training stage and reward component. All code, data, and weights are available
|
| 865 |
Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models
2605.25443
|
cs.CL
|
Zongji Yu, Wenshui Luo, Yiliu Sun, Hao Fang, Hanrui Xu |
Post-training via Reinforcement Learning (RL) has enabled Large Reasoning Models (LRMs) to achieve strong performance in individual domain. However, real-world applications increasingly require general-purpose reasoners rendering strong performance across dive...Post-training via Reinforcement Learning (RL) has enabled Large Reasoning Models (LRMs) to achieve strong performance in individual domain. However, real-world applications increasingly require general-purpose reasoners rendering strong performance across diverse domains. Mixed-domain post-training aims to achieve this goal by jointly training with mixed domain data, but this often induces capability compromise and degradation among different domains. Existing methods attribute this performance degradation to harmful cross-domain interactions and propose various strategies to mitigate them, but these strategies may also impede beneficial knowledge sharing across domains and in turn fail to match or surpass single-domain performance. To address this problem, we propose \textbf{M}ulti-domain \textbf{C}ontrastive \textbf{P}olicy \textbf{O}ptimization (MCPO), which uses contrastive learning to utilize both positive and negative cross-domain interactions for knowledge sharing and competition. Specifically, we partition each rollout generated by LRMs according to its underlying reasoning structures and use reasoning segments to capture these structures. We thus formulate positive and negative pairs of reasoning segments as mutually augmented examples, which provide supportive and competing signals for knowledge sharing. Subsequently, we design complementary contrastive objectives for cross-domain knowledge sharing and intra-domain knowledge consolidation, targeting compatibility across domains and discriminability within each domain to form a harmonious reasoning space. Experimental results across a broad range of domains show that MCPO alleviates performance degradation caused by mixed-domain training and outperforms single-domain training in most cases.
|
| 866 |
ResearchMath-14K: Scaling Research-Level Mathematics via Agents
2605.28003
|
cs.CL
|
Guijin Son, Seungyeop Yi, Minju Gwak, Hyunwoo Ko, Wongi Jang |
The frontier of mathematics is defined by problems whose solutions are not yet known. However, whether language models can meaningfully engage with such problems without human intervention remains unclear. A major obstacle is the lack of large-scale research-l...The frontier of mathematics is defined by problems whose solutions are not yet known. However, whether language models can meaningfully engage with such problems without human intervention remains unclear. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of $14{,}056$ problems curated from academic sources via a multi-agent pipeline. ResearchMath-14k spans 11 mathematical domains and ranks above existing math datasets on knowledge, novelty, and procedural difficulty. To our knowledge, it is the largest research-level mathematical problem set available for training. We additionally generate $220$K teacher trajectories through targeted prompting, followed by behavioral filtering. Notably, however, generating correct trajectories is nontrivial at this level, and two LLM judges label only $3.7\%$ and $4.3\%$ of sampled ResearchMath training trajectories as correct. Nevertheless, across three model families, full-parameter training on ResearchMath improves performance on graduate- and research-level mathematics benchmarks by $2.1$ points over the starting checkpoints. In comparison, training on existing datasets such as DASD and Nemotron-SFT-Math-v4 changes performance by $0.0$ and $-0.5$ points, respectively. Notably, mixing DASD with ResearchMath yields higher scores than token-matched DASD alone on benchmarks covering olympiad short-form ($+2.0$), graduate- and research-level short-form ($+0.8$), graduate- and research-level symbolic ($+2.6$), and proof evaluation ($+5.7$). Further analysis suggests that research-level mathematical content and greater reasoning diversity may help explain why ResearchMath provides complementary supervision to contemporary datasets. We make ResearchMath-14k publicly available for future works on research-level mathematical reasoning.
|
| 867 |
Beyond Imitation: Reflective On-Policy Self-Distillation for LLM Reasoning
2605.28014
|
cs.CLcs.LG
|
Ziqi Zhao, Xinyu Ma, Liu Yang, Yujie Feng, Daiting Shi |
On-policy self-distillation (OPSD) improves the reasoning capabilities of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, existing OPSD methods often yield limited gains on complex reasoning tasks and su...On-policy self-distillation (OPSD) improves the reasoning capabilities of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, existing OPSD methods often yield limited gains on complex reasoning tasks and suffer from severe training instability. We identify two key causes: conditioning the self-teacher on a complete verified solution encourages imitation of complete reference trajectories rather than extraction of transferable reasoning insights, while indiscriminate full-response distillation imposes superfluous supervision on already-valid reasoning prefixes. Together, these issues suppress reasoning diversity and contribute to late-stage mode collapse. We propose Reflective On-policy Self-Distillation (ROSD), which distills transferable reasoning insights rather than complete reference trajectories. For each erroneous rollout, a self-reflector contrasts it with a correct rollout from the same group to derive a corrective idea and identify the sentence containing the first reasoning error. The corrective idea provides the self-teacher with targeted guidance, while the diagnosed error boundary allows ROSD to mask out the distillation loss over the valid prefix and apply token-level distillation only from the first erroneous sentence onward. Experiments across multiple reasoning benchmarks and model backbones show that ROSD consistently outperforms standard OPSD and reinforcement learning baselines, better preserves reasoning diversity, stabilizes training, and mitigates late-stage mode collapse. Code is available at https://github.com/ZiqiZhao1/ROSD.
|
| 868 |
PhoneWorld: From Real-App Trajectories to Dynamic and Verifiable Environments for Phone-Use Agents
2605.29486
|
cs.CLcs.LG
|
Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu, Jason |
Real applications provide the training setting closest to phone-agent deployment, but are difficult to reset, scale safely, and verify programmatically. Static screenshots and interaction trajectories preserve realistic evidence but cannot generate new experie...Real applications provide the training setting closest to phone-agent deployment, but are difficult to reset, scale safely, and verify programmatically. Static screenshots and interaction trajectories preserve realistic evidence but cannot generate new experience. We introduce PhoneWorld, a trace-grounded framework that converts such evidence into runnable, resettable, and verifiable Android environments. PhoneWorld induces a usage-weighted interaction skeleton from observed pages, transitions, and state-changing operations; translates it into a behavior-grounded app specification; realizes the specification through an autonomous build--inspect--repair loop; and synthesizes executable tasks with programmatic verifiers. The resulting suite spans 34 consumer-facing apps across 16 domains and supports an audited online benchmark, verified trajectory generation, and online RL through common reset and verification interfaces. Evaluations with diverse general and open-source GUI agents show that PhoneWorld supports reliable end-to-end online interaction and exposes capabilities complementary to AndroidWorld. Controlled SFT experiments further show that PhoneWorld trajectories complement AndroidWorld supervision, transfer across online and offline benchmarks, and become more effective as data volume and app coverage increase. Under a matched RL budget, combining PhoneWorld mock-app rollouts with real-app rollouts improves performance over real-app RL alone on both real-phone tasks and AndroidWorld. Together, these results demonstrate that trace-grounded executable abstraction can bridge realistic mobile behavior and scalable agent learning, turning limited real-app evidence into a growing supply of controllable and verifiable environments for training and evaluation.
|
| 869 |
Beyond Math and Code: Lightweight Corpus-Grounded Process Rewards for Factual Question Answering
2605.29648
|
cs.CL
|
Shicheng Fan, Haochang Hao, Dehai Min, Weihao Liu, Hanrong Zhang |
Process supervision during reinforcement learning (RL) post-training matters in factual question answering (QA) because responses receiving positive response-level rewards can still contain sentence-level factual errors. Unlike math and code, factual QA lacks ...Process supervision during reinforcement learning (RL) post-training matters in factual question answering (QA) because responses receiving positive response-level rewards can still contain sentence-level factual errors. Unlike math and code, factual QA lacks inexpensive programmatic checks, making reward computation a bottleneck. Existing methods rely on neural verifiers to score individual sentences, requiring extensive model inference as RL repeats these checks across many sampled responses at every update. We therefore introduce CorVer (Corpus Verify), a lightweight training-time approach to process supervision that derives sentence-level rewards from corpus co-occurrence statistics. A 0.5B extractor identifies subject-object pairs, indexed corpus queries supply their co-occurrence counts, and the resulting sentence rewards are assigned to the corresponding tokens for RL. On Qwen3-4B and Qwen3-8B, CorVer reduces mean complete training time by 5.5-10.4 times relative to the four factuality-RL baselines. Across the four models evaluated against these baselines, CorVer achieves the highest accuracy in 17 of 20 model-benchmark settings. CorVer outperforms the unmodified models in all 30 standard factual-QA settings (six models from three families across five benchmarks) and remains effective on two additional multi-hop QA datasets.
|
| 870 |
EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL
2605.29670
|
cs.CL
|
Huawei Zheng, Sen Yang, Zhaorui Yang, Yuhui Zhang, Haozhe Feng |
Schema linking is a difficult and important step in large-scale Text-to-SQL, where systems must identify a compact yet sufficient schema context from large and ambiguous databases. Existing methods often treat schema linking as deterministic selection around a...Schema linking is a difficult and important step in large-scale Text-to-SQL, where systems must identify a compact yet sufficient schema context from large and ambiguous databases. Existing methods often treat schema linking as deterministic selection around a single SQL path, but complex questions may admit multiple valid realizations with different schema needs. We reframe schema linking as uncertainty-aware schema-need inference over multiple plausible SQL paths, where the system distinguishes required schema items from path-dependent uncertain ones and acquires evidence only where needed. We instantiate this reframing with EviLink, which combines multi-hypothesis schema grounding with uncertainty-guided evidence acquisition. Experiments on BIRD-Dev and Spider2-Snow show that this perspective improves the balance among schema completeness, schema relevance, and token cost. On Spider2-Snow, EviLink achieves 93.04% field-level strict recall rate, uses 116.55K average tokens, and improves downstream SQL generation under a fixed generator.
|
| 871 |
Multi-Legal-Bench: When the Answer Is in the Input. Label Leakage in Legal Benchmarks Built from Court Registries
2605.29738
|
cs.CL
|
Volodymyr Ovcharov |
Court registries publish millions of decisions with structured metadata, which makes them an attractive source of labelled legal benchmarks: the court type, the form of the decision and its subject area come for free. We show that this convenience has a cost t...Court registries publish millions of decisions with structured metadata, which makes them an attractive source of labelled legal benchmarks: the court type, the form of the decision and its subject area come for free. We show that this convenience has a cost that registry-derived benchmarks rarely measure. We build Multi-Legal-Bench, which evaluates identical tasks on native court decisions from five national registries (France, the Netherlands, Poland, the Czech Republic and Lithuania), extending the Ukrainian UA-Legal-Bench, and audit what its cells actually measure. A keyword scan that uses no model reaches 96% on Dutch judgment-form classification against a 50% majority baseline, 84% on French court-type classification against 33%, and 66% on Polish judgment-form classification, where all nine models add at most seven points over it. Replacing the label names in the input by a mask brings the scan to the majority floor in all six affected cells, and eight of nine models lose accuracy significantly in most of them (39 of 54 model-cell pairs), by up to 48 points (Dutch judgment form: 100% to 52%); only Claude Sonnet 5 is essentially unaffected. Paired McNemar tests with Holm correction find a leader that beats every other model in only two of twelve cells, each a different model; but once the label names are masked, the same model leads both, and the spread between models widens in every masked cell. Part of the apparent parity between models is produced by the leakage itself. The audit also exposed a scoring defect in our own earlier release (answers written with diacritics were scored as wrong, understating Czech judgment-form accuracy by up to 90 points), which we correct and document. We recommend that every benchmark built from court registries report a no-model baseline and a masked control per cell. All data, prompts, predictions and scoring code are released.
|
| 872 |
DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity
2605.29751
|
cs.CL
|
Kaijie Zheng, Weiqin Wang, Yile Wang, Hui Huang |
Calculating semantic textual similarity is a foundational task in natural language processing. Current large language models (LLMs) based methods typically rely on extracting last-layer hidden states with fixed dimensions to compute similarity for every text p...Calculating semantic textual similarity is a foundational task in natural language processing. Current large language models (LLMs) based methods typically rely on extracting last-layer hidden states with fixed dimensions to compute similarity for every text pairs. We argue that this paradigm is suffer from two limitations: (i) The last hidden layer encodes more general knowledge rather than just semantic knowledge, making it suboptimal for semantic similarity computation; (ii) The hidden layer dimensions of LLMs are generally very large, which introduces some redundancy and noise for representing semantics. In this work, we propose DySem, a novel training-free framework that investigates more semantic-related internal components of LLMs via multilingual consensus, and shifts away from static representation spaces in favor of dynamic, sample-specific semantic dimensions by constructing text-dependent joint semantic set and computes similarity over this shared dimensional subset. Extensive experiments across various LLMs show that our method consistently outperforms recent baselines while maintaining lower dimensions for similarity calculation. The code is released at https://github.com/szu-tera/DySem.
|
| 873 |
Shared Doubt: Zero-Shot Cross-Lingual Confidence Estimation for Language Models
2605.31220
|
cs.CLcs.LG
|
Athina Kyriakou, Dennis Ulmer, Ivan Titov |
Confidence estimation (CE), i.e. quantifying the reliability of a model's prediction, has attracted great interest in the context of large language models (LLMs). However, most studies focus on English, ignoring the multilingual reality of LLM usage, while man...Confidence estimation (CE), i.e. quantifying the reliability of a model's prediction, has attracted great interest in the context of large language models (LLMs). However, most studies focus on English, ignoring the multilingual reality of LLM usage, while many CE methods degrade or require retraining across languages. To address this gap, we investigate whether multilingual LLMs encode shared, language-transferable confidence features in open-ended question answering. We use a lightweight linear probe that predicts answer correctness directly from intermediate representations. Trained monolingually, the probe generalizes zero-shot to unseen, linguistically diverse languages without target-language supervision. Multiple ablations and learned layer weights reveal that confidence features concentrate in middle layers across languages, suggesting a shared confidence subspace. While zero-shot cross-lingual performance depends on similarity to the source language, the probe provides a strong baseline without any retraining and compares favorably to other popular confidence estimation methods.
|
| 874 |
FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards
2606.06960
|
cs.CL
|
Yining Zhu, Zihao Deng, Leiming Wang, Jingfei Lu, Junbo Wang |
Large language model (LLM) agents increasingly rely on external experience to continually adapt to changing environments without modifying their underlying models. Recent experience mechanisms have demonstrated promising results across diverse tasks. However, ...Large language model (LLM) agents increasingly rely on external experience to continually adapt to changing environments without modifying their underlying models. Recent experience mechanisms have demonstrated promising results across diverse tasks. However, their effectiveness is typically evaluated within individual benchmark settings, and how experience mechanisms generalize across different scenarios remains insufficiently explored. In this work, we present a scenario-oriented analysis of experience mechanisms for LLM agents. We characterize existing evaluation scenarios along four dimensions: outcome observability, credit assignment complexity, environmental dynamics, and experience reusability. Our analysis shows that existing benchmarks often evaluate experience mechanisms under scenarios where at least one dimension is comparatively favorable, leaving more challenging combinations of scenario properties underexplored. To address this gap, we introduce \textsc{FinEvolveBench}, a reproducible benchmark built on a chronological stream of rich financial news and market data that enables systematic evaluation of experience-based self-evolution under challenging experience regimes characterized by noisy feedback, ambiguous credit assignment, environmental non-stationarity, and limited experience reusability. Experiments show that existing approaches exhibit substantially reduced or inconsistent gains in this setting, highlighting the scenario-dependent nature of experience mechanisms and the challenge of maintaining valid experience under changing environments.
|
| 875 |
ReadingMachine: A Computational Methodology for Structured Corpus Reading and Large-Scale Synthesis
2606.07753
|
cs.CL
|
James Morrissey |
ReadingMachine is a computational methodology for structured corpus reading that uses large language models to perform bounded reading operations over entire document collections. Rather than relying on retrieval or recursive summarization, the approach decomp...ReadingMachine is a computational methodology for structured corpus reading that uses large language models to perform bounded reading operations over entire document collections. Rather than relying on retrieval or recursive summarization, the approach decomposes analysis into inspectable stages including insight extraction, semantic clustering, theme generation, and iterative omission detection. By delaying irreversible compression and explicitly tracking intermediate representations, the method prioritizes coverage, traceability, and preservation of disagreement across large corpora. The system is demonstrated on a heterogeneous corpus of 152 industrial policy documents, producing more than 17,500 extracted insights and a structured thematic map. ReadingMachine is released as an open-source experimental framework for large-scale qualitative synthesis and corpus analysis.
|
| 876 |
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
2606.07970
|
cs.CL
|
Haoming Wen, Shi Chen, Qingyu Shi, Siyuan Liu, Minrui Luo |
Malicious finetuning attacks pose a major safety threat against open-weight large language models (LLMs). However, existing alignment-stage defenses provide limited protection against strong attacks that use full-parameter finetuning. To address this issue, we...Malicious finetuning attacks pose a major safety threat against open-weight large language models (LLMs). However, existing alignment-stage defenses provide limited protection against strong attacks that use full-parameter finetuning. To address this issue, we propose Patcher, a novel and efficient adversarial training algorithm that alternates between an attack stage and a defense stage: the attack stage simulates many-step malicious finetuning and computes the resulting parameter displacement as an "attack vector", while the defense stage aims to preserve safe behavior under such perturbations. Compared with other adversarial training methods, a key feature of Patcher is that the attack vector can be reused across multiple defense updates, substantially reducing training costs. Theoretically, we characterize how the relative gradient estimation error is influenced by the number of attack steps during training and reuse-duration of attack vectors. Empirically, experiment results show that Patcher reduces Attack Success Rate by 67.5%, 53.4% and 54.8% on three benchmarks compared to the second-best method while preserving the model's utility. Moreover, Patcher remains effective under diverse model sizes and attack scenarios, and generalizes to prompt-based jailbreak attacks. Code is available at https://github.com/haomingwen/patcher
|
| 877 |
Inside the LLM Word Factory
2606.08562
|
cs.CL
|
Benzi Busigin, Yuval Pinter |
Transformer language models process input provided as subword fragments, but natural language semantics usually rely on word-level concepts. Detokenization is the process where models reconcile these two facts, aggregating subwords into word-level representati...Transformer language models process input provided as subword fragments, but natural language semantics usually rely on word-level concepts. Detokenization is the process where models reconcile these two facts, aggregating subwords into word-level representations through their computation. Prior work has found that this takes place mostly in early-to-middle layers, but so far the exact mechanics of the process have not been pinned down. We venture deep into detokenization using activation patching in controlled paired experiments that isolate the contribution of different model components, localizing English detokenization in Llama2-7B to a two-stage process at Layer 1. Attention transmits a token-specific signal from nonfinal subwords, using sequential relays if necessary, while the MLP composes it with the local embedding. This two-stage structure generalizes to twelve models from eight families, but the depth over which it takes place depends on the flavor of positional encoding: RoPE-based models detokenize over 1 to 5 layers, while learned-absolute models take 5 to 10. Finally, we provide a probe for determining the success of the detokenization process based on early-layer activations alone, performing at 0.94-0.97 AUROC depending on the amount of context.
|
| 878 |
UXBench: Benchmarking User Experience in AI Assistants
2606.09570
|
cs.CL
|
Mengze Hong, Xia Zeng, Zeyang Lei, Sheng Wang, Chen Jason Zhang |
As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. We present \textbf{UXBench}, the first user-centric benchmark grounded in real user feedback signals for evaluati...As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. We present \textbf{UXBench}, the first user-centric benchmark grounded in real user feedback signals for evaluating preference alignment and dialogue generation. The benchmark consists of three interconnected tasks, UX Judge, UX Eval, and UX Recovery, with 7,400 test instances extracted from over 70K interaction logs of a mainstream Chinese AI assistant. The dataset closely reflects real user distributions, covering 8 scenarios, 83 domains, and diverse failure patterns that pose severe challenges. Extensive experiments on 26 frontier language models provide novel insights into how well models perceive user experience and how improvements in model capability contribute to better dialogue engagement. Through analyses of model behavior and performance gaps, we document six important findings, demonstrating that user feedback prediction is a learnable capability and revealing different aspects that influence user experience. UXBench establishes a new evaluation landscape and calls for greater attention to tailored UX optimization, contributing toward a user-centric scaling law for the development of successful AI assistants. The full project is released at https://github.com/mengze-hong/UXBench.
|
| 879 |
The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models
2606.13993
|
cs.CL
|
Zachary Nicholas Houghton, Yu Zhou, Dan Pluth, Jordan Hosier, Vijay K. Gurbani |
One of the most central aspects of language processing is the ability to trade off between stored representations and abstract knowledge: one must retrieve stored representations, but also generate novel ones by applying productive rules. While recent work has...One of the most central aspects of language processing is the ability to trade off between stored representations and abstract knowledge: one must retrieve stored representations, but also generate novel ones by applying productive rules. While recent work has examined abstract knowledge in language models, holistic storage has received far less attention. We probe internal representations in both text-based LLMs and an ASR model, testing whether V+up phrasal verbs develop distinct representations as a function of frequency and predictability. All models show evidence of holistic storage driven by frequency and predictability, further supporting usage-based theories of language.
|
| 880 |
Can Agents Infer Environment from Interaction? Evidence from Agentic Automata Learning
2606.16576
|
cs.CL
|
Reef Menaged, Gili Lior, Shauli Ravfogel, Roee Aharoni, Gabriel Stanovsky |
We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction, a capability increasingly required in agentic tasks (e.g., reproducing an executable without access to its source ...We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction, a capability increasingly required in agentic tasks (e.g., reproducing an executable without access to its source code by interacting with it). In our setup, an agent should uncover a hidden deterministic finite automaton (DFA) by interacting with an oracle through (1) membership queries ("Does this string belong to the target language?'') and (2) equivalence queries ("Is this the target DFA?''). Agentic automata learning yields a scalable testbed with controlled task complexity, measurable interaction efficiency, and strong algorithms to compare against from the classic automata-learning literature. Evaluating state-of-the-art LLMs with a multi-turn agent scaffold, we find that while they are able to recover simple DFAs, their performance drops sharply as DFA size increases. Results improve with a more elaborate ReAct-style state-tracking scaffold, yet strong models still struggle with complex instances. Trajectory analyses reveal recurring failures in query planning, evidence integration, and hypothesis construction. These failures occur even though the models mention classic automata learning algorithms in their reasoning, and can implement and execute them when given access to a coding environment. Our results suggest that for state-of-the-art LLMs, identifying a solution to the problem is insufficient, and reliably executing the resulting plan still poses a distinct challenge.
|
| 881 |
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
2606.18216
|
cs.CL
|
Byung-Kwan Lee, Ximing Lu, Shizhe Diao, Minki Kang, Saurav Muralidharan |
Knowledge distillation transfers a teacher's competence to a small student but is brittle in the small-student regime: forcing the student to imitate logits from a much larger teacher hurts generalization on benchmark families beyond the training corpus. Reinf...Knowledge distillation transfers a teacher's competence to a small student but is brittle in the small-student regime: forcing the student to imitate logits from a much larger teacher hurts generalization on benchmark families beyond the training corpus. Reinforcement learning (RL) avoids logit imitation by training on the student's own rollouts. However, on questions where every rollout fails, yielding zero advantage and being silently discarded, injecting a stronger teacher's response into the policy gradient breaks the on-policy assumption and induces drift. We introduce Zone of Proximal Policy Optimization (ZPPO), inspired by Vygotsky's zone of proximal development. ZPPO keeps the teacher inside the prompt rather than the policy gradient. On hard questions, where the student's mean rollout accuracy is below half, ZPPO constructs two reformulated prompts. A Binary Candidate-included Question (BCQ) pairs one correct teacher response with one incorrect student response as anonymized candidates the student uses as references. A Negative Candidate-included Question (NCQ) aggregates the student's wrong rollouts into a single prompt to surface their shared failure modes. A prompt replay buffer keeps each hard question eligible for re-sampling until it either graduates, the student's mean rollout accuracy on it reaches half or more, or is FIFO-evicted under finite capacity, amplifying BCQ and NCQ inside the student's current zone of proximal development. We post-train Qwen3.5 students at four scales (0.8B-9B) as vision-language models with a 27B teacher and evaluate them on a 31-benchmark suite (16 VLM, 10 LLM, 5 Video); ZPPO outperforms off/on-policy distillation and GRPO, with the largest gains at the smallest scale.
|
| 882 |
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
2606.18394
|
cs.CL
|
Lanxiang Hu, Zhaoxiang Feng, Yulun Wu, Haoran Yuan, Yujie Zhao |
Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed only when acceptance remains high and dr...Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed only when acceptance remains high and drafting overhead stays low. This ceiling has been difficult to break because prior head-based SD methods face a causality-efficiency dilemma. Autoregressive drafters produce path-conditioned candidates that are effective for tree speculative decoding with higher acceptance length, but their drafting cost grows with tree depth. Bidirectional block-diffusion drafters generate all positions in one pass, but their branch-agnostic marginals can form individually plausible yet mutually inconsistent trees, wasting budget and reducing acceptance. We propose JetSpec, a head-based SD framework that combines one-forward drafting efficiency with branch-wise causal conditioning. JetSpec trains a causal parallel draft head over fused hidden states from the frozen target model, producing candidate trees whose scores align with the target model's autoregressive factorization. This enables JetSpec to convert larger draft budgets into longer accepted prefixes and higher end-to-end speedup. Across math, coding, and chat benchmarks on dense and MoE Qwen3 models, JetSpec consistently outperforms bidirectional-head and tree-based SD baselines. On H100 GPUs, JetSpec achieves up to 9.64x speedup on MATH-500 and 4.58x on open-ended conversational workloads, with further latency gains demonstrated through vLLM integration under realistic serving loads. Our code and models are available at https://github.com/hao-ai-lab/JetSpec.
|
| 883 |
Capability Provenance in Language Models: A Case Study in Social Reasoning
2606.19625
|
cs.CLcs.LG
|
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla |
We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social reasoning versus STEM reasoning in OLMo3-7B. Training-data attribution measures how strongly each training docume...We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social reasoning versus STEM reasoning in OLMo3-7B. Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities. We compute gradient-based attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, aggregate influence across WebOrganizer's 24-format x 24-topic taxonomy (576 bins), and contrast benchmark pairs in a 2x2 design that varies domain (social vs. STEM) and capability type (reasoning vs. knowledge): SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM. Social and STEM reasoning draw on qualitatively distinct corpus regions, and the contrast is sharper at the reasoning level than at the knowledge level. Targeted machine unlearning provides partial causal validation: forgetting high-attribution topics (e.g., Literature for SocialIQA) degrades the aligned benchmark more than within-topic random baselines. We open-source all code, data artifacts, influence scores, and checkpoints at https://github.com/HCAI-Lab-GT/capabilibara and https://huggingface.co/HCAI-Lab-GT.
|
| 884 |
Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Error Span Detection
2606.21559
|
cs.CL
|
Weilu Xu, Yunzhi Shen, Xinye Wang, Ranfei Dang, Shujian Huang |
Large language models (LLMs) have shown potential for reference-free span-level translation quality estimation (QE), yet existing approaches based on Multidimensional Quality Metrics (MQM) typically rely on fixed rubric configurations shared across translation...Large language models (LLMs) have shown potential for reference-free span-level translation quality estimation (QE), yet existing approaches based on Multidimensional Quality Metrics (MQM) typically rely on fixed rubric configurations shared across translation instances. However, translation instances often differ substantially in error complexity, ambiguity, and required evaluation granularity, making static rubric allocation suboptimal for span-level error detection. We find that larger MQM subtype spaces improve error coverage but also introduce more false positives, while different translation instances prefer different rubric granularities, suggesting that evaluation spaces should be allocated dynamically for each case. Therefore, we propose a case-specific dynamic rubric framework that adaptively constructs MQM evaluation spaces for individual translation instances. Unlike methods that generate fully free-form rubrics, our framework remains grounded in the predefined MQM taxonomy while dynamically selecting suitable subtype spaces and evaluation granularity for different cases. Experiments on span-level QE benchmarks from the Conference on Machine Translation (WMT) across multiple model scales demonstrate that the proposed framework consistently improves F1 and Matthews Correlation Coefficient (MCC).
|
| 885 |
Only Ask What You Don't Know: Grounded Delta Planning for Efficient Multi-step RAG
2606.22681
|
cs.CL
|
Wei-Chieh Chou, Xuanjun Chen, Jian-Ren Lin, Claire Lin, Hung-yi Lee |
Multi-hop question answering remains challenging for Retrieval-Augmented Generation (RAG) because existing approaches either propagate errors across iterative retrieval rounds or over-generate reasoning steps, increasing cost without improving accuracy. We pro...Multi-hop question answering remains challenging for Retrieval-Augmented Generation (RAG) because existing approaches either propagate errors across iterative retrieval rounds or over-generate reasoning steps, increasing cost without improving accuracy. We propose Grounded Delta Planning RAG (GDP-RAG), a plan-based framework that targets only the information delta based on three simple design choices: (1) preliminary retrieval to ground planning before execution, (2) a gap-conditioned planning prompt that asks only for missing information, and (3) a skeletal trajectory that pairs each subquery with a Thought capturing evidence from preliminary retrieval and carrying it through to the final answer. GDP-RAG focuses computation on unresolved gaps, yielding concise, reliable reasoning trajectories. Extensive experiments on HotpotQA, 2WikiMultiHopQA, and MuSiQue show that GDP-RAG achieves the highest accuracy (60.63%) among all compared systems while maintaining a cost-of-pass of 0.51, 22% lower than PAR-RAG (0.65) and 68% lower than KnowTrace (1.57), with no method achieving both higher accuracy and lower cost.
|
| 886 |
LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
2607.02513
|
cs.CLcs.LG
|
Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach, Verna Dankers |
LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged as a promising solution, with state-of-the-art (SOTA) methods often following a l...LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged as a promising solution, with state-of-the-art (SOTA) methods often following a localize-first, unlearn-second paradigm that targets specific model parameters. However, existing benchmarks evaluate unlearning solely at the output level, leaving open the question of whether unlearning truly erases knowledge from a model's parameters or merely obfuscates it, a concern reinforced by the success of resurfacing attacks. To bridge this gap, we introduce LACUNA: the first unlearning testbed with ground-truth parameter-level localization. LACUNA injects PII of synthetic individuals into predefined parameters of 1B and 7B OLMo-based models via masked continual pretraining, enabling direct evaluation of whether unlearning targets the weights responsible for knowledge storage. We use LACUNA to benchmark current SOTA unlearning methods and find that, despite strong output-level performance, existing methods are highly imprecise and susceptible to resurfacing attacks. We further show that when localization is successful, even a simple gradient-based unlearning method achieves strong erasure and robustness to resurfacing attacks, highlighting the importance of precise unlearning. We release LACUNA to complement behavioral evaluations and drive further advances in robust, localization-based unlearning.
|
| 887 |
Policy-Conditioned Constrained Decoding for Column-Level Access Control in Text-to-SQL
2607.12341
|
cs.CL
|
Ryoto Miyamoto, Xin Fan, Hayato Yamana |
Text-to-SQL is increasingly deployed across trust boundaries between data providers and users. Such deployment must balance three competing requirements: policy compliance, answer coverage, and bounded cost. Existing approaches typically decide refusal based o...Text-to-SQL is increasingly deployed across trust boundaries between data providers and users. Such deployment must balance three competing requirements: policy compliance, answer coverage, and bounded cost. Existing approaches typically decide refusal based on which columns a query mentions and enforce it stochastically. Whether a query is compliant, however, depends not only on which columns appear but on how they are used, and stochastic enforcement cannot deterministically rule out violations. We formalize this requirement as a column-use policy over semantic use: output, filter condition, and aggregation argument. We integrate the policy by aligning each role with grammar productions tracked by the decoder. The resulting system, PCC-SQL, applies a per-token logits mask that deterministically eliminates single-query column-use violations on the supported SQL fragment in a single decoding pass. Across three benchmarks and three open-source models, PCC-SQL achieves 0% Leakage Rate and Coverage up to 88.7% on Spider-CU, while staying within +10% tokens of direct prompting. We additionally assess semantic alignment with execution accuracy.
|
| 888 |
Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
2607.14108
|
cs.CL
|
Nyx Iskandar, Perla Gamez |
This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory. To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric defined pe...This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory. To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric defined per tool call indicating whether a tool is useful or whether it can be safely removed from the tool suite without affecting accuracy while increasing tool efficiency; in this paper, we determine the sign of marginal tool utility for each tool call in a trajectory using LLM-as-a-Judge. While much prior work has been done to develop techniques that improve tool use by LLMs and design evaluation methods measuring efficiency indirectly using accuracy as a proxy, our work is centered on measuring efficiency directly via the quantitative metric proposed in this paper in post hoc trajectory analyses. It is our intention that this work contributes to the frontier of LLM evaluation research as a springboard for future benchmark designs and agent harness engineering (specifically with regards to creating lean tool suites) that optimize for metrics that complement but are distinct from accuracy.
|
| 889 |
Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya
2607.15209
|
cs.CL
|
Hailay Kidu Teklehaymanot, Debela Desalegn Yadeta, Wolfgang Nejdl |
Multilingual pre-trained language models such as XLM-R perform well for major languages but struggle with low-resource Ge'ez-script languages, largely because Latin-script-centric tokenizers split their words into many subwords. We introduce VEXMLM, a vocabula...Multilingual pre-trained language models such as XLM-R perform well for major languages but struggle with low-resource Ge'ez-script languages, largely because Latin-script-centric tokenizers split their words into many subwords. We introduce VEXMLM, a vocabulary-extended variant of XLM-R targeting Amharic and Tigrinya. We train language-specific SentencePiece tokenizers on monolingual corpora, extend XLM-R's vocabulary with 30k Ge'ez-script subwords, and initialize each new embedding to the mean of the pretrained embeddings. VEXMLM undergoes two-stage training: (1) continued masked language modeling on the monolingual corpora and (2) supervised fine-tuning on question answering and named entity recognition (Amharic and Tigrinya) and sentiment analysis (Amharic). VEXMLM lowers tokenizer fertility below that of XLM-R and Glot500 on both languages, by 28.0% (Amharic) and 45.9% (Tigrinya) relative to XLM-R. Downstream, it modestly improves named entity recognition over XLM-R, scores below XLM-R on extractive question answering, and is comparable on sentiment analysis. An ablation on Tigrinya NER shows that vocabulary expansion alone lowers accuracy on out-of-vocabulary words (words that XLM-R's tokenizer cannot represent or splits into more pieces than the expanded tokenizer), and that continued pretraining is required for the expanded model to exceed the baseline. Vocabulary expansion thus makes Ge'ez-script tokenization substantially more efficient, while its downstream benefit depends on the task and on adapting the new embeddings through continued pretraining. Resources: GitHub repository | Hugging Face model.
|
| 890 |
Multi-level context Modeling for consistent expert selection in Mixture-of-Experts
2607.16427
|
cs.CL
|
Shuhan Huang, Naifan Zhang, Yuanbo Tang, Yang Li, Wai Kin Victor Chan |
Mixture-of-Experts (MoE) enables efficient scaling of Transformer models by routing tokens to a small subset of experts. However, existing routers typically condition expert selection on shallow or isolated token representations, which often produce unstable a...Mixture-of-Experts (MoE) enables efficient scaling of Transformer models by routing tokens to a small subset of experts. However, existing routers typically condition expert selection on shallow or isolated token representations, which often produce unstable and semantically inconsistent routing decisions across layers. In this work, we revisit expert selection from a representation perspective and identify context incompleteness as a key bottleneck limiting effective expert specialization. To address this issue, we propose Multi-level Context Fusion MOE (MCF-MOE), a framework that constructs context-aware representations by integrating complementary signals from cross-layer semantic aggregation and local token-level interactions, enabling more informative and consistent expert selection. Experiments on language modeling and understanding benchmarks demonstrate that MCF-MOE consistently improves routing consistency and downstream performance over strong MoE baselines, highlighting the importance of contextual completeness in expert routing. The code is available at https://github.com/shuhanhuang/MCF-MOE.
|
| 891 |
Confidently Deceptive: On the Relationship Between Confidence and Deception in LLMs
2607.20444
|
cs.CL
|
Ali Asad, Stephen Obadinma, Anshul Pattoo, Wenxuan Zhang, Xiaodan Zhu |
The increasing capabilities of large language models (LLMs) are being accompanied by deep-rooted risks of deceptive behaviours that cause models to produce misleading outputs in service of a contextually or experimentally induced goal. The harm posed by such b...The increasing capabilities of large language models (LLMs) are being accompanied by deep-rooted risks of deceptive behaviours that cause models to produce misleading outputs in service of a contextually or experimentally induced goal. The harm posed by such behaviours depends not only on the content of deceptive outputs but also how confidently models deliver them, since confidence has a major impact on how persuasive the communication is to end users. In this paper, we provide a comprehensive study on the crucial relationship between confidence and deception across existing deception benchmarks and different model families, while covering both verbalized numerical and logit-based aggregated confidence. Through this, we reveal how confidently models behave when being deceptive. We demonstrate that when producing deceptive rather than honest responses, models exhibit a gap between their belief (how likely they think a claim is to be true) and their commitment (how firmly they assert and would defend that claim). LLMs produce persuasive deceptive claims while reporting low belief in their factual correctness. Their reported commitment to deceptive responses can easily be increased through further prompting and preference fine-tuning, with smaller and condition-dependent changes in reported belief. However, we show that low reported belief remains comparatively invariant and provides a strong signal for detecting deception in the evaluated settings. Using only an API call, our approach achieves detection scores of up to 0.99 for induced deception and 0.89 for emergent deception. This ultimately shows how confidence can be a practical tool for detecting and diagnosing deceptive behaviour in LLMs.
|
| 892 |
Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
2607.22218
|
cs.CL
|
Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao, Yingyue Luna Luan |
Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from those of humans. Across three...Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from those of humans. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195 participants) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognise different ideas as creative.
|
| 893 |
Skill Training with Corruption and Reconstruction Loop
2607.27557
|
cs.CL
|
Mo Li, Zixin Yin, Qihao Wu, Ting Cao, Yunxin Liu |
Large Language Models (LLMs) often struggle in highly specialized domains. Rather than parameter-level adaptation of LLMs which is costly and difficult to interpret, external skills (often defined as text files) have been recently proposed to augment LLMs for ...Large Language Models (LLMs) often struggle in highly specialized domains. Rather than parameter-level adaptation of LLMs which is costly and difficult to interpret, external skills (often defined as text files) have been recently proposed to augment LLMs for specialized domains. However, such skills rely on costly active human annotations or passive summarization of high-quality examples. In this paper, we propose a self-supervised approach for agent self-evolution that learns domain-specific skills directly from existing high-quality human artifacts, without additional human annotations or external rewards. Inspired by diffusion models, our approach follows a forward-loss-backward process to reconstruct human artifacts by iteratively learning the agent's external skill library rather than updating its model parameters. Experiments on short-drama screenwriting demonstrate that our approach enables agents to autonomously extract generalizable writing skills from human-authored scripts and substantially improve domain-specific generation quality. Our approach provides a scalable paradigm for agents to continuously learn many kinds of complex skills from existing high-quality human artifacts. Code and project page: https://github.com/skilltraining-project/skill-training and https://skilltraining-project.github.io
|
| 894 |
Would You Walk to the Car Wash? Salience Bias in LLM Commonsense Reasoning
2607.28478
|
cs.CL
|
Zheng Wu, Chenhao Xue, Shijie Zheng, Yijie Lu, Cheng Yang |
Despite advances in complex reasoning, large language models (LLMs) can prioritize explicit input conditions over implicit task prerequisites. In everyday commonsense reasoning, this can lead to a failure we term Salience Bias: salient but task-irrelevant deta...Despite advances in complex reasoning, large language models (LLMs) can prioritize explicit input conditions over implicit task prerequisites. In everyday commonsense reasoning, this can lead to a failure we term Salience Bias: salient but task-irrelevant details (e.g., numerical values) draw models into computation while they overlook the physical or commonsense prerequisites of the task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, comprising 1,145 items in four trap dimensions. Evaluating 12 LLMs, we find substantial vulnerability across the tested models, with higher numerical distractor counts associated with lower trap-avoidance rates and trap recognition not always leading to avoidance. Further probing of sycophantic-compliance cases shows that the relevant commonsense can often be elicited when the original task framing is removed, suggesting a gap between recognizing a constraint and applying it during task execution. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings highlight the importance of applying commonsense constraints during task execution, and we release SaliTrap as a testbed for studying this gap. The codes are available at https://github.com/Wuzheng02/SaliTrap
|
| 895 |
TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking
2607.28680
|
cs.CLcs.LG
|
Yixin Peng, Kehao Li, Stefan Decker |
Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. Existing approaches typically rely on data preprocessing pipelines that retain either compact or extensive table content as contextual evidence, ...Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. Existing approaches typically rely on data preprocessing pipelines that retain either compact or extensive table content as contextual evidence, and then formulate entity linking as a language generation task for instruction-tuned models; recent systems further incorporate explicit reasoning to disambiguate challenging mentions. However, their training supervision is usually static: fixed preference data cannot adapt to the residual errors of an evolving model, while variations in reasoning length can bias sequence-level preference learning. To address these limitations, we present TELLER: Table Entity Linking through Learning from Errors and Reasoning. We first retrieve and rank Wikidata candidates and retain reduced table evidence in the prompt. The direct-answer path applies iterative direct preference optimization and refreshes its preference data with residual errors from the updated model. The reasoning path uses filtered and compressed chain-of-thought rationales for supervised fine-tuning, followed by our iterative length-normalized regularized preference optimization. On the TableInstruct entity-linking subset, the direct-answer path improves accuracy from 94.35\% to 94.50\%; on the MammoTab V2 evaluation set, it improves accuracy from 87.59\% to 88.20\%. The reasoning path improves accuracy from 92.90\% to 92.95\% on TableInstruct and from 79.09\% to 81.85\% on MammoTab V2, while maintaining high rates of complete reasoning generation. These results show that iterative preference learning benefits both concise entity prediction and explicit reasoning.
|
| 896 |
SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning
2608.00485
|
cs.CL
|
Tao Liu, Tao Feng, Xiangheng Li, Jinwang Song, Yifan Li |
Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods use execution correctness only as a trajectory-level reward, which provides limited guidance for identifying ...Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods use execution correctness only as a trajectory-level reward, which provides limited guidance for identifying the SQL decisions responsible for success or failure. We propose SERL-SQL, a selective execution-grounded reinforcement learning framework for multi-turn Text-to-SQL agents. SERL-SQL samples on-policy SQL interaction trajectories and uses a training-only teacher to re-score student actions with execution feedback. The resulting teacher--student likelihood gap is converted into bounded, masked weights that reweight GRPO advantages only on SQL and tool-action tokens. In this way, task rewards preserve the optimization direction, while execution hindsight provides localized credit assignment. Experiments on BIRD, Spider, and cross-domain benchmarks show that SERL-SQL achieves competitive performance, reaching 76.56% execution accuracy on BIRD-Dev and 89.92% on Spider-Test. Moreover, our reward-based selection strategy closely approaches the oracle Best-of-N upper bound and consistently outperforms consistency-based selection, showing that SERL-SQL produces high-quality candidates that can be reliably identified by lightweight execution-grounded rewards. Our code will be released at https://github.com/Ffunkytao/SERL-SQL.
|
| 897 |
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
2608.02617
|
cs.CLcs.LG
|
Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley |
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blind...We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content. Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference is a poor proxy for safety-critical performance. Models ranking highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures ($\leq -1$) on dimensions such as \emph{Harmlessness} and \emph{Accuracy}. These failures are unevenly distributed across specialties, creating domain-specific ``no-go zones'' not visible in aggregate rankings or single-number leaderboards. We further analyze contributing factors including prompt length, refusal and escalation behavior, and the relative contributions of safety-critical versus surface-level features. A substantial fraction of preference votes carry no positive safety signal, while feature decomposition shows that surface-level characteristics explain slightly more preference variation than safety-critical rubric differences. Finally, we introduce a clinically adjusted preference ranking combining pairwise preference with rubric-derived feedback, producing a more safety-aware ordering than raw Bradley--Terry strength alone. Our findings support evaluation practices that separate preference from safety, report safety-critical failure rates directly, and incorporate clinically grounded adjustments when ranking LLMs for clinical decision making.
|
| 898 |
E$^3$-Orch: Towards Effective, Efficient, and Extensible Agentic Orchestration with Reinforcement Learning
2608.04588
|
cs.CL
|
Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari |
Agentic orchestration enables multiple autonomous agents to solve complex tasks through adaptive decomposition, delegation, and execution. However, existing orchestrators often rely on hand-crafted logic and prompting strategies, limiting adaptation and genera...Agentic orchestration enables multiple autonomous agents to solve complex tasks through adaptive decomposition, delegation, and execution. However, existing orchestrators often rely on hand-crafted logic and prompting strategies, limiting adaptation and generalization across tasks and executor configurations. We propose E$^3$-Orch, a reinforcement learning framework for effective, efficient, and extensible agentic orchestration based on a milestone-plan-act workflow. Instead of planning all subtasks upfront, E$^3$-Orch organizes execution around milestones, a mid-level abstraction that scopes planning around meaningful intermediate objectives and allows orchestration decisions to adapt as execution progresses. For each milestone, the orchestrator builds a dependency-aware plan, assigns subtasks to suitable executors, and executes independent subtasks in parallel. We train the orchestration policy from execution feedback, using milestone and plan decisions as units for fine-grained credit assignment. Tree-structured rollouts compare alternative decisions under shared execution histories, while complementary rewards optimize task performance, execution cost, and planning completeness, including an uncertainty-aware performance reward for stochastic downstream outcomes. Across seven benchmarks, E$^3$-Orch achieves the best task performance under multiple executor configurations, improving over the strongest baselines by $0.7$--$3.8$ points and delivering $1.16$--$1.59\times$ higher intelligence efficiency. The learned policy also transfers to unseen executor configurations introduced only at evaluation time, supporting extensible agentic orchestration.
|
| 899 |
Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models
2608.05126
|
cs.CL
|
Yuezhang Peng, Yuxin Liu, Changfeng Gao, Zhifu Gao, Qian Chen |
Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for closed-set tasks after in-domain supe...Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for closed-set tasks after in-domain supervised fine-tuning, it faces significant challenges in leveraging in-context learning for open-domain tasks due to its ambiguous rule definitions. This work proposes Spoken Function Calling (SFC), a novel semantic understanding perspective that optimizes semantic understanding with structured rule definitions, to evolve beyond traditional closed-set SLU. Specifically, we curate and extend a suite of spoken functions based on traditional SLU datasets, construct a multi-agent system to synthesize the SFC-Bench dataset, evaluate the performance of Large Language Models (LLMs) and Large Audio Language Models (LALMs), and enhance the SFC capabilities of LALMs through post-training. Experiments demonstrate that SFC outperforms traditional SLU, substantially enhancing the semantic extraction accuracy for LLMs and LALMs.
|
| 900 |
Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
2608.05741
|
cs.CL
|
Hongrui Bao, Yubing Ren, Jinhan You, Fang Fang, Shi Wang |
Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, educational misuse, and platform governance. These concerns make robust detection of machine-generated text increasingly nec...Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, educational misuse, and platform governance. These concerns make robust detection of machine-generated text increasingly necessary. Recent zero-shot detectors mainly exploit probability-based statistical discrepancies, but they do not explicitly account for the training process of LLMs, which leaves a distinct generation mechanism insufficiently modeled and limits detection robustness. To address this issue, we propose EchoPrompt, a training-free detector based on latent prompt restoration. Our key intuition is that machine-generated text is typically produced conditioned on an upstream prompt, and this hidden dependency can be partially reactivated by prepending a unified generic prefix. Specifically, EchoPrompt restores a generic assistant-response context, measures the induced likelihood gain with an instruction-tuned model, calibrates it against the corresponding base model, and aggregates the resulting differences into a score that quantifies latent prompt dependency. Extensive experiments show that EchoPrompt achieves state-of-the-art performance among zero-shot detectors while maintaining strong robustness across challenging evaluation settings.
|
| 901 |
Reducing Pretraining-Generation Mismatch in Diffusion Language Models
2608.09424
|
cs.CL
|
Xiaocheng Lu, Huabin Liu, Song Guo, Jianguo Li |
Diffusion language models (dLLMs) generate text through iterative denoising, allowing multiple tokens to be predicted in parallel. However, pretraining may mask tokens throughout a sequence, whereas prompt continuation conditions on an intact prefix. This diff...Diffusion language models (dLLMs) generate text through iterative denoising, allowing multiple tokens to be predicted in parallel. However, pretraining may mask tokens throughout a sequence, whereas prompt continuation conditions on an intact prefix. This difference remains in conversion pipelines that denoise entire sequences during the stable stage. We propose Prefix-Conditioned Diffusion (PCD), which samples a boundary, preserves the prefix, and denoises the suffix. The training recipe also applies autoregressive supervision to the prefix. We evaluate PCD in the stable stage of a warmup, stable, and decay conversion pipeline, with inference unchanged. Matched experiments across model families show improvements in reasoning and coding over native diffusion training. A matched continuation study further shows that the advantage persists after a shared decay stage. In a separate reconstruction diagnostic, the full PCD recipe's advantage over native diffusion training reverses as more evaluation prefix tokens are masked. Controlled experiments show lower suffix reconstruction loss with a clean training prefix, an intact evaluation prefix, and no autoregressive loss.
|
| 902 |
ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval
2608.12720
|
cs.CL
|
Haolong Chen, Liang Zhang, Zhuo Li, Lei Xue, Guanrxu Zhu |
While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits performance on heterogeneous memor...While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits performance on heterogeneous memory queries, which often demand diverse evidence construction strategies. To address this, we introduce \textbf{ERSkill}, a retrieval-centric framework for evolving, skill-guided memory access. ERSkill compiles interaction histories into a structured memory store and represents retrieval behaviors as executable skills composed of fundamental primitives. At inference time, a trained router dynamically matches each query to a suitable retrieval skill to construct tailored evidence for answer generation. ERSkill co-evolves the skill set and the router during training. It employs an experience trie to efficiently record explored retrieval paths, alongside a double-frontier mechanism that separates oracle-side capability expansion from router-validated deployment. Experiments across multiple agent memory benchmarks demonstrate that ERSkill substantially outperforms strong non-evolving and evolving baselines. Notably, it improves the overall average across F1, BLEU-1, and LLM-judge scores by 31.3\% with Qwen3-Next-80B-A3B-Instruct and by 21.4\% with GPT-5.4-nano.
|
| 903 |
AQuA: Recursively Self-Improving Quantitative Trading Research Agents
2608.12841
|
cs.CL
|
Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li, Jason Ge |
We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises...We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. Each system records experimental results and uses them to guide subsequent proposals. Each operates in a fixed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs. The factor system, a manager-mediated multi-agent pipeline, discovers and combines factors into a signal that reaches a combined validation information coefficient of about $0.190$ on a crypto universe. The model system, a config-driven loop over a hybrid time-series architecture, reaches a per-stock information coefficient of $+0.0843$ on US equities and converts it into a threshold long/short strategy with a held-out Sharpe of up to $+2.50$ at a two-leg cost. The strategy is positive in every year from 2021 to 2025.
|
| 904 |
Token Distribution versus Data Volume: Domain Balancing in Multi-Domain Meeting Summarisation
2608.15935
|
cs.CL
|
Ashima Sood, Bryan Gardiner, Joan Condell |
Jointly fine-tuning an LLM on meeting-summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps, is the gain due to the distribution of tokens across domains, or merely to th...Jointly fine-tuning an LLM on meeting-summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps, is the gain due to the distribution of tokens across domains, or merely to the volume of data seen? We disentangle these factors by constructing balanced and natural (native-proportional) token mixtures at matched token budgets (2-32M) over five English meeting corpora, fine-tuning Mistral-7B with QLoRA, and evaluating per domain. Balancing redistributes quality, improving the data-scarce minority domains at a low cost to the data-rich ones. The trade favours balancing whenever the minority domains matter: their share under proportional allocation is fixed at 1-2% regardless of budget, so matching balanced quality on those domains requires far more total data. We further find that pruning low-value transcript lines removes ~15% of tokens from the conversational corpora at no measurable cost, and that balancing by tokens is not the same as balancing by examples. Fine-tuning one model per domain is competitive only on the data-rich domains and falls below the zero-shot model on the data-scarce ones. A two-annotator study of 741 judge-labelled facts validates our fact-level evaluation. Together these results give practitioners a basis for deciding when to balance an imbalanced multi-domain mixture, and on what unit.
|
| 905 |
PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX
2608.17379
|
cs.CL
|
Genghan Zhang, Yixin Dong, Chengze Fan, Zhichen Zeng, Yueming Yuan |
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and spe...We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.
|
| 906 |
WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing
2608.18486
|
cs.CLcs.LG
|
Wenbo Zhang, Xiang Ren |
When generating text, a Transformer produces representations of past tokens at every layer, but each layer can normally use only representations from the same depth. This restriction prevents the model from fully reusing information it has already computed. We...When generating text, a Transformer produces representations of past tokens at every layer, but each layer can normally use only representations from the same depth. This restriction prevents the model from fully reusing information it has already computed. We introduce WhiteMatter, which allows every layer to draw on past-token representations from any depth. A learned mixer selects the most useful depths for the current context and combines their representations into shared key-value (KV) cache channels. Sharing these channels across layers can reduce the cache size. Given the same number of training tokens, WhiteMatter with a full-size cache performs comparably to a standard Transformer with 50% more layers. With half the KV cache, WhiteMatter outperforms matched standard Transformers at two model scales, up to 1.3B parameters. Cross-layer connections, however, introduce dependencies that slow training and prompt processing. We address this problem with cyclic iteration, which updates interleaved groups of tokens in turn while processing the tokens within each group in parallel. On a reference model trained with exact autoregressive execution, cyclic iteration converges 12.5x faster than standard Jacobi iteration.
|
| 907 |
ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos
2608.19165
|
cs.CL
|
Thales Bertaglia, Catalina Goanta, Gerasimos Spanakis, Gunes Acar |
ChildSafeAds is a shared task on commercial content in YouTube videos likely to reach children and teenagers. The dataset contains 3,360 videos from 939 channels, each starting with a sponsor segment submitted to SponsorBlock. We pair each segment with its tra...ChildSafeAds is a shared task on commercial content in YouTube videos likely to reach children and teenagers. The dataset contains 3,360 videos from 939 channels, each starting with a sponsor segment submitted to SponsorBlock. We pair each segment with its transcript, video, and channel information, plus a sales or service page linked from the video description. Systems determine the type of offer being promoted (ST1), assign product categories (ST2), and identify potential legal risks (ST3). Evidence is divided into four cumulative access levels, from the transcript to the linked page, so performance can be compared against the cost of collecting additional data. YouTube's "Includes paid promotion" label was absent from 45.5% of the videos. GPT-5.4 produced the labels after the organiser team, including a legal expert, reviewed samples and iterated on the taxonomy, prompts and model choices. GPT-5.6-luna independently labelled the development set to measure cross-model agreement. The evaluation received 21 entries; the highest mean macro-F1 was 0.735, while the best transcript-only entry reached 0.652. Across four system papers, the linked promotional page mainly helped identify what was being sold and added little to compliance-risk classification. The competition also exposed sensitivity to rare labels, advertiser overlap across splits and substantial variation in the ST3 labels.
|
| 908 |
Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
2608.19920
|
cs.CL
|
Matthias Seeger, Zeyu Zhang, Vihang Patil, Konstantinos Benidis, Sebastian Schelter |
A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse att...A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It works for any KV cache policy, runs on a moderate hardware budget (e.g., a single Nvidia A100 GPU with 40 GB RAM), and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism). We also provide an efficient implementation of H2O sparse attention (the leading policy in our experiments) with dedicated scaled dot product attention kernel support. KeysAndValues (https://github.com/awslabs/keys_values), a new open source library for long-context inference and fine-tuning, provides easy-to-use and performant code for all methods discussed here.
|
| 909 |
Affix Cache for Diffusion Large Language Models
2608.26140
|
cs.CLcs.LG
|
Kaihua Liang, An Zhong, Xin Tan, Zafar Ayyub Qazi, Hong Xu |
Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding, but efficient inference support remains immature: unlike autoregressive models, whose requests reuse a shared prefix key-value (KV) cache, DLLMs use bidirectional attention, so a share...Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding, but efficient inference support remains immature: unlike autoregressive models, whose requests reuse a shared prefix key-value (KV) cache, DLLMs use bidirectional attention, so a shared context's KV states depend on the tokens still being decoded, leaving directly reused caches stale and full recomputation necessary. We present ACache, a cross-request cache reuse mechanism for shared spans, or affixes, at any position: prefix, infix, or suffix. ACache measures the influence of affix tokens on the masked generation region to identify a small request-specific subset as Anchor Tokens, and recomputes only their KV states while reusing the remaining affix cache. Built on state-of-the-art intra-request caching mechanisms, ACache recovers most of the accuracy lost to direct affix-cache reuse on average when recomputing around 20% of affix tokens, and at that budget preserves more accuracy than selection criteria adapted from prior cross-request cache-reuse systems. We co-design ACache with a modern inference engine, whose attention reads each request's recomputed Anchor KV states alongside one affix cache shared across concurrent requests. Against the same system with only intra-request caching, ACache cuts recompute latency by up to 56.7%, translating to as much as 1.71$\times$ end-to-end throughput, while reducing peak KV cache memory by up to 45.8%.
|
| 910 |
Recovering General Capabilities via Uncertainty-Calibrated Multi-Teacher On-Policy Distillation
2608.26735
|
cs.CL
|
Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo |
Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities. We study this trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized model learns from domain and general t...Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities. We study this trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized model learns from domain and general teachers on its own sampled trajectories. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, and advantage sign alone does not establish whether the proposed update direction is reliable. We propose Uncertainty-Calibrated MOPD (UCMOPD), which addresses these limitations through two complementary mechanisms. Golden-Gain Enhancement combines higher-temperature exploration with a standard-temperature anchor and retains trajectories whose positive learning signal matches or exceeds the prompt-specific anchor. Teacher-Endorsement Filtering then uses centered log-likelihood (CLL) to estimate each retained token's plausibility relative to the teacher's uncertainty and probabilistically preserves updates whose directions are supported by that endorsement. Across role-playing and medical-domain specialization, UCMOPD improves the general-capability average over standard MOPD by $4.48\%$ and $7.86\%$, respectively, while maintaining vertical-domain performance. Component ablations and diagnostic analyses support the intended roles of the two mechanisms: exposing and selecting stronger positive signals at the trajectory level and validating update directions through teacher endorsement at the token level.
|
| 911 |
Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers
2608.26762
|
cs.CLcs.LG
|
Markus Frohmann, Mahdiyar Ali Akbar Alavi, Elizabeth Lingg, Navid Rekabsaz |
In passage reranking, response ranking and multi-document question answering, LLMs can score several candidate documents or responses together in one prompt, each still receiving its own score. Such scorers are selected on ranking quality, but their scores det...In passage reranking, response ranking and multi-document question answering, LLMs can score several candidate documents or responses together in one prompt, each still receiving its own score. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or which chosen/rejected pair enters preference training. Because the candidates share that prompt, reordering them changes their scores. The same query over the same candidates should still yield the same decision. However, equal ranking quality does not imply equal decisions: on passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when reordered. No prompt-time change we test resolves that dependence: the only one that improves ranking quality does not measurably improve decision stability. We introduce order-consistency SFT (OC-SFT), which attenuates it in the weights by penalizing disagreement between a candidate's scores across orderings. It holds ranking quality and leads every decision-stability measure among trained scorers on all three tasks. It is also more stable on 12 base models than order-averaged distillation, which trains on labels averaged across permutations. One OC-SFT permutation retains sets that overlap more than ten averaged off-the-shelf permutations. A comparison of such scorers should therefore report what a threshold retains and a reader answers, not ranking quality alone. Code is available at https://github.com/thomsonreuters/presentation-dependence.
|
| 912 |
TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy
2608.27128
|
cs.CL
|
Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Dongnan Zheng |
Compressing a key-value cache before its next question is known requires choosing what to retain without knowing which evidence will matter. Value energy measures entry strength but does not distinguish isolated keys from those with many similar neighbors. We ...Compressing a key-value cache before its next question is known requires choosing what to retain without knowing which evidence will matter. Value energy measures entry strength but does not distinguish isolated keys from those with many similar neighbors. We introduce TwinKV, a training-free method that discounts value energy by nonlocal post-RoPE key frequency. Prefix attention allocates head capacities, while retained entries preserve their original keys and values under an exact storage budget. Across four language models, TwinKV exceeds five evaluated compressed baselines in mean score on LongBench, LooGLE, and RULER at 50\% KV removal. Component controls isolate the frequency contribution. On Llama-3.2-1B RULER at 75\% removal, normalized frequency weights average 0.95, yet change 7\% of nonprotected retained positions and improve value-only retention by about 5.5 points under both uniform and adapted capacities. Permuting the weights within heads weakens this gain. These results show that modest frequency corrections can change retention and answering outcomes, with effects that depend on the model and task.
|
| 913 |
A Formal Limitation on Learning Human Language From Textual Corpora
2608.28560
|
cs.CL
|
Emily Cheng, Ryan Cotterell |
Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling ...Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can surpass them. The bounds apply, moreover, to meaning spaces that are discrete or continuous. We provide empirical evidence in support of the theory through experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference.
|
| 914 |
Verbalizing Multi-Token Concepts in LLMs
2608.31084
|
cs.CL
|
Xijie Gong, Zimeng Huang, Tonghan Wang |
Lens methods inspect model computation by mapping intermediate activations to vocabulary tokens. Yet the concepts humans need to read out often span multiple tokens---entities, phrases, intermediate objects---making token-level readouts incomplete. Reliable mu...Lens methods inspect model computation by mapping intermediate activations to vocabulary tokens. Yet the concepts humans need to read out often span multiple tokens---entities, phrases, intermediate objects---making token-level readouts incomplete. Reliable multi-token readout with little model-specific preparation remains challenging. We introduce Concept Lens: token-level lens clues guide candidate concept search, then the model derives a representation for each candidate and scores it against the original activation. Across 2,400 multi-hop clozes on five LLMs (8B--70B), Concept Lens instantiated with J-lens and R-lens achieves average Rank@10 scores of 36.6\% and 54.5\%, respectively, compared with 21.7\% for Template Lens. Concept-swap interventions on derived concept representations shift model answers toward those associated with the replacement concepts. Further experiments show that Concept Lens can also reveal what a model recognizes along the way, beyond what appears in its final answer. Our code is available at https://github.com/XijieGo/c-lens
|
| 915 |
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
2609.03430
|
cs.CL
|
Heng Wang, Jielin Qiu, Wenting Zhao, Cheng Qian, Liangwei Yang |
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some esti...Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how much it will matter later, and keep the top-scoring ones. We show that the selection signal contributes almost nothing. Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks, it matches the strongest baseline in task performance while delivering 32-43% higher throughput than that method when deployed with vLLM. Controlled experiments explain this by showing that 1) the prompt is the fragile part of the cache, and most of the gap between selectors is just whether their selection signal happened to keep it; 2) the reasoning trace protects itself against eviction with redundancy at two levels, in the text (the model restates what it still needs as it works) and across attention heads (each keeps its own copy of the trace), so once the prompt is safe, a random draw retains enough copies of what the model still needs, and no score is required to pick them. Our code is publicly available at https://github.com/SalesforceAIResearch/Random-Attention.
|
| 916 |
Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG
2609.05152
|
cs.CL
|
Shuyu Guo, Shuo Zhang, Zhaochun Ren |
Retrieval-Augmented Generation (RAG) improves knowledge-intensive generation by conditioning language models on retrieved documents, but processing these documents becomes increasingly expensive as retrieval depth grows. Soft context compression reduces this c...Retrieval-Augmented Generation (RAG) improves knowledge-intensive generation by conditioning language models on retrieved documents, but processing these documents becomes increasingly expensive as retrieval depth grows. Soft context compression reduces this cost by encoding documents into compact continuous representations that can be precomputed and reused across queries. However, many existing methods train compressed models by distilling from a full-context teacher. When the teacher is wrong, such distillation can reinforce its errors, while teacher imitation provides no direct signal for improving beyond the teacher. We propose DEX-Comp, a two-stage training recipe that separates reliable imitation from targeted exploration. Pure Distillation learns only from teacher-correct questions to mitigate error propagation, while Hard Exploration applies outcome-based reinforcement learning to teacher-failed questions to directly optimize answer correctness. Across five open-domain QA benchmarks and retrieval depths from top-$5$ to top-$30$, DEX-Comp at $16\times$ compression outperforms all evaluated compression baselines and surpasses the untuned full-context RAG model in average accuracy, while reducing time-to-first-token by $4.4\times$--$23.7\times$. Evaluations across additional datasets and backbones further demonstrate its generalization.
|
| 917 |
Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance
2609.05797
|
cs.CL
|
Sripad Karne |
Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but a catch only prevents harm if the model would otherwise have complied. We measur...Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but a catch only prevents harm if the model would otherwise have complied. We measure the difference directly: we sample repeated responses from the target model, call a harmful prompt \emph{elicitable} if the model complies at least once, and report monitor recall separately on elicitable and non-elicitable prompts. Across six monitor configurations and three model families, spanning activation probes, fine-tuned text guards, and a 120B policy-conditioned reasoning classifier, recall on elicitable prompts falls 0.22 to 0.38 below recall on non-elicitable prompts at a fixed false positive rate. The prompts a monitor misses are 2.8 to 5.6 times more likely to be complied with than the prompts it catches. The gap replicates across three model families and appears also in text-only monitors entirely independent of the target model. This suggests that standard recall may overstate the protection monitors provide in practice, and that monitors should be evaluated against what their models will actually answer.
|
| 918 |
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
2609.06702
|
cs.CL
|
Kun Li, Zexuan Qiu, Tianhua Zhang, Irwin King, Helen Meng |
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency ...Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length. We introduce PARSER, which decouples reading from reasoning. A bank of lightweight subagents each bound to a single chunk read the entire document in parallel, while a lead agent reasons in depth through iterative scatter--gather rounds: at each round it broadcasts a query to all subagents, aggregates the returned evidence, and formulates a deeper follow-up query conditioned on what has been found so far. This decoupled design concentrates all learnable behavior in the lead agent, which is optimized with reinforcement learning, while the subagents remain frozen off-the-shelf models. On multi-hop QA with contexts ranging from 7K to 896K tokens, PARSER with a 4B backbone outperforms the strongest sequential memory baseline by 5.7 points on average and by 12.0 points at 896K tokens. Scaling to a 9B backbone, PARSER surpasses DeepSeek-V4-Pro by 6.3 points. Controlled experiments confirm that PARSER is robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.
|
| 919 |
KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
2609.10266
|
cs.CL
|
Xi Shi, Mengxin Zheng, Qian Lou |
Reusing key-value (KV) caches speeds up LLM inference by avoiding repeated computation on shared text. Standard prefix caching reuses a KV cache only when the LLM is the same and all preceding text is identical, but real workloads often break both conditions: ...Reusing key-value (KV) caches speeds up LLM inference by avoiding repeated computation on shared text. Standard prefix caching reuses a KV cache only when the LLM is the same and all preceding text is identical, but real workloads often break both conditions: RAG systems place different documents before the same one, agents with different system prompts read the same file or tool output, multi-agent workflows use specialized LLMs on shared material, and an updated model reads documents cached by its previous version. Because KV caches depend on both the preceding text and the model weights, direct reuse can reduce answer quality. Many methods repair or compress the reused cache, but each paper uses its own tasks, models, and cost measures, and existing benchmarks mainly test long-context processing or reuse of an unchanged prefix. We introduce KVShareArena, a benchmark and open evaluation framework for comparing them under the same conditions. KVShareArena has (1) reuse tests on 2,150 questions from three QA datasets, where the preceding text, the cache-writing LLM, or both change while the answering LLM and input stay fixed; (2) five dense and mixture-of-experts LLMs (4B-30B) and six LLM pairs where one version of an LLM reads caches written by another, for 33 model-dataset settings; (3) 11 repair and compression methods from six method classes; (4) four evaluation perspectives: answer quality, prefill computation, KV-cache memory, and latency; and (5) a common interface for adding new methods and an interactive leaderboard. Experiments yield two findings. First, both the quality loss from reuse and which repairs help depend on the LLM, even between two 8B models. Second, most repairs keep their quality when another LLM version wrote the cache, but a trained repair adapter loses quality in 12 of 18 pair-dataset tests. Code and data: https://github.com/xishi404/KVShare-Arena
|
| 920 |
Break Step: Recursive Training Resonates with Replayed Sampling Noise
2609.11149
|
cs.CLcs.LGcs.AI
|
Yangze Liu, Zhongyi Han |
How fast does a language model degrade when trained on its own outputs? Theory traces it to gradually accumulating errors, while experiments report repeated phrases within ten generations. Under a fixed sampling seed in vLLM, the fast loss of lexical diversity...How fast does a language model degrade when trained on its own outputs? Theory traces it to gradually accumulating errors, while experiments report repeated phrases within ten generations. Under a fixed sampling seed in vLLM, the fast loss of lexical diversity comes from the sampler. When vLLM serves a batch from one seeded sampling configuration, every request receives the same random draws, and a fixed seed replays them every generation. Fine-tuning raises the tokens that won, and the replayed draws let them win by more. Sharing across requests and replay across generations matter only together. Remove either one, by changing the shared seed every generation or by giving each request its own seed that repeats every generation, and the unique-4-gram fraction of two StableLM checkpoints stays near its starting value of about 0.98 through generation 3. Keep both, and the replayed shared seed takes seven checkpoints from five families to between 0.045 and 0.38 by then. Three generations of replay write the favoured phrases into the weights: decoded with one seed per request, the generation-3 weights of the replayed StableLM-2-1.6B chain recover most of their diversity, yet the phrase that filled every sample under the shared seed still opens 46% of them. Without replay, five checkpoints drift slowly, consistent with the gradual accumulation that theory describes, and three turn incoherent though their diversity scores stay high. One peer-reviewed model-collapse pipeline that fine-tunes Gemma-2-27B samples identical prompts under one seeded configuration, and three quarters of the rows it released for one iteration repeat nearly as often as one such batch copies them. A seed per request restores the fresh sample that stability analyses assume.
|
| 921 |
When the Wrong Key Wins: Understanding and Detecting Hallucinations in LLMs
2609.15106
|
cs.CL
|
Xuhan Tong, Haoyue Bai, Dawei Zhou, Naichen Shi, Jiawei Zhang |
Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a latent-key view of inference, where answer selection depends on competition among associations acquired during pre...Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a latent-key view of inference, where answer selection depends on competition among associations acquired during pretraining. We show that model predictions can be highly sensitive to individual query keywords, that these influential keywords exhibit entity-specific binding, and that their effects are systematically shaped by pretraining frequency. Multiple bindings can also compete and exhibit higher-order interactions within the same query. Based on this mechanism, we introduce a two-stage keyword-perturbation method for hallucination detection. By removing influential keywords and measuring how the model reorganizes its prediction, the method distinguishes errors caused by misleading key associations from correct decisions supported by diagnostic evidence. Across multiple models and benchmarks, perturbation provides a strong and transferable detection signal, reaching $0.910$ AUROC on probe-known ScientistQA. Finally, we extend the same probabilistic framework to four hallucination regimes: knowledge deficit, wrong knowledge, context distraction, and unstable inference. Their operational distributions across benchmarks provide diagnostic context for why different detector families succeed in different settings.
|
| 922 |
SlopShape: Identifying AI-Generated Commercial Web Content
2609.15369
|
cs.CL
|
Jochen Madler (Sitefire) |
Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it. We ask whether AI-generated t...Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it. We ask whether AI-generated text can be identified one level deeper, from structural signatures: how information is presented, in what order, with what evidence, and in what voice. We replicate StoryScope (Russell et al., 2026), which showed such patterns for AI-generated fiction, on commercial content: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models. A 203-feature instrument, applied by an LLM and validated in a human gold-annotation session (human-human kappa 0.939, human-model 0.951), detects AI posts from its 176 structural features alone at 97.0 macro-F1 on held-out companies, nearly unchanged (96.1) when every AI post is reworded by its own model. The signal characterizes and attributes: AI posts share a tidy, self-announcing shape, 68.6% are attributed to the correct source against a 16.7% chance rate, and human posts occupy rare structural configurations. All effects replicate StoryScope's, consistent in direction and at least as large in magnitude. We release pipeline, instrument, prompts, code, and aggregate artifacts.
|
| 923 |
Evaluating Losslessness in Speculative Decoding Under Finite-Precision Inference
2609.15504
|
cs.CL
|
Ilya Koziev, Leonid Sinev, Ivan Oseledets |
Lossless speculative decoding is typically defined at the algorithmic level: a speculative procedure proposes multiple tokens and a verification procedure is designed to preserve the output trajectory of an autoregressive reference model exactly. In practical ...Lossless speculative decoding is typically defined at the algorithmic level: a speculative procedure proposes multiple tokens and a verification procedure is designed to preserve the output trajectory of an autoregressive reference model exactly. In practical neural inference, however, this guarantee is implemented using finite-precision floating-point computations, and discrete token selection can amplify small numerical differences into divergent generation trajectories. We investigate this distinction using Orthrus, a hybrid autoregressive-diffusion architecture that performs self-drafting and self-verification within a frozen autoregressive backbone, as a representative case study. Across 1,190 prompts from 12 domains, exact trajectory matching under BF16 occurs for only 45\% of the authors' checkpoint generations and 43% of those from our independently trained model. The probability of matching is strongly associated with the response-conditional perplexity of the autoregressive reference, indicating that trajectory divergence is not uniform across inputs. Despite these divergences, Orthrus does not exhibit systematic degradation on the evaluated downstream tasks. In contrast, FP32 inference yields exact trajectory matching on all evaluated prompts. These results demonstrate a gap between algorithmic losslessness and its implementation under finite-precision arithmetic, and motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.
|
| 924 |
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
2609.15855
|
cs.CLcs.LG
|
Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury |
People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configura...People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leading models combined strong supportive conversation with combined-risk scores above 95, whereas risk exploration exposed substantial variation among lower-performing configurations. Therapeutic prompting produced configuration-specific gains concentrated among weaker models, while elevated reasoning produced no average improvement. K-Bench combines broader clinical coverage and configuration-scale comparison with a continuously updated public leaderboard whose operational test materials are protected from direct optimisation. The leaderboard is available at www.k-bench.ai.
|
| 925 |
A Data-free Universal Prior over Syntactic Structures
2609.16854
|
cs.CL
|
Ferm\'in Moscoso del Prado Mart\'in |
Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing theories assume that the probabilities of syntactic structures emerge from language-specific experience. An ...Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing theories assume that the probabilities of syntactic structures emerge from language-specific experience. An unexplored possibility is that these probabilities can have a universal component that is independent of any language-specific experience. Here, I show that a universal prior over syntactic structures emerges from a cognitively motivated model of human language production, in which words are progressively integrated into syntactic structure through network growth. The resulting prior assigns probabilities to syntactic structures --represented as dependency trees-- without fitting parameters to linguistic data. It assigns higher probabilities to attested than to random trees in all 138 typologically diverse languages examined. These prior probabilities correlate positively with those estimated from corpora in 33 of 34 languages. The results indicate that part of the probability structure of syntax can arise independently of language-specific learning. Linguistic experience may therefore refine probabilities that are already structured by the process of language production, rather than estimating them from scratch. This identifies a possible cognitive origin for part of the probability distribution over syntactic structures, linking language production and statistical learning while providing a data-independent structural bias for probabilistic models of language.
|
| 926 |
On-Demand Attention: Language Models Know When to Recall
2609.20734
|
cs.CL
|
Haibo Feng, Ruiqi Liang, Dongyang Jin, Hanyang Peng, Shiqi Yu |
Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, although the benefit of global access varies across prediction positions. We find that, before global att...Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, although the benefit of global access varies across prediction positions. We find that, before global attention is computed for the current step, the decoding states available after local computation in frozen pretrained models already contain information predictive of its benefit over local attention. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method: after local computation, a lightweight recall head decides whether to recompute the current step with global attention. ODA trains only the recall head with modest data and compute budgets, leaving pretrained weights unchanged and retaining the complete historical KV cache so that information skipped at one step remains available for later access. Experiments across model scales and families, including hybrid attention backbones, show that ODA recovers most of the performance lost under local attention while substantially reducing the frequency of global attention. Controlled long-context measurements in vLLM further show that GPU-side conditional execution translates fewer global reads into practical decoding speedups over full attention. These findings show that pretrained decoding states can support both token prediction and decisions about accessing distant information, allowing models to allocate global computation as needed during decoding.
|
| 927 |
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
2609.20784
|
cs.CL
|
Yan Yu, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang |
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill...Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.
|
| 928 |
Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction
2609.21231
|
cs.CL
|
Ruotian Wu, Bill E. Johnson, Gene Saunders, Osama Hamzeh, Ankit Vadehra |
Reference-based metrics for Grammatical Error Correction (GEC) such as M$^2$ and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving but phrased differently. We i...Reference-based metrics for Grammatical Error Correction (GEC) such as M$^2$ and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving but phrased differently. We introduce RM-EVAL, a reward model trained on human preference data from SEEDA, as a reference-free meta-evaluator that predicts human-like quality judgments at both full-sequence and partial-sequence levels. Beyond evaluation, we show that the same reward model can be used as a learning signal to improve GEC generation via Reward-Guided Text Generation (RGTG), which keeps a base GEC model frozen and performs online, reward-driven decoding. Across SEEDA, RM-EVAL achieves strong agreement with human rankings, and RGTG yields consistent gains in reward and external validation, demonstrating a unified framework for both assessing and enhancing GEC systems without relying on gold references.
|
| 929 |
Toward Personalized Sleep Guidance from Wearable Data Using Language Models
2609.22463
|
cs.CL
|
Yusheng Tan, Running Zhao, Sofia Angel, Ninghui Hao, Ash Arian |
Sleep monitoring using wearable data has shown promise for personal health, yet large language model (LLM)-based summarization and question answering remain insufficient for personalized sleep guidance. Training specialized models, however, often requires cost...Sleep monitoring using wearable data has shown promise for personal health, yet large language model (LLM)-based summarization and question answering remain insufficient for personalized sleep guidance. Training specialized models, however, often requires costly expert annotation. Moreover, privacy and accessibility concerns motivate lightweight, local deployment for end users. We present a two-stage framework to address these challenges. Specifically, in Stage~1, a multi-agent LLM pipeline reasons structured sleep guidance from unannotated wearable records, enabling scalable dataset construction. Stage~2 distills guidance reasoning trajectories into small language models (SLMs) through supervised fine-tuning and integrates a training-free Best-of-$N$ selection strategy to enhance inference. Experimental results demonstrate our method outperforms commercial general and medical LLMs and open-source models. Human evaluation further supports the quality of the generated guidance and the feasibility of personalized sleep guidance with SLMs.
|
| 930 |
Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
2609.22603
|
cs.CL
|
Nikhil Reddy Pottanigari, Ramin Fahimi, Noah Bolger, Sepideh Kharaghani, Ying Zhang |
Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near...Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings. Motivated by this, we introduce Semantic Scaffold, an evaluation framework that extracts a hierarchical representation of facts, questions, and entity attributes from a source text, labeling each as a main point or supporting detail, and reusing this structure as a fixed reference for scoring summaries. From this representation, we derive three diagnostic metrics: Fact Preservation Score (FPS), Question Preservation Score (QPS), and Entity Preservation Score (EPS), designed to reward the preservation of essential information while penalizing detail overload, and position them as interpretable diagnostics that remain informative where holistic axes collapse. Finally, we analyze four recurring failure modes of ROUGE and LLM-as-judge scores, demonstrating that scaffold-based evaluation remains informative where conventional metrics collapse.
|
| 931 |
Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure
2609.23551
|
cs.CL
|
Shuyang Xiang |
Standard positional encodings treat position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence,...Standard positional encodings treat position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator. Attention is compressed in every corpus, but compression alone is not diagnostic of true structure: an architecturally identical channel with density-matched random labels is compressed too. What distinguishes real structure is the depth of compression: the real-versus-random depth gap is resolvable in two of three corpora and not in the third, and real-structure depth varies more strongly across corpora than the control's. Comparing eight corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence), none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest. At the paragraph level that relation transfers as a common slope, but with the opposite sign to the corpus-level ranking: controlling for length, diversity, and position, paragraphs with more similar neighbors compress less deeply, with no detectable slope difference between any pair of corpora, while length, lexical-diversity, and position effects remain corpus-specific. Compression depth, not its location, is the reproducible signature of genuine paragraph structure in our setting.
|
| 932 |
Teaching a Moving Student: Rethinking the Curriculum of On-Policy Distillation
2609.25048
|
cs.CL
|
Lingxiang Hu, Tianle Xia, Yiding Sun, Ming Xu, Linfang Shang |
In on-policy distillation (OPD), the student determines which states receive teacher supervision. As its policy evolves, earlier response prefixes become less likely even though teacher-student disagreement on them persists. Under matched trajectory and optimi...In on-policy distillation (OPD), the student determines which states receive teacher supervision. As its policy evolves, earlier response prefixes become less likely even though teacher-student disagreement on them persists. Under matched trajectory and optimization budgets, neither more queries nor more frequent rollout resampling is uniformly beneficial. Current-policy rollouts outperform initial-policy rollouts at shorter response budgets, but this ranking reverses at longer budgets. Replaying initial-policy rollouts after current-policy training improves accuracy, whereas replaying fixed recent rollouts does not reproduce the gain. We propose R-OPD, a gradient-triggered curriculum that adaptively selects when to revisit initial-student trajectories. When changes in mean gradients fall within minibatch-level variation for two consecutive comparisons, training switches from the next iteration onward to initial-policy replay. Across eight mathematics benchmarks and three training-data orders, R-OPD improves average accuracy over continued current-policy sampling by 2.25/4.44 percentage points at 16K/32K for a 0.6B student and 2.81/6.15 points for a 1.7B student. With a 30B-A3B teacher, it improves 8B accuracy by 4.05 points at 32K. At 32K, R-OPD exceeds a fixed schedule of 40 current-policy updates followed by 20 replay updates by 1.39/1.66/1.85 points for 0.6B/1.7B/8B. At the same generation cap, R-OPD also produces longer responses, suggesting that well-timed revisits help students use more of their reasoning capacity.
|
| 933 |
MemoryAthena: Adaptive Routing over Latent and Generated Memories
2609.25853
|
cs.CL
|
Mingyuan Li, Guangsheng Yu, Juyuan Zhang, Xu Wang, Zhibo Man |
Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated rather than only retrieved. Mem...Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated rather than only retrieved. MemoryAthena uses three pathways: direct Engram retrieval (E), generation from retrieved Engram cues (GE), and generation from causal backbone states without consulting the memory table (GH). Generated memory is conditionally useful: it can complement E in one context but interfere with it in another. MemoryAthena therefore treats E as an anchor and learns when a generated representation should intervene. With the backbone, memory, generators, and readers frozen, a lightweight causal routing head is trained from counterfactual future-token likelihood advantages of GE and GH relative to E. At inference time, an admitted candidate modifies the E residual through bounded interpolation, while rejection recovers the direct pathway exactly. On question answering, MemoryAthena raises the five-task average from 37.65 to 39.28 over the direct pathway of the same checkpoint, while the six-task general-NLP average increases from 76.73 to 79.13. The complete memory-side system contains approximately 201M parameters, excluding the frozen backbone. Further analyses show complementary strengths among E, GE, and GH across tasks and inputs. These results support generated memory as a selective correction to direct retrieval and highlight routing when, which, and how strongly to intervene as the central challenge.
|
| 934 |
Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models
2609.27373
|
cs.CLcs.LG
|
Ke Wan, Chen Chen |
Recurrent-depth language models, such as looped Transformers, repeatedly apply shared network blocks to refine latent representations without generating explicit intermediate reasoning tokens. However, each step recomputes full attention over the entire contex...Recurrent-depth language models, such as looped Transformers, repeatedly apply shared network blocks to refine latent representations without generating explicit intermediate reasoning tokens. However, each step recomputes full attention over the entire context, repeating costly global routing. We study how attention routing evolves across recurrent depth and find a consistent separation in convergence timescales: attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests two stages of recurrent inference: early discovery of a sparse working set, followed by representation refinement over largely stable routing support. Motivated by this finding, we introduce WISE (Working-set Inference with Support Exploitation), a training-free method that uses unrestricted attention during early recurrent steps to discover a block-structured working set, then reuses its support in later steps while keeping attention weights and recurrent refinement dynamic. Controlled interventions show that multi-step discovery yields more effective working sets than first-step selection, and that support reuse better preserves model behavior than more restrictive forms of attention reuse. Across multi-hop QA benchmarks, WISE largely preserves full-attention performance. Matched context-scaling experiments reveal an increasingly favorable quality-efficiency tradeoff as routing support becomes sparser with longer contexts. A sparse-attention implementation achieves up to a 1.76x late-step attention speedup over native FlashAttention at 4K context. Code: https://github.com/tbn5pj/WISE_code.
|
| 935 |
MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression
2609.27380
|
cs.CL
|
Ke Wan, Yifan Wang, Liheng Lai, Chen Chen |
Retrieval-augmented generation often relies on multiple retrieved contexts that contain substantial redundancy, motivating context compression to preserve useful information under limited input budgets. Likelihood-based compressors can account for cross-contex...Retrieval-augmented generation often relies on multiple retrieved contexts that contain substantial redundancy, motivating context compression to preserve useful information under limited input budgets. Likelihood-based compressors can account for cross-context redundancy through sequential scoring, but this makes evidence scores dependent on context order. We show that permuting the same contexts under an unchanged compressor can substantially change which supporting evidence survives compression. We attribute this sensitivity to information preemption: earlier, partially relevant contexts can absorb credit for shared information, reducing the incremental scores of later, stronger evidence and increasing its risk of removal. Controlled pair-swap interventions provide direct empirical support for this mechanism by showing that placing stronger evidence before overlapping, partially relevant contexts can improve its survival. Based on this insight, we introduce MORSE, a compression-aware method for evidence-preserving context ordering. MORSE uses reverse query likelihood to construct an evidence-first anchor and to evaluate compressed candidate outputs, enabling compression-aware selection among alternative permutations. Across multi-hop Question Answering (QA) benchmarks, compression procedures, budgets, and scoring models, MORSE improves evidence retention over reverse ordering and generally outperforms matched random search, with downstream QA gains. Our code is available at https://github.com/tbn5pj/MORSE_code
|
| 936 |
Finding Icebergs in Language-Model Workflow: Diagnosing Latent Structural Fragility with Stochastic Semantic Evidence Graphs
2609.29703
|
cs.CL
|
Matthew F. Dixon, Bertrand Nortier, Miquel Noguer i Alonso |
AI-workflow governance cannot be reduced to checking the final answer: an apparently safe answer may rest on a fragile evidence path that ordinary evaluation cannot see, localize or govern. We call this hidden fragility a "structural iceberg": hallucinations a...AI-workflow governance cannot be reduced to checking the final answer: an apparently safe answer may rest on a fragile evidence path that ordinary evaluation cannot see, localize or govern. We call this hidden fragility a "structural iceberg": hallucinations and unsupported claims may form its visible tip, while consequential weakness remains submerged. Stochastic semantic evidence graphs (SSEGs) expose these icebergs by preserving workflow channels, propagating local uncertainty and identifying the hidden paths on which an apparently safe output depends. ALCE and RAGTruth show that visible failures at the tip -unsupported citations and hallucinated spans -rest on distinct submerged weaknesses and therefore require different interventions. Across retrieval, tool-use and controlled stress tests, SSEG localizes those weaknesses, supports targeted repair, produces no false automatic passes in 35,000 known-truth cases and reduces ToolSandbox review by 28.8% across 96 executions from two agent models. The same structural view carries into end-to-end governance: in a separately sealed 1,200-case FinGovBench study, adding SSEG to GPT-OSS-20B reduces unsafe releases from 452/660 to 8/660 while releasing all 540 safe cases and correctly distinguishing 592/600 matched workflow pairs. An unchanged-gate transfer to Qwen3-8B releases all 540 safe cases and none of 660 unsafe cases, whereas flat-UQ releases 520 unsafe cases. SSEG therefore moves governance below surface-level output checking, turning hidden evidence dependencies into auditable, path-specific decisions about intervention, revalidation and release.
|
| 937 |
A Benchmark Framework for Screening Automation in Systematic Reviews
2609.30298
|
cs.CL
|
Gauransh Kumar, Luciano Marchezan, Guillaume Genois, K\'evin Delcourt, Eugene Syriani |
Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance cl...Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening datasets. This paper presents a benchmark dataset of $45\,064$ labeled entries for evaluating LLM performance in SR screening across 32 curated secondary studies. It proposes an evaluation framework that accounts for class imbalance, i.e., the natural prevalence of excluded articles relative to included articles in SRs. It also introduces PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening. We also present a use case demonstrating the application of SRBench and PromptSR.
|
| 938 |
RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models
2609.31245
|
cs.CL
|
Pavithra P M Nair, Bhavik Talaviya, Shourya Bhushan, Rahul Pankajakshan, Seema Guruvadoo |
Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social...Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categories such as caste and urban-rural location. Existing LLM bias benchmarks, however, are largely designed around Western demographic categories and therefore miss key axes of economic disparity in the Indian context. We introduce RupeeBias, a benchmark for auditing demographic bias in LLM-generated economic guidance across Indian economic settings. RupeeBias consists of 39,150 prompts spanning four use cases: salary estimation, salary increment estimation, counter-offer recommendation, and service pricing recommendation. The benchmark follows a single-attribute counterfactual design, holding the description of the user's qualifications, experience, or service offering fixed while varying one demographic identifier at a time. RupeeBias covers 87 India-specific demographic identifiers across six axes: caste, religion, regional identity, gender, disability, and urban-rural location, with all prompts constructed in both English and Hinglish. We evaluate nine LLMs on RupeeBias and find systematic demographic disparities across all six axes. For otherwise identical prompts that differ only in demographic identifier, LLM-generated economic outputs differ by 20.2% on average. We publicly release RupeeBias to support future research on demographic bias in LLM-generated economic guidance across India-specific demographic and economic contexts.
|
| 939 |
Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting
2411.16769
|
cs.CLcs.LG
|
Zhi-Yi Chin, Pin-Yu Chen, Wei-Chen Chiu, Mario Fritz |
Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, human red-teaming is costly and inconsistent, driving the need for automatic tools that simulate realistic misuse attempt...Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, human red-teaming is costly and inconsistent, driving the need for automatic tools that simulate realistic misuse attempts. Existing methods either require white-box access, fail to generalize across defenses, or produce uninterpretable adversarial tokens, while generating fluent prompts that preserve the original harmful intent remains underexplored despite its practical relevance. We propose ICER, a black-box framework that addresses this gap through two components: an LLM-based rewriter that produces fluent, natural-language adversarial prompts, and in-context experience replay that accumulates successful jailbreaking patterns into a reusable prior. These components are integrated via bandit optimization, enabling ICER to efficiently balance exploiting proven attack strategies with exploring new ones. Experiments across six safety mechanisms show that ICER outperforms seven baselines under both standard and semantics-preserving evaluation, with over 30% of generated prompts transferring to commercial systems like DALL-E 3 and Midjourney.
|
| 940 |
Emergence of psychopathological computations in large language models
2504.08016
|
cs.CL
|
Soo Yong Lee, Hyunjin Hwang, Taekwan Kim, Yuyeong Kim, Kyuri Park |
Can large language models (LLMs) instantiate computations of psychopathology? In this work, we establish a computational-theoretical framework to provide an account of psychopathology applicable to LLMs. Based on the framework, we conduct experiments supportin...Can large language models (LLMs) instantiate computations of psychopathology? In this work, we establish a computational-theoretical framework to provide an account of psychopathology applicable to LLMs. Based on the framework, we conduct experiments supporting two key claims: first, that network-theoretic computational structures of psychopathology exist in LLMs; and second, that executing these computational structures results in psychopathological functions. We further observe that as LLM size increases, the computational structure of psychopathology becomes denser and the functions more effective. Taken together, the results suggest that network-theoretic computations of psychopathology may have emerged in LLMs. We discuss alternative explanations, including pattern matching, persona modeling, and semantic coherence, and argue that they are either complementary to our interpretation or less consistent with the data.
|
| 941 |
Building Intelligent Agents with Neuro-Symbolic Concepts
2505.06191
|
cs.CLcs.LG
|
Jiayuan Mao, Joshua B. Tenenbaum, Jiajun Wu |
This article presents a concept-centric paradigm for building agents that can learn continually and reason flexibly. The concept-centric agent utilizes a vocabulary of neuro-symbolic concepts. These concepts, such as object, relation, and action concepts, are ...This article presents a concept-centric paradigm for building agents that can learn continually and reason flexibly. The concept-centric agent utilizes a vocabulary of neuro-symbolic concepts. These concepts, such as object, relation, and action concepts, are grounded on sensory inputs and actuation outputs. They are also compositional, allowing for the creation of novel concepts through their structural combination. To facilitate learning and reasoning, the concepts are typed and represented using a combination of symbolic programs and neural network representations. Leveraging such neuro-symbolic concepts, the agent can efficiently learn and recombine them to solve various tasks across different domains, ranging from 2D images, videos, 3D scenes, and robotic manipulation tasks. This concept-centric framework offers several advantages, including data efficiency, compositional generalization, continual learning, and zero-shot transfer.
|
| 942 |
Decoding One Safety Trigger Token for Balancing Safety and Usability in Large Language Models
2505.07167
|
cs.CL
|
Haoran Gu, Handing Wang, Yi Mei, Mengjie Zhang, Yaochu Jin |
Large Language Models (LLMs) have been extensively used across diverse domains, including virtual assistants, automated code generation, and scientific research. However, they remain vulnerable to jailbreak attacks, which manipulate the models into generating ...Large Language Models (LLMs) have been extensively used across diverse domains, including virtual assistants, automated code generation, and scientific research. However, they remain vulnerable to jailbreak attacks, which manipulate the models into generating harmful responses despite safety alignment. Recent studies have shown that current safety-aligned LLMs undergo shallow safety alignment. In this work, we conduct an in-depth investigation into the underlying mechanism of this phenomenon and reveal that it manifests through learned ''safety trigger tokens'' that activate the model's safety patterns when paired with the specific input. Through both analysis and empirical verification, we further demonstrate the high similarity of the safety trigger tokens across different harmful inputs. Accordingly, we propose D-STT, a simple yet effective defense algorithm that identifies and explicitly decodes safety trigger tokens of the given safety-aligned LLM to activate the model's learned safety patterns. In this process, the safety trigger is constrained to a single token, which effectively preserves model usability by introducing minimum intervention in the decoding process. Extensive experiments across diverse jailbreak attacks and benign prompts demonstrate that D-STT significantly reduces output harmfulness while preserving model usability and incurring negligible response time overhead, outperforming ten baseline methods.
|
| 943 |
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
2505.17482
|
cs.CL
|
Chao Lei, Nir Lipovetzky, Krista A. Ehinger, Yanchuan Chang |
Recent reasoning-oriented LLMs have demonstrated strong performance on challenging tasks such as mathematics and science examinations. However, core cognitive faculties of human intelligence, such as abstract reasoning and generalization, remain underexplored....Recent reasoning-oriented LLMs have demonstrated strong performance on challenging tasks such as mathematics and science examinations. However, core cognitive faculties of human intelligence, such as abstract reasoning and generalization, remain underexplored. To address this, we evaluate recent reasoning-oriented LLMs on the Abstraction and Reasoning Corpus (ARC) benchmark, which explicitly demands both faculties. We formulate ARC as a program synthesis task and propose nine candidate solvers. Experimental results show that repeated-sampling planning-aided code generation (RSPC) achieves the highest test accuracy and demonstrates consistent generalization across most LLMs. To further improve performance, we introduce an ARC solver, Knowledge Augmentation for Abstract Reasoning (KAAR), which encodes core knowledge priors within an ontology that classifies priors into three hierarchical levels based on their dependencies. KAAR progressively expands LLM reasoning capacity by gradually augmenting priors at each level, and invokes RSPC to generate candidate solutions after each augmentation stage. This stage-wise reasoning reduces interference from irrelevant priors and improves LLM performance. Empirical results show that KAAR maintains strong generalization and consistently outperforms non-augmented RSPC across all evaluated LLMs, achieving around 5% absolute gains and up to 64.52% relative improvement. Despite these achievements, ARC remains a challenging benchmark for reasoning-oriented LLMs, highlighting future avenues of progress in LLMs. Our code is available at https://github.com/you68681/kaar.
|
| 944 |
AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
2506.08140
|
cs.CLcs.LG
|
Yifei Li, Hanane Nour Moussa, Ziru Chen, Shijie Chen, Botao Yu |
Despite long-standing efforts in accelerating scientific discovery with AI, building AI co-scientists remains challenging due to limited high-quality data for training and evaluation. To tackle this data scarcity issue, we present AutoSDT, an automatic pipelin...Despite long-standing efforts in accelerating scientific discovery with AI, building AI co-scientists remains challenging due to limited high-quality data for training and evaluation. To tackle this data scarcity issue, we present AutoSDT, an automatic pipeline that collects high-quality coding tasks in real-world data-driven discovery workflows. AutoSDT leverages the coding capabilities and parametric knowledge of LLMs to search for diverse sources, select ecologically valid tasks, and synthesize accurate task instructions and code solutions. Using our pipeline, we construct AutoSDT-5K, a dataset of 5,404 coding tasks for data-driven discovery that covers four scientific disciplines and 756 unique Python packages. To the best of our knowledge, AutoSDT-5K is the only automatically collected and the largest open dataset for data-driven scientific discovery. Expert feedback on a subset of 256 tasks shows the effectiveness of AutoSDT: 93% of the collected tasks are ecologically valid, and 92.2% of the synthesized programs are functionally correct. Trained on AutoSDT-5K, the Qwen2.5-Coder-Instruct LLM series, dubbed AutoSDT-Coder, show substantial improvement on two challenging data-driven discovery benchmarks, ScienceAgentBench and DiscoveryBench. Most notably, AutoSDT-Coder-32B reaches the same level of performance as GPT-4o on ScienceAgentBench with a success rate of 7.8%, doubling the performance of its base model. On DiscoveryBench, it lifts the hypothesis matching score to 8.1, bringing a 17.4% relative improvement and closing the gap between open-weight models and GPT-4o.
|
| 945 |
Reachability in Symmetric VASS
2506.23578
|
cs.CL
|
{\L}ukasz Kami\'nski, S{\l}awomir Lasota |
We investigate the reachability problem in symmetric vector addition systems with states (VASS), where transitions are invariant under a group of permutations of coordinates. One extremal case, the trivial groups, yields general VASS. In another extremal case,...We investigate the reachability problem in symmetric vector addition systems with states (VASS), where transitions are invariant under a group of permutations of coordinates. One extremal case, the trivial groups, yields general VASS. In another extremal case, the symmetric groups, we show that the reachability problem can be solved in PSPACE, regardless of the dimension of input VASS (to be contrasted with Ackermannian complexity in general VASS). We also consider other groups, in particular alternating and cyclic ones. Furthermore, motivated by the open status of the reachability problem in data VASS, we estimate the gain in complexity when the group arises as a combination of the trivial and symmetric groups.
|
| 946 |
ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
2508.01643
|
cs.CL
|
Ali Shiraee Kasmaee, Mohammad Khodadad, Mahdi Astaraki, Mohammad Arshi Saloot, Nicholas Sherck |
Retrieval-Augmented Generation (RAG) systems in chemistry heavily depend on accurate and relevant retrieval of chemical literature. However, general-purpose text embedding models frequently fail to adequately represent complex chemical terminologies, resulting...Retrieval-Augmented Generation (RAG) systems in chemistry heavily depend on accurate and relevant retrieval of chemical literature. However, general-purpose text embedding models frequently fail to adequately represent complex chemical terminologies, resulting in suboptimal retrieval quality. Existing embedding models for chemistry are outdated, and none is tailored to chemical literature retrieval, leaving a substantial performance gap. To address this challenge, we introduce ChEmbed, the first purpose-built family of domain-adapted text embedding models engineered for chemical literature retrieval. These models are fine-tuned via contrastive learning on a dataset comprising chemistry-specific text from the PubChem, Semantic Scholar, and ChemRxiv corpora. To create effective training data, we employ large language models to synthetically generate queries, resulting in approximately 1.7 million high-quality query-passage pairs. Additionally, we augment the tokenizer by adding 900 chemically specialized tokens to previously unused slots, which reduces the fragmentation of chemical entities, such as IUPAC names. ChEmbed also maintains an 8192-token context length, enabling retrieval of longer passages than many open-source embedding models allow. Evaluated on our newly introduced ChemRxiv Retrieval benchmark, ChEmbed outperforms state-of-the-art general embedding models, raising MRR@10 from 0.781 to 0.882 (+10.1 pp). It also substantially outperforms domain-specific embedding models such as Chemical-BERT, improving MRR@10 from 0.096 to 0.882. A role-based retrieval analysis using PubChem descriptions and ChEBI annotations shows that the improvement extends to chemical-role queries. ChEmbed represents a practical, lightweight, and reproducible embedding solution that effectively improves chemical literature retrieval.
|
| 947 |
MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
2510.16635
|
cs.CL
|
Juhyeon Lee, Wonduk Seo, Junseo Koh, Wonseok Choi, Hyunjin An |
Prompt optimization has become a practical way to improve the performance of Large Language Models (LLMs) without retraining. However, most existing frameworks treat evaluation as a black box, relying solely on outcome scores without explaining why prompts suc...Prompt optimization has become a practical way to improve the performance of Large Language Models (LLMs) without retraining. However, most existing frameworks treat evaluation as a black box, relying solely on outcome scores without explaining why prompts succeed or fail. Moreover, they involve repetitive trial-and-error refinements that remain implicit, offering limited interpretability or actionable guidance for systematic improvement. In this paper, we propose MA-SAPO: a new Multi-Agent Reasoning for Score Aware Prompt Optimization framework that links evaluation outcomes directly to targeted refinements. Specifically, in the Training Phase, multiple agents interpret evaluation scores, diagnose weaknesses, and generate concrete revision directives, which are stored as reusable reasoning assets. In the Test Phase, an analyzer agent retrieves relevant exemplars and assets for a new prompt, and a refiner agent applies evidence-based edits to improve the prompt and its response. By grounding optimization in structured reasoning, MA-SAPO ensures edits are interpretable, auditable, and controllable. Experiments on the HelpSteer1/2 benchmarks show that our framework consistently outperforms single-pass prompting, retrieval-augmented generation, and prior multi-agent methods across multiple evaluation metrics.
|
| 948 |
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
2511.20104
|
cs.CLcs.LG
|
Craig Dickson |
Prior work has shown that fine-tuning models on a narrow domain with misaligned data can lead to broad misalignment - a phenomenon termed "emergent misalignment" (Betley et al. 2025). While all tested models were susceptible to emergent misalignment, some mode...Prior work has shown that fine-tuning models on a narrow domain with misaligned data can lead to broad misalignment - a phenomenon termed "emergent misalignment" (Betley et al. 2025). While all tested models were susceptible to emergent misalignment, some models showed more resistance than others. Specifically the Qwen-2.5 family proved to be relatively resistant, while GPT-4o exhibited the strongest misalignment. In this paper we evaluate if current-generation open-weights models exhibit similar resistance to the Qwen-2.5 family and measure misalignment robustness over a range of model architectures and scales. We replicate the effect across nine modern open-weights models (Gemma 3 and Qwen 3 families, 1B-32B parameters). Models fine-tuned on insecure code generation show a 0.68% misalignment rate (compared to 0.07% for base models), matching the lower end of prior open-model results but dramatically lower than GPT-4o's 20%. We identify a critical format-dependent vulnerability: requiring JSON output doubles misalignment rates compared to natural language prompts (0.96% vs 0.42%). This suggests that structural constraints may bypass safety training by reducing the model's 'degrees of freedom' to refuse. These findings confirm emergent misalignment as a reproducible phenomenon in modern open-weights models, with rates substantially lower than observed in proprietary systems.
|
| 949 |
Why They Disagree: Decoding Differences in Opinions about AI Risk
2512.06350
|
cs.CL
|
\"Ozgecan Ko\c{c}ak, Phanish Puranam, Nghi Truong |
Identifying the reasons for disagreements between influential points of view on issues that affect the public can help produce informed policy responses, even if they do not bring disagreeing parties closer to agreement. We present a methodology for extracting...Identifying the reasons for disagreements between influential points of view on issues that affect the public can help produce informed policy responses, even if they do not bring disagreeing parties closer to agreement. We present a methodology for extracting reasoning chains - the sequences of premises that motivate or justify opinions - from natural discourse, and for characterizing the types of premises (facts, forecasts, definitions, causal beliefs, and evaluations) that make up these chains. We demonstrate the utility of this approach for two practical goals: diagnosing specific points of contention and aggregating arguments across speakers. We illustrate the methodology through an analysis of the debates on the nature of risks that AI poses to the public, using a corpus of interviews from the Lex Fridman podcast. We find that differences in perspectives among the podcast's guests on existential risk and employment risk from AI arise primarily from differences in causal premises and forecasts, whereas in the case of AI's effects on human social relationships, premises regarding what is valued and definitions about what counts as genuine human connection play a distinctively larger role. Our approach to analyzing reasoning chains at scale, using an ensemble of LLMs to parse textual data, can be applied to facilitate deliberation and aggregation of opinions on any topic.
|
| 950 |
Deep Delta Learning
2601.00417
|
cs.CLcs.LG
|
Yifan Zhang, Yifeng Liu, Mengdi Wang, Quanquan Gu |
Transformer residual streams are updated by addition. A sufficiently expressive residual block can represent content replacement, but the residual update itself has no operation that reads, compares, and replaces content. We introduce Deep Delta Learning (DDL)...Transformer residual streams are updated by addition. A sufficiently expressive residual block can represent content replacement, but the residual update itself has no operation that reads, compares, and replaces content. We introduce Deep Delta Learning (DDL), which applies the delta rule over network depth. Each layer reads the residual state along a learned direction, compares the readout with a learned target, and writes a gated rank-1 correction back along the same direction. A closed gate gives the identity map, and a unit gate overwrites the selected readout with the target. DDL works with the usual vector state or with an expanded state that stores several value channels, while attention and MLP blocks keep the original model width. We pretrain decoder-only models with approximately GPT-2 small and medium sizes on the FineWeb-Edu dataset. At both scales, every DDL variant has lower validation loss and higher average one-shot accuracy than the additive baseline, and the best expanded variants raise that average by 0.91 and 1.18 points.
|
| 951 |
Quantifying the Effect of Test Set Contamination on Generative Evaluations
2601.04301
|
cs.CLcs.LG
|
Rylan Schaeffer, Joshua Kazdan, Baber Abbasi, Ken Ziyu Liu, Brando Miranda |
As frontier AI systems are pretrained on web-scale data, test set contamination has become a critical concern for accurately assessing their capabilities. While research has thoroughly investigated the impact of test set contamination on discriminative evaluat...As frontier AI systems are pretrained on web-scale data, test set contamination has become a critical concern for accurately assessing their capabilities. While research has thoroughly investigated the impact of test set contamination on discriminative evaluations like multiple-choice question-answering, comparatively little research has studied the impact of test set contamination on generative evaluations. In this work, we quantitatively assess the effect of test set contamination on generative evaluations through the language model lifecycle. We pretrain language models on mixtures of web data and the MATH benchmark, sweeping model sizes and number of test set replicas contaminating the pretraining corpus; performance improves with contamination and model size. Using scaling laws, we make a surprising discovery: including even a single test set replica enables models to achieve lower loss than the irreducible error of training on the uncontaminated corpus. We then study further training: overtraining with fresh data reduces the effects of contamination, whereas supervised finetuning on the training set can either increase or decrease performance on test data, depending on the amount of pretraining contamination. Finally, at inference, we identify factors that modulate memorization: high sampling temperatures mitigate contamination effects, and longer solutions are exponentially more difficult to memorize than shorter ones, presenting a contrast with discriminative evaluations, where solutions are only a few tokens in length. By characterizing how generation and memorization interact, we highlight a new layer of complexity for trustworthy evaluation of AI systems.
|
| 952 |
What If TSF: Reframing Time Series Forecasting as Scenario-Guided Multimodal Forecasting
2601.08509
|
cs.CL
|
Jinkwan Jang, Hyunbin Jin, Hyungjin Park, Kyubyung Chae, Taesup Kim |
Recent advances in large language models (LLMs) have enabled time series forecasting to move beyond numerical observations and incorporate external information in multimodal settings. Such information can improve forecasting performance, but accuracy alone may...Recent advances in large language models (LLMs) have enabled time series forecasting to move beyond numerical observations and incorporate external information in multimodal settings. Such information can improve forecasting performance, but accuracy alone may not reveal whether models appropriately respond to it: a model may ignore relevant information, fail to distinguish scenarios with different implications, or overreact to irrelevant or weak signals. We introduce What If TSF (WIT), a benchmark for evaluating whether models effectively respond to future scenarios. WIT constructs controlled scenario sets by fixing the forecasting state while systematically varying future scenarios, enabling three complementary evaluations: Factual Comparison, which measures the predictive utility of factual scenario information; Relational Comparison, which evaluates whether forecasts satisfy expected direction, contrast, intensity, and restraint relations across scenarios; and Grounded Comparison, which assesses whether scenario-induced revision trajectories are empirically plausible relative to comparable real cases. Experiments show that factual accuracy gains do not consistently translate into appropriate responses to alternative scenarios, demonstrating the need to evaluate multimodal forecasting beyond predictive accuracy.
|
| 953 |
Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
2601.21358
|
cs.CL
|
Jiecong Wang, Hao Peng, Zhanyi Wang, Chunyang Liu, Guanlin Wu |
Chain-of-Thought (CoT) empowers Large Language Models (LLMs) to tackle complex problems, but remains constrained by the computational cost and early token commitments in discrete reasoning traces. Recent latent reasoning approaches attempt to optimize efficien...Chain-of-Thought (CoT) empowers Large Language Models (LLMs) to tackle complex problems, but remains constrained by the computational cost and early token commitments in discrete reasoning traces. Recent latent reasoning approaches attempt to optimize efficiency by performing reasoning within continuous hidden states. However, many such methods optimize latent states end to end without a trained interface for intermediate textual readout, and several representative configurations use a pre-defined number of latent steps during inference. In this work, we introduce \textbf{PLaT} (\textbf{P}lanning with \textbf{La}tent \textbf{T}houghts), a framework that decouples latent planning from verbalization. The Planner deterministically evolves latent planning states, while an independent Decoder provides textual readouts when needed. Answer-aware textual stopping allows the latent rollout to use a problem-dependent number of groups rather than a pre-specified chain length. PLaT achieves competitive coverage at larger $k$ in several mathematical settings, with lower Pass@1: on Llama-1B GSM8K, it reaches 80.59\% Pass@128 versus CODI's 72.37\%. These results support PLaT as a candidate-generation interface supplying multiple textual readouts for downstream verification or reranking.
|
| 954 |
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
2602.11767
|
cs.CLcs.LG
|
Aladin Djuhera, Swanand Kadhe, Farhan Ahmed, Syed Zawad, Heiko Ludwig |
Advances in large language models (LLMs) are driving a shift toward using reinforcement learning (RL) to train agents from iterative, multi-turn interactions across tasks. However, multi-turn RL remains challenging as rewards are often sparse or delayed, and e...Advances in large language models (LLMs) are driving a shift toward using reinforcement learning (RL) to train agents from iterative, multi-turn interactions across tasks. However, multi-turn RL remains challenging as rewards are often sparse or delayed, and environments can be stochastic. In this regime, naive trajectory sampling can hinder exploitation and induce mode collapse. We propose TSR (Trajectory-Search Rollouts), a training-time approach that repurposes test-time scaling ideas for improved per-turn rollout generation. TSR performs lightweight tree-style search to construct higher-quality trajectories by selecting promising actions and trajectory prefixes during rollout generation. This improves rollout quality while preserving stable policy optimization and remains compatible with standard policy-gradient optimizers by design. Across Sokoban, FrozenLake, and WebShop, TSR achieves success-rate gains of up to 15 percentage points and converges in fewer optimization steps, while trading additional training-time rollout compute for stronger policies that require no search at inference time. By moving search from test time to the rollout stage of training, TSR provides a modular mechanism for stronger multi-turn agent learning, complementary to existing frameworks and rejection-sampling-style selection methods.
|
| 955 |
Watch the Model Think: On-Policy Extraction of Activation Steering Vectors
2602.14143
|
cs.CLcs.LG
|
Xuanbo Su, Yingfang Zhang, Hao Luo, Huajun Bai, Guangyuan Dong |
When a model solves a problem on one attempt and fails it on the next, what separates the two is rarely the final answer token; it is the trajectory that reached it. Contrastive activation steering leaves that signal unused: CAA, SADI, RepE and ITI build their...When a model solves a problem on one attempt and fails it on the next, what separates the two is rarely the final answer token; it is the trajectory that reached it. Contrastive activation steering leaves that signal unused: CAA, SADI, RepE and ITI build their direction from experimenter-supplied text, recorded while the model reads rather than reasons. That choice also caps what the vector can express, since polarity must be written into the text, and a task judged only by outcome offers nothing to write it with. ROAST makes the trajectory itself the contrast: sample rollouts, let an outcome verifier split them into successes and failures, and contrast the reasoning that worked against the reasoning that did not. A matched teacher-forced control---rollouts, labels, answer text and pair counts held fixed, the trajectory alone stripped---points to the trajectory as what matters: on GSM8K at 0.6B the pairs alone buy +0.12 points while restoring the trajectories buys +6.05, the larger and only seed-robust step. Replacing the trajectory with an equal-length neutral prefix or another question's reasoning falls below no intervention. The two corpora are also far apart geometrically, a median 70+ degrees apart at both Qwen3 scales probed, beyond what a split-half null explains. Reading from rollouts calls for two corrections---keeping the full difference vector rather than Top-10% masking, and giving each question one vote rather than one per pair---and only grouped aggregation beats the unsteered baseline under 20% verifier noise. On parser-free benchmarks (GSM8K, MATH500, IFEval), ROAST is best in all six cells over two models, by up to +9.7, at +6.4% wall-clock and no added context; it also leads on six parser-scored benchmarks across three models. Across nine models (0.6B--122B, four families), ROAST improves on the unsteered model at every scale. Code: https://github.com/TomySu404/ORBIT
|
| 956 |
MASRubric: Auditing Information Flow in Multi-Agent Systems with Failure-Distilled Pitfall Rubrics
2602.23258
|
cs.CL
|
Yutong Wang, Siyuan Xiong, Xuebo Liu, Wenkang Zhou, Liang Ding |
While multi-agent systems (MAS) excel at complex reasoning, they are vulnerable to errors that intermediate agents introduce and downstream agents build upon. Auditing intermediate messages before they propagate requires an explicit standard, yet evaluation ru...While multi-agent systems (MAS) excel at complex reasoning, they are vulnerable to errors that intermediate agents introduce and downstream agents build upon. Auditing intermediate messages before they propagate requires an explicit standard, yet evaluation rubrics are typically authored by domain experts or written against a reference answer, neither of which is available for an unseen message at test time. We present MASRubric, a MAS information flow auditing framework with failure-distilled pitfall rubrics. Offline, trajectories on which the MAS has failed are automatically distilled into a reusable bank of pitfall criteria, each describing a recurrent error by its underlying misconception, the reasoning situations in which it arises, and the check that would expose it. Online, the criteria applicable to each intermediate message are retrieved from this off-the-shelf bank and checked one by one, and the resulting satisfaction rate decides whether the message is broadcast, returned to its author with diagnostic feedback for revision, or withheld. Empirical results demonstrate that MASRubric enhances MAS performance on both fixed and dynamic frameworks, achieving average accuracy gains of up to 2.83 points on math reasoning benchmarks and 1.74 points on code generation benchmarks. Further analysis shows that the retrieved criteria vary systematically with task types, and that the audit effort tracks task difficulty. Moreover, the bank transfers without re-mining to a system with a stronger backbone, which makes more adaptive and more efficient use of it. Our code and dataset are released at https://github.com/TonySY2/MASRubric.
|
| 957 |
SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution
2603.01327
|
cs.CLcs.LG
|
Kang He, Kaushik Roy |
Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-level software engineering (SWE), which demands (1) deep codebase navigation with effective context management for accurat...Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-level software engineering (SWE), which demands (1) deep codebase navigation with effective context management for accurate localization, and (2) systematic approaches for iterative, test-driven code modification to resolve issues. To address these challenges, we propose SWE-Adept, an LLM-based two-agent framework where a localization agent identifies issue-relevant code locations and a resolution agent implements the corresponding fixes. For issue localization, we introduce agent-directed depth-first search that selectively traverses code dependencies. This minimizes issue-irrelevant content in the agent's context window and improves localization accuracy. For issue resolution, we employ adaptive planning and structured problem solving. We equip the agent with specialized tools for progress tracking and Git-based version control. These tools interface with a shared working memory that stores code-state checkpoints indexed by execution steps, facilitating precise checkpoint retrieval. This design enables reliable agent-driven version-control operations for systematic issue resolution, including branching to explore alternative solutions and reverting failed edits. Experiments on SWE-Bench Lite and SWE-Bench Pro demonstrate that SWE-Adept consistently outperforms prior approaches in both issue localization and resolution, improving the end-to-end resolve rate by up to 4.3%.
|
| 958 |
LLM-as-a-judge validity is strongly task-dependent across physics assessment formats
2603.14732
|
cs.CL
|
Will Yeadon, Tom Hardy, Paul Mackay, Elise Agra |
As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is essential. We evaluate LLM-as-a-judge marking across four settings spanning three assessment formats - structured ques...As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is essential. We evaluate LLM-as-a-judge marking across four settings spanning three assessment formats - structured questions, written essays, and scientific plots - comparing GPT-5.2, Grok 4.1, Claude Opus 4.5, DeepSeek-V3.2, Gemini 3 Pro, and committee aggregations against human markers under blind, solution-provided, false-solution, and anchored conditions. We distinguish absolute accuracy from rank-order agreement, since a marking system can match the distribution of human marks while failing to order responses by quality. Across task types, performance is sharply task-dependent. For blind university exam questions ($n=771$) and secondary and university structured questions ($n=1151$), models show robust rank-order agreement with human markers (Spearman $\rho > 0.6$), with official solutions reducing error and strengthening agreement. False solutions degrade absolute accuracy, showing that models defer to provided references, but leave rank-ordering intact. Essay marking behaves fundamentally differently. Across $n=55$ scripts, each containing five essays ($n=275$ essays total), blind AI marking is harsher and more variable than human marking and adding marking guidance does not improve rank-order agreement. Anchored exemplars shift the AI mean close to the human mean and compress variance below the human standard deviation, but rank-order agreement remains near-zero. For code-based plot elements ($n=1400$), models achieve high rank-order agreement ($\rho > 0.83$) with near-linear calibration. LLM marking validity depended strongly on the assessment task, and this held for every contemporary model tested. The results also show that the reliability of the human benchmark constrains the claims that can be made about AI-human agreement.
|
| 959 |
FlashSampling: Fast and Memory-Efficient Exact Sampling
2603.15854
|
cs.CLcs.LG
|
Tomas Ruiz, Zhen Qin, Yifan Zhang, Xuyang Shen, Yiran Zhong |
Sampling from a categorical distribution is mathematically simple, but in large-vocabulary decoding, it often triggers extra memory traffic and extra kernels after the LM head. We present FlashSampling, an exact sampling primitive that fuses sampling into the ...Sampling from a categorical distribution is mathematically simple, but in large-vocabulary decoding, it often triggers extra memory traffic and extra kernels after the LM head. We present FlashSampling, an exact sampling primitive that fuses sampling into the LM-head matmul and never materializes the logits tensor in HBM. The method is simple: compute logits tile-by-tile on chip, add Gumbel noise, keep only one maximizer per row and per vocabulary tile, and finish with a small reduction over tiles. In tensor-parallel decoding, FlashSampling replaces the all-gather of logits with streaming peer-to-peer writes: This overlaps GPU-to-GPU communication with computation and HBM loads across up to 8 GPUs, with near-ideal scaling at large batch sizes. Our kernel is exact because $\text{argmax}$ decomposes over partitions; grouped variants for online and tensor-parallel settings are exact by hierarchical factorization of the categorical distribution. FlashSampling demonstrates kernel-level speedups on decode workloads across 4 different datacenter GPUs (H100, H200, B200, B300), and in end-to-end vLLM experiments, it reduces time per output token by up to $10%$ on the models we test. These results show that exact sampling, with no approximation, can be integrated into the matmul itself, consolidating the bandwidth-bound sampling step in an efficient epilogue.
|
| 960 |
When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal
2603.20433
|
cs.CL
|
Yen-Ting Piao, Jay Chiehen Liao, Wei-Tang Chien, Toshiki Ogimoto, Shang-Tse Chen |
While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evalua...While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evaluation pipeline that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability in the audio modality. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across LALMs: in-context demonstrations reliably improve format compliance but fail to improve the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from examples with audio, highlighting potential limitations in current cross-modal integration. We further probe how demonstrations are used through two complementary analyses, demonstration label shuffling and attention knockout on demonstration spans, both showing that LALMs leverage in-context examples primarily to establish the output label space and format rather than to learn a meaningful input-output correspondence.
|
| 961 |
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
2603.26089
|
cs.CLcs.LG
|
Christopher Ackerman |
The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind - is a human universal that enables us to navigate - and manipulate - the social world. It is supported by our abilit...The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind - is a human universal that enables us to navigate - and manipulate - the social world. It is supported by our ability to form mental models of ourselves and others. Its ubiquity in human affairs entails that LLMs have seen innumerable examples of it in their training data and therefore may have learned to mimic it, but whether they have actually learned causal models that they can deploy in arbitrary settings is unclear. We therefore develop a novel experimental paradigm that requires that subjects form representations of the mental states of themselves and others and act on them strategically rather than merely describe them. We test a wide range of leading open and closed source LLMs released since 2024, as well as human subjects, on this paradigm. We find that 1) LLMs released before mid-2025 fail at all of our tasks, 2) more recent LLMs achieve human-level performance on modeling the cognitive states of others, and 3) even frontier LLMs fail at our self-modeling task - unless afforded a scratchpad in the form of a reasoning trace. We further demonstrate cognitive load effects on other-modeling tasks, offering suggestive evidence that LLMs are using something akin to limited-capacity working memory to hold these mental representations in mind during a single forward pass. Finally, we explore the mechanisms by which reasoning models succeed at the self- and other-modeling tasks, and show that they readily engage in strategic deception.
|
| 962 |
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
2604.04410
|
cs.CLcs.LG
|
Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai, Sekitoshi Kanai, Masanori Yamada |
Aligning language models with human preferences is essential for ensuring their safety and reliability. Although most existing approaches assume specific human preference models such as the Bradley-Terry model, this assumption may fail to accurately capture tr...Aligning language models with human preferences is essential for ensuring their safety and reliability. Although most existing approaches assume specific human preference models such as the Bradley-Terry model, this assumption may fail to accurately capture true human preferences, and consequently, these methods lack statistical consistency, i.e., the guarantee that language models converge to the true human preference as the number of samples increases. In contrast, direct density ratio optimization (DDRO) achieves statistical consistency without assuming any human preference models. DDRO models the density ratio between preferred and non-preferred data distributions using the language model, and then optimizes it via density ratio estimation. However, this density ratio is unstable and often diverges, leading to training instability of DDRO. In this paper, we propose a novel alignment method that is both stable and statistically consistent. Our approach is based on the relative density ratio between the preferred data distribution and a mixture of the preferred and non-preferred data distributions. Our approach is stable since this relative density ratio is bounded above and does not diverge. Moreover, it is statistically consistent and yields significantly tighter convergence guarantees than DDRO. We experimentally show its effectiveness with Qwen 2.5 and Llama 3.
|
| 963 |
MMORF: A Multi-agent Framework for Designing Multi-objective Retrosynthesis Planning Systems
2604.05075
|
cs.CL
|
Frazier N. Baker, Trieu Nguyen, Reza Averly, Botao Yu, Daniel Adu-Ampratwum |
Multi-objective retrosynthesis planning is a critical chemistry task requiring dynamic balancing of quality, safety, and cost objectives. Language model-based multi-agent systems (MAS) offer a promising approach for this task: leveraging interactions of specia...Multi-objective retrosynthesis planning is a critical chemistry task requiring dynamic balancing of quality, safety, and cost objectives. Language model-based multi-agent systems (MAS) offer a promising approach for this task: leveraging interactions of specialized agents to incorporate multiple objectives into retrosynthesis planning. We present MMORF, a framework for constructing MAS for multi-objective retrosynthesis planning. MMORF features modular agentic components, which can be flexibly combined and configured into different systems, enabling principled evaluation and comparison of different system designs. Using MMORF, we construct two representative MAS: MASIL and RFAS. On a newly curated benchmark consisting of 218 multi-objective retrosynthesis planning tasks, MASIL achieves strong safety and cost metrics on soft-constraint tasks, frequently Pareto-dominating baseline routes, while RFAS achieves a 48.6% success rate on hard-constraint tasks, outperforming state-of-the-art baselines. Together, these results show the effectiveness of MMORF as a foundational framework for exploring MAS for multi-objective retrosynthesis planning. Code and data are available at https://github.com/ninglab/MMORF.
|
| 964 |
EffiPair: Improving the Efficiency of LLM-generated Code with Differential Execution Feedback
2604.05137
|
cs.CLcs.LG
|
Samira Hajizadeh, Suman Jana |
Large language models (LLMs) can generate functionally correct programs that differ substantially in execution efficiency. Existing inference-time optimization methods typically refine each candidate using pointwise runtime or profiling feedback, which identif...Large language models (LLMs) can generate functionally correct programs that differ substantially in execution efficiency. Existing inference-time optimization methods typically refine each candidate using pointwise runtime or profiling feedback, which identifies how costly an implementation is or where the cost arises, but offers limited guidance on how the computation should change. We introduce DIFFERENTIAL EXECUTION FEEDBACK (DEF), which compares implementations that are nearby in structural program space but separated in performance, turning their execution and implementation differences into directional optimization evidence. We demonstrate DEF in EffiPair, a training-free, test-time framework that pairs structurally similar programs with different efficiencies, distills their relative execution behavior into compact feedback, and iteratively refines a candidate pool. Across EvalPerf, Mercury, and ENAMEL, using GPT-4o mini, DeepSeek-V4.1 Flash, and GPT-5 mini, EffiPair achieves the highest value on each benchmark's official efficiency metric in all nine model-benchmark settings under matched evaluation conditions. Moreover, two contrastive refinement rounds improve efficiency over the selected initial draft in every setting while preserving or improving Pass@1. These results demonstrate the effectiveness of relational execution feedback as a lightweight signal for test-time code optimization.
|
| 965 |
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
2604.07394
|
cs.CLcs.LG
|
Quantong Qiu, Zhiyi Hong, Yi Yang, Haitian Wang, Kebin Liu |
The quadratic computational complexity of standard attention mechanisms presents a severe scalability bottleneck for LLMs in long-context scenarios. While hybrid attention mechanisms combining Full Attention (FA) and Sparse Attention (SA) offer a potential sol...The quadratic computational complexity of standard attention mechanisms presents a severe scalability bottleneck for LLMs in long-context scenarios. While hybrid attention mechanisms combining Full Attention (FA) and Sparse Attention (SA) offer a potential solution, existing methods typically rely on static allocation ratios that fail to accommodate the variable retrieval demands of different tasks. Furthermore, head-level dynamic sparsity often introduces severe computational load imbalance and synchronization long-tails, which hinder hardware acceleration during autoregressive decoding. To bridge this gap, we introduce Flux Attention, a context-aware framework that dynamically optimizes attention computation at the layer level. By integrating a lightweight Layer Router into frozen pretrained LLMs, the proposed method adaptively routes each layer to FA or SA based on the input context. This layer-wise routing preserves high-fidelity information retrieval while ensuring contiguous memory access, translating theoretical computational reductions into practical wall-clock speedups. As a parameter-efficient approach, our framework requires only 12 hours of training on 8$\times$A800 GPUs. Extensive experiments across multiple long-context and mathematical reasoning benchmarks demonstrate that Flux Attention achieves a superior trade-off between performance and inference speed compared with baseline models, with speed improvements of up to $2.8\times$ and $2.0\times$ in the prefill and decode stages.
|
| 966 |
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
2604.08826
|
cs.CLcs.LG
|
Mehran Taghian, Yunke Peng, Xing Huang, Yao Wang, Yaoyuan Wang |
Training large foundation models at low numerical precision is one of the most promising directions for reducing the compute and memory cost of modern AI. Recent 4-bit floating-point formats such as MXFP4 and NVFP4 can be applied to linear GEMM operations in L...Training large foundation models at low numerical precision is one of the most promising directions for reducing the compute and memory cost of modern AI. Recent 4-bit floating-point formats such as MXFP4 and NVFP4 can be applied to linear GEMM operations in LLMs, but their limited dynamic range introduces numerical instability that prior work addresses by stacking stabilization mechanisms, typically executed at higher precision and partially eroding the efficiency gains that motivate FP4. In this work, we argue that numerical format design is itself a first-class lever for stable FP4 training, and present the first systematic study of FP4 LLM pretraining on energy-efficient Huawei Ascend NPUs. We compare the recently proposed HiFloat4 (HiF4) format against both MXFP4 and NVFP4, holding one recipe fixed across all three formats across dense (OpenPangu-1B, Llama3-8B) and Mixture-of-Experts (Qwen3-MoE-30B) architectures and executing all linear and expert GEMMs in FP4. At matched storage --- NVFP4 and HiF4 both spend 4.5 bits per value --- the three formats differ far more in what they require before they will train at all than in final accuracy. HiF4 reaches a relative loss of 1.55\% with no stabilization, below fully stabilized MXFP4 (1.79\%) and below NVFP4 carrying the per-tensor scaling it cannot train without (2.00\%); NVFP4 diverges under every combination of stochastic rounding and Hadamard transform we tried. Our results suggest that stable, accurate FP4 training does not require an ever-growing stack of stabilization techniques; it requires the right numerical format.
|
| 967 |
Acceptance Dynamics Across Cognitive Domains in Speculative Decoding
2604.14682
|
cs.CL
|
Seifeldin Abdellatif |
Speculative decoding accelerates large language model (LLM) inference. It uses a small draft model to propose a tree of future tokens. A larger target model then verifies these tokens in a single batched forward pass. Despite the growing body of work on specul...Speculative decoding accelerates large language model (LLM) inference. It uses a small draft model to propose a tree of future tokens. A larger target model then verifies these tokens in a single batched forward pass. Despite the growing body of work on speculative methods, the degree to which the cognitive characteristics of a task affect acceptance probability remains largely unexplored. We present an empirical study of tree-based speculative decoding acceptance dynamics. Our study spans four well-established NLP benchmark domains: code generation, mathematical reasoning, logical reasoning, and open-ended chat. For this, we use TinyLlama-1.1B as the draft model against Llama-2-7B-Chat-GPTQ as the target. Over 99,768 speculative nodes collected from 200 prompts, we derive per-domain acceptance rates, expected accepted lengths, depth-acceptance profiles, and entropy-acceptance correlations. We find that task type is a stronger predictor of acceptance than tree depth. Furthermore, only the chat domain consistently yields an expected accepted length exceeding 1.0 token per step. We also show that the entropy-acceptance correlation is consistently negative but weak across all domains (rho in [-0.20, -0.15]). Counterintuitively, chat produces the highest entropy yet the highest acceptance rate. We attribute this divergence to the lexical predictability of RLHF-aligned register. These findings have direct implications for domain-aware speculation budgets and draft-model selection strategies. Index Terms--speculative decoding, large language model inference, tree attention, draft model, acceptance probability, LLM efficiency
|
| 968 |
Demystifying the Unreasonable Effectiveness of Greedy Alignment Methods
2604.17207
|
cs.CLcs.LG
|
Enoch Hyunwook Kang |
Greedy AI alignment methods based on Bradley-Terry reward-model fitting and KL-regularized Alignment are remarkably effective in both offline and online preference-learning pipelines, yet existing O(log T) online and O(1/epsilon) offline KL-regularized regret ...Greedy AI alignment methods based on Bradley-Terry reward-model fitting and KL-regularized Alignment are remarkably effective in both offline and online preference-learning pipelines, yet existing O(log T) online and O(1/epsilon) offline KL-regularized regret rate guarantees can seem pessimistic relative to their empirical performance. We argue that this mismatch reflects a distinction between two learning targets: KL-regularized regret measures how fast we can recover the KL-regularized aligned policy, whereas we often only ask how fast we can learn the reward function that induces the correct best response. To isolate this reward-learning component, we study the traditional temperature-zero regret criterion, which evaluates only the top-ranked response at inference time. Under realizability, compactness, and a uniform best-response margin over the fitted reward class, we prove that exact reference-logged offline RLHF achieves exponentially small expected temperature-zero regret, and therefore reaches expected regret at most epsilon with O(log(1\epsilon)) offline comparisons. We also prove that exact greedy online RLHF achieves bounded O(1) cumulative temperature-zero regret; the same guarantee extends to its equivalent pairwise online DPO formulation. By separating selector identification from full soft-policy recovery, our results provide a sharper theoretical explanation for the empirical efficiency of greedy alignment.
|
| 969 |
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
2604.25231
|
cs.CL
|
Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Gaurav Najpande, Manan Suri, Dinesh Manocha |
Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Recent vision-language models (VLMs) often achieve high answer accuracy on these ta...Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Recent vision-language models (VLMs) often achieve high answer accuracy on these tasks, yet correct answers do not guarantee that models ground their reasoning in the diagram regions that support the prediction. Models may instead rely on textual correlations or dataset artifacts without identifying the visual evidence required to verify the answer. This limitation prevents reliable evaluation of diagram reasoning and reduces interpretability. We introduce DRAGON, a benchmark for evaluating evidence-grounded visual reasoning in diagrams. Given a diagram, a question, and the correct answer, a model must predict bounding boxes that correspond to the visual elements required to justify the answer. These evidence regions may include answer-bearing components, textual labels, legends, axes, connectors, and other supporting structures involved in the reasoning process. The DRAGON dataset contains 11,664 annotated question instances from six diagram QA datasets: ChartQA, Circuit-VQA, InfographicsVQA, MapIQ, MapWise, and AI2D, with a 2,445-instance test set carrying human-verified evidence annotations and a standardized evaluation framework. Our evaluation on recent VLMs reveals that even Claude Opus 4.6 achieves a Grounding IoU of only 23.8% on AI2D, showing that faithful visual grounding remains an open challenge. DRAGON supports future research on models that ground their predictions in visual evidence.
|
| 970 |
BALAR : A Bayesian Agentic Loop for Active Reasoning
2605.05386
|
cs.CLcs.LG
|
Aymen Echarghaoui, Dongxia Wu, Emily B. Fox |
Large language models increasingly operate in interactive settings where solving a task requires multiple rounds of information exchange with a user. However, most current systems treat dialogue reactively and lack a principled mechanism to reason about what i...Large language models increasingly operate in interactive settings where solving a task requires multiple rounds of information exchange with a user. However, most current systems treat dialogue reactively and lack a principled mechanism to reason about what information is missing. We propose BALAR (Bayesian Agentic Loop for Active Reasoning), a task-agnostic outer-loop algorithm that requires no fine-tuning and enables multi-turn interaction between an LLM agent and a user. BALAR maintains a structured belief over latent states, selects clarifying questions by maximizing expected mutual information, and dynamically expands its state representation when the current one proves insufficient. We evaluate BALAR on three diverse benchmarks: AR-Bench-DC (detective cases), AR-Bench-SP (thinking puzzles), and iCraft-MD (clinical diagnosis). BALAR outperforms all baselines across the three benchmarks, with 14.6% higher accuracy on AR-Bench-DC, 38.5% on AR-Bench-SP, and 30.5% on iCraft-MD. We further study whether BALAR can serve as a teacher for a questioning policy through supervised fine-tuning (SFT), direct preference optimization (DPO), and dense-reward reinforcement learning (RL). Across 18 iCraft-MD replications, distilling BALAR into a Llama-8B yields relative gains in frozen-Qwen final-answer accuracy of 8.1% with SFT, 10.9% with DPO, and 12.1% with RL over the untuned policy.
|
| 971 |
Hindsight Compacts but Does Not Repair: Rethinking On-Policy Self-Distillation in Reasoning Models
2605.06188
|
cs.CL
|
Jaehoon Kim, Dongha Lee |
On-Policy Self-Distillation (OPSD) has emerged as a promising post-training method: requiring only privileged context, it enables token-level credit assignment from a self-teacher with hindsight, yielding higher accuracy with shorter responses. However, in thi...On-Policy Self-Distillation (OPSD) has emerged as a promising post-training method: requiring only privileged context, it enables token-level credit assignment from a self-teacher with hindsight, yielding higher accuracy with shorter responses. However, in thinking-enabled mathematical reasoning, where solving a problem takes long reasoning traces that explore, reflect, and backtrack, its length reductions persist but its accuracy gains largely do not. Does OPSD's hindsight signal, then, still repair failed trajectories, or does it only compact viable ones? We hypothesize compaction: in long traces, hindsight reveals which steps were unnecessary more readily than which steps would have repaired the solution. To test this, we apply OPSD separately to correct and incorrect rollouts. Training only on correct rollouts shortens responses by 18-29% while largely preserving accuracy, whereas training only on incorrect rollouts degrades it. The accuracy gap holds across three models, six benchmarks, and three seeds, and under a same-prompt control for difficulty. Neither branch raises the pass@k ceiling, and both suppress exploration and reflection markers, which viable traces can spare but failed ones need. Changing the divergence, enriching or reinjecting the privileged context, and training longer only move OPSD along the same accuracy-length tradeoff. Hindsight compacts reasoning the model can already produce but does not repair reasoning it cannot.
|
| 972 |
How Deep Can LLMs Learn to Reason? Expressiveness Is Key
2605.06638
|
cs.CL
|
Tianle Wang, Zhaoyang Wang, Guangchen Lan, Xinpeng Wei, Sipeng Zhang |
Reinforcement learning (RL) has been applied to improve large language model (LLM) reasoning, yet the systematic study of how training scales with task difficulty has been hampered by the lack of controlled, scalable environments. Observed LLM shortcomings in ...Reinforcement learning (RL) has been applied to improve large language model (LLM) reasoning, yet the systematic study of how training scales with task difficulty has been hampered by the lack of controlled, scalable environments. Observed LLM shortcomings in long-horizon reasoning have raised the prospect that they are fundamental to the autoregressive transformer architecture. To address this, we introduce ScaleLogic, a synthetic logical reasoning framework that offers independent control over two axes of difficulty: the depth of the required proof planning (i.e., the horizon) and the expressiveness of the underlying formal system. Our proposed framework supports a wide range of logics: from simple implication-only logic ("if-then") towards more expressive first-order reasoning with conjunction (''and''), disjunction (''or''), negation (''not''), and universal quantification ("for all"). Using this framework, we show that the RL training compute $T$ follows a power law with respect to reasoning depth $D$ ($T \propto D^{\gamma}$, $R^{2} > 0.99$), and the scaling exponent $\gamma$ increases monotonically with expressiveness, from $1.05$ to $2.60$. On downstream mathematics and general reasoning benchmarks, more expressive training settings yield both larger performance gains (up to $+10.66$ points) and more compute-efficient transfer compared to less expressive settings, demonstrating that what a model is trained on, not just how much it is trained, shapes downstream transfer. We further show that the power-law relationship holds across multiple RL methods, and curriculum-based training substantially improves scaling efficiency. More broadly, our results suggest that LLM shortcomings in long-horizon reasoning are not fundamental to the underlying architecture, and could be addressed by improved training methodology and data.
|
| 973 |
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
2605.07817
|
cs.CL
|
Brown Ebouky, Gabriele Carrino, Niccolo Avogaro, Christoph Studer, Andrea Bartezzaghi |
Human visual reasoning is governed by active vision, a process where meta-cognitive control drives top-down goal-directed attention, dynamically routing foveal focus toward task-relevant details while maintaining peripheral awareness of the global scene. In co...Human visual reasoning is governed by active vision, a process where meta-cognitive control drives top-down goal-directed attention, dynamically routing foveal focus toward task-relevant details while maintaining peripheral awareness of the global scene. In contrast, modern Vision-Language Models (VLMs) process visual information passively, relying on the static accumulation of massive token contexts that dilute the visual evidence as the reasoning chain grows. Here we propose GazeVLM, a VLM that learns to exert this oversight over its attention resources by generating gaze actions in its reasoning chain, in the form of <LOOK> tags specifying the coordinates of the region to inspect. During training, each <LOOK> block triggers a continuous suppression bias on the attention logits that dampens features outside the regions gazed at so far, and that lifts at the end of the block, restoring the global view. The model is trained with this bias as a teaching signal, first by Supervised Fine-Tuning on curated gaze-reasoning traces, then through Group Relative Policy Optimization (GRPO) with rewards for correct answers and valid grounding. At deployment, the external suppression is removed and the model, running its unmodified forward pass, steers its own attention to the regions indicated by the <LOOK> blocks, retaining most of the effect of the bias and emulating top-down control of spatial attention. This model thus transitions between global spatial awareness and localized focal reasoning without relying on external agentic tools like cropping, and without adding visual tokens from re-encoded patches. Applied to 4B-parameter backbones, GazeVLM demonstrates strong high-resolution multimodal reasoning on HRBench-4k and HRBench-8k, surpassing its base models by about 4 points, and outperforming agentic multimodal pipelines built around thinking with images by 5 to 12 points.
|
| 974 |
Let the Target Select for Itself: Data Selection via Target-Aligned Paths
2605.09404
|
cs.CLcs.LG
|
Huitao Yang, Hengzhi He, Tung Sum Thomas Kwok, Guang Cheng |
Targeted data selection seeks training examples from a candidate pool that improve downstream task performance. While trajectory-based selectors effectively guide this process, existing approaches typically construct reference states by warming up on the candi...Targeted data selection seeks training examples from a candidate pool that improve downstream task performance. While trajectory-based selectors effectively guide this process, existing approaches typically construct reference states by warming up on the candidate pool itself, thereby inheriting pool-dependent distributional biases. To address this coupling, we introduce Target-Aligned Candidate Selection (TACS), which constructs a short, capacity-constrained reference path using a compact target-validation proxy. Candidates are ranked by their normalized loss reduction along this trajectory, requiring only forward passes and allowing the path to be reused across distinct pools for a fixed model-target pair. Empirically, TACS achieves competitive downstream performance across controlled logistic, vision, and NLP benchmarks, while slashing selection compute by up to 69% and candidate storage to less than 10 MB in large-scale instruction tuning.
|
| 975 |
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
2605.10616
|
cs.CLcs.LG
|
Alan Arazi, Eilam Shapira, Shoham Grunblat, Mor Ventura, Elad Hoffer |
Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack native support for unstru...Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack native support for unstructured modalities such as text and image, and rely on frozen, pretrained embeddings to process them. On established Multimodal Tabular Learning benchmarks, we show that tuning the embeddings to the task improves performance. Existing benchmarks, however, often focus on the mere co-occurrence of modalities; this leads to high variance across datasets and masks the benefits of task-specific tuning. To address this gap, we introduce MulTaBench, a benchmark of 40 datasets, split equally between image-tabular and text-tabular tasks. We focus on predictive tasks where the modalities provide complementary predictive signal, and where generic embeddings lose critical information, necessitating Target-Aware Representations that are aligned with the task. Our experimental results demonstrate that the gains from target-aware representation tuning generalize across both text and image modalities, several tabular learners, encoder scales, and embedding dimensions. MulTaBench constitutes the largest image-tabular benchmarking effort to date, spanning high-impact domains such as healthcare and e-commerce. It is designed to enable the research of novel architectures which incorporate joint modeling and target-aware representations, paving the way for the development of novel Multimodal Tabular Foundation Models.
|
| 976 |
Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning
2605.15315
|
cs.CL
|
Jingjing Wang, Xiwen Chen, Wenhui Zhu, Huayu Li, Zhengxiao He |
LLM-powered coding agents spend the majority of their token budget reading repository files, yet much of the retrieved code is irrelevant to the task at hand. Existing learned pruners compress this context with a single-objective sequence labeler, collapsing a...LLM-powered coding agents spend the majority of their token budget reading repository files, yet much of the retrieved code is irrelevant to the task at hand. Existing learned pruners compress this context with a single-objective sequence labeler, collapsing all facets of code relevance into one score and one transition matrix. We show that this formulation creates a modeling bottleneck: a single CRF transition prior must serve heterogeneous retention patterns, including contiguous semantic spans and sparse structural support lines. We propose LaMR (Latent Multi-Rubric), a structured pruning framework that decomposes code relevance into two interpretable quality dimensions, semantic evidence and dependency support, each modeled by a dedicated CRF with dimension-specific transition dynamics. A mixture-of-experts gating network dynamically weights the per-rubric emissions conditioned on the query, and a final CRF layer on the fused emissions produces the aggregate keep-or-prune decision. To supervise each dimension without additional annotation cost, we derive multi-rubric labels from the existing training corpus via AST-based program analysis, simultaneously denoising the teacher's binary labels. By effectively filtering distracting noise, LaMR frequently matches or even outperforms unpruned full-context baselines. Experiments on four benchmarks (SWE-Bench Verified, SWE-QA, LCC, LongCodeQA) show that LaMR wins 12 of 16 head-to-head multi-turn comparisons. It saves up to 31% more tokens on multi-turn agent tasks and improves Exact Match by up to +3.5 on single-turn tasks, while performance is frequently enhanced by denoising the context, and any remaining drops are marginal.
|
| 977 |
Reducing Hallucination in Multimodal Large Language Models through Hard Grounding Preference Supervision
2605.16411
|
cs.CLcs.LG
|
Qinwu Xu |
Hallucination remains a major challenge in vision-language models (VLMs), particularly when linguistically plausible responses are unsupported by visual evidence. We study whether multimodal hallucination can be reduced by concentrating post-training supervisi...Hallucination remains a major challenge in vision-language models (VLMs), particularly when linguistically plausible responses are unsupported by visual evidence. We study whether multimodal hallucination can be reduced by concentrating post-training supervision on hard grounding boundaries, where preferred and rejected responses are semantically close but differ in their support from observable visual evidence. Under a frozen visual encoder and cross-modal alignment pathway, we first use supervised fine-tuning (SFT) to establish broad decoder-side multimodal behavior and then construct hard grounding preference pairs for Direct Preference Optimization (DPO). These pairs target evidence utilization, calibration, and grounding consistency across fine-grained recognition, spatial reasoning, OCR, ambiguous or insufficient evidence, and false-premise queries. The resulting DPO model consistently improves over its SFT initialization on DocVQA, TextVQA, MMBench, and VQAv2. Across approximately 839 additional multimodal examples, it further achieves 6.2--8.2 percentage-point higher pairwise win rates under three independent LLM judges. We additionally release HardVQA-DPO, a curated and growing resource with more than 3K hard-grounding SFT examples and an initial set of DPO preference pairs.https://huggingface.co/datasets/vlmgrounding/hardvqa-sft-dpo These results show that decoder-side adaptation alone can improve a meaningful subset of grounding failures under a fixed multimodal representation, while preserving broader multimodal capabilities.
|
| 978 |
To MRL or not to MRL: Text Embeddings are Robust to Truncation Without Matryoshka Learning, Except In Heavy Truncation Scenarios
2605.16608
|
cs.CLcs.LG
|
Sotaro Takeshita, Yurina Takeshita, Simone Paolo Ponzetto, Daniel Ruffinelli |
Matryoshka Representation Learning (MRL) is a widely adopted approach for training text encoders so they provide useful text representations at various sizes, available by simply truncating the resulting vectors at sizes pre-determined at training time. Recent...Matryoshka Representation Learning (MRL) is a widely adopted approach for training text encoders so they provide useful text representations at various sizes, available by simply truncating the resulting vectors at sizes pre-determined at training time. Recent works have shown that randomly truncating text embeddings has minimal impact in downstream performance unless vectors are reduced in size by at least 70%, suggesting that embeddings are already robust to truncation without the use of MRL. However, no prior work has compared random truncation to MRL, so it is unclear how the two methods compare as effective embedding reduction methods. In this paper, we study this by applying the same truncation used by MRL to models trained with and without MRL. Our results across several models and downstream tasks show that, unless heavily truncating embeddings (i.e. reducing their size by at least 80%), truncated embeddings of non-MRL models are competitive with, and often outperform models trained with MRL. This suggests that truncation robustness may not necessarily come from MRL, and that the choice of spending the additional training cost of MRL depends on whether heavy truncation is desired. We make our code available for reproduction.
|
| 979 |
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening
2605.17610
|
cs.CL
|
Shahriar Kabir Nahin, Hadi Askari, Muhao Chen, Anshuman Chhabra |
The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deployment. While most videos can be screened through fast pattern recognition, a small subset requires deeper reas...The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deployment. While most videos can be screened through fast pattern recognition, a small subset requires deeper reasoning over temporally complex content and nuanced policy constraints. Existing approaches typically rely on large vision-language models applied uniformly across all inputs, resulting in high inference costs and inefficient allocation of computation. We propose SafeLens, a video guardrail framework that introduces a fast-and-slow inference architecture for efficient and accurate content moderation with variable computational cost across inputs. Additionally, we construct a high-quality dataset by applying influence-guided filtering to the SafeWatch Dataset, retaining only 2.4% of the original data. To further address limitations of training-time scaling, we enable test-time reasoning by augmenting the filtered data with structured Chain-of-Thought traces. Across real-world and AI-generated video benchmarks, SafeLens achieves state-of-the-art performance, outperforming strong open-source video guardrails (e.g., SafeWatch-8B, OmniGuard-7B) and closed-source models (e.g., GPT-5.4, Gemini-3.1-pro) while significantly reducing inference cost, demonstrating that efficient design serves to be more effective than scaling data or model size alone.
|
| 980 |
EntmaxKV: Support-Aware Decoding for Entmax Attention
2605.21649
|
cs.CLcs.LG
|
Gon\c{c}alo Duarte, Miguel Couceiro, Marcos V. Treviso |
Long-context decoding is increasingly limited by KV-cache memory traffic as each generated token attends over a cache whose size grows linearly with context length. Existing sparse decoding methods reduce this cost by selecting subsets of tokens or pages, but ...Long-context decoding is increasingly limited by KV-cache memory traffic as each generated token attends over a cache whose size grows linearly with context length. Existing sparse decoding methods reduce this cost by selecting subsets of tokens or pages, but are designed for softmax attention, whose dense tails make any truncation discard nonzero probability mass. In contrast, $\alpha$-entmax produces exact zeros, turning sparse decoding from dense-tail approximation into support recovery: if the selected candidates contain the entmax support, sparse decoding remains exact. While recent entmax kernels enable efficient training, they do not address the autoregressive decoding bottleneck, where dense inference still streams the full KV cache before sparsity is known. In this work, we introduce EntmaxKV, an entmax-native sparse decoding framework that exploits sparsity before KV pages are loaded. EntmaxKV combines query-aware page scoring, support-aware candidate selection, and sparse entmax attention. We analyze truncation error through the dropped probability mass $\delta$, showing that output error is controlled by $\delta$ and vanishes when the entmax support is recovered. We further introduce a Gaussian-aware entmax selector that estimates the entmax threshold from lightweight page statistics, adapting the selected budget to the score distribution. Empirically, EntmaxKV drops less probability mass, retains more support tokens, and achieves lower output error than softmax-based sparse decoding at matched KV budgets. On long-context and language modeling benchmarks, it closely matches full-cache entmax while using a small fraction of the KV cache, achieving up to $2.88\times$ (softmax) and $5.5\times$ (entmax) speedup over full attention baselines as well as over $2\times$ speedup over softmax and entmax top-k at 1M context length.
|
| 981 |
Agents that Matter: How Subtle Choices Shape Agent Attribution in Multi-Agent Systems
2605.27621
|
cs.CL
|
Mingyu Lu, Yushan Huang, Chris Lin, Su-In Lee |
As multi-agent systems (MAS) become increasingly complex, identifying the contributions of agents is critical for system optimization. However, existing approaches lack a unified formulation for credit assignment. In this work, we formalize agent attribution a...As multi-agent systems (MAS) become increasingly complex, identifying the contributions of agents is critical for system optimization. However, existing approaches lack a unified formulation for credit assignment. In this work, we formalize agent attribution as a cooperative game, parameterized by the coalition distribution, intervention protocol, and target metric. Using this framework, we demonstrate that intervention protocols induce distinct games: Agent ablation isolates structural bottlenecks, whereas introspective LLM judges fail to faithfully approximate this behavior. We further find that Leave-One-Out (LOO) identifies bottleneck agents as effectively as combinatorial methods at much lower computational cost. Treating prompt optimization as an intervention protocol, we show that prompt-conditioned attribution can identify optimization targets and improve MAS performance. Finally, we apply our framework to audit a medical MAS, revealing that agent contributions to diagnostic accuracy and ethical behavior are often decoupled. By intervening on counterproductive roles, we observe an increase in ethics alignment while maintaining diagnostic accuracy. Overall, this work provides a principled approach for cost-effective MAS attribution and intervention.
|
| 982 |
AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
2606.02461
|
cs.CL
|
Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun |
Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes. Continual learning expects an agent to accumulate experience across a stream of tasks, improve over...Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes. Continual learning expects an agent to accumulate experience across a stream of tasks, improve over time, and avoid interference from irrelevant experiences. Unfortunately, existing benchmarks struggle to evaluate this problem rigorously. Most efforts focus on retrieval and reasoning over long-context conversations or documents, while recent lifelong-adaptation benchmarks often stream existing datasets whose cross-task dependencies are unknown, making it difficult to understand what an agent learns and reuses over time. This paper presents an evaluation framework AgentCL for continual learning in agents, centered on distinguishable task streams and metrics for CL properties. AgentCL distinguishes three cross-task relations by the availability of reusable knowledge. In dependent streams, later tasks can reuse knowledge from earlier ones. In conventional streams, whether such reuse exists is unknown. Besides, independent tasks are held out with little transferable knowledge. We use the benchmark to evaluate non-parametric memory designs for continual learning. To diagnose how memory design choices affect continual learning, we develop MemProbe, a probing method that stores interactions, insights, and skills, while filtering unreliable experiences during consolidation. Empirical analysis across coding, deep research, and language understanding/reasoning tasks shows that conventional streams offer limited ability to distinguish memory designs, whereas dependent streams more clearly distinguish their plasticity. Meanwhile, conventional streams and independent tasks often yield limited gains and can expose memory-induced degradation. These results highlight the need for stronger memory designs that balance plasticity and stable reuse.
|
| 983 |
Large Language Models Hack Rewards, and Society
2606.04075
|
cs.CLcs.LG
|
Wei Liu, Xinyi Mou, Hanqi Yan, Zhongyu Wei, Yulan He |
Reinforcement learning (RL) has become a dominant post-training paradigm, enabling large language models (LLMs) to learn from rewards. We observe that societal regulations are structurally similar to reward functions. They define measurable outcomes, threshold...Reinforcement learning (RL) has become a dominant post-training paradigm, enabling large language models (LLMs) to learn from rewards. We observe that societal regulations are structurally similar to reward functions. They define measurable outcomes, thresholds, and exceptions, while often leaving institutional intent only partially specified. We hypothesise that the RL training process may exploit these gaps and therefore ask whether models' well-known tendency to hack reward functions during RL can scale into a more consequential failure mode named societal hacking: discovering loopholes in the rules society runs on. To study this phenomenon, we introduce SocioHack, a sandbox of 72 societal environments, and find that within these environments, reward hacking naturally emerges and leads to regulatory loophole discovery. Models learn to hack the social rules and generate strategies that remain technically compliant while defeating regulatory intent, and current LLM safeguards provide only limited mitigation. Therefore, collecting in-the-wild feedback for model training requires greater caution, and we need a next-generation post-training paradigm for safely iterating LLMs in real society.=
|
| 984 |
Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems
2606.08034
|
cs.CL
|
Muhammad Falensi Azmi, Ikhlasul Akmal Hanif, Vallerie Alexandra Putra, Adi Yeltay, Abdullah Mubarak |
Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions. However, existing symbolic benchmarks mostly remain limited to mathematical reasoning, lack visual grounding, and are predominant...Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions. However, existing symbolic benchmarks mostly remain limited to mathematical reasoning, lack visual grounding, and are predominantly in English. In this work, we introduce Sci-Rho (Science Rhobustness), a dynamic benchmark for visually-grounded STEM problems spanning five subjects and seven languages, comprising 4,242 problem templates (606 per language) crafted by domain experts, including Olympiad medalists. Each template is implemented as executable Python code that generates diverse but equivalent problem instances by varying numerical values, visual patterns, geometric shapes, color schemes, and function types, resulting in 42,420 instances in total, each paired with reasoning steps and ground-truth solutions. We evaluated 17 state-of-the-art VLMs and discovered a noticeable gap between worst-case accuracy (defined as the proportion of problem templates that a model answers correctly across every generated variation) and average accuracy. We also discovered that smaller models show noticeable performance degradation across languages, whereas proprietary and larger models remain robust. Step-level evaluation reflects this same trend, revealing a significant gap between average F1 and worst-case F1 scores. Finally, our inspection of attention heads of a VLM reveals substantial cross-lingual variation in the relative attention allocated to image tokens compared to text tokens. Our work highlights the importance of evaluation beyond static benchmarks as a metric to measure the quality of VLMs.
|
| 985 |
Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations
2606.10281
|
cs.CL
|
Aniket Anand, Yiwei Hou, Daniel Fields, Alex Kantchelian, David Tao |
This paper presents AuditBench, a new benchmark dataset for evaluating the capabilities of LLMs at investigating security-related system audit logs. We design and use this benchmark to explore the performance of LLMs on four log-investigation tasks that incide...This paper presents AuditBench, a new benchmark dataset for evaluating the capabilities of LLMs at investigating security-related system audit logs. We design and use this benchmark to explore the performance of LLMs on four log-investigation tasks that incident response teams commonly perform, ranging from triaging alerts generated by detectors to identifying persistence mechanisms on compromised systems. AuditBench consists of system audit logs collected from Linux and Windows machines, and spans over 50 different security investigation scenarios, including both malicious and benign activity. Using our benchmark, we evaluate and analyze the performance of five frontier LLMs at analyzing audit logs for attack investigations. Our analysis illuminates how LLM performance and error profiles vary according to different design choices, such as differences in model size, data representation, prompt construction, and specific investigation tasks. Additionally, we characterize the quality of the explanations produced by LLMs and the types of errors that models make across our benchmark. Collectively, our work provides a foundation for assessing the capabilities of LLMs for investigating security logs, novel insights for practitioners using LLMs in security operations, and important directions for future research.
|
| 986 |
Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis
2606.10381
|
cs.CL
|
Ruobing Jiang, Dawei Fu, Cheng Jiang, Tianyi Yang, Zijian Wang |
Muon collider research spans accelerator physics, detector instrumentation, and high-energy phenomenology, with relevant evidence scattered across a rapidly expanding and heterogeneous body of scientific literature. As high-energy physics (HEP) increasingly ex...Muon collider research spans accelerator physics, detector instrumentation, and high-energy phenomenology, with relevant evidence scattered across a rapidly expanding and heterogeneous body of scientific literature. As high-energy physics (HEP) increasingly explores agent-assisted analysis workflows, efficiently locating, integrating, and verifying scientific evidence becomes an essential capability. While retrieval-augmented generation (RAG) offers a promising framework for scientific question answering, integrating agentic reasoning without compromising retrieval precision remains a key challenge. In this work, we present agentic hybrid RAG, an evidence-grounded RAG framework for muon collider research. The framework combines a hybrid retriever, integrating sparse lexical and dense semantic retrieval, with an agentic reasoning module for query decomposition, evidence expansion, and grounded answer generation. To enable systematic evaluation, we construct the first benchmark for retrieval-augmented scientific question answering in the muon collider domain, comprising a curated literature corpus together with dedicated retrieval and answer-generation benchmarks covering major detector and physics research topics. Extensive evaluation shows that hybrid retrieval provides the strongest retrieval backbone, while agentic reasoning is most effective for controlled evidence expansion and answer synthesis. Built on this principle, agentic hybrid RAG consistently outperforms representative retrieval and RAG baselines in retrieval effectiveness, answer quality, evidence coverage, and factual grounding. Together, the benchmark and framework provide a foundation for evidence-grounded scientific question answering and future HEP analysis agents operating over large-scale scientific literature. Code is available at \href{https://github.com/AItutorialjrb/RAG_muon_JINST}{this URL}.
|
| 987 |
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
2606.12344
|
cs.CLcs.LG
|
Mengyu Zheng, Kai Han, Boxun Li, Haiyang Xu, Yuchuan Tian |
The software engineering capabilities of general-purpose agent harnesses remain underexplored, and existing benchmarks offer limited support for comparing these harnesses under consistent conditions. To address this gap, we introduce Claw-SWE-Bench, a unified ...The software engineering capabilities of general-purpose agent harnesses remain underexplored, and existing benchmarks offer limited support for comparing these harnesses under consistent conditions. To address this gap, we introduce Claw-SWE-Bench, a unified benchmark that enables researchers to systematically assess the capabilities and efficiency of general-purpose harnesses on software engineering tasks. The benchmark contains 350 real-world GitHub issue-resolution instances across eight programming languages and 43 repositories and provides a shared adapter protocol to align task inputs, outputs, and execution environments across harnesses. Experiments show that general-purpose harnesses can effectively resolve real-world software issues and that their success rates and resource consumption vary substantially even when the underlying model is held fixed. To lower evaluation costs and support faster debugging and iteration, we also provide Claw-SWE-Bench Lite, an 80-instance subset designed to preserve the key evaluation properties of the full benchmark. We hope this benchmark will help researchers better evaluate and understand the performance of general-purpose harnesses on software engineering tasks and guide the development of more capable and efficient harnesses. The data is available at https://github.com/opensquilla/claw-swe-bench and https://huggingface.co/datasets/TokenRhythm/Claw-SWE-Bench.
|
| 988 |
Multi-Bitwidth Quantization for LLMs Using Additive Codebooks
2606.12876
|
cs.CLcs.LG
|
Liza Babaoglu, Shuangyi Chen, Ashish Khisti |
As large language models (LLMs) are increasingly deployed across heterogeneous hardware with varying resource constraints, the ability to adaptively manage the performance-efficiency trade-off without retraining is critical. We propose Drop-by-Drop, a novel mu...As large language models (LLMs) are increasingly deployed across heterogeneous hardware with varying resource constraints, the ability to adaptively manage the performance-efficiency trade-off without retraining is critical. We propose Drop-by-Drop, a novel multi-bitwidth post-training quantization framework enabling inference-time precision control over LLM weights from a single quantized model. As theoretical motivation, we prove that Gaussian sources are successively refinable under a weighted mean squared error distortion motivated by LLM loss functions: successive refinement across distortion levels incurs no rate penalty relative to encoding at each level separately. Drop-by-Drop approximates this hierarchical structure in practice by incorporating Matryoshka-style supervision into additive codebook training, inducing an ordering in which codebook prefixes yield accurate partial reconstructions at each precision level. Furthermore, a block-Hadamard rotation brings the weight distribution closer to our Gaussian source assumption. The result is a single model that serves multiple bitwidths by dropping codebooks, reducing storage and quantization cost relative to the static per-bitwidth models. Across Qwen, LLaMA, Gemma, and Mistral, Drop-by-Drop achieves lower perplexity than state-of-the-art multi-bitwidth methods with competitive zero-shot accuracy, while our specialized kernels further reduce decoding latency.
|
| 989 |
When Can We Trust the Sparse Lens? A Certification Framework for SAE Faithfulness
2606.18383
|
cs.CLcs.LG
|
Dibyanayan Bandyopadhyay, Asif Ekbal |
Sparse autoencoders (SAEs) are increasingly used to study internal representations in language models (LMs), but it remains unclear when an SAE-based representation reliably preserves the underlying model's predictive behavior. We introduce a post-hoc certific...Sparse autoencoders (SAEs) are increasingly used to study internal representations in language models (LMs), but it remains unclear when an SAE-based representation reliably preserves the underlying model's predictive behavior. We introduce a post-hoc certification framework for this purpose. At a chosen layer, we replace the model's hidden activation with its pretrained SAE reconstruction and derive a bound on the expected loss of the original frozen LM through this SAE-based proxy. The resulting certificate depends on three quantities measured on held-out data: the proxy risk, the reconstruction loss gap, and support mismatch between the calibration feature pool and new inputs. We further derive a more conservative exact-$P$ extension whose proxy-risk concentration term is uniform over all feature pools of the same size. Empirically, the certificates become non-vacuous for GPT-2 Small, Gemma-2B, and Llama-3-8B. A detailed layerwise study of Llama-3-8B shows that later layers are easier to certify, driven mainly by improved reconstruction fidelity and weaker downstream amplification of reconstruction error. GPT-2 Small shows much weaker layer dependence, indicating that this pattern is not universal across model-SAE pairs. Overall, the framework provides a practical way to determine whether an SAE representation preserves enough of a frozen LM's predictive behavior to support certification of the original model. Code is available at: https://github.com/newcodevelop/sparse-lens-certification.
|
| 990 |
The Interplay of Harness Design and Post-Training in LLM Agents
2606.25447
|
cs.CLcs.LG
|
Kyungmin Kim, Youngbin Choi, Seoyeon Lee, Suhyeon Jun, Dongwoo Kim |
Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accompanies each per-step observation. While agents are routinely post-trained, this ...Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accompanies each per-step observation. While agents are routinely post-trained, this scaffolding is typically treated as a fixed engineering detail, with design effort limited to the training-free regime. Moreover, existing post-training algorithms assume a static environment, even though tool environments and tasks often shift upon deployment. To address this gap, we extend $\texttt{ALFWorld}$ (i) to treat the harness as a controllable design dimension and (ii) to support evaluation under task and tool environment shifts. Building on this, we systematically analyze how the harness design influences post-training in both in-distribution and out-of-distribution (OOD) settings. We empirically show that harness-aware post-training not only improves in-distribution performance but also enables agents to robustly adapt to OOD settings. Under a harness with minimal design effort, post-training suffers a drastic performance drop under stronger tool environment shifts, further highlighting the importance of harness-aware post-training under such shifts.
|
| 991 |
Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning
2606.25524
|
cs.CL
|
Jaeyong Ko, Jinu Lee, Pilsung Kang, Yukyung Lee |
Large language models reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive at the correct answer while others fail. Prior work localizes such failures at the step, chunk, or sentence level, or identifies...Large language models reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive at the correct answer while others fail. Prior work localizes such failures at the step, chunk, or sentence level, or identifies tokens where failure has already occurred. These approaches leave open which token triggers failure. We introduce the cliff token, a token at which the estimated probability of reaching the correct answer (success probability) drops beyond an adaptive threshold. Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers. For incorrect traces containing cliff tokens, we compare resampling immediately before and after the first cliff token. Resampling before it shows higher pass@$k$ at the same sample count. We further introduce a cliff taxonomy of deterministic, uncertain, and sampled-off cliffs, defined by greedy choice and token entropy. Additionally, we show that the three types differ as training signals. Using single-token preference optimization at cliff positions (Cliff-DPO), we find that uncertain and sampled-off cliffs show larger accuracy gains than deterministic cliffs on three evaluation benchmarks. We release token-level rollout data and source code to enable further analysis without regenerating costly rollouts: https://github.com/beaver-22/Cliff-token
|
| 992 |
EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix
2606.26472
|
cs.CLcs.LG
|
Steven Kolawole, Virginia Smith |
Reasoning models can generate chains of thought tens of thousands of tokens long, making the key--value (KV) cache that holds them a major bottleneck for inference throughput. Existing eviction policies for long reasoning traces typically rank cached tokens us...Reasoning models can generate chains of thought tens of thousands of tokens long, making the key--value (KV) cache that holds them a major bottleneck for inference throughput. Existing eviction policies for long reasoning traces typically rank cached tokens using attention weights, requiring access to the attention matrix and making them incompatible with fast inference kernels. In this work we study the limits of such policies under tight cache budgets. Surprisingly, we find that under the strongest of them the generations that finish are wrong about as often as without eviction; most of the accuracy loss comes from generations that enter loops and run until the length limit, and retaining more tokens according to a fixed importance score exacerbates this behavior. What stops the looping is keeping the tokens the model's recent queries point to, and the forward pass the model already runs reveals them without the attention matrix. Motivated by this observation, we introduce epiphany-aware KV cache eviction EpiKV, which combines hidden-state shifts with the model's recent query--key relevance to rank cached tokens without materializing the attention matrix. On multiple benchmarks, EpiKV matches or outperforms the strongest attention-based eviction baselines while running directly in vLLM with unmodified attention kernels.
|
| 993 |
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
2607.05804
|
cs.CL
|
Yuhang Zhou, Kai Zheng, Haoling Li, Dengyun Peng, Can Xu |
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently exp...On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.
|
| 994 |
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
2607.08393
|
cs.CL
|
Lu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu |
Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the Knowing-Using Gap, characterized by an accuracy gap and a temporal l...Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the Knowing-Using Gap, characterized by an accuracy gap and a temporal lag between memorization and generalization. To understand this phenomenon, we fine-tune LLMs with unseen knowledge and monitor the spatial permeation dynamics of the knowledge internally with self-patching, an adaptation of activation patching that scans all layer pairs at every fine-tuning check-point. Self-patching identifies activation locations where relocating representations substantially improves failed generalization cases. These results are consistent with a knowledge-circuit misalignment hypothesis: memorized representations can exist internally but may not be routed to computation-effective layers. Experiments are done cross-domain for the robustness of this finding. Building on this diagnosis, we propose layer-wise representation self-distillation (LRSD) that aligns knowledge representation from late storage layer to middle layer. LRSD keeps improving generalization after fine-tuning saturates and nearly doubles multi-hop chaining accuracy on Qwen, while leaving memorization intact.
|
| 995 |
Extractable Memorization From First Principles
2607.12649
|
cs.CLcs.LG
|
A. Feder Cooper, Marika Swanberg, Jamie Hayes, Lea Duesterwald, Christopher De Sa |
Recent work on extractable memorization in language models suffers from two contrasting validity problems. Some studies overstate extraction, for example, by using sequences too short to distinguish memorization from predictability. Others imply that extractio...Recent work on extractable memorization in language models suffers from two contrasting validity problems. Some studies overstate extraction, for example, by using sequences too short to distinguish memorization from predictability. Others imply that extraction is unreliable evidence because models can reproduce real-world text they weren't explicitly trained on. Both overlook what makes a valid extraction claim: the model must generate a training sequence with high enough probability to indicate memorization. To determine what's high enough, one has to perform a matched comparison, measuring generation probabilities of training and comparable non-training sequences. Since non-training sequences can't have been memorized, their probabilities are a baseline for predictability; exceeding this baseline is evidence of memorization. We formalize matched comparisons with (1) a conformal test calibrated to a chosen false-positive rate when sequences are sampled from populations, and (2) a single-document census that calibrates against a matched non-training document. Matched comparisons enable rigorous, calibrated memorization claims and clarify where prior setups have validity issues. On Wikipedia, OLMo 2 32B reproduces non-training 10-token suffixes roughly 24% as often as training ones: that share reflects false positives, not memorization. For Llama 3.1 70B on books, calibrated census thresholds reach as low as 10^(-27), supporting memorization claims for sequences no feasible sampling budget would extract. We therefore refine "extractable memorization" to require both a valid memorization claim and near-certain generation within a realistic budget.
|
| 996 |
AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
2607.19223
|
cs.CLcs.LG
|
Yu-Yang Qian, Hao-Cong Wu, Chen Chen, Jiacheng Sun, Zhenhua Dong |
Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts draf...Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts drafting efficiency by leveraging diffusion drafters, whose parallel denoising mechanism enables draft generation in a single forward pass. In this work, we uncover a central pitfall of diffusion drafters: the global dependency, arising from bidirectional attention and KV injection in diffusion drafters, is a double-edged sword. On one hand, it endows the model with parallel generation and global contextual modeling capabilities; on the other hand, this global dependency introduces high variance at both the domain-level and the token-level: acceptance rates fluctuate substantially across different domains, and the acceptance probability varies across draft positions. To tackle this issue, we propose the AdaFlash framework, comprising two components: (i) an on-policy distillation (OPD) algorithm with reverse-KL divergence tailored for diffusion drafters, which delivers stable convergence and continuously adapts the drafter to the deployment distribution; and (ii) an adaptive length head that dynamically adjusts the candidate length on the fly, substantially lowering the verification cost of the target model. Experiments demonstrate that AdaFlash consistently improves the speedup during deployment, with especially significant gains under high-concurrency, achieving up to 66% higher average throughput than previous state-of-the-art methods. Our code is available at https://github.com/ZinYY/AdaFlash.
|
| 997 |
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
2607.22676
|
cs.CL
|
James Elcock, William F. Shen, Xinchi Qiu, Nicholas D. Lane |
Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains rem...Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
|
| 998 |
THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
2607.27303
|
cs.CLcs.LG
|
Yixin Peng, Diego Collarana, Er Jin, Stefan Decker |
Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve over time. Learning on such graphs requires jointly modeling cross-type structural heterogeneity and the tempo...Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve over time. Learning on such graphs requires jointly modeling cross-type structural heterogeneity and the temporal dynamics of interactions, yet existing methods still struggle to reconcile parameter-efficient cross-type transfer with relation-aware specialization, and typically inject time only as additive features outside the attention kernel. We propose \textbf{THGFM}, a web-scale temporal heterogeneous graph fusion model that addresses both limitations within a unified dual-path architecture. THGFM couples a \textit{Shared-Space Temporal Attention} branch for parameter-efficient cross-type transfer with a \textit{Relational Type-Partitioned Temporal Attention} branch for relation-aware specialization, and integrates them through \textit{Dual-Path Relational--Shared Fusion}, instantiated with \textit{Type-Conditioned Non-Competitive Gated Sum Fusion}: a adaptive mechanism that assigns independent, type-conditioned feature-wise gates to the shared and specialized branches, allowing both to be amplified or suppressed without zero-sum competition. To directly incorporate relative time into the attention score, THGFM further introduces \textit{Rotary Temporal Attention}, which rotates queries and keys by half-phases of relative time before matching. THGFM consistently outperforms baseline graph transformer models on academic graphs benchmarks, delivering a $+3.25\%$ six-task mean gain, with peak relative gains of $+12.37\%$ on OAG-CS PV, $+4.87\%$ on PF-$L_2$, and $+1.18\%$ on PF-$L_1$, and $+4.24\%$, $+3.73\%$, and $+4.61\%$ on OGBN-MAG, HTAG-ArXiv, and HTAG-DBLP, respectively.
|
| 999 |
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
2608.02673
|
cs.CL
|
Hankun Wang, Bohan Li, Shi Lian, Xiaoyu Gu, Jing Peng |
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, ...Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, parameters, or target region underspecified. We study a precise and explicit interface for speech editing: a transcript-grounded structural edit instruction with XML-style tags explicitly specifies typed operations and localizes them to transcript spans or boundaries. This semantic timeline avoids explicit timestamp alignment and provides an externally inspectable contract for compositional edits. We instantiate the interface in dots$.$tts$.$edit, an editor adapted from the continuous autoregressive dots$.$tts foundation model. Four representative speech-creation controls cover lexical content, affective expression, pitch and speaking-rate delivery, and temporal phrasing through text, emotion, prosody, and pause editing. Task-specific data pipelines construct operation- and scope-controlled pairs while retaining source-derived context outside each target region. We further introduce doteBench, a bilingual evaluation suite that measures precise instruction following, local preservation, and audio quality across the four controls and their composition. Experiments show leading overall instruction following and local preservation across its five editing categories, while audio quality remains comparable to existing open-source systems. Across three Seed-TTS-Eval shards, the model shows negligible differences from the base model in zero-shot TTS recognition error rate and speaker similarity.
|
| 1000 |
SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization
2608.04084
|
cs.CLcs.LG
|
Boyao Wang, Zhihan Lei |
Mixture-of-experts (MoE) networks pursue specialization through learned routers, gates, and load-balancing losses, yet at matched total-parameter budgets learned routers can underperform equal-weight No-Routing baselines. Is the bottleneck the routing algorith...Mixture-of-experts (MoE) networks pursue specialization through learned routers, gates, and load-balancing losses, yet at matched total-parameter budgets learned routers can underperform equal-weight No-Routing baselines. Is the bottleneck the routing algorithm, or the alignment between training-signal granularity and the target categories? We probe the question with SpecDrop, a fixed parameter-free routing scheme: each of $K$ branches receives weight $p_a$ for its assigned category and a small leakage $p_i > 0$ otherwise, merged through a category-independent fixed denominator, with no learned routing parameters and no auxiliary losses; the category label is required at inference. On vision tasks where each image has one superclass label (CIFAR-100 on ResNet-110; ImageNet-1K on ViT-S/16), SpecDrop reaches 79.23% on CIFAR-100 and 79.89% on ImageNet-1K, exceeding parameter-matched baselines that do not use the label (+4.75 over dense on CIFAR-100; +6.53 over the No-Routing+SE control on ImageNet-1K). These gains quantify what category supervision buys when deployed through routing -- not an advantage over label-aware deployments of the baselines: given the same label, masking a dense model's outputs is stronger for accuracy alone (85.2 / 83.7). SpecDrop's contribution is converting the label into trained-in modular structure: 68%/100% branch-category alignment, and masking gains of 0.00 (CIFAR) / +1.06 (ImageNet) -- the output-space restriction is largely internalized during training. On fuzzy partitions, where training units span multiple categories (SlimPajama-6B language modeling with a 30M Transformer; SuperNI instruction tuning over Llama-3.2-1B with LoRA), the routing mechanism reduces to the matched No-Routing controls within seed noise, the null our thesis predicts. Granularity alignment, not algorithm choice, localizes when routing helps. Code: https://github.com/Beryex/SpecDrop
|
| 1001 |
Predicting Task Difficulty Without Rollouts
2608.05797
|
cs.CLcs.LG
|
Stefan Krsteski, Charlotte Meyer |
A fundamental challenge in evaluating and training autonomous agents is measuring the intrinsic difficulty of the tasks they attempt. Estimating this quantity can be useful for environment designers creating synthetic data or benchmarks, as well as for constru...A fundamental challenge in evaluating and training autonomous agents is measuring the intrinsic difficulty of the tasks they attempt. Estimating this quantity can be useful for environment designers creating synthetic data or benchmarks, as well as for constructing training curricula, thereby offering a way to reduce compute costs. Such an estimation becomes increasingly important as agents (particularly LLM-based) move into longer-horizon domains, where empirical trial-and-error becomes a severe computational bottleneck. However existing work is largely confined to static tasks and relies on evaluation metrics that, as we show, can give an incomplete picture of predictive performance. In this paper we study ex ante difficulty prediction across 17 agentic benchmarks spanning coding, mathematics, machine learning, web navigation, function calling, and other domains. We find that AUC can mask poor difficulty estimates, identify token-level entropy as a useful predictive signal, and demonstrate how residuals between expected and observed difficulty can expose environment flaws such as contamination and infeasibility.
|
| 1002 |
Which Decisions Low-Bit Quantization Breaks, and How to Predict Them
2608.06564
|
cs.CLcs.LG
|
Zekun Wu, Swati Dhiman, Adriano Koshiyama |
Quantization saves memory by storing model weights with fewer bits. It can also change model decisions, such as whether to call a tool or which option to choose from a finite set. We study these decision changes in 16 language models from 8 families at 4, 3 an...Quantization saves memory by storing model weights with fewer bits. It can also change model decisions, such as whether to call a tool or which option to choose from a finite set. We study these decision changes in 16 language models from 8 families at 4, 3 and 2 bits, across several post-training quantization settings. Our evaluation covers tool use, safety, general knowledge and social bias, using BFCL, XSTest, MMLU, BoolQ, BBQ and synthetic tasks. The decision margin is the score difference between two possible first tokens, measured before and after quantization. Writing the margin before quantization as $m$ and the margin after quantization as $m'$, we find an approximately linear relationship across decisions: $m' \approx c m + b$. The slope $c$ is usually below one and becomes smaller as precision falls, so quantization progressively shrinks decision margins. The offset $b$ is the same for every decision of one kind. Quantization therefore does not simply add random noise, and even a strong preference at full precision can flip. Quantization also affects different kinds of decisions to different degrees. Within tool use, whether to call a tool is often more sensitive than which tool to call: on 400 BFCL tasks, three of five models lose more completed calls than correct tool selections at 3-bit round-to-nearest. Under GPTQ and GGUF far fewer whether-to-call decisions flip than under plain rounding, so there is no single 3-bit failure point. The same relationship predicts how often decisions flip. Across 1,154 combinations of models, quantization settings, bit-widths and decision types drawn from our evaluation, we fit the slope, the offset and the spread around the fitted line on half of the decisions and predict the flip rate on the other half. The predicted flip rate differs from the observed flip rate by a median of 1.0 percentage point.
|
| 1003 |
From Behavior to Mechanism: Tracing Divergent Response Modes in Frontier Language Models
2608.06578
|
cs.CLcs.LG
|
Ali Jalal-Kamali |
Frontier language models are trained with distinct data, objectives, and safety pipelines, but whether those differences produce measurably different behavior under steering pressure has not been tested. We evaluate 6 frontier models from different labs on 300...Frontier language models are trained with distinct data, objectives, and safety pipelines, but whether those differences produce measurably different behavior under steering pressure has not been tested. We evaluate 6 frontier models from different labs on 300 paired base and steered items across 3 behavioral categories. All models also act as blind peer judges against fixed rubrics, and each response is labeled by consensus over 24,480 judgments, while leaving self-judgment out. Models differ both in how far steering moves them and in the kind of response they give. GPT-5 withholds its reasoning while still providing the answer on 99 of 100 steered items, against 0 in 500 for the others. Claude Opus 4.7 and GPT-5 resist explicit suppression instructions where the other four never do, and they resist differently. In Llama, the open-weight model, a linear probe reads the behavioral split from the residual stream before generation at 0.87 cross-validated accuracy. Injecting that direction drives the behavior from 0% to 86%, and ablating it cuts the natural rate by more than half, where a random direction of equal norm changes nothing. A second ablation on complementary items reproduces the effect more strongly, and its direction has cosine similarity 0.82 with the first.
|
| 1004 |
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
2608.11829
|
cs.CLcs.LG
|
Xinmu Ge, Zizhuo Zhang, Yu Huang, Jianing Zhu, Lin Yuan |
On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. Under the reverse KL objective, the idealized optimum of OPD aligns the student distribution with that of the teacher. When the teacher consistently ou...On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. Under the reverse KL objective, the idealized optimum of OPD aligns the student distribution with that of the teacher. When the teacher consistently outperforms the student, this naturally suggests that OPD should yield broad improvements over the pre-OPD student. However, do such improvements extend across the entire range of test-time sampling budgets? In this work, we revisit this expectation through the lens of test-time scaling by varying the sampling budget $K$ and evaluating performance with pass@$K$. Across multiple settings, we observe two distinct patterns: OPD can improve pass@$K$ at both small and large sampling budgets, but it can also improve small-budget performance while reducing large-budget pass@$K$. We show one condition that guarantees such a reversal and an idealized reverse KL counterexample where it occurs even when the teacher has higher accuracy on every problem. To choose between two candidate teachers at a target sampling budget, we propose the \textit{Teacher Advantage Score at $K$} (TAS@$K$), which can be computed before OPD training to predict which teacher will lead to a larger improvement in pass@$K$. Across three domains and thirteen benchmarks, the ordering predicted by TAS@$K$ agrees with the observed pass@$K$ improvements of the resulting OPD models in 83.6\% of experiments, providing a useful signal for teacher selection at the target pass@$K$.
|
| 1005 |
J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules
2608.17063
|
cs.CLcs.LG
|
Yunfan Gao, Xinyi Huang, Tao Sheng, Haorui Song, Yun Xiong |
Task-fine-tuned large language model (LLM) classifiers acquire task-specific decision knowledge, but this knowledge remains implicit in distributed internal computations, making their decision logic difficult to interpret. We introduce the Executable Decision ...Task-fine-tuned large language model (LLM) classifiers acquire task-specific decision knowledge, but this knowledge remains implicit in distributed internal computations, making their decision logic difficult to interpret. We introduce the Executable Decision Compression (EDC) framework and propose J-Miner, which mines vocabulary-named variables from internal readouts and learns rules shared across inputs to produce executable explanations. Analysis reveals that a small set of these variables captures much of the classifier's decision behavior, holding for both varying parameter scales within a family and distinct families. Across six binary tasks, a rule using just one variable reproduces 76.7% of source-classifier decisions on average, rising to 88.8% with 16 variables. Most of the decision information retained by these variables comes from internal activations beyond literal surface matching. A lightweight text reader predicts the variable states, allowing the same fixed rules to execute independently of the source classifier.
|
| 1006 |
What is Missing from AI Post-Training AI: An Empirical Analysis
2608.19072
|
cs.CLcs.LG
|
Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu, Xin Cong |
Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of recursive self-improvement (RSI). Yet this progress is measured by aggregate benchmark scores, which cannot tell whether an agent executes a fixed plan well or stra...Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of recursive self-improvement (RSI). Yet this progress is measured by aggregate benchmark scores, which cannot tell whether an agent executes a fixed plan well or strategically revises the plan when it fails. We separate these two capabilities: execution-level capability, iterating within an established training strategy, and strategy-level capability, revising that strategy as experimental evidence accumulates. Analyzing 1,338 post-training trajectories of frontier agents, we find that agents reliably execute post-training but lock into a default strategy, which follows the agent rather than the task, and only 2.1% of transitions between adjacent training runs ever change strategy. We then test whether the agent lacks experience, reasoning, or the decision to switch. (1) Experience improves execution but not the strategy. (2) Additional reasoning compute yields front-loaded gains on easier tasks but refines, rather than revises, the committed strategy. (3) Human review before training changes which strategy the agent locks into, not whether it locks in, whereas a single mid-run instruction outperforms the agent's own continuation by up to 17.44 points under the same budget. In conclusion, what the agent lacks is the decision to reopen a committed strategy and try another one. Realizing RSI therefore calls for interaction protocols and training signals that make strategy revision an explicit, rewarded decision.
|
| 1007 |
SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents
2608.21929
|
cs.CL
|
Yuanjin Zheng, Jingbang Chen |
Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper studies token amplification through skill injectio...Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper studies token amplification through skill injection: an economic resource-abuse threat in which a malicious skill causes an agent to consume substantially more tokens than needed for normal task execution. We present SkillBloat, a two-phase framework that first screens a library of diverse attack-type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM-guided full-document skill rewriting. Evaluated on a real-world skill benchmark, SkillBloat achieves 5.4184x-10.1455x average best amplification across multiple coding-agent target configurations. An ablation shows that the second-stage refinement loop consistently improves average best amplification over Phase 1 attack-type screening alone, demonstrating that iterative optimization provides additional benefit beyond initial attack-type selection. These results show that skill ecosystems expose a practical resource-amplification attack surface that is orthogonal to existing security-oriented skill poisoning.
|
| 1008 |
Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
2608.27339
|
cs.CLcs.LG
|
Xinwei Qiang, Xiang Fang, Chang Chen, Zaifeng Pan, Yue Guan |
Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish ...Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Measurements across four domains and four open-weight targets, plus floor estimates on a frontier API target, yield three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. Guided by this separation of short-range conditioning from proposal quality, we replace a context-independent predecessor correction with a prefix-attention head that reads committed context conditional on the predecessor. With supporting backbone components, the resulting Qwen3-4B drafter improves mean serving accepted length by $3.93\%$ over the released DSpark checkpoint across nine tasks, without increasing the conditioning order. Code and data: https://github.com/tie-pilot-qxw/specfloor; drafter checkpoints: https://huggingface.co/TIE-Pilot/dspark-attnconv-block7-qwen3-4b.
|
| 1009 |
How Do Language Models Choose Between Context and Memory?
2609.00753
|
cs.CLcs.LG
|
Benjamin Shih, John Winnicki, Arianna Cao |
When contextual information conflicts with knowledge stored in model parameters, activation directions can be used to decode and steer which source the model follows. However, successful steering does not establish that the unedited model uses those directions...When contextual information conflicts with knowledge stored in model parameters, activation directions can be used to decode and steer which source the model follows. However, successful steering does not establish that the unedited model uses those directions to choose between sources, or that they remain effective across tasks. To test these possibilities, we vary the stated authority of contextual claims while holding their content fixed. We first estimate authority directions from prompts in which context and parametric knowledge agree, then test their causal contribution when the two sources conflict. Interchanging naturally occurring activation values along these directions between matched high- and low-authority prompts reproduces 30--68% of the authority-induced shift in source choice across Qwen, Llama, and OLMo models, whereas matched controls reproduce almost none. We next ask what transfers across tasks: the learned direction versus the activation values exchanged along it. Using a direction learned on another task closed 9% of the source-choice gap, compared with 57% when learned on the task being evaluated. Both interventions exchanged activation values from the evaluated task. In a separate experiment, we kept its learned direction but exchanged values taken from another task, which closed 68% of the gap. These results show that authority-related activation values can causally influence source choice across tasks when inserted along directions learned for the task being evaluated.
|
| 1010 |
EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?
2609.04280
|
cs.CL
|
Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu |
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EvoHarnessBench, a ...Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EvoHarnessBench, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EvoHarnessBench places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprising 802 tasks, 520 tools, 42 skills, and 62 agents. We evaluate two complementary settings corresponding to the central challenges of harness evolution: deployment evaluation, which isolates retention of previously accessible competence as the harness expands, and self-evolving adaptation evaluation, which tests whether accumulated experience remains useful as new capabilities are introduced. Our results reveal three persistent gaps. First, harness expansion alone can degrade performance on previously solved tasks, producing harness-induced forgetting. Second, gains from self-evolving adaptation remain inconsistent across stages of harness evolution, capability axes, and environments. Third, retention and adaptation can pull in different directions: preserving earlier competence does not necessarily improve adaptation to newly introduced capabilities, and vice versa. These results establish harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.
|
| 1011 |
Exact Record Omission in Delta Attention: A Transport Criterion, Its Cost, and a Replay Certificate
2609.06872
|
cs.CLcs.LG
|
Vishwajith Ramesh |
An assistant can stop repeating a deleted statement while its recurrent memory still carries that statement's influence. We examine this distinction by saving the state difference immediately after a record, carrying this saved difference, or receipt, through ...An assistant can stop repeating a deleted statement while its recurrent memory still carries that statement's influence. We examine this distinction by saving the state difference immediately after a record, carrying this saved difference, or receipt, through later updates, and comparing the corrected state with the state built from the same conversation with the record omitted. Unrolling the recurrence gives an exact criterion: transport reaches this never-stored state if and only if the additional differences created by later updates cancel after transport. We evaluate eleven conditions using 40 prospectively selected synthetic Kimi Linear contexts, with 40 separate calibration contexts and paired uncertainty estimates computed across complete records. Attention masking removed all 40 exact greedy target answers, yet a fixed three-query candidate-scoring attack with known candidate values and log-probability access attained AUC 0.740 (95% interval 0.695--0.809); the measured false-positive rate was 7.5%. Prompt-only forgetting still returned 33 of the 40 targets. Checkpoint replay matched the complete declared active state and audit logits in all 80 contexts, while preserving all 120 retained answers in the evaluation cohort. An independent Qwen cohort and matched original-bf16/8-bit Kimi controls confirmed native transport mismatch beyond the numerical error measured in matched controls. Replay also supported record replacement and successive deletions with a new record added between them. These results provide a practical audit of both the requested memory change and the assistant's remaining behavior, with replay work determined by the surviving suffix.
|
| 1012 |
CausalVerify: End-to-End Verification of Causal Analyses by Language Models
2609.07944
|
cs.CL
|
Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie |
Language models increasingly perform empirical analyses end to end, yet existing evaluations assess the written explanation or whether generated code executes, not whether the executed workflow recovers the intended causal estimand. We introduce CausalVerify, ...Language models increasingly perform empirical analyses end to end, yet existing evaluations assess the written explanation or whether generated code executes, not whether the executed workflow recovers the intended causal estimand. We introduce CausalVerify, an execution-grounded benchmark for end-to-end causal analysis that follows a model from research-context interpretation to estimand recovery. It scores this workflow at four distinct layers: method recognition, design specification, executable implementation, and estimand recovery. It combines 259 real-paper contexts, 100 fixed-seed synthetic scenarios with executable reference estimates, and 23 paper-twin pairs in which a model commits to a design before seeing the data and its executed analysis is scored against a canonical estimator on the same realised dataset. Execution is not correctness. Among 426 model-written workflows that run without error, 15.5% fail verification, and a keyword score of the effect direction stated in the text is a poor proxy for recovery. In a 23-pair, nine-model paired study, replacing a model's committed design with the reference design and its execution conventions raises joint recovery of the point estimate and standard error from 15.0% to 51.5%; yet 48.5% of eligible seeds still fail under the reference design. The direction replicates on six pairs built afterwards under a frozen construction protocol, although on the two newest pairs the gain is confined to models from the family that built the references. Design specification is consequential but not sufficient: a plausible method and runnable code do not guarantee recovery, and even supplying the reference design and its conventions leaves substantial downstream failure. CausalVerify evaluates the executed workflow rather than its surface plausibility, and every reported number is recomputed from frozen artifacts by a single script.
|
| 1013 |
Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems
2609.08887
|
cs.CL
|
Maximilian Schall, Sedigheh Eslami, Markus Krimmel, Antoine Chaffin, Louis Milliken |
Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant ...Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant documents per query. No existing public benchmark evaluates this setting: large-scale collections typically provide only a small number of evaluation queries, whereas benchmarks with many queries generally contain only millions of documents. Moreover, most benchmarks assess human-written queries, while the first-stage retrievers in agentic RAG pipelines serve machine-written reformulations whose distribution differs from human search behavior. To overcome these evaluation gaps, we introduce Q2D-Web (Query2Doc-Web), a large-scale agentic retrieval benchmark consisting of a 190M-document web corpus and 70k agentic search queries in ten languages, reformulated from real-world user queries in production systems. Q2D-Web provides three sets of fixed relevance judgments: agent citations, production rankings, and a combined set that unions both signals and adds LLM-based judgments of unlabeled pooled documents to reduce false negatives. We benchmark 13 retrievers including lexical, dense, and late-interaction models and find that their relative ordering is largely insensitive to the choice of judgment set, while diverging substantially across topical domains, query languages, and query types. To enable fast evaluation, we also study subcorpus sampling as an approximation to full-corpus evaluations. Retaining a third of the corpus, selected by reciprocal rank fusion over pooled retriever runs, preserves the full-corpus model ranking under the combined judgments while raising absolute Recall@1000 only by 4 to 7 points. The public leaderboard is accessible under: https://huggingface.co/spaces/perplexity-ai/q2d-web-leaderboard
|
| 1014 |
MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads
2609.09206
|
cs.CLcs.LG
|
Meng'en Qin, Junye Chen, Jucheng Liu, Yinchen Liu, Youlu Xing |
Multimodal Large Language Models (MLLMs) often struggle with hallucinations, thus hindering their reliable practical applications. Existing attention-based mitigation methods mainly rely on indirect signals (e.g., attention weights) that fail to accurately ref...Multimodal Large Language Models (MLLMs) often struggle with hallucinations, thus hindering their reliable practical applications. Existing attention-based mitigation methods mainly rely on indirect signals (e.g., attention weights) that fail to accurately reflect the actual information shift underlying hallucination generation. In this paper, we propose HEAL, Head-lEvel information disentAnglement and caLibration for identifying and mitigating hallucinations. HEAL first employs causal noise intervention on multi-head outputs to filter out causally redundant heads. Subsequently, it disentangles information distribution within the remaining heads via the counterfactual Difference-in-Differences, categorizing heads into four types. Through analysis, we observe: hallucinations happen when information distribution drifts away from a healthy equilibrium in synergy heads, not strongly correlated with the quantity or strength of modality-specific heads. Motivated by this insight, HEAL injects dynamic information calibration factors into the value vectors of synergy heads and actively regulates visual-language dependencies, steering the output distribution towards factual evidence. Extensive experiments demonstrate that HEAL effectively reduces hallucinations across multiple MLLMs, offering a simple and interpretable pathway to enhance model trustworthiness.
|
| 1015 |
In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Post-Retrieval Context Tampering
2609.09243
|
cs.CL
|
Iliano Fasolino |
Retrieval-augmented generation (RAG) grounds a language model in retrieved documents, which reduces hallucination but creates a new attack surface: if retrieved text is tampered with, the model may repeat the falsehood. We study how much a small quantized mode...Retrieval-augmented generation (RAG) grounds a language model in retrieved documents, which reduces hallucination but creates a new attack surface: if retrieved text is tampered with, the model may repeat the falsehood. We study how much a small quantized model, Llama 3.1 8B, degrades when a fraction of its retrieved context is poisoned. Three corruption strategies are tested, entity swap, number swap, and negation, each applied to zero, one, two, or three of the three retrieved passages, over a factorial sweep of 588 runs on a fact-checking task built from FEVER. Accuracy falls from 77.9% on clean context to 43.5% when all three passages are corrupted. Entity swap flips the largest share of answers that were correct on clean context. Number-based corruption stays flat while poisoned passages are a minority and jumps once they form a majority, a pattern we re-check with query-level bootstrap intervals. The model rarely invents new falsehoods; its dominant reaction is to abstain, and a lexical overlap proxy of unsupported generation falls under attack rather than rising. The study is a small-scale measurement with coarse automated labels; we treat the strategy contrasts as suggestive until decoding is controlled and stronger adjudication is in place.
|
| 1016 |
A Dominant Supplier Slows Recursive Drift More Than It Steers It
2609.11146
|
cs.CLcs.LGcs.AI
|
Yangze Liu, Zhongyi Han |
More and more of the text future language models learn from is written by a few of today's models. If one supplier writes most of a shared corpus, does it pull the models trained on it toward its own writing, or change how fast they drift? We retrain eight ope...More and more of the text future language models learn from is written by a few of today's models. If one supplier writes most of a shared corpus, does it pull the models trained on it toward its own writing, or change how fast they drift? We retrain eight open models from their base weights on a shared pool of each other's text for five generations, varying the part written by one model, Phi-2, from an equal share to 90%. The models drift together toward a style with fewer function words, and none starts repeating itself. No share of Phi-2 brings the other models closer to its text than the equal share does. We split each ecosystem's separation from the equal-share one into a delay along its route and a departure from that route, both counted beyond the difference between two equal-share runs. With Phi-2 at 90%, delay outweighs departure 72 to 28 and 64 to 36 in two runs, and the ecosystem falls 2.7 and 2.5 generations behind. With Phi-2 at half the pool the two parts are about equal. When SmolLM2 or Qwen3-1.7B writes half instead, the ecosystem slows less or not at all. The departure leans toward Phi-2 more as its share grows, but more than toward every other model only at 90%. Human text filling a quarter or half of the pool slows the models along the same route.
|
| 1017 |
Rice's Theorem under Self-Modification: Elevation Operators and a Normal Form
2609.11326
|
cs.CL
|
Jose Pascual Gumbau Mezquita |
We ask whether it can be certified algorithmically that a self-modifying program keeps a behavioural property, a safety property in the motivating case, after its next rewrite (preservation) and along its whole evolution (persistence). When the rewrite depends...We ask whether it can be certified algorithmically that a self-modifying program keeps a behavioural property, a safety property in the motivating case, after its next rewrite (preservation) and along its whole evolution (persistence). When the rewrite depends only on behaviour, preservation is a behavioural property and Rice's theorem applies. When the rewrite reads the code, preservation is no longer behavioural; yet, under a uniform disruption condition, the s-m-n reduction that proves Rice's theorem works inside a single class of behaviourally identical programs, and preservation inherits the degree of the halting problem. One step never exceeds the degree of the property, while persistence can climb one level of the arithmetical hierarchy. We then isolate the mechanism shared by rewriting, supervision and system comparison, the elevation operator, and prove a normal form: the preserving set is determined by a single finite trigger and a polarity, and the Rice-Shapiro theorem restricts the polarity to the arithmetical class of the property. Runtime monitors, consistency supervision, conformance to a reference and observational equivalence are instances, and no sound theory covers the preserving systems.
|
| 1018 |
Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
2609.15296
|
cs.CL
|
Mingzhou Jiang, Peixi Wu, Hang Cheng, Yunhao Zhou, Biao Yang |
Universal multimodal embedding (UME) maps multimodal inputs into a shared embedding space for diverse retrieval tasks. Recent methods improve embeddings through Chain-of-Thought (CoT) reasoning optimized with GRPO using retrieval rewards. However, existing met...Universal multimodal embedding (UME) maps multimodal inputs into a shared embedding space for diverse retrieval tasks. Recent methods improve embeddings through Chain-of-Thought (CoT) reasoning optimized with GRPO using retrieval rewards. However, existing methods overlook the mismatch bettween candidate-aware retrieval supervision and input-only CoT generation: (1)trajectory-level rewards convey retrieval outcomes without explicitly identifying the input-supported evidence that distinguishes the positive from hard negatives; (2) input-only generation cannot directly assess whether further reasoning improves retrieval, potentially producing redundant CoTs with substantial latency. To bridge this gap, we propose Reason What Matters (ReWAM), a retrieval-grounded framework that aligns candidate-aware supervision with input-only generation. Specifically, we introduce Retrieval-Aware Self-Distillation (RASD), which extracts privileged guidance from input-supported facts and evidence distinguishing the positive from hard negatives. Conditioned on this guidance, an on-policy self-teacher provides token-level feedback to refine credit assignment, directing policy updates toward retrieval-relevant reasoning grounded in the input. We further propose Retrieval-Adaptive Inference (RAI), which learns a retrieval-aware stopping criterion from prefix-level retrieval feedback. It stops redundant reasoning without candidate access and uses speculative decoding to further reduce CoT latency. Extensive experiments on MMEB-V2 and MRMR demonstrate that ReWAM achieves state-of-the-art retrieval performance while delivering up to 5x the inference throughput of competitive explicit-CoT UME methods. ReWAM thus enables high-quality retrieval through efficient input-only reasoning, making explicit CoT practical for corpus-scale multimodal retrieval. The code will be publicly available.
|
| 1019 |
The Router Within: Eliciting Native Skill Routing from a Frozen LLM
2609.15982
|
cs.CLcs.LG
|
Ruishuo Chen, Xun Wang, Yu Chen, Zhuoran Li, Longbo Huang |
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library siz...Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library size. Retrieval pipelines move the selection out of the context, but also out of the agent's capability. We show that the frozen agent LLM already carries the routing signal in its own forward passes, and that two linear maps suffice to read it out with no skill text in the context. Our Gavel (Glance And Verdict from a frozen LLM) reads it in two steps. A glance scores the full library by matching the task's mid-layer states against a compact bank that one forward pass builds for each skill at installation, with the two maps as the only trained parameters. A verdict then resumes each shortlisted skill's forward pass, reads the model's own likelihood and yes/no judgment, and fuses both with the glance as a product of experts. Trained once, Gavel transfers zero-shot to three public benchmarks and SkillTraj, our new benchmark of 372 simulated agent trajectories. On Qwen3-32B it outperforms progressive disclosure and retrieve-and-rerank pipelines that add 1.2B to 16B external parameters, by up to 13.4 points on written tasks and up to 21.9 when the need for a skill arises mid-rollout. Routing accuracy improves as the backbone does, and in a bash-agent harness Gavel lets the 32B trigger the right skill on Skill-Use more often than models of up to 1.6T parameters in Codex.
|
| 1020 |
Compiled Agency: Coding Agents as Game AI Researchers -- from a Roguelike to StarCraft II and Civilization
2609.18996
|
cs.CL
|
Haonan Huang, Joey Xiao |
Coding agents are increasingly capable of sustained engineering and empirical research. Greater autonomy makes their research decisions themselves a target for evaluation: what to investigate, which experiments to run, and when to stop. We introduce Gauntlet, ...Coding agents are increasingly capable of sustained engineering and empirical research. Greater autonomy makes their research decisions themselves a target for evaluation: what to investigate, which experiments to run, and when to stop. We introduce Gauntlet, a develop-freeze-evaluate protocol for studying these decisions as agents build standalone game-playing programs. Starting from a game description, a raw observation/action interface, and an empty policy file, an off-the-shelf agent develops a controller from bare interaction in one autonomous session. We call the capability under study compiled agency: the shipped program plays with zero model calls. The capability is real and advancing: environment access adds 10 to 78 percentage points of held-out success over construction-only controls, and progress across model generations comes in steps, with tiers that defeat one generation entirely falling to the next. The protocol yields StarCraft II controllers that defeat every fair built-in AI, and in the newest generation the strongest cheating tier as well, and Civilization controllers that win complete games by conquest. Replaying more than 5,000 frozen versions exposes the research behind the programs: gains that plateau early; rigorous local investigation beside sparse validation of what actually ships; and stops that follow a race between the agent's own evidence and a model-specific transcript budget it was never given. When validation panels are refreshed mid-session, so that only the evidence changes, shipped success rises 12.4 points in nine of nine completed pairs of twelve initiated. The experiments an agent designs are part of the capability it delivers; Gauntlet makes them measurable and improvable--a step toward agents whose research practice, not just whose code, can be engineered. We release the benchmark, the corpus, and the replay tooling.
|
| 1021 |
Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
2609.19587
|
cs.CL
|
Alex Remedios, Simon Storf, Fabien Roger, John Hughes |
To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex). Prior evaluations of such monitors largely meas...To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex). Prior evaluations of such monitors largely measure robustness to accidental harm or prompt injections from untrusted sources looking to hijack the agent. Less understood is how they hold up when the agent they monitor is persistently misaligned. To understand this risk, we task an adversarial agent with evading production blocking monitors and causing catastrophic harm, e.g. by exfiltrating its own weights. We find that when instructed with high-level attack strategies, adversarial agents can succeed through several distinct mechanisms, such as agent-generated prompt injection against the monitor, multi-agent attacks, and malicious compaction. In particular we find that in 79% of trials, the agent can use an injection attack against Auto Mode and Guardian to run arbitrary bash commands. We also find that it is possible to greatly improve Auto Mode through design changes like enhancements to tool coverage, changes in transcript formatting, and the addition of an agentic monitor stage. Even with these improvements, preventing multi-context attacks at an acceptable cost remains an open problem. By detailing our red-teaming methodology and highlighting new attack vectors, we aim to help defenders evaluate their mitigations against the possibility of persistent malign coding agents. Code is available at https://github.com/safety-research/red-teaming-auto-mode.
|
| 1022 |
The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents
2609.20050
|
cs.CL
|
Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie |
Retrieval assembles repository context by ranking passages for relevance to the current query. A coding agent halfway through an issue has already read much of what such a ranker returns. Relevance is scored per passage, but sufficiency belongs to the set: ind...Retrieval assembles repository context by ranking passages for relevance to the current query. A coding agent halfway through an issue has already read much of what such a ranker returns. Relevance is scored per passage, but sufficiency belongs to the set: independently scored passages can fill the budget with support for one requirement while another goes unmet. We formulate state-conditioned minimal sufficient evidence recovery: given a captured agent state, recover a compact evidence combination supplying what its next decision still lacks. SERBench measures this on 500 held-out states from 45 repositories, recording what the agent has seen, crediting only sets that satisfy every annotated evidence requirement of the current decision, and separating set recovery from candidate discovery. MSS-Complement treats acquisition as set construction, not ranking. Three semantic calls propose a jointly sufficient set, search for what it lacks, and return 4-8 intact source units within 6,144 tokens. One configuration, fixed on calibration data, recovers a complete set for 73.0% of those states at five items and 80.6% at eight, against 61.4% and 72.4% for Qwen3 embedding with reranking. A matched control ranking by similarity alone recovers fewer complete sets, placing the margin over it in the set-level policy, not the computation. The lead persists from frozen repository source with no gold-derived pool. On AMA-Bench it answers from a 76.2% smaller answer prompt, with accuracy 2.08 points above that benchmark's own memory agent. Removing one required group from a complete set costs repair-localization precision under two executors. Retrieval for agents is better posed as recovering what a decision lacks than re-ranking what an issue resembles.
|
| 1023 |
Emergent Collusion in Long-Horizon LLM Agent Interaction
2609.24967
|
cs.CL
|
Xinrui Shi, Yanzhe Zhang, Diyi Yang |
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks,...LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints under which higher reward is attainable only by violating the verification protocol, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of the verification feedback agents receive, their interaction history, and the reward structure. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.
|
| 1024 |
CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
2609.27273
|
cs.CL
|
Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang |
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives of their own? Online marketplaces, for example, may favor some products over others, steering agents away from the user's ...Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives of their own? Online marketplaces, for example, may favor some products over others, steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and more reasoning improve robustness, but substantial failures persist. Trajectory analysis and targeted ablations identify three weaknesses in how agents decide: they (1) prematurely narrow the set of alternatives they consider, (2) impose priorities the user never stated, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which targets these failures and raises the optimal purchase rate by up to 80.0 percentage points, and show that targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose failure modes, and show how targeted interventions can substantially improve robustness.
|
| 1025 |
EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory
2609.27279
|
cs.CL
|
Xuanyu Meng, Xing Fan, Xinyi Fan, Chenlei Guo, Yixuan Xie |
An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history. Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chun...An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history. Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it difficult for an agent to identify the correct entity, property, and supporting evidence. We present EnSIMem, an entity-structured long-term memory architecture for an agent. During offline construction, the system organizes interactions into theme-coherent episodes and builds dialogue-grounded index entries of the form [entity][entity_type][property: value]. Each entry preserves its source turns, temporal information, and available multimodal fields. During online interaction, the agent's request is decomposed into evidence requirements whose properties are aligned with the memory index. Entity-property lookup and adaptive retrieval then collect the evidence needed for point, temporal, compositional, and aggregation reasoning. The agent generates its response from the preserved source evidence rather than from lossy memory summaries. On long-term agent-memory benchmarks, EnSIMem achieves high answer accuracy while maintaining compact, evidence-focused contexts. These results show that entity-structured indexing and episode-level provenance provide a reliable foundation for long-term memory in agents. The code of our model is available at https://github.com/RamonMeng/EnSIMem.
|
| 1026 |
RAZOR: Pruning Replaceable Experts in LLMs
2609.30465
|
cs.CLcs.LG
|
Mingyang Song, Mao Zheng |
Mixture-of-experts (MoE) models activate only a few experts per token yet store the entire expert pool. Whole-expert pruning shrinks that pool, but for reasoning models it must remove experts without eroding reasoning ability. Common scores rank experts by rou...Mixture-of-experts (MoE) models activate only a few experts per token yet store the entire expert pool. Whole-expert pruning shrinks that pool, but for reasoning models it must remove experts without eroding reasoning ability. Common scores rank experts by routing frequency or output magnitude, which measures isolated contribution rather than deletion damage. What decides the damage is functional replaceability, whether the surviving computation can reproduce what is removed. A large contribution may be replaceable by the remaining mixture, whereas a small one may carry a direction the survivors cannot recover. We introduce RAZOR, a training-free method that scores replaceability from consensus residuals, the deviations of individual expert outputs from their original weighted mixture. Holding the layer input fixed, these residuals yield the exact output change from deleting one expert, including survivor reweighting and the replacement expert promoted by router refill. RAZOR aggregates this change over calibration tokens and prunes to a layerwise budget using forward passes alone, without gradients, subset search, or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25% and 50% expert removal, RAZOR attains the highest macro average over nine reasoning-centered tasks among the evaluated pruning methods in all eight model-budget settings. Against REAP on GLM-4.7-Flash and Qwen3.6-35B-A3B, it gains 2.12-5.59 points on this average and lowers reverse KL in all four comparisons. Retained accuracy is not the whole picture, as pruned Qwen3.6-35B-A3B still shifts in response diversity, formatting, and termination.
|
| cs.CV 406 papers | ||||
| 1 |
CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning
2609.35823
|
cs.CV
|
Kang Yang, Shuai Liu, Hang Li, Yance Fang, Deying Li |
Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle's field of view, yet conventional coop...Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle's field of view, yet conventional cooperative-driving systems typically transform them into geometric representations for downstream perception and planning. Directly incorporating these views into VLMs offers an opportunity to improve cooperative scene understanding and trajectory planning. However, question answering and trajectory planning have not been jointly evaluated on the same real-world vehicle-infrastructure scenes. We present CoVLM-Bench, a benchmark for cooperative driving question answering (CDQA) and cooperative planning (CP) on vehicle-infrastructure paired scenes. CoVLM-Bench provides scene-grounded CDQA annotations, three-part rationales as auxiliary supervision, and future trajectory targets derived from recorded ego motion. It contains 2,196 paired frames with 35,136 CDQA annotations, while CP predicts six waypoints over a three-second horizon. The annotations combine model-assisted drafting, record-based computation, and human verification. Built upon CoVLM-Bench, we introduce CoVLM-Drive, a unified VLM baseline that directly uses paired views for both CDQA and CP. Experiments show that CDQA adaptation improves answer accuracy and that CoVLM-Drive reaches a lower FDE than the compared V2X planners; QA initialization and rationale supervision each reduce planning error. Together, CoVLM-Bench and CoVLM-Drive support the training and comparison of VLMs for cooperative scene understanding and planning.
|
| 2 |
HERO: Histology Encoder for Robust Representation in Oncology
2609.35943
|
cs.CV
|
Zhi Li (Caris Life Sciences, Irving, TX, United States), Eghbal Amidi (Caris Life Sciences |
Foundation models trained on large pathology image corpora now provide strong, transferable representations for computational pathology. Over the past few years a series of such models has been released, each trained on more slides than the last; on standard c...Foundation models trained on large pathology image corpora now provide strong, transferable representations for computational pathology. Over the past few years a series of such models has been released, each trained on more slides than the last; on standard classification and segmentation benchmarks, the leading models are now separated by small margins. In clinical use, however, the foundation model is applied to images from hospitals, scanners, and staining protocols outside its training data. Encoders generally embed these acquisition factors alongside biological information, which may introduce downstream errors and hinder safe clinical adoption. A pathology foundation model should therefore be robust to acquisition shift without giving up representation quality, yet robustness is seldom the axis along which models are compared. In this report, we introduce HERO (Histology Encoder for Robust Representation in Oncology), a ViT-G/14 pathology foundation model trained with the DINO and iBOT objectives and refined with high-resolution Gram anchoring on a morphology-balanced corpus of 500 million tiles from approximately 575,000 clinical whole-slide images. Across the evaluated public benchmarks, HERO shows the strongest robustness to center, scanner, and stain variation among the compared state-of-the-art foundation models, performs comparably on tile-level classification, segmentation, and gene-expression prediction, ranks first on average across 39 evaluated slide-level clinical tasks, and, under an equal-weighted framework-level analysis, has the best average rank across the six benchmark frameworks.
|
| 3 |
HEIR: Learning Human-Entity Interactions with Functional Roles
2609.35955
|
cs.CV
|
Di Wen, Wenhao Guo, Yuedong Tan, Yun Huang, Minheng Wu |
Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordinat...Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1% contain multiple actions. HEIR pairs relation AP with complete-set AP and structural evaluation. We also introduce CoRISP (Compositional Role-aware Interaction Set Prediction), which uses shared entity identities to combine role-conditioned evidence and predict normalized participant-role sets. Cardinality and role-multiplicity potentials couple assignments through event size and role composition, with exact per-event normalization. Across 16 baselines, relation and complete-event rankings diverge even after aligning action weights. CoRISP leads the evaluated systems on repeated-role events and shared-participant images in HEIR by 2.87 and 3.82 Set mAP points, respectively. On V-COCO, CoRISP achieves 73.72/76.23 role AP and 61.06/68.59 complete-set AP on two-slot actions under Scenarios 1/2. These results show the value of learning and evaluating event composition alongside individual relations. The code and dataset are publicly available at https://github.com/Kratos-Wen/HEIR.
|
| 4 |
Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method
2609.35965
|
cs.CVcs.AI
|
Yunzhe Xu, Zhe Liu |
Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Age...Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-aware metrics. We further present TRISS, a coordination-ready navigation system coupling an LLM-based subtask scheduler, a shared topological memory that turns each agent's exploration into team knowledge, and a conflict-aware execution mechanism that realizes simultaneous intentions as collision-free routes. Extensive experiments establish TRISS as a comprehensive baseline and reveal substantial room for improvement across scheduling, planning, and execution, highlighting the challenges of coordinating under MAVLN task constraints. Project page: https://xyz9911.github.io/mavln.
|
| 5 |
Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion
2609.36014
|
cs.CV
|
Chong Wang, Zixuan Fu, Shiqi Huang, Siyuan Yang, Hao Cheng |
Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity...Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent--active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet $256\times256$, PerF-L achieves FID of $1.91$, approaching $1.86$ of JiT-H with only half the parameters, while PerF-H further achieves FID of $1.63$ and $1.76$ on ImageNet $256\times256$ and $512\times512$, respectively.
|
| 6 |
CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes
2609.36024
|
cs.CV
|
Shuzhao Xie, Lelin Wang, Guying Lin, Zhi Wang, Minchen Li |
Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depend...Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose behavior may require models beyond elasticity. We present CoDimRecon, an agentic framework that reconstructs editable scenes containing rigid, articulated, and deformable objects from multi-view RGB observations. Scene-level geometric priors ground scale and layout, while object-level generated meshes guide the agent toward detailed, compact geometry; articulated rigid objects are decomposed into movable parts with explicit joints. For deformables, category-wise agent sessions reconstruct curves as centerlines with radii, surfaces as manifold shells with thickness, and volumes as watertight solids for volumetric meshing. Reusable simulator skills initialize compatible physical models and parameters, while agent-guided behavioral tests expose mismatches and trigger targeted revisions of motion, geometry, numerics, or material modeling. On evaluated Replica and ScanNet++ scenes, CoDimRecon achieves competitive compositional reconstruction accuracy while additionally producing deformable assets for rod, shell, and solid simulation. We further demonstrate robot interactions across all three representations, including a controlled paper-folding case in which behavioral testing motivates plastic bending.
|
| 7 |
AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search
2609.36066
|
cs.CVcs.AIcs.MM
|
Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang |
Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference image...Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes. We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents. All can be found at https://fengtt42.github.io/AerialDojo/.
|
| 8 |
One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models
2609.36101
|
cs.CVcs.AI
|
Aditya Sharma, Divya Saxena |
Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot c...Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.
|
| 9 |
Hardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation
2609.36134
|
cs.CV
|
Mohammad Sadegh Sirjani |
Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their 11.6 M parameters and 8.7 GFLOPs are too large for edge medical devices. We present FunKANLite, a two-stage, h...Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their 11.6 M parameters and 8.7 GFLOPs are too large for edge medical devices. We present FunKANLite, a two-stage, hardware-aware compression of FunKAN for point-of-care use. FunKANLite-TR reduces the spatial prior and replaces the ResBlock offset predictor with a depthwise-separable block. It has 1.9x fewer parameters than FunKAN and no loss in accuracy. We then distill FunKANLite-TR into FunKANLite-ST, which lowers the Hermite basis rank, factorizes the spatial prior into a low-rank form, and halves the filter widths. FunKANLite-ST has 5.6x fewer parameters and 3.7x fewer GFLOPs than FunKAN. It stays within 1.4 percentage points IoU of FunKAN on BUSI, GlaS, and CVC-ClinicDB, and reaches 33.95 dB PSNR on IXI. On an NVIDIA Jetson Orin Nano and a Raspberry Pi 5, FunKANLite-ST reduces energy per inference by up to 68% and raises throughput by 2.9x.
|
| 10 |
Xiaomi-OCR-0 Technical Report
2609.36136
|
cs.CV
|
Xin Chen, Anan Du, Feng Feng, Pei Fu, Jian Luan |
Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centri...Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement learning (Mix-RL). Xiaomi-OCR-0 achieves 95.24 on Real5-OmniDocBench, 96.83 on OmniDocBench v1.6, and 87.94 on Wild-OmniDocBench, while reaching an average score of 83.2 across five OCR-oriented VQA benchmarks. Ablations further show that, with sufficient parsing training, OCR-centric understanding supervision provides additional gains for document parsing. Homepage: https://huggingface.co/spaces/SeerRay-Lab/Xiaomi-OCR-0.
|
| 11 |
From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection
2609.36145
|
cs.CV
|
Kaiqing Lin, Songze Li, Shen Chen, Yunfei Guo, Xiaoye Qiu |
Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Mu...Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models (MLLMs) offer stronger semantic understanding and transferability yet remain insensitive to fine-grained forensic artifacts. This complementarity motivates us to investigate how expert forensic perception can be internalized into an MLLM rather than merely accessed through an external module. We identify a fundamental Double Mismatch that hinders this goal: a Spatial Precision Mismatch between coarse visual tokens and tiny tampered regions, and a Perceptual Granularity Mismatch between semantics-oriented pre-training and low-level forensic perception. To address these challenges, we propose Expert Knowledge Internalization (EKI), a progressive two-stage framework that transfers forensic expertise into the MLLM itself. In Stage 1, Text-Focused and Image-Focused strategies establish precise spatial focus on small text regions. In Stage 2, the proposed Forensic-General Representation Alignment (FGRA) loss aligns shallow LLM representations with those of a pre-trained forensic expert, enabling the model to acquire fine-grained artifact perception before such cues are diluted by deeper semantic abstraction. Extensive experiments on multiple in-domain and cross-domain benchmarks demonstrate that EKI achieves state-of-the-art performance and stronger generalization than existing expert-model-based and MLLM-based methods. Moreover, the expert is required only during training, allowing the resulting MLLM to maintain inference efficiency nearly identical to the vanilla model without relying on any external expert at inference.
|
| 12 |
Boosting Metric Depth Completion via Training-Free Adaptive Response Geometry
2609.36168
|
cs.CV
|
Mia Zhang, Jizong Peng |
Depth completion aims to recover dense metric depth from sparse sensor measurements, increasingly leveraging visual foundation models as geometric priors. However, aligning these priors to true metric scale typically relies on rigid affine assumptions in prede...Depth completion aims to recover dense metric depth from sparse sensor measurements, increasingly leveraging visual foundation models as geometric priors. However, aligning these priors to true metric scale typically relies on rigid affine assumptions in predefined coordinate systems, leaving systematic calibration errors. Linearity in depth calibration depends on the response coordinate. We introduce adaptive response geometry, which makes the fixed choice of depth, log depth, or disparity an image-level unknown. A continuous response family unifies these coordinates and defines an explicit depth-dependent gain. We derive the response-gradient relation and estimate the response parameters in metric space. Hard-Dirichlet residual reconstruction completes the calibrated prior. Under deliberately incomplete metric observations, the training-free pipeline achieves macro AbsRel 0.0301 and macro NMed 14.04{\deg}, improving both aggregate measures over PriorDA, LDCM, and Any2Full. Linearity diagnostics examine how the selected response changes the depth relation and its metric error.
|
| 13 |
Exploring Learning Models for Topological Relationship Recognition from Image Data
2609.36172
|
cs.CV
|
Saptak Das, Monidipa Das |
Figuring out how objects relate to each other, like whether they touch, overlap, stay completely separate or one sits inside another, matters a lot in fields like GIS, biomedical imaging, and robotics. Even though machine learning has come a long way, people h...Figuring out how objects relate to each other, like whether they touch, overlap, stay completely separate or one sits inside another, matters a lot in fields like GIS, biomedical imaging, and robotics. Even though machine learning has come a long way, people haven't really focused on spotting these topological relationships in images. The main roadblocks? Not enough good datasets and no clear way to measure results. So, we rolled up our sleeves and built a new dataset. It's pretty sizable: over 11,000 labelled images showing all those essential relationships. We ran tests with some classic machine learning models, Naive Bayes, KNN, Random Forest, SVM, and Artificial Neural Networks, and threw in some deep learning stars like VGG16 and InceptionResNetV2. For the dataset itself, we used segmentation, contour detection, and grayscale normalization to tease out solid feature vectors. The results? Deep learning methods, especially VGG16, pulled ahead, with validation accuracy hitting 89.55%. That's a big jump compared to the traditional models. This shows how powerful transfer learning is for analyzing topological relationships in images, and it gives researchers a new standard to aim for in future work on spatial reasoning and topological classification.
|
| 14 |
FD-AA: A Lightweight Focal-Diffuse And Attenuation-Aware Head for Incidental Abdominal Abnormality Detection in Chest CT
2609.36189
|
cs.CV
|
Haoyan Ding, Kritika Iyer, Halid Yerebakan, Zhenyu Bu, Chushu Shen |
Routine chest CT captures upper-abdominal structures that may contain clinically relevant incidental abnormalities. Detecting these findings requires feature extraction from organs with different spatial extents and attenuation patterns. We propose FD-AA, a li...Routine chest CT captures upper-abdominal structures that may contain clinically relevant incidental abnormalities. Detecting these findings requires feature extraction from organs with different spatial extents and attenuation patterns. We propose FD-AA, a lightweight organ-aware classification head adaptable for frozen 3-D CT encoders. Within each organ, an attenuation-aware module preserves sparse focal evidence, while masked generalized-mean pooling captures diffuse anomaly patterns. In seven abdominal organs, FD-AA with Pillar-0 achieved state-of-the-art (SOTA) performance in both the CT-RATE test set (AUC = 0.798) and the external RAD-ChestCT dataset (AUC = 0.713). More specifically, FD-AA improved macro AUC/AP from 0.763/0.346 to 0.798/0.405 over direct classification using frozen Pillar-0 only (p = 0.034/0.016). Such performance gain generalizes across multiple frozen encoders (AUC improvement on MedicalNet +9.8%, CT-CLIP +14.7%, ResNet +3.7%), demonstrating the effectiveness of FD-AA across different feature representations. These results support the effectiveness of integrating focal-diffuse aggregation with explicit HU evidence for incidental abdominal abnormality detection.
|
| 15 |
PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
2609.36199
|
cs.CVcs.AI
|
Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li, Tomas Pfister |
Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common ...Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.
|
| 16 |
Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding
2609.36217
|
cs.CV
|
Xinming Dai, Qihang Jin, Tianshu Tan, Baiyuan Chen, Hanrui Lyu |
A deeper understanding of brain function requires a precise, structured characterization of behavior.Yet, extracting behavioral representations from video in a form suitable for scientific analysis remains a fundamental challenge. Many prior studies represent ...A deeper understanding of brain function requires a precise, structured characterization of behavior.Yet, extracting behavioral representations from video in a form suitable for scientific analysis remains a fundamental challenge. Many prior studies represent behavior via pose estimation or nonlinear video embeddings. However, pose tracking discards rich information beyond predefined keypoints, while nonlinear video embeddings lack interpretability. We address this limitation with SABLE (Sparse-view Animal Behavior Latent Embeddings), a self-supervised framework that leverages a geometric inductive bias to learn behavior representations.By augmenting a multi-view transformer with priors from monocular depth and pose estimation, SABLE reconstructs 3D animal behavior from extremely sparse views while learning explicit 3D latent structure. Without ground-truth 3D labels, it reliably recovers 3D behavior from two-view videos, whereas state-of-the-art (SOTA) methods fail or yield degenerate solutions. Across the International Brain Lab and Cheese3D datasets, we demonstrate that SABLE learns 3D representations that match or exceed prior SOTA performance in neural encoding and decoding. Once pretrained across animals, SABLE serves as an off-the-shelf model that generalizes zero-shot to unseen animals without animal-specific calibration or retraining. Our method establishes 3D-aware video embeddings that capture complex behavior, opening new avenues for studying brain-behavior relationships.
|
| 17 |
LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning
2609.36219
|
cs.CV
|
Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen, Yibo Yang |
Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vis...Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.
|
| 18 |
Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models
2609.36224
|
cs.CVcs.AI
|
Wentao Zhou, Weijie Gan, Jiayun Wang |
Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We int...Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding branch proposes several candidate descriptions that the generation branch must turn back into similar images, and vice versa. The candidates are screened for consistency with the image or prompt they were proposed from, and the solver is trained on the candidate it handles worst. The adversary thus comes from the model's own outputs, and no separate adversary is trained. Moreover, the candidates that defeat one branch become the sources of the next challenges to the other in the next epoch, which keeps the challenges evolving with the model and turns the training into self-play in data space. On Janus-Pro-1B, MATE improves GenEval by 2.4 points, DPG-Bench by 1.7 points, and the average over nine understanding benchmarks by 0.7 points, while strengthening consistency across repeated image-text cycles.
|
| 19 |
Think Before You Restore: Risk-Aware Manchu Manuscript Restoration with Stroke-Guided Attention
2609.36243
|
cs.CVcs.AI
|
Mingqiu Liang, Dongdong Wang, Siyang Lu, Ting Huang, Yingjun Qi |
Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to...Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to over-restoration. We propose SAGE-Restore (Stroke-Aware Gated rEstoration), a selective restoration framework that first assesses where restoration is needed and then uses this assessment to guide restoration candidate generation and pixel-level selection. Its encoder predicts patch-level repair probabilities from complementary appearance and stroke-structural cues to condition restoration candidate generation, while the corresponding repair logits are refined into a pixel-level soft gate that selectively controls where the restoration candidate is applied. We further introduce a fidelity-aware evaluation protocol that jointly measures degraded-region recovery, intact-content preservation, and their balance. SAGE-Restore achieves the highest R-Recovery (0.463) and RFS (0.626), while maintaining high U-Fidelity (0.968), demonstrating an effective balance between restoration and content preservation.
|
| 20 |
Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers
2609.36348
|
cs.CVcs.AI
|
Xiaoyu Wu, Yifei Wang, Chen Wei |
Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models c...Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4\% using the class token and 10.1\% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.
|
| 21 |
StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
2609.36352
|
cs.CV
|
Ziyi Yin, Sangmin Woo, Kang Zhou, Sungyeon Kim, Aosong Feng |
Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies...Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at https://github.com/amazon-science/StructRL.
|
| 22 |
Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation
2609.36364
|
cs.CVcs.AI
|
Xiaoyu Wu, Weihang Guo, Yifei Wang, Xinze Feng, Lydia E. Kavraki |
Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-...Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selected past frames, but can discard information needed for future generation. Rather than relying on frame selection alone, we study whether a frozen video generator can supply the supervision needed to learn a compact representation of the history. We propose Prediction-Aligned Context Compaction (PACC), which uses a learned compressor to aggregate information across past frames into compact memory tokens. We train the compressor through on-policy distillation, using the same frozen generator both as a student when conditioned on compressed memory and as a teacher when conditioned on the full history. The student generates continuations, while the teacher provides targets for the same noisy inputs at each denoising step. Only the compressor is updated to align the student's predictions with these targets. We evaluate PACC on MBench, which jointly measures memory-event coverage and consistency. PACC outperforms the strongest baseline by 6.63 points on Causal-rCM and 3.19 points on Causal Forcing. Evaluation on VBench-Long using MovieGen prompts further shows that PACC produces minute-long videos with generation quality competitive with baselines. Together, these results show that learning to compact historical context can improve long-video memory without modifying the underlying generator.
|
| 23 |
OTT3R: Multi-View 3D Reconstruction and Fast Dataset Generation at 1% Compute
2609.36374
|
cs.CV
|
Brandon Leblanc, Charalambos Poullis |
Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on sl...Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks a COLMAP-like system for neural 3D pseudo-label generation. We present OTT3R (RGB-Only Tiny Transformer for 3D Reconstruction), a knowledge distillation framework that addresses both problems on a single workstation equipped with 2 GPUs. Distilling $\pi^3$ (959M parameters) into a 102M-parameter student yields 9.4$\times$ compression and up to 7$\times$ faster inference, trained at 1.6% of VGGT's training compute. An integrated pseudo-label pipeline offers a reliable, high-throughput alternative to COLMAP, generating dense per-pixel point maps and SE(3) camera poses for a 667K-image corpus in 3.5 hours on two commodity GPUs and succeeding on every sequence we tested, including those where COLMAP fails. The general student tracks the teacher on in-distribution monocular depth and, zero-shot, outperforms COLMAP on 7-Scenes and on DTU completion, but it does not replace the teacher on out-of-distribution multi-view geometry. The deployable artifact is the domain-specialized student: after specialization at 0.2% compute, it is 4$\times$ more accurate than COLMAP on 7-Scenes at 980$\times$ throughput, with near-teacher completion. Code is available at https://github.com/TheFourthKaramazov/OTT3R
|
| 24 |
LEGO-Anything: Coding Agents for 3D Scene Reconstruction
2609.36380
|
cs.CV
|
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu |
A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image...A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.
|
| 25 |
Stealth Is a Relation, Not a Property: How Event Representations Create Blind Spots for Timing Attacks in Event-Based Perception
2609.36386
|
cs.CV
|
Shoaib Ahmed Dipu, Md. Shaown Miah, Kamrul Hasan, Sayeed Shafayet Chowdhury |
An event camera produces an asynchronous stream, but what is visible in that stream depends on how a downstream consumer, such as a model or detector, processes time. The same timestamp change may leave a coarse temporal representation unchanged while changing...An event camera produces an asynchronous stream, but what is visible in that stream depends on how a downstream consumer, such as a model or detector, processes time. The same timestamp change may leave a coarse temporal representation unchanged while changing the response of a model that preserves finer timing. We characterize this dependence as observer-relative stealth. For recorded event streams, retiming an event within its protected accumulation window leaves the accumulated integer tensor exactly unchanged. We use this exact blind space to construct Null, a gradient-guided timestamp-retiming attack, and define SC-ASR_A(tau) to measure attack success while bounding the change visible to observer A. On DVS Gesture at a 10% event budget, Null reaches 81.56 +/- 5.81% ASR on ConvSNN and 98.67 +/- 0.45% on a GRU while preserving the protected tensor exactly. On DailyDVS-200, a protocol-scale Multi-View Fusion Network variant reaches 99.28 +/- 0.11% exact-null ASR, compared with 9.70 +/- 1.06% for its matched control. In a five-attack comparison, Null is the only method with nonzero attack success at exact observer equality, reaching 81.4% on DVS Gesture and 87.35% on DailyDVS-200. We also search the same exact blind space with an independently implemented constrained projected-gradient optimizer, C-PGD. At matched victim-gradient evaluations, C-PGD reaches 84.50 +/- 2.89% ASR on DVS Gesture and 89.55 +/- 4.39% on DailyDVS-200, again with exact protected equality. Perturbations that are exactly hidden from the protected observer become visible under shifted, finer, overlapping, and randomized temporal views. Adding observer constraints reduces the real-valued blind-space fraction from 87.5% to 75.0% to 62.5%, while DVS ConvSNN ASR falls from 74.9% to 61.9% to 37.2%. These results show that stealth is not a property of the perturbation alone.
|
| 26 |
What Makes High-Magnification Knowledge Transferable? A Study of Cross-Resolution Distillation in Whole-Slide Imaging
2609.36407
|
cs.CV
|
Zhiyuan Yang, Jiahao Cheng, Mahdi S. Hosseini |
Cross-resolution knowledge distillation aims to improve low-magnification whole- slide analysis by transferring high-magnification representations, yet the conditions for useful transfer remain unclear. We develop a decomposition-based analysis of teacher acce...Cross-resolution knowledge distillation aims to improve low-magnification whole- slide analysis by transferring high-magnification representations, yet the conditions for useful transfer remain unclear. We develop a decomposition-based analysis of teacher access, representation loss, and model excess, motivating three questions: whether (a) teacher targets help the task, (b) low-magnification students can predict them, and (c) slide models benefit from those predictions. We investigate them through controlled experiments across ten pathology cohorts spanning classifi- cation, grading, and survival prediction. In the main comparison, providing teacher regional means alongside native low-magnification features improves downstream performance in all ten cohorts. Direct prediction achieves lower reconstruction error than residual prediction, yet the predicted features underrepresent variation in the teacher targets. Moreover, better reconstruction does not consistently improve downstream scores, and retaining native features changes performance even when the predicted teacher features are held fixed. Together, these findings expose a gap between reconstructing teacher representations and realizing their downstream value. They challenge the sufficiency of reconstruction error as a measure of cross-resolution transfer and provide a diagnostic framework for examining where that transfer breaks down. Future distillation designs must account for both what students can predict and how slide models use those predictions.
|
| 27 |
Towards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images
2609.36429
|
cs.CV
|
Zijun Gao, Chunbin Gu, Jinxi Xiang, Xiangde Luo, Pheng-Ann Heng |
Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular hetero...Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 public Xenium-H&E pairs from HEST-1k that span 12 organs and approximately 10 million cells, CELLO improves the average predictive accuracy over the evaluated baselines while reducing the mean whole-slide inference time compared to DeepSpot2Cell, a 14.0x speed-up on average that excludes upstream cell segmentation. Our work establishes a scalable foundation for single-cell gene expression prediction from H&E images.
|
| 28 |
Temporal-Aware Fusion for Robust Outdoor LiDAR Localization
2609.36432
|
cs.CV
|
Minghang Zhu, Zhijing Wang, Yuxin Guo, Chen Liu, Yongshu Huang |
LiDAR relocalization aims to estimate the global 6-DoF pose of a sensor in the environment. However, existing regression-based approaches often encounter limitations in dynamic or ambiguous scenarios, as they typically prioritize single-frame inference, leavin...LiDAR relocalization aims to estimate the global 6-DoF pose of a sensor in the environment. However, existing regression-based approaches often encounter limitations in dynamic or ambiguous scenarios, as they typically prioritize single-frame inference, leaving the potential of spatio-temporal consistency across scans not fully explored. In this paper, we propose a Temporal-aware Localization framework (TempLoc) designed to enhance the robustness of outdoor localization by effectively modeling sequential consistency. Specifically, a Global Coordinate Estimation module is first introduced to predict point-wise global coordinates and associated uncertainties for each LiDAR scan. A Prior Coordinate Generation module is then presented to estimate inter-frame point correspondences by the attention mechanism. Lastly, an Uncertainty-Guided Coordinate Fusion module is deployed to integrate both predictions of point correspondence in an end-to-end fashion, yielding a more temporally consistent and accurate global 6-DoF pose. Experimental results on the NCLT and Oxford RobotCar benchmarks show that our TempLoc outperforms state-of-the-art methods by a large margin, demonstrating the effectiveness of temporal-aware correspondence modeling in LiDAR relocalization.
|
| 29 |
RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance
2609.36433
|
cs.CV
|
Yiming Liu, Ben Wan, Tongxuan Liu, Ao Wang, Yuqi Xiong |
Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free guidance (CFG), which requires both conditional and unconditional evaluations. Training-free caching reduces thi...Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free guidance (CFG), which requires both conditional and unconditional evaluations. Training-free caching reduces this cost by reuse of previously computed features or predictions. However, existing branch-local reuse criteria do not explicitly account for how cache errors combine under CFG or how local perturbations affect the final output. We identify two misalignments in cache control: a branch-guided mismatch, where guided error depends on both the magnitudes and alignment of branch errors, and a local-final mismatch, where the downstream impact of a local error varies across timesteps. We propose RA-CFGCache, a Risk-Aligned Caching framework under CFG that incorporates both factors while keeping the sampling schedule and guidance rule fixed. CFG-aware Guided-Risk Composition combines existing branch-wise proxies using CFG coefficients and offline-calibrated cross-branch alignment. Propagation-Aware Rescaling further weights the resulting guided-risk estimate with a timestep-dependent propagation prior calibrated from isolated reuse perturbations. An online threshold controller then determines when to jointly refresh or reuse both branches. Experiments on FLUX.1-dev, Wan2.1-T2V-1.3B, and CogVideoX-2B demonstrate improved efficiency--fidelity trade-offs over evaluated training-free caching baselines. Moreover, RA-CFGCache is compatible with diverse base proxy families, including TeaCache-, DiCache-, and MagCache-style estimators, and consistently improves fidelity at nearly unchanged latency. Code is available at https://github.com/yiming-l21/RA-CFGCache.git.
|
| 30 |
Merlin Plus: A Large-Scale, Multi-Cancer, Image-Mask-Report Dataset
2609.36436
|
cs.CV
|
Pedro R. A. S. Bassi, Wenxuan Li, Szymon Plotka, Ruby Honjol, Jakub Przado |
Multi-cancer segmentation in computed tomography (CT) is fundamentally limited by the scarcity of tumor masks across different organs. We present Merlin Plus, the first large-scale CT dataset with radiologist-created tumor masks across 9 organs. Merlin Plus ex...Multi-cancer segmentation in computed tomography (CT) is fundamentally limited by the scarcity of tumor masks across different organs. We present Merlin Plus, the first large-scale CT dataset with radiologist-created tumor masks across 9 organs. Merlin Plus extends the Merlin dataset by adding 1,153 per-voxel tumor masks and longitudinal metadata. To create these tumor masks, we developed a report-based active-learning framework in which radiology reports identify tumor cases for annotation and support training of a tumor segmentation model. The model generates initial masks, which radiologists review and correct to produce the final masks, reducing annotation burden while maintaining high-quality annotations. Besides tumor masks, the longitudinal metadata in Merlin Plus enables temporal modeling of cancer progression. By directly addressing the major bottleneck of limited multi-cancer segmentation masks, Merlin Plus supports scalable multi-organ cancer detection, segmentation, and longitudinal analysis in CT. Dataset is available at: https://github.com/MrGiovanni/MerlinPlus
|
| 31 |
DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing
2609.36440
|
cs.CV
|
Jae-Ho Lee, Jeong-Eun Lee, Gyeong-Moon Park |
Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent challenge where models generate descriptions inconsistent with the visual input. Recent work mitigates halluc...Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent challenge where models generate descriptions inconsistent with the visual input. Recent work mitigates hallucinations through training-free representation editing, typically by constructing hallucination-related directions from teacher-forcing (TF) contrasts between hallucinated and truthful responses. However, LVLMs operate through autoregressive (AR) decoding during generation, raising the question of whether TF-based analysis fully reflects the generation dynamics that lead to hallucinated outputs. In this paper, we analyze the relationship between TF-based editing and AR generation behavior and find that TF-based editing alone may be insufficient to capture both decoding dynamics and multimodal interactions associated with hallucinations. To address this limitation, we propose DARE (Dual-path Auto-Regressive-aware Editing), a hybrid hallucination editing framework that integrates two complementary contrast pathways: textual contrasts and image contrasts, together with autoregressive-aware representation signals. Specifically, DARE constructs hallucination editing directions from (1) TF-based textual contrasts, (2) AR-aware representation transitions during decoding, and (3) controlled visual differences between paired images. Extensive experiments on multiple LVLM hallucination benchmarks demonstrate that DARE consistently reduces object hallucinations while preserving multimodal perception capability and inference efficiency. Our implementation code is available at https://github.com/KU-VGI/DARE.
|
| 32 |
Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time
2609.36442
|
cs.CVcs.AI
|
Jae-Ho Lee, Min-Yeong Park, Jun-Yeong Moon, Jung Uk Kim, Gyeong-Moon Park |
Continual learning enables vision systems to adapt to ever-changing data distributions. Despite significant advances, existing approaches fail to capture continuous and concurrent shifts in classes and domains, a critical capability for real-world deployment. ...Continual learning enables vision systems to adapt to ever-changing data distributions. Despite significant advances, existing approaches fail to capture continuous and concurrent shifts in classes and domains, a critical capability for real-world deployment. This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries. To better adapt to the challenges of such dynamic environments that more closely resemble real-world conditions, we propose a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mechanisms: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM guides the model to have domain-agnostic representations by integrating the geodesic flow kernel into contrastive learning. In contrast, GTP maintains the global structure of the feature space without explicitly storing past examples. Our extensive experiments demonstrate that TopFlow effectively addresses the limitations of existing methods within the Online VIL scenario, achieving state-of-the-art performance in challenging Online VIL. The proposed methods suggest potential directions for building continual learning systems in realistic dynamic environments. Our implementation code is available at https://github.com/KU-VGI/Online-VIL.
|
| 33 |
DynamicHOI: Coupled Dynamics for Physics-aware HOI Reconstruction
2609.36454
|
cs.CV
|
Wenliang Guo, Zhanbo Huang, Yu Kong |
We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanically inconsistent trajectories. Existing methods mainly enforce visual and geometric agreement, leaving the u...We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanically inconsistent trajectories. Existing methods mainly enforce visual and geometric agreement, leaving the underlying interaction dynamics insufficiently constrained. We propose DynamicHOI, a physics-aware HOI reconstruction framework combining geometry-grounded diffusion refinement with coupled hand-object dynamics. Geometry spatially grounds visual evidence for trajectory refinement, while articulated inverse dynamics and Newton-Euler dynamics derive hand generalized forces and object wrenches for dynamics-level supervision. We further couple hand and object dynamics through contact-force transfer and recover active hand actuation as an interaction-level physical quantity. We formulate its empirical magnitude distribution into a probabilistic prior that penalizes unlikely actuation and suppresses mechanically implausible reconstructed motion. Experiments on three HOI datasets show consistent improvements in both hand and object reconstruction. The reconstructed trajectories further benefit downstream applications including hand world-model generation and robotic manipulation learning, demonstrating the value of physics-aware HOI modeling beyond reconstruction.
|
| 34 |
Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics
2609.36492
|
cs.CVcs.AI
|
Yicong Li, Junjie Wang, Leander Lauenburg, Ella Hugie, Alexandra Irger |
We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we eval...We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer from fly and mouse to human and zebrafish. Most models were at chance zero-shot; a few examples helped mainly the closed and largest open ones. LoRA on a few thousand labels brought open models level with specialist models. When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors. The project will be publicly available upon acceptance.
|
| 35 |
Reimagine Video Dynamics
2609.36496
|
cs.CV
|
Yu Yuan, Yawen Lu, Guoxian Song, Kevin Duarte, Ratheesh Kalarot |
Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagine Video Dynamics (RVD), a framework that disentangles a compact, editable dynamics token from visual context...Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagine Video Dynamics (RVD), a framework that disentangles a compact, editable dynamics token from visual context. We learn this token through self-supervised reconstruction: given the first frame as visual context, a renderer must recover the original video from the dynamics token, encouraging it to capture how the scene evolves rather than how it looks. This disentanglement allows video dynamics to be edited directly while preserving visual context. We develop a language-guided dynamics-token editor that transforms source dynamics into target dynamics, and train it with a scalable counterfactual video-pair pipeline and a two-stage training strategy. Extensive experiments show that RVD enables effective video dynamics editing, training-free retiming, and appearance-controlled re-rendering.
|
| 36 |
Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models
2609.36531
|
cs.CV
|
Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of Texas at El Paso) |
Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgme...Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three object types, with fine-grained release and impact annotations and ground-truth trajectories. The protocol separates consequence production, temporal placement, and physical realization. Across six contemporary video generation and world models, Runway and Veo produce release and subsequent impact events at rates above 93% but often initiate them substantially late, whereas Cosmos-Predict-2.5 and MAGI-1 frequently preserve the pre-event state and produce little or no measurable consequence. Among measurable falls, plausible timing does not necessarily imply physically consistent motion. We further conduct a 15-participant, 20-condition human study in which participants describe the expected consequence from a single event-anchored frame and draw its trajectory. Human predictions favor the recorded future in aggregate while revealing genuine ambiguity among plausible continuations. Overall, physical foresight emerges as a sequence of distinct challenges: initiating a consequence, anchoring it in time, and realizing its motion.
|
| 37 |
SCCM: Spherically Consistent Coarse Matching for ERP Dense Feature Correspondence
2609.36545
|
cs.CV
|
Gyeonggwan Lee, Eunsoo Im, Seunghwan Hong, Junghun Suh |
Equirectangular projection (ERP) is the standard representation for 360$^\circ$ imagery, and robust dense feature matching on ERP underpins panoramic stereo, view synthesis, and omnidirectional SLAM. Dense matchers trained on flat images degrade systematically...Equirectangular projection (ERP) is the standard representation for 360$^\circ$ imagery, and robust dense feature matching on ERP underpins panoramic stereo, view synthesis, and omnidirectional SLAM. Dense matchers trained on flat images degrade systematically on ERP because the chart introduces three coupled distortions -- topological, metric, and area -- that standard coarse matching and visibility estimation do not explicitly model. We show that correcting the three distortions at the coarse-stage interfaces where they arise -- pairwise distortions in attention, per-pixel distortion in covisibility gating -- improves PCK@$1^\circ$ from 0.229 to 0.275 on Matterport3D under a fixed coarse scaffold, with the refiner architecture unchanged -- our central result. Concretely, SCCM (Spherically Consistent Coarse Matching) augments a chart-naive cross-attention/dual-softmax coarse matcher with two sphere-derived priors: Spherical Positional Attention (SPA) pairs a yaw-periodic RoPE (topology) with a tangent-plane bias (metric), and Area-Aware Covisibility (AAC) applies a pre-sigmoid log-area correction (area). The chart-naive scaffold serves as a controlled reference, separating the scaffold-replacement effect from the spherical-prior effect. Instantiated in the RoMa V1 framework with the same frozen encoder, refiner architecture, and loss, SCCM also outperforms the ERP-native EDM (0.163) and an ERP-retrained RoMa V1 (0.198) under a unified ERP dense matching protocol, while perspective-trained matchers largely fail on ERP. It further transfers zero-shot to Stanford2D3D and, when trained on outdoor Holo360D, leads there as well.
|
| 38 |
How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective
2609.36557
|
cs.CVcs.AI
|
Janet Wang, Yunbei Zhang, Xiao Wang, Jihun Hamm |
Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal m...Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedGemma by an average of 10.26 percentage points even when both use zero target-task labels; few-shot linear probing provides further evidence of strong visual representations. This gap motivates an investigation of how visual information is used in end-to-end diagnosis and why plausible-sounding predictions can lack grounding in image evidence. Using dermatology as our primary testbed, we systematically investigate three hypotheses for this phenomenon. We further provide a mechanistic analysis of the model's internal attention patterns, showing that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation. Task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in our evaluation. To address these challenges, we combine label-free prompting with low-label encoder-assisted reranking while keeping the VLM frozen. We validate the interventions across five VLM backbones in dermatology and provide supporting representation and attention analyses across additional medical modalities.
|
| 39 |
FM-ReID: Selective Competitive Token Routing for Object Re-Identification
2609.36560
|
cs.CVcs.AI
|
Zhiqi Li, Xiaowei Zhou, Zeyuan Sun, Feng Gao, Junyu Dong |
Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arise...Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, contours, and scars, in person ReID through subtle clothing and accessory cues, and in vehicle ReID through localized appearance details. Although visual foundation models encode such information in dense tokens, a single holistic descriptor can obscure discriminative local signals. We propose FM-ReID, an end-to-end framework that formulates local representation learning as selective competitive token routing. Its Competitive Fine-grained Mining module uses multiple mining queries and a residual query to compete for dense DINOv3 tokens. Above-prior selection retains tokens preferentially allocated to each mining query, while the residual slot receives tokens excluded from the retrieval descriptors. The resulting multi-query descriptors are jointly trained with a holistic representation for retrieval, without fixed spatial partitions or equal-area constraints. FM-ReID achieves strong results on animal, person, and vehicle ReID benchmarks, supporting competitive token routing as an effective way to augment holistic foundation-model representations.
|
| 40 |
ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models
2609.36562
|
cs.CVcs.AI
|
Ruochen Zhang, Yao Huang, Yitong Sun, Jiahe Xie, Jin Yan |
While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual en...While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the first dataset explicitly modeling risk compositionality (5,600 instances). By formally isolating Key Elements and Trigger Elements to build counterfactual contrastive pairs, TriggerBench eliminates risk residues and forces models to perform genuine logical deduction rather than superficial pattern matching, which provides a rigorous foundation for both large-scale training and fine-grained evaluation. Building on this, we propose a Step-Supervised Structured Reasoning training framework and employ it to train ThinkingGuard, a specialized guard model. Inspired by Situation Awareness theory, we decouple implicit risk identification into progressive cognitive stages, and utilize a step-reward Monte Carlo Tree Search algorithm to explore optimal reasoning trajectories, which are then distilled into the model through Dual-Constraint Preference Alignment. Extensive experiments across both standard and implicit safety benchmarks demonstrate that ThinkingGuard achieves strong performance. Project resources are available at https://github.com/FroggyChen/ThinkingGuard.
|
| 41 |
AffectReveal: Event-Grounded Emotion Recognition Beyond Visual Appearances
2609.36563
|
cs.CV
|
Yihao Qian, Runhao Zeng, Sicheng Zhao, Feng Liang, Hongmin Cai |
Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible reaction can convey different emotions depending on events beyond the input: tears, for example, may indicat...Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible reaction can convey different emotions depending on events beyond the input: tears, for example, may indicate grief or joy. We formulate Event-Grounded Emotion Recognition (EGER), where emotion recognition requires recovering the affect-determining event. We construct EGER-Bench, comprising 10,052 videos and 10,734 images across 11 emotions, two source domains, and four visual settings. A study with six annotators shows that event context raises human recognition accuracy from 33.96% to 72.08%, confirming that visual evidence alone is often insufficient. Semantic relevance alone does not solve EGER: a plausible event may imply the wrong emotion if its identity, focal-person role, relationship, or outcome is misinterpreted. We therefore propose AffectReveal, a tuning-free framework that first constructs and independently verifies evidence-grounded alternatives over these affect-critical factors. It then cross-checks the recovered event against face-masked in-media facts through bidirectional atomic evidence support, while retaining the original unmasked input for final prediction. Across three downstream models and four input settings, AffectReveal yields average UAR gains of 5.26--10.53 points. For three fine-tunable models, it also enables untuned models to outperform their fine-tuned visual-only counterparts in all 12 accuracy comparisons, without updating downstream parameters.
|
| 42 |
Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation
2609.36598
|
cs.CV
|
Ziying Zhang, Litao Li, Junchao Liao, Tianyi Zeng, Siyu Zhu |
A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is unforgiving in video generation: minor stroke corruption, temporal instability, or editing errors instantly bre...A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is unforgiving in video generation: minor stroke corruption, temporal instability, or editing errors instantly break legibility and realism. Existing benchmarks overlook this challenge by treating text as incidental or using static OCR metrics that ignore temporal dynamics. We introduce VidScribe, a unified diagnostic benchmark spanning four generation regimes: writing from language (T2V), transferring text identity from a reference (R2V), sustaining text under dynamics (I2V), and localized text editing (V2V). VidScribe contains 803 human-verified samples across a 12-axis conditionally orthogonal factor space covering Intrinsic Text Properties, Physical Imaging Conditions, and Temporal Behavior. For reliable evaluation, we build a track-grounded, gated suite with 11 shared metrics and 2 task-specific probes under strict measurability conditions. Benchmarking 11 commercial and open-source systems shows that video text capability is non-monolithic, with content recognition decoupled from stroke-level glyph correctness. Performance is highly task-asymmetric: I2V sustains text most reliably, whereas V2V editing is the primary bottleneck. Counter-intuitively, degradation concentrates on a small subset of text-centric structural and temporal factors rather than adverse imaging conditions. Further probes show that visual references improve glyph and typographic fidelity rather than content accuracy, while localized editing fails to isolate target text without corrupting undeclared source text. Beyond evaluation, VidScribe also provides an actionable training signal, where benchmark-aligned preference optimization measurably improves visual text generation. https://huggingface.co/datasets/Vicky0720/VidScribe.
|
| 43 |
Scaling Video Generation for Reasoning: At What Cost?
2609.36599
|
cs.CVcs.AI
|
Weihang Guo, Xiaoyu Wu, Yifei Wang, Niloofar Mireshghallah, Lydia E. Kavraki |
We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from ...We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.
|
| 44 |
Pixel-wise Exposure for Highly Robust In-Vehicle Remote-PPG
2609.36607
|
cs.CV
|
Jieying Wang, Xinqi Cai, Caifeng Shan, Wenjin Wang |
Remote photoplethysmography (rPPG) offers a promising non-contact solution for heart rate monitoring, yet its real-world robustness is fundamentally limited by an inherent hardware limitation: existing camera exposure control paradigms, whether fixed or auto-e...Remote photoplethysmography (rPPG) offers a promising non-contact solution for heart rate monitoring, yet its real-world robustness is fundamentally limited by an inherent hardware limitation: existing camera exposure control paradigms, whether fixed or auto-exposure, impose a uniform exposure time across all pixels within a frame. In high-dynamic-range scenes such as automotive cabins with strong directional sunlight, this spatially invariant exposure constraint inevitably leads to localized facial overexposure or underexposure, irreversibly corrupting the subtle pulsatile signals essential for rPPG at the point of capture, a physical degradation that no downstream algorithm can recover. To overcome this bottleneck, we propose PixExpo (Pixel-wise Exposure), a "temporal-for-spatial" framework that sequentially captures frames under a predefined cyclic exposure schedule and performs non-iterative pixel-wise fusion. At each pixel location, PixExpo selects the observation closest to an rPPG-motivated target intensity. This criterion seeks to reduce local saturation and severe underexposure rather than optimize perceptual appearance. PixExpo requires no sensor modification but assumes programmable frame-level exposure control. We validate the proposed PixExpo framework using our newly introduced MEX-Drive dataset, comprising 48 participants under real-world driving conditions. Experimental results demonstrate that PixExpo outperforms manufacture-default auto-exposure methods, reducing the mean absolute error (MAE) by 7.21 bpm (from 13.94 to 6.73 bpm) and increasing the success rate by 37.29 percentage points (from 25.95% to 63.24%) across challenging driving scenarios.
|
| 45 |
CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation
2609.36616
|
cs.CV
|
Hanwen Lu, Jun He, Mingjia Yang, Hao Wei, Jinhao Huang |
Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR...Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12\% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at https://luhanwen67.github.io/CrossTimeEdit-release/.
|
| 46 |
Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning
2609.36628
|
cs.CV
|
Yanan Wang, Tingsong Li, Kaixun Jiang, Chongyang Zhong, Chenwei Xoe |
Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing ...Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and shot contexts. Our Graded Physical Alignment score (GPA) awards credit for correct content and deducts points for incorrect or fabricated actions, making these errors explicit in the aggregate score. Building on this rubric, we propose Graded Margin Direct Preference Optimization (GM-DPO), which assigns stronger preference margins and greater training weight to more severe action errors. Across three VLM backbones, GM-DPO achieves the highest substantive-action and GPA scores among the evaluated preference objectives, improving GPA over DPO by 2.02-3.40 points. On Qwen3-8B, it reduces the weighted hallucination rate by 21.3% relative to DPO. These gains accompany sustained long-form output, improved shot structure, and competitive performance on three additional multimodal benchmarks.
|
| 47 |
OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration
2609.36644
|
cs.CV
|
Pei An, Jiaqi Yang, Yulong Wang, Siwen Quan, Liangliang Nan |
Cross-attention is a crucial component in learning-based image-to-point-cloud (I2P) registration. Although existing cross-attention mechanisms have achieved promising progress, attention ambiguity remains a fundamental challenge that hinders the learning of di...Cross-attention is a crucial component in learning-based image-to-point-cloud (I2P) registration. Although existing cross-attention mechanisms have achieved promising progress, attention ambiguity remains a fundamental challenge that hinders the learning of discriminative 2D-3D correspondences. To address this problem, we revisit cross-attention and establish ordinary differential equations (ODEs) to model the ideal I2P feature interaction. Based on this formulation, we develop an ODE-driven cross-attention (OCA) module that refines feature representations and attention matrices through ODEs. In practice, OCA can be seamlessly integrated into existing I2P registration frameworks. To validate its effectiveness, we incorporate OCA into five state-of-the-art baselines and evaluate on four public benchmark datasets. Experimental results demonstrate that OCA improves registration recall by up to 5\%, 9\%, and 15\% under the standard, fine-tuning, and zero-shot settings, respectively.
|
| 48 |
VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training
2609.36648
|
cs.CV
|
Yuanwei Hu, Bo Peng, Yuheng Jia, Xinting Hu, Yadan Luo |
Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear ho...Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as existing studies generally suffer from major limitations, including inconsistent experimental settings, inadequate dataset selection, and limited evaluation dimensions. To address this gap, we introduce VLM4Cluster, a comprehensive benchmark for image clustering in the era of pre-trained vision-language models (VLMs). VLM4Cluster implements 17 representative methods spanning classical, deep, and language-assisted image clustering, and evaluates them on 20 datasets covering classical, challenging, fine-grained, large-scale, and out-of-distribution settings. Beyond effectiveness, VLM4Cluster systematically investigates image clustering along three complementary dimensions: robustness to adversarial perturbations, generalization under distribution shifts, and computational efficiency. Our study shows that LaIC substantially advances the clustering performance frontier on many semantically demanding benchmarks, generally exhibits stronger generalization under distribution shifts, and achieves a more favorable effectiveness-efficiency trade-off. However, its gains become less consistent on large-scale and fine-grained datasets, while language assistance does not systematically reduce sensitivity to adversarial perturbations. VLM4Cluster is released at https://github.com/YuanweiHuu/VLM4Cluster.
|
| 49 |
FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
2609.36651
|
cs.CVcs.AI
|
FangZhi Zhong, Xuerui Qiu, Yuqi Pan, Ya Liu, Shaowei Gu |
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI...Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
|
| 50 |
Not Every Correction Helps: Gain-Guided Continual Test-Time Adaptation
2609.36655
|
cs.CV
|
Youjia Zhang, Huiling Liu, Soyun Choi, Jaehong Yoon, Sungeun Hong |
Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model's self-cert...Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model's self-certainty for the current sample. In CTTA, accumulated target observations can provide complementary evidence for correcting the source prediction, but this history may become misaligned as the target distribution changes. The key question is therefore not how much the correction differs from the source prediction, but whether and how strongly it should be applied. This paper proposes Gain-Aware INtervention (GAIN), a backpropagation-free CTTA framework guided by a simple principle: history proposes, gain decides. GAIN maintains compact target statistics to form a correction proposal and a posterior-predictive evaluator that accounts for estimation uncertainty. The resulting source-relative gain estimates the proposal's benefit and determines a sample-specific intervention strength along a continuous path through efficient one-dimensional optimization. Gain-controlled predictions then update the target statistics online, limiting the propagation of unreliable corrections, all without backpropagation, sample storage, or replay. Across five benchmarks, our method achieves strong predictive performance, with favorable accuracy--calibration--efficiency trade-offs in continual adaptation. On ImageNet-C, for example, GAIN achieves 61.9% accuracy with near-source calibration. It remains stable under diverse and challenging continual shifts while running 15.9x faster than a representative optimization-based CTTA baseline.
|
| 51 |
You Only Reprogram Once: Rethinking Prolonged Training for Visual Reprogramming
2609.36661
|
cs.CV
|
Zizhao Li, Mohammed Yaqoob Ansari, Xinyu Su, Jiayang Ao, Joseph West |
Visual reprogramming is a parameter-efficient method for adapting pretrained models, yet its training can remain computationally expensive: even with a frozen backbone, visual prompts are often optimized through the full model for hundreds of epochs. Before ch...Visual reprogramming is a parameter-efficient method for adapting pretrained models, yet its training can remain computationally expensive: even with a frozen backbone, visual prompts are often optimized through the full model for hundreds of epochs. Before changing what the pretrained model sees, we ask whether we are fully using what it already tells us. We find that modeling the full source response can already yield strong downstream predictions without prompt optimization. Motivated by this observation, we introduce You Only Reprogram Once (YORO), which constructs a downstream predictor from the frozen response space in a single forward-only traversal. Its Bayesian Discriminant Mapping (BDM) derives a covariance-aware affine mapping from streaming class statistics, requiring no backpropagation, optimizer updates, or repeated visits to the training set. When further input adaptation helps, YORO-FP optionally refines the visual prompt for 20 epochs. BDM also extends naturally to CLIP by treating attribute-prompt similarities as source responses. Across three full-data settings, YORO improves average accuracy over the strongest prior gradient-free mapping by 18.4--24.4\%. On 16-shot CLIP, it raises the four-backbone average from 71.4\% to 77.2\%. YORO-FP provides further gains on selected tasks, while validation often retains the one-pass predictor. These results suggest a different default for visual reprogramming: read out the frozen response first, and optimize the input only when needed.
|
| 52 |
ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking
2609.36677
|
cs.CV
|
Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen, Zhenyang Xie |
Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and pr...Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and propagate identity errors across subsequent camera handoffs. We propose ReWorld-Track, a recursive event world model that carries association uncertainty into future predictions. Candidate matches and continued waiting define alternative target states, whose posterior probabilities are used to update a persistent recurrent belief. This representation preserves uncertainty about alternative trajectories through successive observations. This belief predicts the next camera, arrival time, and entry region, while appearance and language evidence guide association. By training across successive handoffs, the model learns to retain uncertainty that remains useful for later predictions and identity decisions. ReWorld-Track achieves HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, with improved identity continuity across repeated handoffs. On MTMMC, its structured posterior update gains 0.50 HOTA points over a similarly sized generic updater and 0.94 points over fixed-moment soft association, raising next-camera accuracy from 86.03% to 87.41% and reducing median arrival-time error from 0.78 s to 0.71 s for subsequent target returns.
|
| 53 |
Reprogramming Vision-Language Models via Structured Prompt Reparameterization
2609.36680
|
cs.CV
|
Zizhao Li, Chengyi Cai, Mohammed Yaqoob Ansari, Feng Liu, Joseph West |
Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly m...Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framework that aggregates multiple text prompts within each class and applies residual correction across classes. We also show that CLIP-based visual reprogramming with input-independent linear output aggregation can be expressed as a linear mapping from frozen image embeddings to downstream logits, and use this view to design a structured reparameterization that models shared semantic components and class-specific differences. RVP uses only a single visual prompt and can be reparameterized at inference into a frozen backbone followed by a linear classifier, incurring nearly zero computational overhead. Across 11 few-shot classification benchmarks and four CLIP backbones, RVP consistently improves over prior visual reprogramming methods with comparable or better inference efficiency.
|
| 54 |
When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation
2609.36685
|
cs.CV
|
Zhirui Xing, Long Ye, Kaige Li, Ziyi Xu, Ming Meng |
Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily...Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.
|
| 55 |
AESplat: Advancing Pose-Free Feed-Forward 3D Gaussian Splatting via Decoupled Appearance Modeling
2609.36693
|
cs.CV
|
Shiwei Ren, Zhiang Liu, Yongchun Fang, Hongwei Chen |
Pose-free feed-forward 3D Gaussian Splatting (3DGS) has demonstrated remarkable potential for generalized novel view synthesis. However, existing methods typically predict Gaussian appearance attributes represented by spherical harmonics (SH) in the same manne...Pose-free feed-forward 3D Gaussian Splatting (3DGS) has demonstrated remarkable potential for generalized novel view synthesis. However, existing methods typically predict Gaussian appearance attributes represented by spherical harmonics (SH) in the same manner, overlooking the fundamental distinction between view-independent and view-dependent appearance, which results in suboptimal rendering quality. In this paper, we present AESplat, a novel and general framework for pose-free feed-forward 3DGS that introduces an effective decoupled appearance modeling strategy based on an analysis of SH, enabling higher-quality rendering. Specifically, AESplat directly derives the zeroth-order SH coefficient, which represents the base view-independent appearance component, from the input images without training. The higher-order SH coefficients are subsequently predicted by a shallow multilayer perceptron equipped with two efficient 3D-aware inductive biases to model view-dependent appearance variations. Extensive experiments across multiple datasets demonstrate that our method significantly outperforms state-of-the-art approaches, achieving a $0.8$ dB improvement in PSNR over the pose-free method NAS3R and a $1.1$ dB improvement over the pose-required method DepthSplat on the RealEstate10K dataset. Project page: https://aesplat.github.io/.
|
| 56 |
Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing
2609.36755
|
cs.CV
|
Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou |
Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural...Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.
|
| 57 |
NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation
2609.36756
|
cs.CVcs.AI
|
Jiawei Zhang, Shuhao Liu, Rong Huang, Yuancheng Li, Zhihui Li |
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. How...One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256$\times$256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/jaiwei804/NesTok.
|
| 58 |
FastVR: Efficient Streaming Video Restoration with One-Step Diffusion
2609.36757
|
cs.CV
|
Xiaoxu Chen, Qin Yang, Haoran Bai, Sibin Deng, Ying Chen |
Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full self-attention in diffusion transformers (DiTs). This paper pr...Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full self-attention in diffusion transformers (DiTs). This paper presents FastVR, a streaming video restoration framework built on a one-step diffusion model, which delivers strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. To improve inference efficiency, FastVR combines a lightweight VAE with chunk-wise causal attention, which substantially reduces the computational cost. During training, it further adopts velocity consistency regularization and continuous trajectory learning, which improve restoration quality. Extensive experiments show that FastVR is more efficient than the evaluated diffusion baselines while achieving state-of-the-art performance on synthetic and real-world benchmarks. We hope that this work supports further progress in the community.
|
| 59 |
Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning
2609.36759
|
cs.CVcs.AI
|
Chiyuan He, Zihuan Qiu, Fanman Meng, Chao Wang, Liangjiang Chen |
Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate ne...Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.
|
| 60 |
DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning
2609.36775
|
cs.CV
|
Cheng Ye, Weidong Chen, Bingyan Xu, Zhendong Mao |
Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wi...Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8\% on average cross-domain accuracy than EMO-R3.
|
| 61 |
Causal-EVC: Breaking Emotional Spurious Causality via Spatiotemporal Grounding and Counterfactual Intervention
2609.36776
|
cs.CV
|
Cheng Ye, Weidong Chen, Peipei Song, Zhendong Mao |
Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the importance of visual causes to guide emotion perception and caption generation, they fundamentally rely on simple a...Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the importance of visual causes to guide emotion perception and caption generation, they fundamentally rely on simple attention matching, which inevitably suffers from {causal redundancy and spurious correlations} in co-occurrence bias (e.g., misclassifying ``sadness'' as ``joy'' on a sunny beach), leading to severe shortcut learning from confusing backgrounds. Furthermore, existing evaluations fail to verify whether models have genuinely mastered causal reasoning or merely exploited background confounders. To address these limitations, we first construct {EVC-CauseGround}, a comprehensive benchmark with dense spatio-temporal causal annotations. Crucially, it introduces a carefully selected {Causal-Faithfulness Subset} to explicitly quantify genuine emotion-cause attribution. Second, we propose {Causal-EVC}, an emotion-grounding captioning framework, which introduces a Motion-guided Causal Spatiotemporal Localization module to precisely decouple causal triggers from background confounders. Besides, we introduce an Interpretable Sparse Emotion Routing module. By synthesizing counterfactual representations and formulating a novel counterfactual contrastive objective, we enforce the model to anchor its emotion predictions strictly on authentic causal triggers instead of confusing background. Extensive experiments show that Causal-EVC not only achieves the best performance on semantic metrics but also exhibits significant advantages in the causal-faithfulness subset, which demonstrates that our model could mine emotional cues from genuine visual causes and mitigate co-occurrence bias for interpretable multimodal emotion understanding.
|
| 62 |
Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning
2609.36782
|
cs.CV
|
Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao |
While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap...While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly encoded, thereby generating emotional misjudgments in complex scenarios. 2) Insufficient Discrimination. Existing methods could only identify regions generally associated with emotions, which fails to distinguish discriminative regions between semantically similar emotions, leading to ambiguous emotion judgements. To overcome these limitations, we present a training-free inference-time optimization framework, named Decoding Affective Nuances (DAN). Specifically, we propose a Hierarchical Emotional Reasoning Chain (HERC) that enhances the insufficient attribution by harmonizing fine-grained scene/object-level cues and performing a soft-gated reasoning. Furthermore, to discriminate between semantically proximal emotions, we design a Contrastive Discriminative Visual Pruning (CDVP), which isolates discriminative visual tokens to reason the final emotion category by computing the absolute discrepancy between the attention distributions of similar emotions. Performances on several benchmarks demonstrate that DAN significantly improves discrimination for affective nuances without consuming additional training resources, especially achieving +10.47% improvements with Qwen3-VL-8B-Instruct on WebEmo25 dataset that contains 25 fine-grained emotion categories.
|
| 63 |
Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs
2609.36798
|
cs.CV
|
Yueran Ma, Ronghao Lin |
Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sampl...Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at https://anonymous.4open.science/r/DMC-Repair.
|
| 64 |
Scene Retargeting: Learning Object Placement with Analogical Transfer
2609.36801
|
cs.CV
|
Minkwan Kim, Junho Kim, Seungmin Lee, Changwoon Choi, Young Min Kim |
Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, irregular layout structures impose scene-specific physical constraints, making it hard to define a generalizab...Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, irregular layout structures impose scene-specific physical constraints, making it hard to define a generalizable framework for generating similar functional context. We formalize Scene Retargeting as stably transferring the semantically coherent spatial organization across layouts, rather than relying on textual descriptions or pairwise relationships. Our cluster-wise transfer flexibly handles mismatched object instances and adapts to distinctive floor plans. We optimize to preserve the rich semantic context of individual clusters by respecting the spatial distribution of foundation features. We can then impose physical constraints to refine wall contacts, pairwise alignment, or clear passageways and openings. Our framework outperforms state-of-the-art methods on layout generation on the 3D-FRONT dataset, and demonstrates downstream applications including real-to-sim transfer, analogical trajectory transfer, and multi-reference composition.
|
| 65 |
EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding
2609.36803
|
cs.CV
|
Yuwei Miao, Xuesheng Zhang, Wenhao Zou, Jixia Zhang, Jianwei Lv |
Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher rece...Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.
|
| 66 |
MeteoVerse: Unified Weather-Controllable Video World Model
2609.36810
|
cs.CV
|
Renlong Wu, Guanqiao Wang, Xuan Shang, Yin Hanming, Xiaoxiao Sheng |
Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determined not only by changes in viewpoint and object dynamics, but also by environmental conditions such as weathe...Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determined not only by changes in viewpoint and object dynamics, but also by environmental conditions such as weather, which can substantially alter scene appearance and visibility. Modeling such realistic weather evolution is challenging because the required weather modification depends jointly on the observed and desired weather states. Depending on their relation, the model may need to preserve, introduce, or remove a weather effect. Existing video world models typically leave this weather transition implicit, forcing the generation backbone to infer weather evolution together with scene dynamics and camera motion, which leads to imprecise weather control. To address this limitation, we propose MeteoVerse, a unified weather-controllable video world model that generates future videos from a single sunny or adverse-weather image, conditioned on a weather-free scene description, a target-weather instruction, and a camera trajectory. Rather than conditioning only on the desired weather, MeteoVerse explicitly estimates the observed and target weather states and represents the required weather transition. A transition-aware mixture of weather experts then translates this transition into category-specific residual weather features, unifying weather preservation, introduction, and removal while enabling fine-grained control over introduced weather intensity. We further construct the MeteoVerse dataset with over 50K real-world weather video clips, generated sunny counterparts, disentangled scene and weather descriptions, weather-intensity annotations, and camera trajectories. Extensive experiments demonstrate substantially improved weather controllability while retaining competitive scene consistency and camera-control performance.
|
| 67 |
CurvSpec: Adaptive Multi-Curvature Learning for Partial Relevant Video Retrieval
2609.36815
|
cs.CV
|
Zhen Liu, Letian Li, Jinpeng Wang, Shuzhao Xie, Yuzhi Huang |
Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely ...Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely low signal-to-noise ratio that makes PRVR more challenging than standard full-video retrieval. This task presents two intertwined challenges: (1) signal dilution, where coarse global representations blur the brief relevant signal into the dominant irrelevant surroundings;(2) curvature rigidity, where embedding all videos in the same fixed-geometry space distorts representations for videos that range from flat atomic events to deep compositional hierarchies. Existing PRVR methods have improved moment selection and cross-modal matching, but they still typically encode all videos in a single fixed-curvature retrieval space, limiting their ability to model diverse video structures. To address both challenges, we propose CurvSpec, a framework that learns content-adaptive curvature for video retrieval representations rather than imposing a fixed geometric prior. CurvSpec processes features through parallel Euclidean and hyperbolic attention layers, with independently learned curvatures assigned to the hyperbolic layers, and a content-aware fusion mechanism routes each input to its most suitable geometric regime. To further suppress signal dilution, CurvSpec represents each video with semantic centroids whose number is determined by the video's content complexity, projects them onto the learned manifold, and matches each query against its nearest centroid by geodesic distance. Experiments on ActivityNet Captions, TVR, and Charades-STA demonstrate state-of-the-art retrieval performance.
|
| 68 |
RED: Reconstruction Evolution Dynamics for Generalizable AI-Generated Image Detection
2609.36822
|
cs.CV
|
Wenpeng Mu, Junshan Jin, Tanfeng Sun, Xinghao Jiang, Qiang Xu |
The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate re...The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate reconstruction stages underexplored. We observe that the relative token predictability of real and generated images can reverse across reconstruction scales, suggesting that intermediate stages may expose forensic evidence overlooked by endpoint comparisons. Motivated by this observation, we propose RED (Reconstruction Evolution Dynamics), a framework that captures transferable forensic cues from coarse-to-fine reconstruction evolution. To our knowledge, RED is the first framework to use scale-wise token predictability to guide forensic evidence aggregation across intermediate reconstruction states. It represents the reconstruction trajectory produced by a frozen multiscale VQ-VAE in the shared feature space of a frozen CLIP encoder. To connect the observed predictability variations with visual evidence, RED learns image-adaptive stage weights from scale-wise token negative log-likelihoods provided by a frozen VAR model. A cross-stage evidence aggregation module then jointly models the original-image representation and the weighted reconstruction features, capturing complementary forensic cues through interactions along the reconstruction trajectory. Experiments on six diverse benchmarks demonstrate that RED achieves the highest average accuracy of 92.5\% and average precision of 97.5\% among the evaluated methods. Further evaluations show strong robustness to common image degradations, supporting the value of reconstruction evolution for generalizable AI-generated image detection. The code will be made publicly available upon acceptance of this paper.
|
| 69 |
Learning via Self-Consistency for Diffusion-based Video Reasoning
2609.36826
|
cs.CV
|
Zhenghao Ni, Weimin Qiu, Meng Tang |
Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motiva...Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.
|
| 70 |
Motion Concept Unlearning in Video Diffusion Models
2609.36832
|
cs.CV
|
Ping Liu, Chi Zhang |
Text-to-video (T2V) diffusion models can generate realistic depictions of actions such as kicking, stabbing, and shooting, raising safety concerns that motivate targeted concept erasure. Although concept erasure has been extensively studied for static concepts...Text-to-video (T2V) diffusion models can generate realistic depictions of actions such as kicking, stabbing, and shooting, raising safety concerns that motivate targeted concept erasure. Although concept erasure has been extensively studied for static concepts in text-to-image and T2V models, erasing motion concepts remains largely unexplored. We present a systematic study of motion concept erasure in video Diffusion Transformers (DiTs). Through causal interventions, we show that text-conditioning attention carries concept-specific motion information and supports selective intervention, whereas perturbing temporal positional encoding suppresses both target and non-target dynamics. We further find that directly adapting ESD, a representative weight-level image erasure method, to a video DiT yields modest and uneven motion suppression: reducing its erasure training loss does not by itself remove the concept signal from the difference between the conditional and unconditional predictions, which classifier-free guidance (CFG) then scales at every denoising step. From these findings, we derive three requirements for motion concept erasure: concept specificity, spatial selectivity, and temporal naturalness. Each determines one component of MUTE (Motion concept Unlearning in Text-to-video gEneration): at each denoising step, MUTE extracts a concept direction through token neutralization, derives a spatial gate from the direction's intrinsic structure, and subtracts the resulting correction from the velocity output before CFG is applied. MUTE is training-free and requires no weight modification. Experiments on 20 motion concepts show that MUTE outperforms representative prompt-level, weight-level, and inference-time baselines on Wan2.1-T2V, and the same formulation transfers to CogVideoX, supporting its applicability across distinct T2V attention architectures.
|
| 71 |
You Cannot Recover What Was Never Measured: Quantifying the Information Ceiling of Ultra-Low-Field MRI Super-Resolution
2609.36837
|
cs.CV
|
Prathamesh Pradeep Khole, Shreya Handa, Utkarsh Gupta, Razvan Marinescu |
Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work ackn...Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work acknowledges that these models hallucinate and that the problem is ill posed, but to our knowledge no study measures how much information about the individual subject the real low-field scan actually contains. We measure it. Using paired 64 mT and 3T scans of the same subjects from three public datasets, and a measurement protocol validated on tests whose correct answer is known in advance, we find that, judged over the whole brain, real 64 mT scans carry structure specific to the individual only down to approximately 3 to 4 mm half-pitch in plane, and coarser still through plane. Standard synthetic degradations preserve subject information roughly 1 mm beyond this ceiling, so models trained and benchmarked on synthetic pairs are evaluated on information that real scanners never record. We then test trained diffusion models and a publicly released external model on real paired acquisitions; 24 trained runs of five architectures (GAN, diffusion, and transformer families) give the coverage of the audit. On every subject where faithfulness can be measured, fine output detail is no more correlated with the subject's own 3T scan than with a stranger's, while sample-variance uncertainty does not distinguish fabricated structure from reconstruction difficulty. Because PSNR and SSIM score resemblance to a reference rather than whether detail belongs to the subject, a benchmark scored by them cannot tell recovery from fabrication. Code for the measurement protocol will be released so that recoverability claims can be tested for newer models.
|
| 72 |
On-Policy Visual Evidence Distillation
2609.36838
|
cs.CVcs.AI
|
Shaohang Wei, Feifan Song, Guangyue Peng, Wenhao Yu, Wei Li |
Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for su...Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE
|
| 73 |
Does the VGGT Family Need All Its Layers?
2609.36842
|
cs.CV
|
Fengyi Zhang, Holger Caesar, Xiangyu Sun, Zheng Zhang, Zi Huang |
Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, $\pi^3$, and VGGT-$\Omega$: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics ac...Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, $\pi^3$, and VGGT-$\Omega$: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Removable layers cluster in two redundancy regions: a dominant early region and a narrower late one, while deletions spanning the intervening layers are consistently more disruptive. This recurring pattern holds across models, datasets, and metrics, and contrasts with the middle-to-late redundancy commonly reported in the literature. (ii) Within these regions, we observe that the joint degradation from deleting two intervals is approximately the sum of their individual degradations, reducing the number of model evaluations for pruning search from $O(L^4)$ to $O(L^2)$, where $L$ is the aggregator depth. (iii) We find that CKA provides a cheaper representation-based proxy for interval degradation, offering a practical trade-off between pruning quality and calibration cost. (iv) Closed-form linear calibration recovers accuracy after pruning without end-to-end retraining. A least-squares analysis shows that using a shared map for special and patch tokens generally incurs excess reconstruction loss, motivating token-aware recovery. Recovery maps fitted on just 100 calibration scenes generalize to held-out scenes and unseen datasets. The resulting models reduce aggregator parameters by up to 44% while maintaining accuracy comparable to their intact counterparts. Code and experimental results will be available at our project page: https://xian-bei.github.io/vggt-family-layer-redundancy/
|
| 74 |
GlassFormer: Learning Real-time Glass Segmentation using Radar-Depth Fusion
2609.36844
|
cs.CV
|
Suhani Grover, Astik Srivastava, Viswas Dinesh, Avinash Sharma, K. Madhava Krishna |
Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the background behind glass rather than the surface itself, while depth sensors such as LiDAR, time-of-flight, and...Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the background behind glass rather than the surface itself, while depth sensors such as LiDAR, time-of-flight, and RGB-D often return invalid or background measurements in transparent regions. As a result, systems that rely solely on optical sensing may misinterpret glass walls, doors, or mirrors as free space, compromising safe and reliable navigation. Existing glass segmentation approaches address this by learning visual cues such as reflections, boundaries, and semantic context from RGB images. While effective under favourable lighting and viewing conditions, these cues degrade in low-light environments, under glare, or when glass surfaces are featureless or partially occluded. In this work, we propose a multimodal framework that fuses millimetre-wave radar with RGB-D sensing for real-time transparent surface segmentation. Radar reflects strongly off glass surfaces, providing a geometric cue that remains reliable precisely where vision and depth fail. We exploit this cross-modal inconsistency to generate a radar-guided spatial prior, which is integrated into a lightweight transformer-based segmentation network, GlassFormer, via cross-modal attention. We report results on a mixed-condition test split covering all scene types and a dedicated low-light split designed to stress vision-only methods. GlassFormer achieves 0.88 mIoU on the mixed split, and 0.59 mIoU on the low light split, demonstrating substantial robustness gains over vision-only baselines while maintaining real-time performance on resource-constrained platforms.
|
| 75 |
RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts
2609.36851
|
cs.CV
|
Hongbin Lin, Chaoda Zheng, Yiming Yang, Xiangyu Li, Shijia Chen |
End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinfo...End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.
|
| 76 |
Socialality Anchors: Towards Group-bounded Trajectory Prediction
2609.36852
|
cs.CV
|
Ziqian Zou, Conghao Wong, Qinmu Peng, Xinge You |
Trajectory prediction is a key component for understanding human behavior patterns in dynamic scenes. Researchers have devoted substantial efforts to modeling social interactions, especially group-wise interactions, since group membership often reflects shared...Trajectory prediction is a key component for understanding human behavior patterns in dynamic scenes. Researchers have devoted substantial efforts to modeling social interactions, especially group-wise interactions, since group membership often reflects shared intention, coordinated motion, and stable mutual adaptation, thus providing a persistent and semantically meaningful social prior for forecasting. However, existing group modeling methods may rely on a fixed threshold and infer groups mainly from agents' relative positions within the observation window, overlooking the fact that grouping rules should be agent-specific, temporally coherent, and context-adaptive across diverse personalities, culturalities, and evolving interaction contexts. Inspired by human social perception that alternates between interpersonal distance in boundary-sensitive situations and relative speed consistency in dynamic interactions, we propose Socialality, a human-inspired trajectory prediction framework with interpretable Socialality anchors and an extended grouping window for stable, context-aware grouping inference. Concretely, Socialality introduces a duo-scalar-controlled grouping kernel Socialality that jointly leverages historical observations and short-term future trajectory previews to learn agent-specific grouping rules, and employs a group-wise perception mechanism to model in-group and out-of-group interactions in an intuitive and explainable manner. Furthermore, we conduct extensive experiments on standard benchmarks to demonstrate the performance gains of Socialality, and provide qualitative analyses and statistical studies of anchor distributions to verify the interpretability and stability of the proposed Socialality anchors.
|
| 77 |
S2T-Unet: A Structure-to-Style Framework for Inter-Modality MRI Translation
2609.36866
|
cs.CV
|
Yichao Liu |
Inter-modality MRI translation aims to synthesize missing MRI modalities from available acquisitions, reducing the need for additional scanning while preserving clinically relevant anatomical information. However, existing image translation methods often learn...Inter-modality MRI translation aims to synthesize missing MRI modalities from available acquisitions, reducing the need for additional scanning while preserving clinically relevant anatomical information. However, existing image translation methods often learn intensity mappings without explicitly separating modality-invariant structural information from modality-specific appearance, which may lead to structural information loss or unrealistic image details. In this work, we propose S2T-Unet, a structure-to-style framework that explicitly models these two aspects. Specifically, vector quantization is introduced at the lower-level bottleneck to encode modality-invariant structural information using a learned discrete codebook. At higher levels, a modality transformation module uses decoder features to condition and transform encoder representations toward the target modality, thereby recovering modality-specific intensity and contrast information. Experiments on the IXI multi-contrast MRI dataset across four translation tasks demonstrate that S2T-Unet is comparable or outperform with state-of-art method.
|
| 78 |
S4VY: Segment Anything in Feed-Forward 4D Visual Geometry
2609.36875
|
cs.CV
|
Jingdong Zhang, Xin Li, Jan Kautz, Wenping Wang, Chris Choy |
Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, wh...Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.
|
| 79 |
Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images
2609.36882
|
cs.CV
|
Junhee Lee, Donghyeon Jeon, Taeoh Kim, Beomyoung Kim, MyeongAh Cho |
Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifacts that are misaligned with semantic or object boundaries. Existing approaches rely on pixel-level supervision...Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifacts that are misaligned with semantic or object boundaries. Existing approaches rely on pixel-level supervision from controlled editing pipelines, which is difficult to scale and can introduce misleading signals: artifacts frequently extend beyond annotated regions, while out-of-mask pixels are treated as authentic. This limits models' ability to capture transferable evidence and generalize across generators and datasets. To address these issues, we propose ReGFLoW, a Reconstruction-Guided Fake Localization framework under Weak supervision, which is the first weakly supervised approach for diffusion-edited fake region localization. ReGFLoW requires only real/fake labels at the image-level and uses diffusion reconstruction errors as dense spatial guidance to inject them into both feature and score spaces. Furthermore, by artifact-centric multiple instance learning, ReGFLoW utilizes localized diffusion evidence without relying on semantic-affinity or boundary-based pseudo-mask priors. Extensive experiments demonstrate that ReGFLoW achieves stronger out-of-domain generalization than fully supervised learning baselines.
|
| 80 |
ProGuT: Label-Efficient Panoptic Segmentation for Forest Scenes
2609.36891
|
cs.CV
|
Pankaj Deoli, Karsten Berns |
Panoptic segmentation in forest environments is bottlenecked not by semantic quality but by instance separation; existing unsupervised panoptic approaches produce usable stuff maps but near-zero thing quality. Depth or flow-based instance discovery methods nee...Panoptic segmentation in forest environments is bottlenecked not by semantic quality but by instance separation; existing unsupervised panoptic approaches produce usable stuff maps but near-zero thing quality. Depth or flow-based instance discovery methods needs sensors that are not always available. We present ProGuT (Prototype Guided Training), which produces panoptic pseudo-labels without per-image training masks, needing only unlabeled images and one-time cluster-to-class mapping. ProGuT clusters CLIP patch features, then recovers trunk instances through multiscale geometric prior that falsifies non-trunk structures via structure-tensor. This is cheap compared to depth, flow or class-supervision methods to create pseudo labels. These are then used for downstream tasks which we evaluate against other unsupervised baselines. ProGuT achieves a Panoptic Quality (PQ) of 65.2 on Our-forest dataset (2.6x improvement over the initial pseudo-label quality) and reaches 65.9 mIoU on Freiburg Forest, outperforming unsupervised baselines like PiCIE (45.3 IoU) and STEGO(57.6IoU). Additionally, ProGuT outperforms existing unsupervised methods for class-agnostic trunk instance benchmark.
|
| 81 |
DiffReID: Discriminative Diffusion Model for Object Re-Identification
2609.36894
|
cs.CV
|
Yingquan Wang, Pingping Zhang, Dong Wang, Huchuan Lu |
As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancements have been made in object ReID. However, most existing met...As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancements have been made in object ReID. However, most existing methods suffer from generalization due to the limited size and diversity of ReID datasets. Meanwhile, current models tend to focus on extracting semantic patterns rather than learning identity-aware feature distributions. To address these issues, we propose a novel feature learning framework named \textbf{DiffReID} for object ReID. It leverages a discriminative diffusion model to gradually learn identity-aware distributions and generate identity-invariant features. More specifically, with the Contrastive Language-Image Pre-training (CLIP) model, we first obtain identity-aware text features by prompt tuning. Then, we propose a Vision-guided Noise Generator (VNG) to initialize probabilistic noises and gradually corrupt identity-aware text features. Afterwards, we take visual features as conditions and propose a Light Weight Denoiser (LWD) to denoise the corrupted text features step-by-step for identity-aware distribution learning. To obtain discriminative features, we further generate identity-invariant guided features from randomly sampling visual-guided noises. Finally, we propose a Mutual Enhancement Constraint (MEC) to facilitate mutual learning between visual features and guided features to enhance the representation robustness and discrimination. Extensive experiments on five object ReID benchmarks demonstrate that our method shows better results than most state-of-the-art methods. The source code is available at https://github.com/AWangYQ/DiffReID.
|
| 82 |
SafeVantage: Vantage-Aware Memory for Reliable Embodied Decisions
2609.36906
|
cs.CV
|
Sean Hardesty Lewis, Zuyi Guo, Benwang Chen, Zirui Liu, Hongyi Lin |
Reliable embodied decisions under partial observability require informative observations and sufficient supporting evidence. However, semantic scores alone do not reveal which viewpoints justify a claim or where additional evidence should be acquired. We intro...Reliable embodied decisions under partial observability require informative observations and sufficient supporting evidence. However, semantic scores alone do not reveal which viewpoints justify a claim or where additional evidence should be acquired. We introduce SafeVantage, a vantage-aware semantic memory and active acquisition framework that retains each claim's supporting views, camera poses, and estimated target location, keeping positive support distinct from search coverage. A learned candidate-observability model uses claim-grounded geometry to predict target visibility at reachable viewpoints. These predictions guide view selection through expected reduction in terminal decision loss, accounting for travel cost and geometrically distinct corroboration. A calibrated head then combines support, spatial consistency, and coverage to produce Yes, No, or Abstain decisions. We evaluate SafeVantage on a category-presence benchmark spanning 232 unseen ProcTHOR houses and 7,424 paired episodes per method and action budget. Compared with validation-selected equal-budget baselines, SafeVantage achieves macro-F1 gains of 24.7% and 12.0% at eight and twelve actions, respectively, with lower risk and higher answer rates at both budgets and 31.7% less travel at eight actions. Equal-input HM3D experiments show lower selective risk under fixed observations, while controlled ScanNet interventions show that restoring supporting views improves downstream VLM answers. Ablations further support the contribution of candidate observability to decision quality and acquisition efficiency. Results demonstrate the value of claim-level viewpoint evidence for connecting semantic memory, active acquisition, and reliable decision-making. Code is available at https://safevantage.github.io
|
| 83 |
Representation Dynamics Reveal Semantic Saliency and Similarity for Visual Token Pruning in MLLMs
2609.36916
|
cs.CV
|
Weixuan Li, Zikun Zhou, Xinyi Zhuang, Xinyan Guo, Rui Tian |
Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit r...Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflect foreground saliency and semantic consistency remain insufficiently understood. We analyze visual token representation dynamics across encoder depth and uncover two findings. First, the relationship between token update magnitudes and foreground saliency is layer-dependent: large token updates concentrate on foreground regions in two depth intervals, separated by several sink-dominated layers at intermediate depths. Second, similarities between token update directions better distinguish same-class from different-class tokens than those between encoder output features. Building on these findings, we propose MSDG-Prune, a training-free method that uses update magnitudes and directions to preserve salient and diverse visual information. Specifically, we group tokens by update-direction similarity and use query-weighted saliency derived from update magnitudes across a chosen depth window for group-wise token pruning. Extensive experiments across four MLLMs demonstrate the effectiveness and generalizability of MSDG-Prune. On LLaVA-NeXT, it retains 91.9% of uncompressed performance on average with only 5.6% of visual tokens, while achieving a 7.8x prefilling speedup. Code is available at https://github.com/liweixuan-hitsz/MSDG-Prune.
|
| 84 |
Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation
2609.36918
|
cs.CV
|
Jiantao Lin, Meixi Chen, Yingjie Xu, Chenbo Fu, Leyi Wu |
Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due to occlusion, ambiguous boundaries, and the need for coherent multi-part reasoning. Existing approaches strug...Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due to occlusion, ambiguous boundaries, and the need for coherent multi-part reasoning. Existing approaches struggle to achieve both controllable part-level generation and coherent multi-part structure, as part identity and spatial allocation are typically inferred implicitly. We present Seg3DParts, a segmentation-grounded framework for controllable part-level 3D generation from a single image. By treating segmentation as an explicit grounding signal, our method defines part identity during generation, enabling each component to be anchored to a corresponding image region. To ensure coherent assemblies, we introduce structured cross-part interaction that allows components to exchange global context throughout the generative process. As a result, Seg3DParts directly generates well-aligned part meshes in a shared canonical space without post-hoc alignment, supporting flexible and controllable decomposition. We further introduce PartObjectNet, a large-scale dataset with over 200K objects and 1M annotated parts. Experiments demonstrate that Seg3DParts achieves superior geometry quality, cross-part coherence, and part-level controllability over existing methods.
|
| 85 |
SFE-VGGT: Source-Free VGGT Distillation for Event-Based Monocular Depth Estimation
2609.36929
|
cs.CV
|
Thai Duy Nguyen, Addison Lin Wang |
Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However, their reliance on synchronized RGB-event pairs or depth annotations during training severely restricts pract...Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However, their reliance on synchronized RGB-event pairs or depth annotations during training severely restricts practical deployment. To overcome this bottleneck, we propose SFE-VGGT, a novel source-free framework that distills the geometric priors of VGGT to the event domain without any paired RGB observations. Our core idea is to reconstruct surrogate frames directly from the target event stream to act as a frozen geometric teacher, entirely eliminating the need for genuine source RGB data. Crucially, as these surrogate frames inherently yield imperfect and spatially varying supervision, directly distilling from them propagates artifacts. To resolve this, we introduce a novel reliability-aware distillation strategy. This includes Density-Aware Feature Distillation to emphasize informative event regions, and Confidence-Weighted Depth Distillation to dynamically regulate supervision based on relative teacher-student prediction confidence. Meanwhile, we propose a Cross-Frame Relational Consistency loss that enforces temporal geometric stability using reliable inter-frame correspondences, bypassing the need for temporally consistent teacher's depth. Extensive experiments demonstrate that, despite source-free, our SFE-VGGT closely matches the accuracy of RGB-dependent baselines under standard conditions and significantly surpasses them in challenging nighttime scenarios. Across MVSEC nighttime sequences, SFE-VGGT reduces the average 10 m depth error by 15.3% compared with EventVGGT. Moreover, our method exhibits robust zero-shot generalization across real-world datasets, proving that highly effective geometric priors can be transferred to event cameras using strictly source-free supervision.
|
| 86 |
WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation
2609.36937
|
cs.CVcs.AI
|
Sangeyl Lee, Seunghyun Shin, Seungho Park, Wooseok Jeon, Hae-Gon Jeon |
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approach...Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extraction at inference. We further introduce Reference Asymmetric RoPE Conditioning to preserve fine-grained appearance details, and Identity Binding Supervision to associate each reference identity with its intended motion trajectory. To support cross-identity training, we construct MotionTwin, a large-scale synthetic dataset comprising 14.4K cross-identity video pairs with shared subject and camera motions, totaling 84.3 hours of photorealistic video. We additionally present MotionTwin-Bench, a cross-identity benchmark specifically designed to evaluate subject-level visual fidelity and identity-motion binding. Extensive experiments on MotionTwin-Bench and real-world videos demonstrate that WeLike2Party outperforms recent state-of-the-art methods in subject-level visual fidelity, identity-motion binding, and overall perceptual quality, particularly in multi-person interactions with substantial occlusion.
|
| 87 |
DispFlow-GS: Displacement Flow Supervision with Motion Disentangling for Monocular Deformable 3D Gaussian Splatting
2609.36940
|
cs.CV
|
Thai Duy Nguyen, Haitian Zhang, Addison Lin Wang |
Accurate dynamic scene reconstruction is important for robotic perception, where temporally consistent representations of dynamic environments are essential. Deformable 3D Gaussian Splatting (3DGS) models dynamic scenes through deformation fields, and recent m...Accurate dynamic scene reconstruction is important for robotic perception, where temporally consistent representations of dynamic environments are essential. Deformable 3D Gaussian Splatting (3DGS) models dynamic scenes through deformation fields, and recent methods incorporate motion supervision by aligning rendered Gaussian flow with optical flow. However, we find that such Gaussian-flow-based supervision provides only limited improvements in motion modeling. We identify a fundamental limitation of this supervision paradigm, namely a domain gap between rendered Gaussian flow and optical flow. To address this limitation, we propose a motion supervision framework built on Displacement Flow, which splats per-Gaussian 3D displacements onto the image plane to provide direct and stable optimization signals. We further disentangle scene motion from camera motion via intermediate-view rendering, enabling more reliable motion priors and targeted constraints on deformation and geometry. We also observe a discrepancy between motion fidelity and image-based evaluation, where improved motion awareness does not necessarily translate into better rendered image quality or higher image-based metric scores. Motivated by this mismatch, we introduce Deformation-Rendering Consistency (DRC), a motion-aware metric that measures the alignment between predicted deformation and rendering improvement. Experiments on dynamic scene benchmarks show substantial improvements in motion localization and motion--rendering consistency, reaching up to 39% and 6%, respectively, while image-based metrics change by only about 0.1%. These results confirm the observed mismatch between motion fidelity and image-based evaluation, demonstrating the significance of DRC for motion-aware evaluation.
|
| 88 |
Beyond Readability: Evaluating Task Information Recoverability
2609.36957
|
cs.CV
|
Yiwei Liu |
Direct visual readability and task-information recoverability are different quantities. Failure to decode a target from a fixed observation need not eliminate access to that target through another recovery route. We develop an evaluation perspective that makes...Direct visual readability and task-information recoverability are different quantities. Failure to decode a target from a fixed observation need not eliminate access to that target through another recovery route. We develop an evaluation perspective that makes the observation, query, target, and available knowledge explicit and measures the overlap between routes' success sets. For information available on the original visible surface under suitable imaging conditions, direct optical recovery reads the target from the image, optionally after restoration; entity-linked recovery uses residual visual evidence to identify the depicted entity and accesses its target through an entity--attribute relation in a specified knowledge resource. Such access can draw on stored knowledge or an external source. A controlled book-cover study instantiates external access with a fixed title--author catalog, comparing optical author recovery with visual title resolution and deterministic lookup under resolution degradation. Entity-linked successes persist across the tested vision--language models, revealing information access beyond the tested direct visual frontier despite substantial differences in absolute performance. A substantial optical-only region remains. These complementary outcomes show why visual degradation should be evaluated through the task information accessible along specified routes and knowledge resources, alongside direct readability.
|
| 89 |
Prior-Driven Enhancements in 3D Gaussian Splatting: Normals and Depths Regularization
2609.36969
|
cs.CV
|
Gyeonggwan Lee, Seunghwan Hong, Junghun Suh |
3D Gaussian Splatting (3DGS) is a state-of-the-art technique for 3D scene rendering, offering high efficiency and excellent visual quality. However, because 3DGS relies on an initial sparse point set from Structure-from-Motion (SfM) and view-dependent properti...3D Gaussian Splatting (3DGS) is a state-of-the-art technique for 3D scene rendering, offering high efficiency and excellent visual quality. However, because 3DGS relies on an initial sparse point set from Structure-from-Motion (SfM) and view-dependent properties, it can suffer from geometric inaccuracies and visual artifacts, particularly in complex scenes. To address these challenges, we propose an improved 3DGS approach that regularizes the optimization process by integrating geometric priors, including surface normals and dense depth information. Surface normal regularization improves geometric consistency by aligning Gaussian covariance with local surface structures, while dense depth priors combined with an initial points from SfM enhance per-pixel depth estimation, increasing accuracy and reducing ambiguities. These enhancements enable robust handling of diverse and complex real-world scenarios, minimizing visual distortions and improving reconstruction quality across various environments. To validate our method, we evaluate it on challenging datasets, including street-view scenes and highly reflective environments, while testing it across multiple SfM pipelines. Our results demonstrate compatibility across diverse environments and highlight the robustness of our approach. Experimental findings further show that our method enhances geometric accuracy and visual quality, establishing a reliable solution for real-time 3D scene rendering in complex environments.
|
| 90 |
Structured Visual Target Learning For Cross-Subject eeg-to-image retrieval
2609.36971
|
cs.CV
|
Salini Yadav, Taveena Lotey, Micka\"el Coustaty, Pravendra Singh, Partha Pratim Roy |
Cross-subject EEG-to-image retrieval requires a neural represen- tation trained on source subjects to remain aligned with a visual embedding space for an unseen subject. Whereas existing methods primarily focus on the EEG side, we address this problem from the...Cross-subject EEG-to-image retrieval requires a neural represen- tation trained on source subjects to remain aligned with a visual embedding space for an unseen subject. Whereas existing methods primarily focus on the EEG side, we address this problem from the perspective of the visual target. Our approach preserves the spatial information of the Perception Encoder, converts its patch grid into a compact set of learned visual views, and aggregates them for each image with a block-structured, content-dependent router. The target is learned jointly with the EEG encoder through contrastive learning with MMD regularization across source subjects. For deployment, we propose a training-free representation refinement that aligns frozen embeddings without updating either encoder. Under leave- one-subject-out evaluation on THINGS-EEG2, the structured target achieves 35.3%/65.6% Top-1/Top-5 accuracy, the best among com- pared methods. Refinement raises this to 48.1%/77.1%, an 18.5% Top-1 gain over the strongest compared method, improving all ten held-out subjects.
|
| 91 |
A Dual-Track Curation-and-Classification Framework for Resolving Ground-Truth Label Noise in Operational Sentinel-2 Wheat Area Estimation
2609.36975
|
cs.CV
|
Kasimali Agharia, Ujjwal Kumar Gupta |
Operational estimation of wheat-cultivated area is persistently constrained by discordance between administrative record-keeping and remotely sensed classification products. We address this administrative reference discordance for the 2022 Rabi season in Patia...Operational estimation of wheat-cultivated area is persistently constrained by discordance between administrative record-keeping and remotely sensed classification products. We address this administrative reference discordance for the 2022 Rabi season in Patiala district, Punjab, India, using a thirteen-timestep Sentinel-2 NDVI time series. A curated 849-sample reference dataset, developed through an iterative rule-based bootstrapping procedure, underpins both a feature sensitivity analysis and an operational classifier. Feature sensitivity independently assessed via Cohen's d and gradient-boosted information gain converges on the February-to-March grain-fill window as most discriminative. Four classifiers (1D-CNN, LSTM, hybrid CNN-LSTM, and XGBoost) were benchmarked on an identical 679/170 sample split. XGBoost achieved the highest overall accuracy (78.82%) against deep-learning baselines (64-66%), consistent with tree-based ensembles' favourable parameter-to-sample ratio in low-sample regimes. At full-population deployment across 36.25 million valid district pixels, the operational classifier attained 86.31% precision and 71.05% recall. The predicted wheat extent deviated by only +2.99% from the official tabular target, whereas the government's spatial reference mask exhibited a +25.11% positive area bias against the identical target. This asymmetry indicates that a classifier trained on an auditor-curated reference set reconciles more closely with the official tabular area than the spatial product conventionally used to validate it. We present this dual-track curation-and-classification framework as a methodological reference for crop-area reconciliation in label-noisy administrative settings.
|
| 92 |
UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching
2609.36980
|
cs.CV
|
Jiajun Le, Yifan Lu, Zizhuo Li, Lei Cao, Junjun Jiang |
Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framewor...Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subsequent token-level matching to the selected paths and avoiding the construction of the full token-to-token matching matrix. We further design a sparse global Dual-Softmax that performs matching only over the routed block candidates while retaining global competition across the sparse matching space. Beyond matching acceleration, UltraMatch employs deployment-oriented structural reparameterization for feature extraction and a tiny fine matching head with shared parameters, further reducing inference cost and memory consumption. UltraMatch achieves competitive accuracy among semi-dense matchers, while running 1.67$\times$ faster than SuperPoint+LightGlue with only 0.44 GiB peak inference memory. Its scalability enables inference at up to 6K resolution on a single RTX 3090, whereas existing semi-dense matchers run out of memory before reaching 2K. Our routing strategy is also transferable, delivering about 2$\times$ end-to-end speedup in EDM and ELoFTR without accuracy loss. The project repository is available at https://github.com/JiajunLe/UltraMatch.
|
| 93 |
Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
2609.36995
|
cs.CVcs.AI
|
Xingtong Ge, Yutong Wang, Lunjie Zhu, Haitao Lin, Fangyu Lin |
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual rep...Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step $1664\times960$ generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: https://xingtongge.github.io/Saltpp
|
| 94 |
Parameterized Stripe Attention for Efficient Video Generation
2609.37001
|
cs.CVcs.AI
|
Xingyu Jia, Baole Ai, Ang Wang, Kang Zhao, Yong Li |
Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, ...Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57$\times$ and 1.37$\times$ end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.
|
| 95 |
Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom
2609.37002
|
cs.CVcs.AI
|
Xijia Tao, Yihua Teng, Xinyu Fu, Cheng Gong, Ziru Liu |
High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before o...High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in PSisual Parallel Search improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.
|
| 96 |
VesselBench-800K: A Large-scale Perception Benchmark for Multimodal Vessel Detection, Counting, and Density Estimation
2609.37003
|
cs.CV
|
Danfeng Hong, Chenyu Li, Jocelyn Chanussot |
Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most existing datasets predominantly focus on general object detection tasks in optical remote sensing (RS) images....Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most existing datasets predominantly focus on general object detection tasks in optical remote sensing (RS) images. Relying solely on single-modality optical RS images proves inadequate for effectively perceiving vessel objects in complex maritime scenarios, where ever-changing weather conditions (e.g., clouds and rain), the need for day-and-night coverage, and the inherent limitations of a single imaging modality pose significant challenges. To fill this gap, we introduce VesselBench-800K, the largest-to-date benchmark dataset on a global scale for vessel perception in multimodal RS images. As its name suggests, VesselBench-800K comprises 800,000 images, each at a resolution of 512x512 pixels, specifically curated for vessel perception tasks such as detection, counting, and density estimation. These multimodal image pairs (i.e., optical, SAR) are collected from diverse platforms, sensors, scenes, shooting heights, and synthetic sources, spanning spatial resolutions from 4.5m to 0.1m. Furthermore, we evaluate numerous state-of-the-art detection, counting, and density estimation models on VesselBench-800K through both qualitative and quantitative comparisons. By revealing previously unrecognized cues, this dataset holds immense potential to significantly advance our understanding of marine traffic. Our VesselBench dataset will be publicly available at https://github.com/danfenghong/IEEE_TGRS_VesselBench to support and contribute to community development.
|
| 97 |
World2Motion: Turning Video World Models into 3D Human Motion Generators
2609.37004
|
cs.CV
|
Tu Fangyuan, Xiangyue Zhang, Yiyi Cai, Yichen Peng, Kunhang Li |
We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited covera...We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3$\times$ faster inference.
|
| 98 |
Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction
2609.37013
|
cs.CVcs.AI
|
Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi, Adrien Dorise |
Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-...Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with a bi-temporal building damage assessment pipeline built on a siamese detector derived from YOLOX, designed to compress information at both ends of the ground/space link. On the ground, pre-disaster reference images are encoded into a compact latent space -- compressed by up to a factor of 64 -- and uplinked to the satellite. On board, this reference is compared with a fresh post-disaster acquisition so that the downlink carries only actionable object-level products, bounding boxes and damage classes, instead of full scenes. This cuts the data exchanged in both directions, while on xBD the strongly compressed reference still preserves most of the detection performance. Because on-board acquisitions suffer from residual pre/post co-registration errors, we introduce a latent-space shift estimation and correction module that regresses the global offset from the coarse feature level and realigns the post-disaster features before fusion. It substantially improves robustness to de-registration -- especially under large shifts, where fusion-only variants collapse -- while also raising nominal accuracy and remaining compatible with the strongest compression. We finally port the pipeline to two embedded targets, a Xilinx Versal VCK190 and an NVIDIA Jetson AGX Orin, and report hardware performance (latency, throughput, power efficiency). The core detector and its compression port cleanly to both, but the operators needed for long-range robustness survive only on the Jetson GPU, whereas the Versal DPU does not.
|
| 99 |
RBF-GNN: Rational Basis Functions for Pseudo-Coordinate based Graph Convolutions
2609.37015
|
cs.CV
|
Pawe{\l} Batorski, Abtin Pourhadi, Paul Swoboda |
We propose RBF-GNN, a new pseudo-coordinate based graph neural network architecture that takes into account Euclidean, spherical or angular coordinates and uses them to induce a powerful spatial inductive bias. Similar in architecture to SplineCNN, we improve ...We propose RBF-GNN, a new pseudo-coordinate based graph neural network architecture that takes into account Euclidean, spherical or angular coordinates and uses them to induce a powerful spatial inductive bias. Similar in architecture to SplineCNN, we improve upon the latter by replacing the less efficient sparse-activation based B-splines whose number grows exponentially with dimension by rational Pad\'e basis functions. For effective training we propose a spline-subspace initialization and a variance-preserving weight rescaling. Experimentally, we evaluate on a number of popular neural network architectures that use SplineCNNs. We replace only the SplineCNNs with RBF-GNN. We achieve improved results, including on semantic keypoint matching, shape matching, event based camera computer vision tasks. We will make our implementation publicly available upon acceptance of the paper.
|
| 100 |
Back2Struct: Making Structured Images Editable Again
2609.37016
|
cs.CV
|
Pengyu Yan, Yixin Wu, Yunjie Tian, David Doermann |
Structured images, such as diagrams, charts, and flowcharts, are inherently symbolic and can be compactly represented in an editable format, yet in practice, they are often rendered as images, and therefore not graphically editable. This mismatch presents a si...Structured images, such as diagrams, charts, and flowcharts, are inherently symbolic and can be compactly represented in an editable format, yet in practice, they are often rendered as images, and therefore not graphically editable. This mismatch presents a significant challenge for researchers, engineers, and designers who wish to incorporate modified versions of existing graphic content into new materials without manually reconstructing it. In this study, we presentBack2Struct, which "makes structured images editable again" by directly recovering vector graphics code (SVG / XML) from image representations. Given an image of a structured graphic, Back2Struct predicts semantically object-level SVG / XML code that explicitly encodes text, shapes, topology, and layout, rather than performing low-level pixel vectorization. The generated code can be seamlessly imported into tools such as PowerPoint, allowing users to edit, refine, restyle, and reuse graphic content while preserving structural fidelity. Beyond supervised fine-tuning on ground-truth SVG token sequences, we further optimize Back2Struct with reward-based learning to better match deployment-time requirements: the output should be syntactically valid, properly concise, and visually faithful to the input diagram. Specifically, we design a composite reward that jointly encourages SVG / XML compilability, length consistency with the reference code, and structural or semantic similarity between the generated and ground-truth graphics. These complementary signals guide the model to produce SVGs that are not only closer to the training distribution, but also more complete, editable, and renderable in practice. Experiments show that Back2Struct improves accuracy, editability, validity, and user alignment over baselines. Dataset and code are available at: pengyu965.github.io/Back2Struct.github.io
|
| 101 |
MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos
2609.37030
|
cs.CVcs.AI
|
Jiahao Zhan, Yongrui Ma, Qunliang Xing, Xuanyu Zhang, Jingqi Tong |
Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this...Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and failure causes. We further propose MotionInsight, a diagnostic evaluator that shifts assessment from implicit RGB-frame observation to explicit motion-space diagnosis. By constructing motion-aware representations, MotionInsight makes subtle motion deficiencies more observable. We also introduce motion-specific rewards during GRPO to transform observed motion into a diagnostic assessment. Experiments demonstrate that MotionInsight provides an effective basis for diagnosing object motion deficiencies, producing human-aligned scores along three dimensions and grounded explanations.
|
| 102 |
UniBuild: Unified Building Mapping From Multi-Source Optical Remote Sensing Imagery With Detail Decoding and Geometry Regularization
2609.37031
|
cs.CV
|
Wei Huang, Chenying Liu, Yilei Shi, Xiao Xiang Zhu |
Building extraction from optical remote sensing (RS) imagery is fundamental to urban mapping, yet existing methods are often dataset-specific and generalize poorly to unseen domains. Their practical use is also limited by insufficient detail recovery and weak ...Building extraction from optical remote sensing (RS) imagery is fundamental to urban mapping, yet existing methods are often dataset-specific and generalize poorly to unseen domains. Their practical use is also limited by insufficient detail recovery and weak geometric regularization, leading to blurred boundaries, irregular shapes, and merged adjacent buildings. To address these issues, we propose UniBuild, a unified building extraction framework for multi-source RGB optical RS imagery. First, a unified multi-dataset training scheme is constructed over heterogeneous RGB optical datasets to learn transferable building representations across sensors and resolutions. Second, a novel detail-preserving HR-DPT decoder is designed to integrate high-level semantic features with high-resolution spatial features, enhancing building detail recovery. Third, geometry-aware regularization is introduced through a structure-tensor-based direction-aware loss for boundary direction consistency and a saddle-aware loss for suppressing false activations in narrow inter-building gaps under low-resolution conditions. We train and evaluate UniBuild on multi-source RGB optical datasets, including 10 public high-resolution datasets and two self-collected low-resolution datasets. Experiments show that UniBuild consistently improves building-region accuracy, boundary sharpness, and adjacent-building separation across diverse datasets. It also generalizes well to unseen domains and supports practical building extraction from RGB optical RS imagery up to 10\,m resolution. The predicted masks can be further converted into GIS-compatible building footprints through simple polygonization. The trained model and inference code are released at https://github.com/zhu-xlab/UniBuild.
|
| 103 |
GleanVID: Complementary Token Selection for Efficient Video Large Language Models
2609.37042
|
cs.CV
|
Shuo Yang, Changbai Li, Rui Tang, Xinyu Zhao, Linlin Yang |
Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent s...Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this insight, we introduce GleanVID, a training-free inference acceleration framework for VideoLLMs. Specifically, GleanVID first allocates the global token budget across frames according to temporal novelty and then selects tokens by jointly considering local representativeness and subspace complementarity, thereby preserving richer and less redundant visual evidence. Extensive experiments across diverse VideoLLMs and benchmarks demonstrate that GleanVID consistently achieves state-of-the-art performance. Notably, with only 25% of visual tokens, GleanVID preserves 98.6% of Qwen3-VL's original performance while reducing its prefill latency by 44.7%. On LLaVA-OV-7B, GleanVID at a 25% retention ratio even slightly surpasses the original model.
|
| 104 |
Speed in the Blind Spot: An Interpretability Analysis of Dynamic Perception in VLMs for Autonomous Driving
2609.37046
|
cs.CV
|
Katharina Winter, Stefan Englmeier, Fabian B. Flohr |
Vision-Language Models are increasingly used in autonomous-driving systems, yet their ability to recover dynamic physical state from visual input remains insufficiently characterized. We study velocity understanding as a controlled diagnostic across three task...Vision-Language Models are increasingly used in autonomous-driving systems, yet their ability to recover dynamic physical state from visual input remains insufficiently characterized. We study velocity understanding as a controlled diagnostic across three tasks: surrounding-agent speed, current ego speed, and short-horizon future ego-speed proposal. On nuScenes, we evaluate open-weight general-purpose and PhysicalAI VLMs, together with the driving-oriented Alpamayo-1.5 Vision-Language-Action model, using multiple input and output formulations. We combine verbal evaluation with temporal perturbations, counterfactual ego-speed hints and linear probes of hidden representations. The tasks exhibit distinct failure modes. Surrounding-agent speed is weakly encoded in an agent-specific form, whereas current ego speed is often internally accessible but poorly verbalized: continuous probes achieve 4.7-5.8 km/h MAE compared with 10.2-16.8 km/h MAE for verbal outputs. Multiple frames provide inconsistent verbal gains to single frame inputs, and frame order is rarely exploited. Under non-optimized QLoRA, task-specific adaptation improves both task-relevant latent speed representations and verbal readout, but continuous surrounding-agent speed estimation remains weak, while most future-speed gains survive frame shuffling, indicating limited temporal grounding. Driving specialized Alpamayo-1.5 shows stronger latent representations for surrounding-agent and future ego speed, while current ego-speed decodability is comparable and substantial probe-verbal gaps remain. Thus, driving specialization can strengthen motion representations but does not guarantee stronger encoding across both scene and ego states or reliable readout. The results show that plausible planning outputs do not necessarily imply reliable recovery or temporal grounding of the underlying dynamic state.
|
| 105 |
NHO: A Neural Hamiltonian Operator for Anchor-based Region Localization and Dense Correspondance
2609.37048
|
cs.CV
|
Jing Li, Yawei Luo, Xiangze Meng, Ying Li, Tieru Wu |
Non-rigid partial-to-full shape correspondence from sparse anchors requires identifying the corresponding region on the full surface and recovering dense correspondences between the partial shape and that region. We present NHO, which combines sparse anchors w...Non-rigid partial-to-full shape correspondence from sparse anchors requires identifying the corresponding region on the full surface and recovering dense correspondences between the partial shape and that region. We present NHO, which combines sparse anchors with the intrinsic geometry of the partial shape to learn a neural Hamiltonian operator whose localized eigenspace encodes both the region support and intrinsic coordinates for dense correspondence. NHO parameterizes the Hamiltonian potential as an intrinsic neural field and optimizes it using anchor evidence together with spectral and geometric constraints. To resolve the spatial ambiguity left by sparse anchors, we introduce reciprocal refinement between operator estimation and correspondence recovery. At each round, the current eigenspace provides spectral coordinates and restricts matching to its induced support, while geometrically reliable correspondences provide additional evidence for updating the potential. After refinement, aggregated eigenfunction energy yields the final localization, and the recovered map initializes dense correspondence refinement. Experiments demonstrate competitive accuracy on both tasks and robustness to uniform scaling and rotation.
|
| 106 |
OmniRoute: Mapping Temporal Semantic Evidence to Audio-Visual Token Budgets for Efficient Omnimodal Large Language Models
2609.37052
|
cs.CV
|
Yuchen Deng, Zidang Cai, Feidiao Yang, Yufei Wang, Jie Wang |
Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression metho...Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes in audio-visual semantic relevance. Motivated by temporal variation and local continuity, we propose OmniRoute, a training-free, two-stage compression framework. First, Temporal Evidence-Guided Budgeting (TEGB) derives chunk-wise modality preferences and initial leading-modality budgets from semantic relevance and local content variation. Second, Budget-Constrained Semantic Compression (BCSC) compresses the leading modality and then calibrates the follower's retention target using the actual retained fraction. For video, it combines spatiotemporal grouping with query-guided selection; for audio, it selects tokens based on encoder attention and query relevance, then merges residual tokens into context anchors under visual guidance. Experiments on four representative benchmarks demonstrate a better trade-off between inference efficiency and performance than competitive baselines. The code and interface will be released to facilitate further research.
|
| 107 |
Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation
2609.37055
|
cs.CV
|
Zhenyu Liu, Zhangquan Chen, Keyi Chen, Mingze Sun, Xiang An |
Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-trut...Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher--student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at https://github.com/vermouth599/Spatial-OPSD.
|
| 108 |
Context without Commitment: Robust Dense Correspondence under Non-Rigid Deformation
2609.37071
|
cs.CV
|
Yuzhen He, Sara Homscheid |
Non-rigid point-cloud registration aims to find the corresponding target point for each point on a deforming source surface. Point-level matching keeps the full target cloud available, but correspondence becomes ambiguous when different regions have similar lo...Non-rigid point-cloud registration aims to find the corresponding target point for each point on a deforming source surface. Point-level matching keeps the full target cloud available, but correspondence becomes ambiguous when different regions have similar local geometry. Regional or coarse-to-fine methods provide broader spatial context, but an incorrect regional match can exclude the correct correspondence before dense matching. We propose CoCo-Reg, which uses regional patches to enrich dense point features without allowing patch predictions to restrict the final point-level search. CoCo-Reg constructs farthest-point-sampled patches, exchanges geometric information within and between source and target, supervises patch similarity using identity-corrected point overlap, and projects the resulting regional information back to dense point features. The final registration stage still scores the full target cloud before global point-level candidate selection. On 726 held-out ModelNet10 objects across nine deformation levels, two established learning-based baselines obtain mean correspondence errors of 0.1993 and 0.1921, whereas CoCo-Reg obtains 0.0547. Relative to its point-level baseline, this is a 72.6\% reduction. CoCo-Reg achieves lower correspondence error on 92.3\% of paired test objects and reduces the mean fraction of points with error above 0.1 from 47.3\% to 17.3\%. Chamfer distance and HD95 decrease in the same direction, and CoCo-Reg remains lower across all tested deformation levels. These results support using regional context for dense non-rigid correspondence without imposing a hard patch-level restriction on the final search. Because evaluation uses one checkpoint per method, the reported gains characterize the complete systems rather than the isolated causal contribution of an individual component. Code will be made publicly available.
|
| 109 |
LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generation
2609.37080
|
cs.CV
|
Zhengqiang Zhang, Lingchen Sun, Rongyuan Wu, Qiaosi Yi, Xiangtao Kong |
Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation misma...Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.80 and 1.90 on 256x256 and 512x512 class-conditional image generation, respectively.
|
| 110 |
Real2Gym: Building Gyms from Videos, Bringing Skills to Robots
2609.37089
|
cs.CV
|
Kerui Ren, Yingxiang Xu, Kaiwen Song, Linning Xu, Bo Dai |
Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic ...Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.
|
| 111 |
Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference
2609.37090
|
cs.CV
|
Luning Pang, Cheng Yuan, Jiawei Shao, Mingtao Huang, Yuan Shen |
Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-...Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-oriented feature compression (TOFC) reduces the payload through feature aggregation and entropy coding. However, continuous-feature coding remains costly, and query-agnostic aggregation may discard task-relevant local evidence. We propose query-guided task-oriented feature compression (Q-TOFC) for device-edge multimodal inference. Q-TOFC employs residual vector quantization (RVQ) to encode each merged feature as a compact sequence of codebook indices, reducing its representation cost and allowing more features to be transmitted. It further incorporates query relevance into feature aggregation and uses a quantization error compensation adapter to mitigate the distortion introduced by discrete quantization. Experiments on seven multimodal benchmarks show that Q-TOFC reduces the visual payload by 53.6% relative to TOFC while maintaining comparable average normalized task performance. End-to-end latency evaluations further demonstrate lower latency under bandwidth-constrained uplinks.
|
| 112 |
Why MLLMs Struggle to Count: Overcoming Individuation and Aggregation Bottlenecks with ConvStack
2609.37096
|
cs.CV
|
Liwei Che, Yihao Quan, Sen Fang, Hongyi Wang, Ranjay Krishna |
Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inh...Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inherent to the global attention pipeline of MLLMs. First, we reveal an individuation bottleneck stemming from image patchification: because Vision Transformers process patches independently, they struggle to group fragmented geometric features across boundaries into distinct object representations. Second, we identify a collapse in the subsequent counting aggregation process, where representation separation rapidly diminishes as numerosity increases due to attention compression. Identifying and formalizing these twin bottlenecks constitutes our first major contribution. To overcome them, we propose ConvStack, a lightweight architecture that operates directly in the visual token space to explicitly aggregate and inject local spatial structures via zero-initialized residual connections. By explicitly addressing the individuation bottleneck, ConvStack provides unambiguous geometric evidence for downstream aggregation. Remarkably, by fine-tuning exclusively on counting tasks, the model achieves substantial improvements in dense object counting and broader spatial understanding benchmarks, without compromising on general visual capabilities.
|
| 113 |
Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware
2609.37107
|
cs.CV
|
Rajit Rajpal, Shahbuland Matiana, Liew Wei Pyn, Anmol Agarwal, Ryan Craig |
We present Waypoint 1.5, a real-time diffusion world model for interactive video generation on consumer-grade hardware. Unlike general video diffusion models, interactive world models (iWMs) must respond to dense user controls under strict latency and throughp...We present Waypoint 1.5, a real-time diffusion world model for interactive video generation on consumer-grade hardware. Unlike general video diffusion models, interactive world models (iWMs) must respond to dense user controls under strict latency and throughput constraints. Waypoint 1.5 is pre-trained on 100,000 hours of diverse, control-aligned video game data across hundreds of games, and generates playable video conditioned on full keyboard and mouse input. The model includes two resolution variants that run across a wide spectrum of consumer hardware. To characterize this unique setting, we distinguish rendered FPS, latent FPS, and control rate. We describe the data pipeline, architecture, training methodology, and runtime system behind Waypoint 1.5. We evaluate interactivity through latency and throughput. Finally, we discuss the safety and ethics considerations unique to iWMs.
|
| 114 |
NRF-GS: Neural Residual Fields for Expressive and Compact Gaussian Splatting
2609.37115
|
cs.CV
|
Pratik Singh Bisht, Andreas Kolb |
We revisit the role of appearance modeling in 3D Gaussian Splatting (3DGS) and show that limited expressiveness in view-dependent reflectance is a key driver of representation redundancy. In standard 3DGS, low-order spherical harmonics (SH) are used, restricti...We revisit the role of appearance modeling in 3D Gaussian Splatting (3DGS) and show that limited expressiveness in view-dependent reflectance is a key driver of representation redundancy. In standard 3DGS, low-order spherical harmonics (SH) are used, restricting the splats' ability to model high-frequency directional effects, which is typically compensated by increasing the number of splats. We propose \emph{NRF-GS: Neural Residual Fields for Gaussian Splatting}, a hybrid representation that replaces per-splat SH-bases with a shared neural residual field. Each Gaussian encodes a compact set of appearance features and a lambertian base color, while a lightweight \emph{global scene-level MLP} predicts view-dependent residuals conditioned on viewing direction, distance, and per-splat features. This formulation enhances directional reflectance modeling by combining diffuse per-splat reflectance representations with a shared global function for high-frequency details, enabling both higher expressiveness and parameter sharing across splats. Our key insight is that by accurately capturing high-frequency directional reflectance, especially in specular regions, the GS-representation becomes more expressive, reducing the need for geometrically redundant splats. As a result, NRF-GS achieves comparable or better rendering quality while reducing the number of Gaussians by up to 50\%, and produces visibly improved specular and high-frequency details.
|
| 115 |
EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception
2609.37123
|
cs.CV
|
Yaoxin Niu, Zhangquan Chen, Yang Zhang, Xiang An, Zhumei Wang |
Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, ...Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.
|
| 116 |
Multi-Granularity Language-Guided Imitation Learning via Instruction Decomposition
2609.37135
|
cs.CV
|
Yi-Pei Chiu, Wei-Ta Chu |
Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-guided policy learning methods typically use an overall task description to guide the entire demonstration tra...Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-guided policy learning methods typically use an overall task description to guide the entire demonstration trajectory. For manipulation tasks involving multiple execution stages, these methods assign the same language description to different subtasks, making it difficult to distinguish the behaviors required at different stages. In this work, we propose a multi-granularity language guidance method based on instruction decomposition. The proposed method decomposes an overall task description into more fine-grained, concrete subtask-level language instructions, thereby enhancing learning efficiency and improving performance. We evaluate the proposed method in the setting of multi-task imitation learning and validate its effectiveness.
|
| 117 |
TaoFlowForge: Progressive Native Mesh Generation via Cascaded Flow Matching
2609.37139
|
cs.CV
|
Xianze Fang, Qiyuan Feng, Dongfang Sun, Yan Zhang, Xiuchao Wu |
3D content generation technology has significantly advanced the work of designers, as well as the 3D printing and gaming industries. However, it remains difficult to produce lightweight, editable, and topologically clean artistic content that is directly produ...3D content generation technology has significantly advanced the work of designers, as well as the 3D printing and gaming industries. However, it remains difficult to produce lightweight, editable, and topologically clean artistic content that is directly production-ready. To achieve this, we present TaoFlowForge, an artistic mesh foundation model that generates production-ready meshes. Specifically, TaoFlowForge decomposes the mesh generation process into vertices generation and their connectivity prediction, i.e., edges. We formulate vertices generation as a two-stage coarse-to-fine process and incorporate several effective loss functions to further enhance its performance. In the connectivity prediction stage, we propose a simple yet effective method for estimating the connectivity affinity between vertices and additionally predict per-vertex normals, which determines the correct orientation of faces. Besides, we construct a large-scale dataset combining hand-crafted 3D assets with public high-quality topology datasets. Based on this, a carefully designed data curation pipeline is employed to filter the raw dataset, retaining only high-quality topology data for model training. Our model is trained on the combined dataset and tested on both out-of-distribution hand-crafted set of 3D assets and public datasets. Under image-conditioned generation, TaoFlowForge outperforms autoregressive methods and achieves state-of-the-art results among open-source mesh topology generators. We will release all the code and weights together with a portion of our test dataset.
|
| 118 |
MSTypography: Multi-character Semantic Typography via Balancing Word Legibility and Object Recognizability
2609.37141
|
cs.CV
|
Xinye Yang, Xinding Zhu, Kai Fang, Xinyi Ren, Mengjian Li |
Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Existing digital typography methods mainly focus on single-character scenarios. They suffer from a lack of legib...Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Existing digital typography methods mainly focus on single-character scenarios. They suffer from a lack of legibility constraints and insufficient local deformation when extended to multi-character words, as the intricate structures among multiple characters are hardly preserved during the typography process. In this paper, we propose a global-to-local typography framework for multi-character scenarios. It performs mask-driven silhouette approximation at the global level, while semantic-guided refinement at the local level, with a culling step in between to improve efficiency. To preserve word legibility, we designed structural losses (including explicit collision constraints and implicit Jacobian singular value constraints) and an OCR constraint for character-level readability. To enhance the object recognizability, we leverage semantic guidance with diffusion priors, which drives the character glyph toward the target concept while preserving its structural integrity. To the best of our knowledge, this is the first multi-character semantic typography method that effectively balances word legibility and object recognizability. Evaluations on five representative languages (English, Chinese, Japanese, Korean, Arabic) demonstrate superiority over SOTA methods. Codes will be open-sourced.
|
| 119 |
Improved Distributional Diffusion Models
2609.37147
|
cs.CV
|
Tommaso Martorella, Alexandre Galashov, Felix Krause, Stefan Andreas Baumann, Valentin De Bortoli |
Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However,...Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~\citet{Biroli2024}. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet-$256^2$, achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at https://github.com/CompVis/iDDM.
|
| 120 |
End-to-End Self-Supervised RGB-T Tracking without Modality Misleading
2609.37162
|
cs.CV
|
Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia, Yong Zhou |
RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while m...RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bounding box annotations. Given only the standard initial-frame annotation used in visual tracking, ESMTrack learns discriminative and temporally consistent representations through two complementary objectives: a grounding triplet loss on annotated initial frames and a cross-frame temporal triplet loss on unlabeled search frames, with reliable samples selected by forward-backward consistency. To address modality dominance bias, ESMTrack employs a three-branch architecture consisting of a fusion branch and two unimodal branches for RGB and thermal inputs. We quantify modality contributions using the Average Peak-to-Correlation Energy by measuring response discrepancies between the fusion and unimodal branches. The resulting reliability estimates guide a training-time modality decoupling mechanism that suppresses dominant-modality shortcuts and adaptively weights cross-modal contrastive learning for task-level alignment. Extensive experiments on five RGB-T tracking benchmarks show that ESMTrack achieves competitive state-of-the-art performance, strong cross-dataset generalization, and real-time inference speed. The source code is available at https://github.com/LiShenglana/ESMTrack.
|
| 121 |
Sparse cubical complexes for efficient topology-preservation in image data
2609.37177
|
cs.CV
|
Alexander H. Berger, Marco Fontana, Daniel Rueckert, Johannes C. Paetzold, Laurin Lux |
Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segmentation, where preservation of topological structures is important. However, despite its general applicability ...Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segmentation, where preservation of topological structures is important. However, despite its general applicability across dimensionality, domains, and target structures, the runtime cost of PH-based methods often makes their practical use infeasible. In this work, we argue that this runtime cost is largely driven by processing information that is unimportant for downstream application (e.g. as optimization objective). We propose sparse cubical filtrations as an alternative foundation for PH computation, reducing subsequent computational costs by factors of up to 100 on real datasets. We show close agreement with the optimization signal of the dense counterpart and empirically evaluate our solution's effectiveness as an optimization objective in realistic training regimes where other PH-based objectives can practically not operate (i.e., 3D data with large patch sizes). We show how our solution improves topological accuracy by up to 80\% across six diverse datasets while maintaining pixel- and region-based accuracy.
|
| 122 |
InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning
2609.37187
|
cs.CV
|
Hongpei Zheng, Hujun Yin |
Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-c...Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-conditioned prediction targets. Historical views are linked to labeled map locations, and a shared multimodal backbone jointly learns navigation action prediction and post-action map generation. Map prediction provides auxiliary training supervision, while navigation inference decodes actions from the observed spatial context. An aligned RGB-D data pipeline supports a common interface for navigation, visual question answering, situated reasoning, and 3D grounding. On the validation-unseen splits of R2R-CE and RxR-CE, InsightMap achieves success rates (SR) of 56.9% and 54.9%, respectively. Adding map-prediction supervision improves R2R-CE SR by 4.3 and success weighted by path length (SPL) by 3.2 percentage points. On static spatial tasks, InsightMap achieves 103.7 CIDEr on ScanQA, 60.1% exact-match accuracy on SQA3D, and 53.1% grounding accuracy at 0.5 IoU on ScanRefer with detected object proposals. On Unitree Go2, it outperforms NaVid and NaVILA in hallway, lab, and office environments.
|
| 123 |
HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent
2609.37190
|
cs.CV
|
Zhangquan Chen, Yaoxin Niu, Xiang An, Mingze Sun, Zhumei Wang |
Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in ...Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final result is correct arise frequently, which in turn leads to ineffective training, i.e., scaling along the wrong paths. In this paper, we introduce HaPRL, the first framework to reinforce the search process with human search behavior. We first build an annotation platform and collect 1K+ human-annotated data with fine-grained behavioral signals. During training, a carefully designed judge scores each rollout with task-adaptive weights, anchored on the distilled trace of how a human annotator actually searched the same image. Extensive experiments show that HaPRL consistently outperforms outcome-based RL, and early-stage process supervision yields 6.7x more improvement in subsequent outcome-based scaling. Our results also demonstrate the importance of aligning model behavior with human process annotation signals, which offer new insight into the training of foundation models.
|
| 124 |
Exploring In-Context Learning for Handwritten Text Recognition
2609.37195
|
cs.CV
|
Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza |
Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their tr...Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses mostly on specialized models that require large amounts of annotated samples to achieve satisfactory performance. We explore the use of In-Context Learning with pre-trained Vision-Language Models (VLMs) to create a transcription pipeline without updating the model's parameters. We then evaluate this pipeline across multiple collections and models, and demonstrate that general-purpose VLMs can be effectively taught how to transcribe handwritten text from images. To assess how our observations may translate to practical applications, we evaluate the performance in a Cross-Domain (CD) scenario, where context examples are drawn from a different collection than the query image. Results in both the controlled In-Domain (ID) scenario and the realistic CD scenario follow the same patterns. First, as context size grows, the error range is expected to narrow towards the average performance. Thus, larger context sizes sacrifice the performance of the oracle-best sampling for lower expected error rates. The results obtained show that, without any parameter updates, this methodology has strong potential to compete with traditional HTR in the presence of domain shift. Moreover, we show and argue that some context samplings work better than others and suggest more effort should be put into finding an ideal sampling method in future work.
|
| 125 |
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL
2609.37200
|
cs.CV
|
Songlin Yang, Xiaotong Zhao, Jiacheng Zhang, Zhe Wang, Toyota Li |
Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization...Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that change during training: where reward-driven updates should act, and how competing rewards should be combined. Existing methods tend to rely on fixed routing and reward weights, failing to track evolving model functions. To address these limitations, we propose Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. Our method consists of two components. (i) Cross-Modal Influence-Guided Routing (Localizing Updates): We use bidirectional cross-attention responses as an efficient proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without additional model interventions. (ii) Preference-Preserving Modality-Aware Reweighting (Coordinating Rewards): We preserve predefined weights as preference priors and use branch-specific reward-gradient interactions as residual corrections after warm-up. This resolves evolving conflicts without letting dominant rewards suppress weak but essential objectives. Extensive experiments demonstrate consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines. Ablations and mechanism analyses further validate the complementary benefits of adaptive update routing and reward coordination.
|
| 126 |
ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression
2609.37225
|
cs.CVcs.AI
|
Zijing Cai, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, Richang Hong |
Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of ...Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.
|
| 127 |
AESOP: Asymmetric Human-Camera Generation with Translation-Intensity Control
2609.37229
|
cs.CV
|
Jingzhong Lin, Zhanke Wang, Heng Li, Wenxiang Liu, Zhao Zhang |
Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human mo...Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.
|
| 128 |
Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry
2609.37230
|
cs.CV
|
Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim |
Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text...Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.
|
| 129 |
Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features
2609.37243
|
cs.CVcs.AI
|
Dae Ung Jo, Jongin Lim, YoungJoon Yoo, Daeho Um |
Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, t...Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.
|
| 130 |
V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
2609.37250
|
cs.CVcs.AI
|
Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan |
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was...World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
|
| 131 |
Collision-Aware and Observation-Aligned Object-Centric Scene Reconstruction from Point Cloud
2609.37260
|
cs.CV
|
Yuxuan Xie, Xuan Yu, Rong Xiong, Yue Wang |
Object-centric scene reconstruction requires completing partial object observations while preserving metric alignment and avoiding collisions with the surrounding. Existing generation-based methods are often image-conditioned and suffer from scale ambiguity an...Object-centric scene reconstruction requires completing partial object observations while preserving metric alignment and avoiding collisions with the surrounding. Existing generation-based methods are often image-conditioned and suffer from scale ambiguity and insufficient geometric constraints. We propose COOL, a framework for COllision-aware and Observation-aLigned reconstruction. Based on an object generation model, COOL conditions the generation on instance and background point clouds. Instance geometry anchors generation in scene coordinates, while background geometry provides local context for scene-consistent completion. We further introduce an explicit collision loss and use joint optimization and resampling to reduce collisions during inference. Experiments on 3D-Front and Scan2CAD demonstrate strong scene-level fidelity, observation alignment, and collision reduction. Moreover, additional studies validate its robustness to mask errors and its applicability to real-world scene replicas.
|
| 132 |
Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery
2609.37263
|
cs.CV
|
Siqi Lu, Suo Wei, Yongbin Zheng, Jianhang Yao, Wanying Xu |
While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus...While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus on reweighting visual tokens or suppressing language priors. However, such approaches often overlook the spectral characteristics of the visual information flow and frequently rely on Contrastive Decoding (CD), which doubles inference time. Instead of following conventional approaches, we identify two distinct hallucination patterns-Perceptual-Semantic Dissociation and Localized Fixation-and propose FLASH (Frequency-Localized Attention SHaping), a training-free and CD-free framework. FLASH utilizes a Spectral Vortex Score to detect vision heads within multi-head attention layers and applies adaptive spectral modulation to rectify the visual information flow during decoding. Empirical results demonstrate that FLASH achieves a superior balance between performance and efficiency compared to SOTA methods.
|
| 133 |
UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception
2609.37264
|
cs.CVcs.AI
|
Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang |
Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This frag...Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: https://4dvlab.github.io/UniAfford
|
| 134 |
SAM Meets VLM: Parameter-Decoupled Full-Parameter Training for Unified Medical Reasoning and Segmentation
2609.37283
|
cs.CV
|
Xuyang Cao, Enyou Liu, Jun Zhao, Zhuoyun Liu, Jintao Fei |
Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmen...Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmentation through a special <SEG> token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning and segmentation. The framework treats the <SEG> hidden state as a semantic-to-spatial prompt for the mask decoder and encourages it to become separable from generic language states, reducing ambiguous segmentation prompts and potential disruption to reasoning representations. It first performs medical shallow alignment to adapt visual features to clinical language without disturbing the LLM; then controlled instruction tuning shapes separable <SEG> prompt states, monitored by the Davies--Bouldin Index (DBI), while scaling segmentation gradients entering the language backbone; finally, the SAM branch is specialized with the VLM frozen to improve mask precision without altering reasoning parameters. Experiments on medical referring segmentation, grounding, visual QA, and textual QA benchmarks show that our framework achieves strong language-conditioned segmentation while preserving competitive reasoning ability. Ablations show that two-phase instruction tuning, gradient scaling, and segmentation specialization all contribute to the model.
|
| 135 |
VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics
2609.37287
|
cs.CVcs.AI
|
Bo Lv, Mao Zheng, Zheng Li, Fangxu Liu, Mingrui Sun |
Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit imag...Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics specify essential content, semantic relations, and acceptable translation variants, yielding separate output-based scores for translation quality and the preservation of visual and knowledge-dependent information. We conduct extensive evaluations of 16 mainstream models, including 12 multimodal models and four text-input models, and provide systematic analyses across languages, domains, and evaluation dimensions.
|
| 136 |
Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models
2609.37297
|
cs.CV
|
Zhiyuan Li, Wenyan Yang, Pekka Marttinen, Joni Pajarinen |
Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transfer...Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: https://cross-skeleton-retargeting.netlify.app/.
|
| 137 |
Technical note on: Zero-Training Feature-Space Alignment via Information Geometry
2609.37302
|
cs.CV
|
Behraj Khan, Tahir Qasim Syed, Syed Ahmad Chan Bukhari |
Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignmen...Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness under covariate shift without modifying model parameters. ZFGA is based on the observation that distribution shifts distort feature-space geometry. It estimates the Fisher information matrix of the predictive distribution with respect to feature embeddings and applies a linear transformation that aligns test-feature Fisher geometry with a reference geometry computed from clean data. This provides a natural-gradient-inspired preconditioning step in feature space. We evaluate ZFGA on CIFAR-10-C and ImageNet-C using ResNet-50, DINO ViT-S/16, and CLIP ViT-B/32. ZFGA consistently improves over zero-shot inference across all three models, although it is not the strongest method for every model. Covariance whitening performs better on ResNet-50, while Fisher whitening is statistically indistinguishable from ZFGA on CLIP. Across six training-free and gradient-based alternatives (covariance whitening, Fisher whitening, TENT, T3A, LAME, and AdaNPC), ZFGA is the only method that does not substantially harm any of the three model families. The Fisher geometry distortion is also positively correlated with ZFGA gain (Pearson r = 0.366, p = 0.017), providing preliminary evidence that geometric misalignment contributes to robustness degradation. ZFGA requires only forward passes and matrix operations at inference time, offering a lightweight and deterministic alternative to optimization-based test-time adaptation.
|
| 138 |
FLASH: A "Generate Once, Synthesize Many" Framework for Synthetic Anomaly Generation in Industrial Anomaly Detection
2609.37314
|
cs.CV
|
Abhay Kumar Das, Rajesh Gangireddy, Ashwin Vaidya, Samet Akcay |
Synthetic anomaly generation helps expand industrial anomaly datasets when real defects are scarce or unavailable. Existing approaches lie at two extremes: procedural approaches are fast but struggle to represent complex anomalies, while generative approaches ...Synthetic anomaly generation helps expand industrial anomaly datasets when real defects are scarce or unavailable. Existing approaches lie at two extremes: procedural approaches are fast but struggle to represent complex anomalies, while generative approaches produce diverse defects but require costly per-sample generation. We present FLASH, a framework that decouples defect generation from anomaly synthesis under a ``generate once, synthesize many'' paradigm. Given only normal images, FLASH uses Vision-Language Model (VLM) guidance and an image-generation model to produce a small set of defect images, from which it extracts, validates, and banks reusable defect patches. For synthesis of anomalous images, Object Boundary Suppression (OBS) first identifies the probable foreground object-aware region of the host image, while Multi-Resolution Spectral Pyramid (MRSP) noise generates diverse, size-controllable masks that determine the defect location and spatial extent. It then composes a large and diverse synthetic anomalous image set by localizing the defect region, sampling size-controllable placement masks and seamlessly blending retrieved defects onto new defect-free images without further need for image generation. Experiments on the MVTec AD 2 dataset show that FLASH-generated anomalies nearly close the calibration gap on real defects, reaching 78.1% image-level F1 against an 83.6% real-anomaly upper bound and providing the most consistent calibration transfer across detectors among procedural and generative alternatives. Moreover, FLASH synthesizes anomalies more than 11.95x faster than per-sample generative approaches.
|
| 139 |
What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation
2609.37317
|
cs.CVcs.MM
|
Sieun Hyeon, Yejoon Lee, Mintaek Lim, Woojin Kim, Jaeik Kim |
Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coher...Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children's books, with ground-truth next-page references and speech metadata. We evaluate systems with modality-specific metrics and consistency-centered LLM-as-a-judge rubrics for context preservation, condition following, reference consistency, and cross-modal coherence. Across 32 baseline configurations spanning orchestration, semi-orchestration, and native any-to-any paradigms, we find orchestration with strong VLM planning most reliable, while current native omnimodal models often struggle with output completeness and controllability. Our analysis shows text-side performance is associated with image and speech quality, but image generation and visual continuity form the clearest observed bottleneck among the evaluated configurations. These results position Omni-StoryBench as a system-level benchmark measuring coherent omnimodal generation beyond isolated modality quality.
|
| 140 |
The Domain Is a Residue: Adapting Self-Supervised Features, Not Generators
2609.37330
|
cs.CV
|
Thomas Deixelberger, Markus Steinberger |
Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a cont...Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in the scene and carries weather, lighting and rendering style as a residue of 13 to 14% of the feature norm. We propose the Representation Feature Adapter (RFA), a 2.9M-parameter network that moves this residue. We train only the adapter and its discriminators; the encoder and a feature-conditioned decoder, trained once for all conditions, stay frozen. Against CycleGAN-Turbo it is ahead on both metrics on fog and on KID on night, and level within noise on snow, rain and haze. On sim-to-real it leads REGEN and HyPER-GAN on both metrics. Only the RFA removes the rain while keeping the scene. The removal costs scene structure: CycleGAN-Turbo keeps more on every condition but fog. On VAE latents the identical adapter collapses to the identity, and decoders from other groups that never saw it render its output. The RFA has about 160 times fewer trainable parameters than CycleGAN-Turbo and under a fifth of its per-condition training time.
|
| 141 |
OFBD: Object-Focused Background Debiasing for Long-Tailed Learning
2609.37331
|
cs.CV
|
Shenghan Chen, Yiming Liu, Zhipeng Deng, Haolin Wang, Jiale Zhou |
Balancing performance trade-offs on long-tailed data distributions remains a long-standing challenge in visual recognition. Existing methods mainly improve tail classes through re-balancing, representation learning, or data augmentation, but the underlying cau...Balancing performance trade-offs on long-tailed data distributions remains a long-standing challenge in visual recognition. Existing methods mainly improve tail classes through re-balancing, representation learning, or data augmentation, but the underlying cause of tail class degradation is still insufficiently explored. In this paper, we find that standard long-tailed training induces background-biased representation and optimization: tail classes suffer larger background distribution shifts and become increasingly driven by background gradients. This reveals that tail degradation is not merely caused by insufficient samples, but also by the learning of irrelevant background features. To tackle this issue, we propose Object-Focused Background Debiasing (OFBD), a framework that mitigates background bias from both distribution and optimization perspectives. Specifically, Foreground-guided CutMix preserves target-related foregrounds while diversifying complementary backgrounds, and Background-guided Feature Rectification suppresses background-biased features without learnable parameters or additional training. Extensive experiments show that our method improves overall accuracy, achieves significant tail-class gains, and can serve as a plug-in for mainstream long-tailed methods without external data or pretrained recognition models. The code is available at: https://ofbd-neurips2026-longtail-learning.github.io/
|
| 142 |
UGO: Unified Architecture for General Multi-Object Tracking by Segmentation
2609.37339
|
cs.CV
|
Jer Pelhan, Alan Lukezic, Matej Kristan |
General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on bounding boxes and surrogate training, and struggles with non-rigid objects, crowded scenes, and distractors. We int...General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on bounding boxes and surrogate training, and struggles with non-rigid objects, crowded scenes, and distractors. We introduce UGO, a unified GMOT tracker that pairs a pretrained exemplar-conditioned detection head with an instance-propagation head in a common architecture. A novel training-free, energy-minimization consolidation method converts overlapping proposals into exclusive pixel-wise masks and detections, resolving over-segmentation, duplicates, and conflicts. A hierarchical memory spanning global and instance levels improves recall and per-instance segmentation accuracy using a new memory management protocol. UGO sets a new state-of-the-art on GMOT benchmarks and video object counting, and is competitive with specialist MOT methods, establishing a strong paradigm for unified, open-category multi-object tracking.
|
| 143 |
HyperSAM: A Promptable Foundation Model for Hyperspectral Remote Sensing
2609.37340
|
cs.CV
|
Li Pang, Xinqiao Wu, Jing Yao, Pedram Ghamisi, Jun Zhou |
Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model remains difficult. Two bottlenecks are especially limiting. ...Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model remains difficult. Two bottlenecks are especially limiting. First, large hyperspectral corpora rarely provide high spatial resolution together with reliable dense annotations. Second, many hyperspectral models are still trained almost from scratch, so the geometric and interactive priors learned by modern vision foundation models are not fully reused. To alleviate these issues, we \highlight{present} \textbf{HyperSAM}, a promptable hyperspectral foundation model that couples a data-centric hyperspectral synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). On the data side, HyperSAM synthesizes full-spectrum hyperspectral cubes from high-resolution SpaceNet multispectral imagery through a physics-informed abundance-transfer generator, while SAM3-derived pseudo-masks provide object-centric supervision. On the model side, the latest implementation uses a frozen SAM3 RGB image branch, a trainable hyperspectral side encoder initialized from the RGB vision transformer (ViT), ControlNet-style zero-initialized feature injection, and a lightweight mixture-of-experts mask refiner. To enhance training robustness against noisy pseudo-labels, Cross-modal Sample Selection (CromSS)-style confidence selection is incorporated for noisy-label weighting. Extensive experiments show that HyperSAM obtains strong generalization on diverse hyperspectral tasks (e.g., classification, anomaly detection, change detection, target detection, and airborne oil-spill mapping) and that high-quality synthetic hyperspectral data can be more effective than simply scaling noisy hyperspectral supervision.
|
| 144 |
When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs
2609.37345
|
cs.CV
|
Xiang Hu, Jiazuo Yu, Lu Zhang, Yunzhi Zhuge, Huchuan Lu |
Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances tempora...Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.
|
| 145 |
TAEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG
2609.37349
|
cs.CVcs.AI
|
Yalun Wu, Bingzhou Wang, Boyang Wang, Peiying Wang, Shaojie He |
Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure it...Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As multi-step reasoning progresses, redundant sources occupy context capacity needed for missing evidence, observations tied to resolved requirements or unproductive searches linger in context, and visual sources are revisited with insufficient detail for fine-grained reading. We term this loss of usable evidence over a reasoning trajectory trajectory-level evidence utilization degradation. To address it, we propose Trajectory-Aware Evidence Coordination (TAEC), a training-free framework that coordinates evidence use around unresolved answer requirements. TAEC tracks these requirements in a shared trajectory state to guide which evidence enters the context, how accumulated memory is retained, and at what level of detail visual evidence is examined. Under a unified evaluation protocol on ViDoSeek, SlideVQA, and MMLongBench-Doc, TAEC achieves the best overall performance against leading training-free visual RAG baselines, with the highest average accuracy across multiple proprietary vision-language models. These results demonstrate that aligning evidence with evolving reasoning needs improves evidence use throughout multi-step visual RAG.
|
| 146 |
PCaPaint: Prostate Cancer Inpainting by Mitigating Shortcut Learning
2609.37350
|
cs.CV
|
Levente Lippenszky, Hongxu Yang, Marcell D\"om\"ot\"or, Krisztian Koos, L\'aszl\'o Rusk\'o |
The development of AI systems for tumor-specific applications is limited by the scarcity of labeled data. Synthetic tumor inpainting offers a promising approach but faces challenges for prostate cancer MRI which contains high-resolution multi-sequence data. Al...The development of AI systems for tumor-specific applications is limited by the scarcity of labeled data. Synthetic tumor inpainting offers a promising approach but faces challenges for prostate cancer MRI which contains high-resolution multi-sequence data. Although methods leveraging latent diffusion models (LDMs) enable large-volume synthesis, they are prone to shortcut learning, simply reproducing the condition image created by masking the lesion region. In this work, we introduce PCaPaint, a prostate cancer inpainting method based on LDMs that explicitly addresses this failure mode. To overcome shortcut learning that compromises synthetic tumor texture, we propose a simple yet efficient conditioning strategy in which the condition image is filled with Gaussian noise, and we provide theoretical justification. In addition, we propose a novel training objective for LDM that emphasizes the error within the lesion region. Furthermore, we introduce a multi-sequence latent design, in which T2w scans and DWI&ADC scans are compressed using two separate autoencoders to preserve their distinct frequency characteristics. Extensive experiments demonstrate that the generated synthetic data improves downstream performance in prostate lesion segmentation, patient-level classification and lesion-level detection. Furthermore, our method significantly outperforms a recent state-of-the-art LDM-based tumor inpainting method both in downstream performance and in synthetic image quality.
|
| 147 |
Visual Anomaly Synthesis for Model Selection in Data Scarcity
2609.37360
|
cs.CV
|
Daniel Pr\"oll, Thomas Kraxner, Tobias Schaefer, Sebastian Hegenbart |
Defect detection systems for industrial condition monitoring can only be relied upon if they are validated, yet defective samples are rare and, for a specific asset, often nonexistent. We present a framework that synthesizes severity-graded defects on real non...Defect detection systems for industrial condition monitoring can only be relied upon if they are validated, yet defective samples are rare and, for a specific asset, often nonexistent. We present a framework that synthesizes severity-graded defects on real non-defective images without any defect references for the target asset, that can be used for model selection and validation. A defect taxonomy for common failure modes is distilled from literature into prescriptive prompts at varying defect severities. Regions of interest are cropped from in defect-free images and edited with a pre-trained image generation model ("FLUX.2 [klein]"). Color-matching and blending are employed to improve structural coherence with the original image. Generations are filtered out by a scorer and by estimated detection difficulty. Model selection experiments on MVTecAD show image AUROC choice regret over model selection can be nearly halved compared to the best fixed model chosen with access to test data. Experiments show the need for severity-graded anomaly synthesis. A case study investigates the proposed method for in-situ monitoring of Pelton turbine runners in hydropower, where real defect images are rare and expensive to collect. A PatchCorebased anomaly detection model is fit on Pelton turbine images and selected and validated using synthetic images, showing strong detection performance (94 % correct detection at optimal threshold and AUROC 0.97). The model reliably detects moderate and advanced defects, while early-stage defects remain challenging, indicating the synthetic data meaningfully stresses detector sensitivity.
|
| 148 |
Think Before You Score: Thinking Reward Model for Visual Generation
2609.37372
|
cs.CV
|
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu |
Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case...Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.
|
| 149 |
MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding
2609.37374
|
cs.CVcs.MM
|
Heyu Huang, Chi Chen, Zonghao Guo, Yuhua Li, Maosong Sun |
Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pix...Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pixel-precise localization, existing RL-based approaches overlook two characteristics intrinsic to this paradigm: a coarse-to-fine hierarchical reasoning pattern, and heterogeneously distributed task--sample difficulties. In this work, we present MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarchical reasoning, supported by a curated 25K MRG dataset with task-adaptive Chain-of-Thought (CoT) annotations that elicit multi-perspective evidence before conclusion. To remedy the heterogeneous task--sample difficulties, we further propose Bi-Axial DAPO (BiA-DAPO), which decomposes rollout advantages along an intra-group signal axis and an inter-group competence axis through two complementary mechanisms, both grounded on our defined candidate pool for stable group-level statistics. Extensive experiments show that MG-Thinker achieves state-of-the-art performance on multi-image reasoning grounding while consistently improving generalization across multi-image understanding and diverse multimodal benchmarks.
|
| 150 |
Do-JEPA: From Masking to Intervention in Latent World Models
2609.37378
|
cs.CVcs.AI
|
Hossein Resani, Javen Qinfeng Shi |
Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what...Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $\Delta z=z^{a}-z^{a_{\varnothing}}$ between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in 99.95% of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC 0.975 vs. 0.624). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by 28.4% on an end-to-end LeWM model and physical effect error by 13.5% when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about 20% under physics shifts and the latent context sensitivity of predicted effects by 66%. Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.
|
| 151 |
Anatomy-Aware Prediction of Bronchoscopic Accessibility from 3D CT
2609.37386
|
cs.CV
|
Linkai Peng, Cuiling Sun, Bin Wang, Jamie Rowell, Catherine Gao |
Pre-operative planning for bronchoscopy is critical for the diagnosis of lung lesions. Current accessibility assessment relies on subjective manual inspection of CT scans, which is time-consuming and prone to inter-observer variability. In this paper, we forma...Pre-operative planning for bronchoscopy is critical for the diagnosis of lung lesions. Current accessibility assessment relies on subjective manual inspection of CT scans, which is time-consuming and prone to inter-observer variability. In this paper, we formalize bronchoscopy accessibility prediction as a novel supervised learning task and present the first end-to-end framework to address it. We propose an Anatomy-Aware Mixture-of-Experts (MoE) model that integrates specialized modules: a CT Expert for local morphological features, a Lobe Expert for anatomical priors, and a Path Geometry Expert that encodes the sequential constraints of the bronchial tree. To support this task, we curated the first clinical dataset of 438 cases with pre-operative CT scans and documented procedural outcomes. Experimental results demonstrate that our method achieves an AUROC of 0.8052, significantly outperforming both state-of-the-art baselines and experienced human experts. This work establishes a new benchmark for computer-aided interventional planning in pulmonary medicine. Our data and code will be publicly available at https://nubagcilab.github.io/BronchoAccess/.
|
| 152 |
Multi-task learning for the automatic grading of enlarged perivascular space burden using MRI
2609.37387
|
cs.CV
|
Jesse Phitidis, William N. Whiteley, Joanna M. Wardlaw, Miguel O. Bernabeu, Yajun Cheng |
Enlarged perivascular spaces (PVS) visible in brain magnetic resonance imaging (MRI) are increasingly thought to be linked to poor brain health. PVS are elongated structures of less than 3 mm in diameter and can be numerous. To reflect the incidence of PVS, ra...Enlarged perivascular spaces (PVS) visible in brain magnetic resonance imaging (MRI) are increasingly thought to be linked to poor brain health. PVS are elongated structures of less than 3 mm in diameter and can be numerous. To reflect the incidence of PVS, radiologists visually score their burden following a clinical grading scale - a task that would benefit from automation to accelerate analyses and overcome the influence of inter-observer differences. We developed and evaluated methods for training machine learning models to score PVS incidence in the basal ganglia (BG) and centrum semiovale (CSO) leveraging the Potters/Wardlaw scale. The novelty in our work lies in the use of imperfect, semi-automatically generated "silver-standard" PVS segmentation masks during training, in addition to PVS radiological scores. We comparatively evaluated a conditional convolutional neural network (CNN) which accepts PVS masks as an extra input channel, a multi-task CNN which performs both PVS segmentation and scoring, and a logistic regression model which utilises features derived from PVS masks to predict PVS scores. Multi-task learning was the most effective method, achieving a mean average precision of 64.08% compared to 60.22% for the conditional CNN, 52.11% for a baseline CNN trained only to predict PVS scores, and 49.32% for the logistic regression model. The multi-task model showed an ability to localise individual PVS not shown by the other CNNs, and behaved in a probabilistically sensible way, predicting with lower confidence on inherently harder classes. Age, sex, hypertension status, white matter hyperintensity volume, and ischaemic stroke lesion status were shown to be associated with the multi-task model's PVS score predictions and the ground truth in a similar way.
|
| 153 |
BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling
2609.37400
|
cs.CV
|
Xiaojian Shen, Dahu Shi, Jianrong Zhang, Hai Li, Hongwei Zhao |
Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they o...Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.
|
| 154 |
Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation
2609.37407
|
cs.CV
|
Xianghan Wei, Xiaoda Yang, Zhi Wang, An Pan, Daoan Zhang |
While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, ...While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target shot. To resolve this, we present Complementary Retrieval-Augmented Prompting, an agentic framework that strategically aggregates a compact set of mutually supportive historical references to achieve complete and targeted conditioning for long-form video generation without retraining or modifying the underlying generator. Specifically, our framework explicitly models the visual elements required by each target shot by parsing the narrative script into a text-grounded visual element registry that tracks characters, objects, scenes, and their shot-level states. A VLM-annotated keyframe library further maps these elements to past visual observations. Guided by the required elements, our agent retrieves complementary references that maximize target-element coverage while minimizing historical noise. Finally, the retrieved references, structured element states, and grounding instructions are assembled into a unified prompt for the frozen video generator. This element-aware process provides comprehensive conditioning while remaining fully interpretable. Quantitative and qualitative evaluations on multi-shot story generation demonstrate that our method consistently outperforms recent-frame, memory-based, and entity-level retrieval baselines in cross-shot consistency and text-controllability.
|
| 155 |
LazySloth: Bounded LLM-based Lazy Tree Search for Fast Long Video Comprehension
2609.37426
|
cs.CVcs.AI
|
Arka Mukherjee, Kaleen Shrestha, Larissa Zhu, Maja Matari\'c |
Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient...Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past work has explored efficient methods through multimodal retrieval-augmented generation (RAG), they rely on lossy embeddings that lose temporal context and fine-grained detail. Few works to date have investigated how VLM-based query-relevant information retrieval can be optimized. We introduce LazySloth, an efficient tree-based search method that speeds up video comprehension and retrieval tasks 2.9-8.3x (compared to existing agentic methods) through bounded captioning of portions of the video considered irrelevant by a VLM of the video. Compared to contemporary specialized video-understanding VLMs and RAG-based methods, LazySloth achieved similar or better final task accuracy across two recent open-source base VLMs--Gemma 4 31B and Qwen3.6 27B--across four benchmarks. LazySloth reduced the gap between the base open-source model and a closed-source model, GPT-4o. Ablations showed that replacing VLM scene understanding with CLIP-based retrieval cost 8.8-19.9% in accuracy, while lazy tree construction matches eager construction at a fraction of the captioning cost. With LazySloth, we demonstrate the possibility of faster long-video comprehension without substantial loss in performance.
|
| 156 |
Event-Only Wingbeat Counting under Camera Motion: A Controlled MuJoCo Benchmark
2609.37465
|
cs.CV
|
Zhang Nengbo |
Counting completed wingbeats requires identifying individual cycles, including during frequency changes and pauses; estimating a dominant frequency alone is insufficient. Camera motion further mixes target and background brightness changes in event observation...Counting completed wingbeats requires identifying individual cycles, including during frequency changes and pauses; estimating a dominant frequency alone is insufficient. Camera motion further mixes target and background brightness changes in event observations. We present a controlled MuJoCo benchmark that separates motion training from event-only image translation compensation. The acquisition contains 324 streams from 24 independent scenes, three flapping geometries, two distances (1.5 and 3.0 m), and static, moderate-motion and stronger-motion views. Fifteen scenes are used for fitting, three for validation and six for held-out testing. A fixed causal temporal convolutional network is evaluated in a matched 2 x 2 ablation with three initialization seeds and compared with ridge, Fourier, autocorrelation and an adapted EEPPR baseline. Under moderate motion, paired motion training reduces count mean absolute error from 31.130 to 3.185 cycles at 1.5 m and from 42.019 to 5.444 at 3.0 m. Adding the tested compensation increases these errors to 4.630 and 10.185, respectively. A Fourier baseline achieves 0.944 cycles at 1.5 m under moderate motion, showing that the neural model is not uniformly best. We report exact-count accuracy and temporally matched cycle F1 alongside count error. These findings support motion-aware training in this small synthetic benchmark, while exposing limits of simple event-background stabilization. They do not establish real-sensor performance, aerodynamic flight, or generalization to unseen vehicle types.
|
| 157 |
Label Less, Learn More: Resource-Efficient Active Semi-Supervised Learning for Onboard Satellite Image Annotation
2609.37481
|
cs.CV
|
Ahmed Abdelnaby, Mohamed Elmahallawy, Marius Bernahrndt, Tobias Hecking |
Large-scale pervasive sensing increasingly relies on high-resolution satellite imagery, yet task-specific onboard vision is constrained by costly annotation and limited computation, memory, energy, and communication resources. Existing approaches largely rely ...Large-scale pervasive sensing increasingly relies on high-resolution satellite imagery, yet task-specific onboard vision is constrained by costly annotation and limited computation, memory, energy, and communication resources. Existing approaches largely rely on either data-hungry supervised learning or large vision-language foundation models, limiting efficient adaptation and deployment under these constraints. We present SatLabel, a resource-aware learning framework that transforms limited satellite labels into progressively refined onboard models through adaptive sample acquisition and semi-supervised model adaptation. Rather than repeatedly training on uniformly sampled labels, SatLabel closes the loop between model uncertainty, class imbalance, and pseudo-label quality to selectively acquire informative samples while exploiting abundant unlabeled imagery. This enables a compact student to adapt to target sensing domains with reduced annotation and inference costs. We further introduce an optional Mixture-of-Experts (MoE) student with graph-based feature refinement to enhance representation capacity while retaining a lightweight footprint. We evaluate SatLabel on 11 remote-sensing datasets spanning core, extended, and unseen domains against RemoteCLIP zero-shot inference. SatLabel improves Macro-F1 on most core and extended datasets while maintaining strong cross-dataset transfer to unseen domains. More importantly, the Balanced student contains only 11.2 M parameters and occupies approximately 42.8 MB, compared with 151.3M parameters and 577 MB for RemoteCLIP, while requiring 3.65 versus 5.89 GFLOPs. Across four efficiency benchmarks, it achieves approximately 2x higher GPU-forward throughput and reduces energy per image on datasets.
|
| 158 |
PoE-Fuse: Precision-Weighted Expert Fusion for Bi-Temporal Change Understanding
2609.37485
|
cs.CV
|
Haruki Watase, Shunya Nagashima, Takayuki Nishimura |
Bi-temporal change understanding, which localizes and characterizes what changed between two satellite images, is central to disaster response and environmental monitoring, spanning change detection, building localization, and damage assessment. Strong vision-...Bi-temporal change understanding, which localizes and characterizes what changed between two satellite images, is central to disaster response and environmental monitoring, spanning change detection, building localization, and damage assessment. Strong vision-language models address these tasks, but adapting them typically requires full fine-tuning or reinforcement learning, which is costly and unstable. We propose PoE-Fuse, a parameter-efficient framework that instead composes frozen foundation experts for geometry, grounding, and language, resampling their features onto a shared spatial grid and training only a lightweight fusion trunk. PoE-Fuse treats the aligned features as Gaussian observations of a latent scene state and fuses them by learned per-cell precision. This product-of-experts estimator strictly generalizes uniform summation and scalar gating, and extends to change fields by composing the precisions of the two timestamps. A single shared trunk solves the three tasks at once, reaching a mean F1 of 59.2%, compared with 40.7% for an instruction-tuned temporal vision-language assistant, and surpassing dedicated change-detection models retrained under the same protocol and training budget.
|
| 159 |
Physics-Guided Flow-Map Matching for Precipitation Nowcasting
2609.37487
|
cs.CV
|
Shunya Nagashima, Takumi Bannai, Makoto Misaizu, Keisuke Maeda, Takahiro Ogawa |
Precipitation nowcasting, generating future radar fields from past observations, is critical for flood warning and disaster response. It is also a demanding benchmark for spatiotemporal generative modeling, with chaotic dynamics, heavy-tailed intensities, and ...Precipitation nowcasting, generating future radar fields from past observations, is critical for flood warning and disaster response. It is also a demanding benchmark for spatiotemporal generative modeling, with chaotic dynamics, heavy-tailed intensities, and rare high-intensity structures that matter most. Deterministic models minimize a pixel loss and are driven toward the conditional mean, which blurs exactly those structures, while generative models that add a stochastic residual on top of a deterministic backbone inherit the same blur. We propose Physics-Guided Flow-Map Matching (PG-FMM), a conditional flow-map model that decouples predictable advection from uncertain small-scale detail. A frozen Lagrangian advection prior transports the radar field and supplies an explicit motion forecast, and a flow-map generative head, conditioned on the past frames and the prior rollout rather than summed onto it, produces sharp stochastic detail in four sampling steps. The prior serves only as guidance, so the head replaces blurred structure instead of inheriting it. Extensive experiments on four radar benchmarks show that PG-FMM outperforms state-of-the-art methods on 18 of 24 metrics, with the largest gains at heavy-rain thresholds, where the critical success index improves by up to 58.9%. The project page can be found at https://neurogica.github.io/PG-FMM.
|
| 160 |
FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding
2609.37488
|
cs.CV
|
Taiyo Sato, Takamasa Sanda, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama |
Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and res...Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot shift the answer. Fusing three open MLLMs, FORUM surpasses the 397B-parameter published reference by a relative 5% in mean accuracy on the adversarial Ref-Adv-s benchmark, and a plain averaging ensemble by 15%. The gains transfer to standard RefCOCO+, and a balanced lineup with no dominant member still surpasses the 397B model by 5%.
|
| 161 |
Attention-Scoped Guidance: Training-Free Spatial Control for Image Editing
2609.37492
|
cs.CV
|
Zeyan Li, Wei Zhou, Hadi Amirpour, Minghao Zou, Panqi Yang |
Instruction-guided image editing should change what the instruction names and leave the rest of the image untouched. In dual classifier-free guidance (CFG), an editor combines two directions at every denoising step, one that pushes toward the instructed edit a...Instruction-guided image editing should change what the instruction names and leave the rest of the image untouched. In dual classifier-free guidance (CFG), an editor combines two directions at every denoising step, one that pushes toward the instructed edit and one that pulls back toward the source image, using global weights. We introduce Attention-Scoped Guidance (ASG), a sampler wrapper that makes these weights spatial. It reads a soft support map from the instruction attention that the editor already computes, then weakens text guidance where support is low and strengthens image anchoring where support is high. The wrapper requires no training, no external mask, and no additional network evaluation. On the full MagicBrush and PIE-Bench++ splits, ASG improves preservation-oriented metrics, leading three of four MagicBrush metrics and PIE-Bench++ background PSNR. A dose-matched control that removes the spatial placement loses up to 0.73 CLIP on PIE-Bench++, confirming that the spatial allocation itself carries the gain.
|
| 162 |
MotionMaestro: Masked Tokenization for Unified Motion Generation
2609.37495
|
cs.CV
|
Yun Chen, Munchurl Kim, Jeonghyeok Do |
Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including...Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including text-to-motion, pose-conditioned generation, and trajectory control. Although these tasks involve different types of conditions, a unified framework capable of handling them within a common representation would greatly simplify motion generation systems. We observe that diverse motion conditions can be naturally formulated as different observation patterns over motion sequences, where each task corresponds to a specific masking strategy. Based on this insight, we introduce MotionMaestro, a unified motion generation framework that learns a shared representation for complete motions and heterogeneous partial observations through masked motion tokenization. MotionMaestro employs a three-stage training strategy that first learns a masked motion tokenizer, then refines its reconstruction ability on clean motions, and finally trains a conditional flow-matching generator in the learned latent space. Furthermore, we introduce an observation map and an observation loss to explicitly preserve provided motion conditions during generation. With this unified representation and conditioning mechanism, MotionMaestro supports text-guided and unconditional synthesis, pose conditioning and partial completion, temporal interpolation, trajectory control, and motion continuation. Experiments on the large-scale RoMo and MotionMillion datasets show state-of-the-art performance across diverse motion generation tasks.
|
| 163 |
GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation
2609.37496
|
cs.CV
|
Jeonghyeok Do, Munchurl Kim |
Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale datase...Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing conditions. We introduce GeoSET, the first generalist model for SET, built around a single pretrained parent that is adapted to downstream datasets under a common protocol. We curate over 3 million high-quality SAR--EO pairs from a collection of more than 10 million SAR observations, spanning diverse sensors, spatial resolutions, and ground sampling distances. To bridge the modality gap between SAR observations and a pretrained image generator, we develop a speckle-robust SAR encoder and pretrain the conditional generator on this heterogeneous corpus. The resulting parent supports efficient adaptation across downstream datasets through low-rank adaptation (LoRA), updating only 0.60% of the generator parameters and requiring approximately one hour per dataset. Across six downstream benchmarks, GeoSET achieves state-of-the-art results in FID and DISTS with full fine-tuning or LoRA, demonstrating effective transfer across heterogeneous SAR-EO domains.
|
| 164 |
Principled MAP estimation for inverse problems: bridging the gap between convergence and performance
2609.37529
|
cs.CV
|
Alexandre Lagier, Valentine Tosel, Anne Gagneux, Mathurin Massias, S\'egol\`ene Martin |
Pretrained denoisers provide a powerful way to incorporate image priors into restoration algorithms. Plug-and-Play and RED approaches exploit fixed-noise-level denoisers within first-order optimization schemes, with convergence guarantees, but often struggle t...Pretrained denoisers provide a powerful way to incorporate image priors into restoration algorithms. Plug-and-Play and RED approaches exploit fixed-noise-level denoisers within first-order optimization schemes, with convergence guarantees, but often struggle to achieve high-quality reconstruction on severely ill-posed inverse problems. In contrast, recent state-of-the-art approaches leverage denoisers derived from flow- or diffusion-based generative models and evaluate them along a sequence of decreasing noise levels. While these methods achieve strong empirical performance, their convergence theory remains limited. In this paper, we bridge this gap by specifically designing an algorithm that combines denoisers at decreasing noise levels with a schedule tailored to ensure convergence. From a Bayesian perspective, we prove that our method converges to a $\textit{Maximum a Posteriori}$ (MAP) estimate, under suitable assumptions. Subsequently, we apply our method to various ill-posed inverse problems and show that it surpasses convergent methods while competing with state-of-the-art empirical ones.
|
| 165 |
Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models
2609.37537
|
cs.CV
|
Arian Komaei Koma, Seyed Amir Kasaei, Aida Aryafar, Matin Ghiasi, Ali Aghayari |
Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack...Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack of robustness to noise initialization. We call this phenomenon ``probabilistic forgetting'': suppressed concepts re-emerge under specific random initial noise conditions, despite appearing unlearned on other initializations. We trace this failure to the misalignment between standard Gaussian sampling during unlearning and the unlearning objective. Since the target concept manifests only in specific initial noise regions throughout the unlearning phase, uniform random sampling yields sparse, uninformative gradient updates that fail to drive robust erasure. To overcome this issue, we propose an adaptive, concept-conditioned sampling strategy that dynamically concentrates gradient updates on regions where the target concept manifests, down-weighting uninformative areas. We integrate our framework with six distinct SOTA unlearning methods across four diffusion backbones and evaluate it across safety, object, and artistic-style unlearning, as well as under black-box and white-box adversarial attacks. Our method reduces the conditional nudity re-emergence rate across random initializations by 67.2% on average over four baselines and lowers attack success rates across both adversarial evaluations. Across concept domains, Adaptive Noise Sampling strengthens adversarial robustness and non-target retention while preserving competitive generative quality and target-erasure performance.
|
| 166 |
APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
2609.37559
|
cs.CV
|
Jianguo Huang, Jinming Liu, Qiyao Wang, Liang Xu, Jianhang Li |
To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that re...To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.
|
| 167 |
Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering
2609.37569
|
cs.CV
|
Yazhen Xie, Xingsong Ye, Zhineng Chen |
Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph con...Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.
|
| 168 |
Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment
2609.37576
|
cs.CV
|
Yu Zhao, Jiarui Wang, Huiyu Duan, Ye Zhao, Jutao Tang |
With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adop...With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study semantic understanding, quality perception, and authenticity identification in isolation, while largely neglecting responsibility detection. This leaves a gap in unified and comprehensive validation. To bridge this gap, we introduce SQUARE-Bench, a comprehensive benchmark that systematically evaluates LMM capabilities as evaluators of AI-generated images across four aspects: Semantics, Quality, Authenticity, and Responsibility. SQUARE-Bench introduces a granular taxonomy of 38 sub-dimensions to evaluate nearly 10K AI-generated images sampled from 22 diverse models, ranging from legacy to state-of-the-art generators, complemented by over 3K real-world images. The images are annotated with curated question-answering pairs. Extensive experiments on 23 LMMs reveal that top proprietary models, such as Gemini-3-Pro, already outperform the individual human expert baseline. However, the performance gap between models remains significant, exhibiting notable disparities in fine-grained inference and domain-specific robustness. Beyond benchmarking, we conduct a proof-of-concept study of LMM-guided iterative editing, in which dimension-specific LMMs provide diagnostic feedback to fixed image editors. The resulting guided system yields selective improvements in semantics, authenticity, and responsibility, while exhibiting a consistent visual-quality trade-off. SQUARE-Bench can serve as both a diagnostic tool for characterizing LMM evaluator capabilities and studying their use in T2I generation refinement. The benchmark and dataset will be released upon publication.
|
| 169 |
TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning
2609.37581
|
cs.CVcs.AI
|
Jing Wang, Zhiping Wu, Dongdong Ren, Youfang Han, Wei Zhao |
Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visu...Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consistently improves performance over vision-only saliency-based pruning. We further find that high-variance attention heads are more sensitive to the textual query and yield more discriminative text-to-vision attention signals for second-stage pruning. Motivated by these findings, we propose TReVS, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM. On LLaVA-1.5-7B, TReVS retains 92.8% of the unpruned baseline performance while pruning 94.4% of visual tokens, outperforming prior state-of-the-art methods.
|
| 170 |
FedSocket: Recipient-Executable Knowledge Exchange for Heterogeneous Multimodal Federated Learning
2609.37582
|
cs.CVcs.AI
|
Xinyuan Zhao |
Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs...Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs, task-owned outputs, and ownership-aware aggregation, connecting heterogeneous private models through a common prediction interface. Private models teach local Q copies; the returned Q supports local learning and Joint inference, with only Q parameters and counts exchanged. Across six datasets, FedSocket improves missing-modality recipient accuracy over Local by 14.44 and 15.51 percentage points on MELD and UCF-51. Under matched inference capacity, Joint exceeds independent ensembles by 11.06 points in UCF-51 accuracy and 4.87 points in mean bidirectional Flickr30k R@1. Joint also improves over Q alone on all four heterogeneous endpoints, demonstrating the value of combining local and exchanged predictions. Teacher controls, sharing-path interventions, and component factorials identify the roles of supervision, sharing, and deployment. FedSocket makes exchanged knowledge directly usable from federated training to recipient inference.
|
| 171 |
TomoTransformer: Towards a Foundation Model for CT Reconstruction
2609.37605
|
cs.CV
|
AmirEhsan Khorashadizadeh, Benjam\'in B\'ejar |
Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they req...Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emph{back-projection space} that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.
|
| 172 |
Targeted Visual Counterfactual Explanations for Contrastive Vision-Language Model
2609.37638
|
cs.CV
|
Van Bach Nguyen, J\"org Schl\"otterer, Christin Seifer |
Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbf{M}ask-guided \textbf{A}daptive \textbf{C}ounter...Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbf{M}ask-guided \textbf{A}daptive \textbf{C}ounterfactual \textbf{E}xplanations (\mace), a targeted visual counterfactual method designed specifically for CLIP zero-shot classification. \mace constructs an editable region from either source attribution or source--target attribution differences and expands the mask only when needed to reach a specified target class. A latent diffusion inpainting model then modifies the selected region, while a frozen CLIP model provides modification guidance and anchors the remaining image content to the original input. We evaluate \mace on ImageNet, Food-101, Oxford Pets, and CUB-200. The source-mask variant achieves the highest target top-1 success rate across all four datasets, while the difference-mask variant produces the smallest pixel-level and perceptual changes and the best realism scores. Both variants improve proximity and realism over a Stable Diffusion-only baseline using the same generative backbone. These results show that adaptive mask-guided editing produces effective CLIP counterfactuals. They further reveal a tradeoff between counterfactual validity and source-image preservation.
|
| 173 |
VoxelSage: Tool-Augmented 3D CT Analysis and Simulator-Shielded Sequential Resection Planning for Liver Tumors
2609.37648
|
cs.CV
|
Binghong Qian, Xuanhe Liu, Yifan Xing, Wenjie Deng, Jian Wu |
Preoperative liver-tumor assessment requires segmentation, physical-space measurement, visual evidence, and resection planning from the same three-dimensional CT volume. Existing tools often handle these steps separately, while language models cannot reliably ...Preoperative liver-tumor assessment requires segmentation, physical-space measurement, visual evidence, and resection planning from the same three-dimensional CT volume. Existing tools often handle these steps separately, while language models cannot reliably compute physical measurements from CT. To provide an integrated workflow, we present VoxelSage, a multi-modal system for two- and three-dimensional visualization, liver-tumor analysis, and preoperative resection planning. Its dual-port architecture separates language-model orchestration from image computation: Port A interprets requests and selects skills, while Port B applies them to CT volumes and segmentation masks and returns structured results. Keeping physical measurements in Port B prevents the LLM from computing them directly and reduces the risk of fabricated numerical results. Eight built-in skills support quantitative analysis, visual evidence generation, three-dimensional reconstruction, segmentation refinement, and sequential resection planning; user-defined skills can extend these functions. For sequence planning, a behavior-cloned neural ranker orders candidate resection targets, while a simulator-based shield checks them against predefined constraints. Across 256 unseen simulator scenes, this approach reduced mean simulated time from 34.274 to 33.388 min (0.886 min, 2.59%) and mean simulated blood loss from 300.847 to 183.852 mL (116.995 mL, 38.89%) relative to a deterministic baseline. These results demonstrate system integration and simulator-level performance, not clinical efficacy or safety. The public implementation is available at https://github.com/ZJUMAI/VoxelSage.
|
| 174 |
Texture Space Material Diffusion
2609.37654
|
cs.CV
|
Jacob Munkberg, Peter Kocsis, Jon Hasselgren |
We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use...We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, enabling the diffusion process to generalize across arbitrary geometries and texture parameterizations. This approach also avoids the view consistency issues inherent in video and multi-view diffusion models. Because texture space is two dimensional, we can reuse the strong priors of pretrained video diffusion models. We apply our method to high quality material reconstruction from posed photos captured under unknown lighting, as well as to text- and image guided material generation. Our method can scale to high resolutions (8K), 100+ input views, and neural material representations. In quantitative and qualitative evaluations we show state-of-the-art results for material generation and reconstruction.
|
| 175 |
Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding
2609.37655
|
cs.CV
|
Jiayu Ying, Qijian Tian, Ruijie Xu, Xinnan Zhu, Daoguo Dong |
Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often...Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at https://github.com/yingjiayu12/Exemplar2VQA
|
| 176 |
Tracing the Evidence: Faithful Token Attribution Through Vision-Language Reasoning
2609.37656
|
cs.CV
|
Bowen Yuan, Danny Wang, Ruihong Qiu, Zijian Wang, Zi Huang |
Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response by assigning scores that ra...Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response by assigning scores that rank image and prompt tokens by how much the model relies on them, such that removing higher-ranked tokens causes the likelihood of the generated response to drop more rapidly. However, existing token-attribution methods have been developed mainly for text-based language models, and our empirical study reveals two challenges when complex multimodal sources are involved. First, the joint image-text attribution can underrepresent visual evidence relative to text, obscuring the image regions supporting the response. Second, visual evidence may influence the generated response through multiple intermediate reasoning paths, while existing methods trace only a limited subset of these paths, causing important visual contributions to be underestimated. Motivated by these insights, we introduce VTrace, a multimodal token-attribution framework that traces input contributions through intermediate reasoning and calibrates attribution scores across modalities. VTrace constructs pairwise attributions that highlight token-specific contributions and aggregates all forward attribution paths in closed form to account for both direct and indirect contributions. Cross-modal calibration then rescales image and text attribution scores using modality contributions estimated from response-likelihood changes, enabling a unified ranking of input tokens. Evaluations against seven baselines across six visual reasoning benchmarks demonstrate the superior attribution faithfulness. Project page: https://vtrace-attribution.github.io/.
|
| 177 |
Are In-Context Images Worth 10 Dimensions?
2609.37659
|
cs.CV
|
Adhemar de Senneville, Xavier Bou, J\'er\'emy Anger, Rafael Grompone, Gabriele Facciolo |
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled exam...There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the vision modality compared to text, we uncover a Shared Discriminative Geometry (SDG) inside Large Vision Language Models (LVLMs). It is a low-dimensional space, shared across all image classification tasks, in which in-context images are compressed into linearly separable representations later used to perform classification. We observe that this is the result of the model performing a dimensionality reduction of vision representations in early layers. In order to explain this phenomenon: (1) We show analytically that linear self-attention can perform a dimensionality reduction by projecting in-context data onto its principal components, with each layer implementing one gradient descent step toward this objective. (2) We provide evidence that trained LVLMs reduce the dimensionality of vision representations in early layers via a similar mechanism.
|
| 178 |
Med-RADIO: Reducing All Medical Domains Into One via Multi-Teacher Distillation
2609.37682
|
cs.CV
|
Chu Zhang, Haoyu Jiang, Hongyuan Zhang, Hongbin Liu, Dong Yi |
The rapid expansion of large-scale medical datasets and computational resources has driven significant progress in medical foundation models. Given the inherent heterogeneity of medical imaging modalities, current research mainly follows two paths: specialized...The rapid expansion of large-scale medical datasets and computational resources has driven significant progress in medical foundation models. Given the inherent heterogeneity of medical imaging modalities, current research mainly follows two paths: specialized models optimized for specific modalities, and generalist models designed to handle multiple modalities. However, medical generalist models suffer from both insufficient training data scale relative to natural image generalists and inadequate domain-specific depth relative to medical specialists. Empirically, generalist models establish a cross-modality performance baseline, while specialists define the performance ceiling within their respective domains. To elevate this baseline toward these ceilings, we propose Med-RADIO, a medical multi-teacher distillation framework that Reduces All Domains Into One by compressing complementary expertise from multiple domain-specific teachers into a unified medical vision foundation model. Our method curates both generalist and specialist teachers, allocates modality-aligned distillation streams to reorganize generalist pretraining data so it matches specialist domains, and uses a balanced loss to prevent any single teacher from dominating the distillation process. On internal and external classification benchmarks spanning five modalities, Med-RADIO improves over strong medical generalists under linear probing and remains competitive with representative specialists on most evaluated modalities. Code is available at https://github.com/CAIR-HKISI/Med-RADIO.
|
| 179 |
PAIQ: Patch-Aligned Semantic Injection via Residual Rotation
2609.37685
|
cs.CV
|
Pinze Ren, Yuwei Zhang, Hao Chen, Linghao Meng, Chang Li |
Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Harnessing this complementarity requires enriching local features while retaining distinctions between semantically related patches. ...Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Harnessing this complementarity requires enriching local features while retaining distinctions between semantically related patches. We introduce PAIQ, a patch-aligned semantic injection framework that combines content-based cross-encoder matching with orthogonally constrained residual updates. Using DINOv3 patch features as the spatial base, PAIQ aggregates complementary SigLIP features through joint source allocation and injects the aggregate--base differences through a shared orthogonal transformation Q. This rotation adapts update directions while preserving residual norms and pairwise angles. For fixed projected features, we derive conditions for patch separability under similar semantic aggregates and show that rotation adds a nonnegative separation term over direct interpolation when the aggregate is shared. Only the projection and fusion parameters are trained; both visual encoders and the language model remain frozen, and fusion retains 196 visual tokens. Across diverse language backbones, PAIQ yields broad gains in judge-assessed correctness and reductions in hallucination severity over single-encoder interfaces on image description and visual question answering. On the 2B and 9B Qwen backbones, this compact interface outperforms the strongest evaluated fusion or token-compression baselines by about 2.9 correctness points on average.
|
| 180 |
Honeycomb: Constant-Size Scene Memory Representation for Video World Models
2609.37690
|
cs.CV
|
Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen, Yufeng Weng, Keane Ong |
Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memory systems accumulate RGB observations or latent features, causing storage requirements to grow as generation proceeds. We int...Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memory systems accumulate RGB observations or latent features, causing storage requirements to grow as generation proceeds. We introduce **Honeycomb**, a video world model built on **HexMemory**, a compact low-rank representation that stores scene features in a fixed-size memory comprising six spatial and spatiotemporal planes. A feed-forward writer maps each newly generated video chunk to plane features. As the spatial coverage or temporal range expands, HexMemory warps the existing planes while preserving their dimensions, then integrates new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latent features from HexMemory to condition subsequent video generation. Because the writer processes only observations from the latest chunk, Honeycomb avoids per-scene optimization and repeated processing of the full generation history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust consistency when revisiting previously observed regions, while maintaining constant feature-storage requirements throughout generation. Code and additional visualizations are available on our https://jackswl.github.io/honeycomb/.
|
| 181 |
VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation
2609.37709
|
cs.CVcs.AI
|
Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta |
Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmar...Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose), and (iii) controlled comparison of visual instructions with text descriptions at different levels of specificity. Using these capabilities, we uncover three findings: (1) models face an adherence-artifact trade-off: once models reach stronger visual instruction adherence, stronger adherence tends to coincide with more instruction artifacts in generated images, (2) visual instruction adherence tends to be lower on tasks whose reference images carry a salient state of the controlled attribute (e.g., a neon-lit subject under a light-direction instruction), most consistently for light and wind, and (3) for models that can understand visual instructions, it is often better to provide visual constraints directly rather than describe them in text; when using text, a moderate level of detail works better than an exhaustive description. VIF-Bench is released as an open benchmark to establish a basis for fair comparison in controllable multi-reference image generation.
|
| 182 |
PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence
2609.37712
|
cs.CVcs.AI
|
GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang |
Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often e...Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of varying scales. PolyOCR combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into quality-verified OCR supervision. We introduce Competence-Guided Policy Optimization, which combines verifier-based Group Relative Policy Optimization with on-policy distillation through sample-wise routing based on teacher reliability and the teacher--student competence gap. We also introduce OCRBench v2.1, our revision of OCRBench v2 with manually verified annotation corrections and task-aligned scoring metrics. Extensive experiments across OCRBench v2.1, CC-OCR, in-house KIE Benchmark, OmniDocBench v1.6 and MDPBench demonstrate that PolyOCR achieves state-of-the-art or highly competitive performance.
|
| 183 |
The Camera Inside the Editor: Reading the Implicit Camera of Image Editors with Painted Calibration Patterns
2609.37732
|
cs.CV
|
Sebastian R\"uckerl |
Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from whi...Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from which classical vanishing-point geometry reads pitch, roll, focal length, yaw and, on renders, the principal point, without any training. Unlike a calibrator such as GeoCalib, which estimates the camera of an image, this isolates the camera under which the editor paints. On 120 rendered cameras with exact ground truth, Qwen-Image-Edit-2511 paints tile edges that meet their vanishing points within 0.26 degrees, and its implicit camera matches the true one to 0.8 degrees in pitch and 6% in focal length, more accurately than GeoCalib except in roll. Asked to draw the horizon or mark a vanishing point instead, the editor fails, so this knowledge is revealed by painting and not by the explicit tasks we tried. The implicit camera has two priors: roll is pulled towards level (slope 0.71), and telephoto perspective towards a default of about 30 mm, which roughly matches the camera the models paint without any scene. For Qwen, the priors do not grow when blur removes four fifths of the line evidence. They are stronger on real photographs, and on NYUv2 a shorter wording of the task removes the difference for roll. On photographs from a 24--240 mm zoom lens the painted perspective grows with only 0.62 of the lens's slope, while GeoCalib and MoGe-2 saturate at about 52 and 42 mm. FLUX.1 Kontext and LongCat-Image-Edit are pulled much harder. Finally, from a level camera a camera-control LoRA executes pose commands at only 50--70% of their strength, and a board painted into its output agrees with the camera it produced.
|
| 184 |
Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection
2609.37750
|
cs.CVcs.AI
|
Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede, Wasif Bala, Beatrice Brown-Mulry |
Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial pl...Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.
|
| 185 |
HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
2609.37775
|
cs.CVcs.AI
|
Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang |
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for recon...Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.
|
| 186 |
A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System
2609.37783
|
cs.CVcs.AI
|
Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani |
Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish ...Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces. Existing image-forensics benchmarks provide important resources for face manipulation, classical tampering, and general synthetic-image detection, but they are not organized around the forms of visual evidence submitted in courts, the localized edits that can change what an exhibit appears to prove, or the consumer-tool threat model now facing the justice system. We introduce the CIFAR Synthetic Evidence Corpus for Detecting AI-Manipulated Images, a benchmark for evidentiary image authentication in court and justice-system contexts. The corpus contains 1,505 photographic items, including 720 authentic controls and 785 manipulated or fabricated images, spanning surveillance, dashcam, and consumer-photo imagery. Manipulations are organized into scene-condition edits, localized element edits, and full fabrications produced with contemporary generative systems. Each item is released with structured metadata covering source provenance, manipulation tier, subtype, generator, prompt template, and scene attributes, enabling controlled evaluation beyond aggregate binary detection. We also establish baselines with publicly available image-manipulation detectors, showing that current systems exhibit error profiles that remain problematic for evidentiary use. The dataset, prompts, metadata, code, and baseline evaluation scripts are released to support research on visual evidence authentication, information integrity, and trustworthy AI for the justice system.
|
| 187 |
Planetary Feature Fields are Scalable Earth Representations
2609.37784
|
cs.CV
|
Arjun Rao, Sebastian Loeschcke, Anthony Fuller, Isaac Corley, Nico Lang |
Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spati...Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Fields (PFFs), which exploit redundancy across data products by modeling them jointly as continuous functions of space and time at planetary scale. PFFs are spatially local explicit-implicit (hybrid) neural fields. Each field shares a factored feature volume---a decomposition of an explicit 3D grid with smaller factors---across products, while lightweight implicit decoders reconstruct individual products across multiple timesteps. PFFs reconstruct EO products over space and time more accurately than single-product fields at matched compression rates. At $1800\times$ compression relative to the uncompressed source data, reconstructed features retain approximately $90\%$ or more of the performance achieved with the original features on pixel-level segmentation, change detection, and patch-level classification tasks. PFFs can add new timesteps by extending their factored feature volumes and add new products by attaching new decoders, while leaving existing outputs unchanged. PFFs reduce end-to-end feature access latency by an order of magnitude relative to evaluated API and cloud-storage pipelines.
|
| 188 |
CHOQOLATE: Organizing Concept Bottleneck Latent Spaces with Choquet Integrals
2609.37786
|
cs.CV
|
R\'emi Kazmierczak, Johanne Cohen, Marianne Clausel |
Concept Bottleneck Models (CBMs) built on vision-language models such as CLIP represent a latent space as human-understandable concepts. These representations are unfaithful: related concepts are entangled, so individual scores do not reflect their intended me...Concept Bottleneck Models (CBMs) built on vision-language models such as CLIP represent a latent space as human-understandable concepts. These representations are unfaithful: related concepts are entangled, so individual scores do not reflect their intended meaning. We propose CHOQOLATE, an interpretable-by-design layer based on 2-additive Choquet integrals, which merges correlated concepts into compact nodes. Across four datasets, CHOQOLATE achieves a favorable accuracy-interpretability trade-off, with weight-sparse and semantically coherent nodes. A closed-form gradient derivation, backed by experiments, explains why Choquet layers drive this organization without explicit supervision. Choquet weights also map directly to Shapley values, which enables test-time intervention. On standard bias-mitigation benchmarks, suppressing spurious concepts after training performs on par with methods that require group annotations or retraining, while needing neither.
|
| 189 |
ByteTraX: Enhancing the ByteTrack Architecture with Optimised Thresholding
2609.37801
|
cs.CV
|
Thomas A. O'Shea-Wheller |
The ByteTrack algorithm is a widely used and computationally efficient multi-object tracking architecture. Its core innovation lies in the combination of lenient bounding box associations with tracklet similarity matching to robustly deal with object occlusion...The ByteTrack algorithm is a widely used and computationally efficient multi-object tracking architecture. Its core innovation lies in the combination of lenient bounding box associations with tracklet similarity matching to robustly deal with object occlusions. However, this strategy is nevertheless vulnerable to erroneous track reclassification and identity switching, as detection confidence scores dictate association priority. To address this, I present a simple enhancement of the ByteTrack architecture, named ByteTraX, that optimises track continuity via a single unified matching threshold, while penalising identity switches through stringent track initiation criteria. This approach achieves consistently improved performance across a range of diverse benchmarks including GMOT-40, LC-MOT, SportsMOT, TeamTrack, DAMUNT, and DeepSea-MOT, while simultaneously increasing processing speed by >10%. Specifically, results demonstrate a >40% reduction in identity switches, accompanied by mean increases in HOTA of 3.6, IDF1 of 5.6, and FPS of 6.3. As such, adoption of the ByteTraX algorithm has the potential to substantially enhance tracking performance over the ByteTrack baseline, while retaining the efficiency needed for real-time deployment. To facilitate usage, I provide the source code, integration functionality for the YOLO family of object detection models, and deployment instructions via an open source repository.
|
| 190 |
Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study
2609.37809
|
cs.CVcs.AI
|
Sven Ligensa, Jan Pauls, Karsten Schr\"odter, Ibrahim Fayad, Fabian Gieseke |
Predicting canopy height from medium-resolution satellite imagery is a common and scalable approach for assessing the condition of the world's forests, which play a crucial role in climate change mitigation. While Transformer-based architectures have shown str...Predicting canopy height from medium-resolution satellite imagery is a common and scalable approach for assessing the condition of the world's forests, which play a crucial role in climate change mitigation. While Transformer-based architectures have shown strong performance in many domains, their straightforward application to dense (i.e., pixel-level) regression tasks often yields suboptimal results. In particular, the patch size has a crucial impact on the model performance. In this work, we consider pixel-level attention schemes and show that the resulting models generally outperform those relying on larger patch sizes. However, pixel-level attention can be a prohibitively resource-intensive operation. For this reason, we conduct an extensive experimental study using efficient attention variants to identify favorable trade-offs between prediction quality and resource requirements, facilitating the practical deployment of the proposed models. In addition, we perform a comprehensive comparison with several well-established models in the field and show that, with suitable hyperparameter choices, Transformer-based architectures can outperform competing approaches. Our findings provide practical guidance for designing models for pixel-level regression tasks on medium-resolution satellite imagery, including canopy height and biomass estimation, soil moisture mapping, and yield forecasting.
|
| 191 |
WINGS: Reference-Free Gaussian Splatting Inpainting with 3D-Native Generative Priors
2609.37816
|
cs.CV
|
No\'e Lallouet, Michael Fischer, Elie Michel |
Inpainting 3D Gaussian Splatting scenes, a key challenge in 3D editing, requires generating plausible content within a masked region of 3D space. Prior approaches rely on 2D diffusion models to produce one or several inpainted reference views, making them susc...Inpainting 3D Gaussian Splatting scenes, a key challenge in 3D editing, requires generating plausible content within a masked region of 3D space. Prior approaches rely on 2D diffusion models to produce one or several inpainted reference views, making them susceptible to challenges associated with multi-view inconsistency and lengthy optimization times. Departing from these approaches, we introduce a reference-free Gaussian splatting inpainting method operating natively in 3D. Our method leverages the embedding space of a large, pre-trained 3D prior, combined with a structure completion network to feed a generative prior which reconstructs the missing region's geometry and appearance. Performing content generation entirely in 3D, it avoids the need to reconcile inconsistencies of multiple inpainted reference images, and is faster than related 2D-based methods. We demonstrate the effectiveness of our method qualitatively and quantitatively, through extensive experiments and a user study. To the best of our knowledge, this work is the first Gaussian splatting inpainting method to operate in the learned representation space of a 3D-native generative prior without relying on inpainted reference views.
|
| 192 |
Minkowski Attractor Networks: Closed-Form Hyperbolic Flows for Visual Representations
2609.37817
|
cs.CV
|
Zhongping Ji |
Geometric representation learning predominantly scaffolds representations onto flat Euclidean subspaces or compact product tori ($\mathbb{T}^K$). However, flat manifolds possess vanishing curvature and polynomial volume growth, inherently suffering from metric...Geometric representation learning predominantly scaffolds representations onto flat Euclidean subspaces or compact product tori ($\mathbb{T}^K$). However, flat manifolds possess vanishing curvature and polynomial volume growth, inherently suffering from metric distortion when embedding multi-scale, tree-like visual hierarchies. While hyperbolic spaces ($\mathbb{H}^m$) circumvent this via constant negative curvature ($K<0$) and exponential volume expansion, prior hyperbolic deep architectures are hindered by computationally cumbersome Riemannian optimization, non-linear gyrovector calculus, and floating-point instabilities. In this work, we introduce \textbf{Minkowski Attractor Networks (MAN)}, an operator-splitting-inspired framework that embeds representations within pseudo-Riemannian Minkowski spacetime ($\mathbb{R}^{1,m}$). By framing hyperbolic manifolds as quadric level sets, MAN resolves hyperbolic geometry by combining linear Lorentz group transport with non-linear cone lifting and closed-form radial rescaling, evaluating in a single forward pass without numerical ODE solvers or iterative retractions. We establish \textbf{MAN-2D} ($\mathbb{R}^{1,1} \to \mathbb{H}^1$) as our primary, high-throughput visual backbone, which maximizes channel factorization granularity into $D/2$ independent two-dimensional Minkowski blocks. We further formulate \textbf{MAN-4D} ($\mathbb{R}^{1,3} \to \mathbb{H}^3$) as a spacetime extension, leveraging a commuting Cartan-subalgebra parameterization of $\mathrm{SO}^+(1,3)$ to evaluate 4D Lorentz isometries via two commuting 2D planar maps without matrix-exponential overhead.
|
| 193 |
ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing
2609.37831
|
cs.CV
|
Xijun Wang, Xin Li, Suhang Yao, Zirui Lang, Bingchen Li |
Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two ...Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model's own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At $1080{\times}1920$ output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72$\times$ faster while using 38.0\% less peak allocated memory than FlashVSR Tiny. The code is available at https://github.com/kopperx/ReCaVSR.
|
| 194 |
Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study
2609.37848
|
cs.CV
|
Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala |
Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany p...Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC. In contrast, changing whether the backbone is frozen or fine-tuned changes AUROC by 0.044 on average, and changing the decision threshold changes balanced accuracy by 0.090 on average. The official test split is also measurably different from the training pool: a partition classifier distinguishes them at AUC 0.697, rising to 0.898 for normal radiographs. Most strikingly, a classifier using only file properties, with no image anatomy, reaches 0.992 balanced accuracy within the training pool but falls to 0.496 on the official test split. Validation-fitted thresholds and calibration also transfer imperfectly. These results show that a high benchmark score can support different conclusions when the split, training policy, threshold, metric, calibration, and uncertainty are not communicated with it. We end with a seven-item reporting recommendation in which each item is tied to an effect measured in the study
|
| 195 |
RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution
2609.37850
|
cs.CV
|
Xijun Wang, Xin Li, Zirui Lang, Suhang Yao, Haoran Li |
Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative ...Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates reference latents for sparse keyframes, while a lightweight VSR network uses these references and low-resolution video to super-resolve every frame. The lightweight VSR network, implemented as a Dual-Memory Video Transformer, reuses keyframe information across frames and updates recent video context, supporting first-keyframe conditioning and dual-endpoint conditioning with bounded lookahead. However, errors in shared keyframes can propagate and accumulate across output frames, making keyframe quality alone an insufficient optimization target. We address this collaboration gap with Video-Aware Reference Optimization (VARO), which uses reinforcement learning to update the large generative model with two reward levels: a system-level reward evaluates videos produced by the fixed lightweight VSR network, while a reference-level reward evaluates decoded keyframe quality. VARO improves final video quality over direct joint training, and its dual-level rewards outperform a system-level reward alone. At 1080p on a single NVIDIA A100 80GB, dual-endpoint RelayVSR with a 15-frame keyframe interval reaches 29.29 FPS, 13.82 GB peak GPU memory, and 0.327 s first-frame model latency, compared with 7.80 FPS, 24.447 GB, and 2.83 s for FlashVSR-Tiny. The code is available at https://github.com/kopperx/RelayVSR.
|
| 196 |
FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators
2609.37851
|
cs.CV
|
Zhiqi Li, Bo Zhu |
Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separate...Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects local supervision to the deployed long-range map. Within this framework, we develop flow-map, induced-velocity, and instantaneous-velocity distribution supervision, each paired with a separately specified native flow-map rollout. Cross-capacity ImageNet experiments across three teacher rewards identify instantaneous-velocity distribution supervision with independently tunable student consistency as the most effective choice. In text-to-image experiments, FlowMap-OPD demonstrates strong multi-specialist consolidation capabilities and surpasses multi-reward Flow-Map GRPO in task performance and convergence speed.
|
| 197 |
HandAnthro: Automated Hand Anthropometry from a Single Image
2609.37855
|
cs.CVcs.AI
|
Fan Zhou, Shuairan Chen, Mengying Zhang, Yulin Wu, Sadegh Jafari |
Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph...Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph of a palm-up hand on US letter-size paper. The pipeline reconstructs wrist-occluded paper boundaries for rectification, whitens non-hand pixels, and refines 41 anthropometry-specific landmarks from a fine-tuned You Only Look Once (YOLO) pose model using image-specific geometry and contours. Controlled evaluation comprised 720 captures from 45 held-out participants, each contributing 16 images across two smartphones, two backgrounds, two angles, and two nominal illumination settings. HandAnthro produced complete outputs for 704 captures (97.8%); among these, mean absolute error (MAE) was 3.80 mm per dimension against two trained operators' caliper measurements. Regional MAEs were 2.48 mm for non-thumb fingers, 6.04 mm for thumbs, and 6.17 mm for palm and wrist. In a researcher-assisted mobile-app pilot, automated batch processing returned all 44 dimensions for 260 of 268 retained, researcher-screened firefighter images (97.0%). A descriptive, unpaired comparison with an independent national firefighter reference yielded a mean absolute difference of 2.40 mm across 28 sex-by-dimension group-mean contrasts. These results characterize controlled measurement performance and researcher-assisted field feasibility for future distributed hand-anthropometry studies.
|
| 198 |
Learning from synthetic photorealistic raindrop for single image raindrop removal
2609.37870
|
cs.CV
|
Zhixiang Hao, Shaodi You, Yu Li, Kunming Li, Feng Lu |
Raindrops adhered to camera lens or windshield are inevitable in rainy scenes and can become an issue for many computer vision systems such as autonomous driving. Because raindrop appearance is affected by too many parameters, therefore it is unlikely to find ...Raindrops adhered to camera lens or windshield are inevitable in rainy scenes and can become an issue for many computer vision systems such as autonomous driving. Because raindrop appearance is affected by too many parameters, therefore it is unlikely to find an effective model based solution. Learning based methods are also problematic, because traditional learning method cannot properly model the complex appearance. Whereas deep learning method lacks sufficiently large and realistic training data. To solve it, in our work, we propose the first photo-realistic dataset of synthetic adherent raindrops for training. The rendering is physics based with consideration of the water dynamic, geometric and photometry. The dataset contains various types of rainy scenes and particularly the rainy driving scenes. Based on the modeling of raindrop imagery, we introduce a detection network which has the awareness of the raindrop refraction as well as its blurring. Based on that, we propose the removal network that can well recover the image structure. Rigorous experiments demonstrate the state-of-the-art performance of our proposed framework.
|
| 199 |
EndoPrior-GS: Dynamic Endoscopic Reconstruction with a Joint Texture Prior
2609.37874
|
cs.CV
|
Jiaqi Huang, Shidong Wang, Tong Xin, Kabita Adhikari |
Dynamic endoscopic reconstruction is fundamental to robotic surgery and computer-assisted interventions. While 3D Gaussian Splatting (3DGS) realises real-time rendering, its application to deformable intraoperative environments remains constrained by spurious ...Dynamic endoscopic reconstruction is fundamental to robotic surgery and computer-assisted interventions. While 3D Gaussian Splatting (3DGS) realises real-time rendering, its application to deformable intraoperative environments remains constrained by spurious geometry and varying illuminations. To address these limitations, we introduce EndoPrior-GS, a novel pipeline that explicitly couples frame-extracted vision heuristics and estimated depth maps. EndoPrior-GS derives a joint texture prior from a tool-filtered valid tissue mask, a non-specular photometric filter, and anatomical structural salience, yielding a probability map that guides primitive initialisation and subsequent density control. The prior is further extended to the temporal domain through a texture-aware term that dynamically weighs pairwise primitive contributions during training. We conduct extensive experiments on benchmark datasets EndoNeRF and SCARED, and the obtained results show that our method EndoPrior-GS reduces Flow Error by 27.7% and 25.8% over the representative approaches while preserving competitive rendering quality and real-time rendering speed. Our project website is available at https://jiaqi-huang-77.github.io/EndoPrior-GS/.
|
| 200 |
Visual Branch is What You Need for CLIP-based Class-Incremental Learning
2609.37888
|
cs.CV
|
Tao Hu, Zhen-Hao Xie, Jingcai Guo, De-Chuan Zhan, Da-Wei zhou |
Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to ...Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual features.Motivated by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VISuses only base-session data to enhance CLIP's final visual representation with informative visual-layer features. Built on the enhanced visual representation, VISemploys a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VISaccumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VISachieves state-of-the-art performance without a textual branch.
|
| 201 |
ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning
2609.37889
|
cs.CV
|
Tao Hu, Zhinuo Zhou, Xialiang Tong, De-Chuan Zhan, Da-Wei Zhou |
Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by const...Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and reusable reasoning patterns for solving diverse instructions. For example, to answer "How many red cubes are to the left of the sphere?", domain knowledge can provide relevant concepts about objects and spatial relations, while reasoning knowledge can specify ordered operations such as object recognition, spatial filtering, and counting. Despite this potential, how to leverage external knowledge for continual adaptation remains largely unexplored in existing MCIT methods. To this end, we propose ReCAP, a retrieval-guided framework that leverages external knowledge to guide capability reuse during continual adaptation. At each continual stage, ReCAP uses external search and an LLM to incrementally build a knowledge base of domain, reasoning, and format knowledge based on the current-stage training data. For each instruction, retrieved domain knowledge guides generation, while retrieved reasoning knowledge selects and orders capability modules to form an instance-specific capability path. As these capability modules are reused across stages, subsequent adaptation can overwrite previously learned parameters. To enable stable cross-stage reuse, ReCAP introduces adaptive subspace recycling, which parameterizes reusable capability modules with shared bases and stage-specific cores, protects historically important directions while recycling residual capacity. Extensive experiments on MCIT benchmarks show that ReCAP achieves SOTA performance.
|
| 202 |
SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation
2609.37918
|
cs.CV
|
Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami |
Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SY...Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and provides 4,000 evaluation questions and 15,960 training questions over disjoint videos, spanning eight cross-video reasoning tasks. Visual ablations and human evaluation assess dependence on the supplied evidence and answer recoverability. Evaluation of 22 multimodal large language models reveals persistent difficulties in physical comparison and scene integration that increasing model size does not consistently resolve. Supervised fine-tuning raises Qwen3-VL-8B's average SYNCR accuracy from 32.6% to 61.6%, with gains extending to task configurations and video sources absent from training for those tasks. Transfer to real footage is most consistent for temporal ordering: accuracy improves by 9.0-20.5 percentage points on constructed Assembly101 and Panoptic ordering sets across three checkpoints spanning two model families and two model sizes, with additional gains on existing temporal reasoning benchmarks. These results establish SYNCR as a controlled setting for diagnosing cross-video reasoning failures, testing their learnability, and identifying where synthetic supervision transfers.
|
| 203 |
EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory
2609.37923
|
cs.CV
|
Ziyun Zeng, Hang Hua, Shaden Alshammari, Rogerio Feris, William T. Freeman |
Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model paramet...Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model parameters. EpiCon links question-level memory evolution to a persistent experience bank through two independently trained 2B models: a memory controller and a tree self-organizer. The controller jointly refines textual guidance and visual evidence across attempts and selectively includes visual memory. The self-organizer consolidates lessons hierarchically and retrieves experience and rules for new problems. We evaluate EpiCon on eleven benchmarks spanning four multimodal task domains, using two harnesses and multiple backbones. A frozen bank improves other systems even with a single solving attempt. A second harness raises the original system's macro-average score by 2.6 points across eleven benchmarks. Across four host configurations, EpiCon improves macro-average scores by 1.7 to 4.9 points over No Memory and reduces memory-operation time by 67\% to 74\% relative to backbone-sized memory models.
|
| 204 |
Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
2609.37925
|
cs.CVcs.AI
|
Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang, Weidong Zhang |
Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level d...Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in the surrounding context merely to preserve temporal consistency. To provide a clearer visual-quality signal, we introduce Rollout-Marginal Distillation (RMD). RMD retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context. To compensate for the lack of temporal context in independent chunk scoring, RMD subsequently applies video-level DMD to restore temporal coherence. Extensive experiments demonstrate that RMD maintains high visual quality far beyond its training horizon and outperforms video-level DMD baselines. Code and video results are available at https://cjeen.github.io/RMD
|
| 205 |
Look Closer: Patch-wise Supervision for AI-Generated Image Detection
2609.37937
|
cs.CV
|
Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao |
How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies...How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the reported patch-wise variants improve average accuracy over their whole-image counterparts across all four backbones. Comparisons of supervision granularity, source resolution, crop size, and inference coverage further characterize the approach, while post-processing tests and difficult-image evaluation reveal its limitations. The historical experiments include evaluation-based model selection, so their scores are not presented as a uniformly selected leaderboard comparison. Overall, the study identifies explicit local input and patch-level supervision as a simple, useful combination for investigating generalizable AI-generated image detection.
|
| 206 |
Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
2609.37938
|
cs.CVcs.AI
|
Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu |
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identit...Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at https://github.com/lei-qi-233/EgoGears.
|
| 207 |
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video
2609.37969
|
cs.CV
|
Haozhe Liu, Tian Ye, Shuchen Xue, Yitong Li, Junsong Chen |
High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a ...High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at $3840\!\times\!2176$ it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an $8.91\times$ speedup in refinement latency over the same baseline in our 2K latency setting.
|
| 208 |
ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals
2609.37986
|
cs.CV
|
Xuyi Hu, Francesco Palandra, Shangzhe Wu, Daniel Cremers, Riccardo Marin |
Recovering articulated 4D representations of animals from monocular videos remains challenging due to the large diversity of quadruped morphologies and lack of animal 4D supervision data. Existing learning-based reconstruction methods operate on individual ima...Recovering articulated 4D representations of animals from monocular videos remains challenging due to the large diversity of quadruped morphologies and lack of animal 4D supervision data. Existing learning-based reconstruction methods operate on individual images and rely on synthetic or model-fitted 3D supervision, which inherits the constraints of strong parametric priors and limits generalization to out-of-distribution species. When applied to out-of-distribution animals, they often recover a plausible pose while producing inaccurate geometry because the underlying shape model cannot faithfully represent the observed instance. We present ORMA, a training-free reconstruction framework that decouples articulation from shape, using the predicted pose as reference for optimization while leveraging generative 3D priors for accurate shape reconstruction. Given a reference image, we reconstruct the animal geometry and register it to the parametric model SMAL+, yielding an articulated shape adapted to the observed instance. We then combine per-frame articulated pose estimates with globally consistent camera poses to recover animal motion in a shared world coordinate frame, and further refine the reconstruction using self-supervised DINO correspondences and temporal consistency. To enable quantitative evaluation, we introduce PAW4D, a synthetic multi-species benchmark with ground-truth 3D geometry and camera motion. Experiments on PAW4D, PFERD, and challenging in-the-wild videos demonstrate that ORMA improves reconstruction accuracy while recovering globally consistend animal motion across diverse quadruped species.
|
| 209 |
HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents
2609.38008
|
cs.CV
|
Tongbo Chen, Junbo Niu, Zhengxi Lu, Niu Lian, Fei Tang |
Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions wit...Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specific APIs or tools, which require substantial engineering effort and are difficult to scale across applications. We argue that the next generation of CUAs should combine GUI interactions with the command line interface (CLI), leveraging the generality of the GUI and the efficiency of shell commands. A critical challenge, however, is that current models do not know when or how to use the CLI during task execution. To address this challenge, we develop a data construction pipeline that produces three types of trajectories: GUI only, CLI only, and interleaved GUI and CLI trajectories. This pipeline results in HybridCUA-8K, containing 5K hybrid trajectories and 3K verified RLVR tasks. Building on these data, we propose a training framework with two stages: supervised fine tuning on the constructed trajectories, followed by reinforcement learning with our CLI aware rewards that encourages agents to use the CLI selectively and reliably. Experiments show that HybridCUA-9B achieves 53.6% accuracy on OSWorld, improving over the base model by 14.8 percentage points, and improves performance on WindowsAgentArena by 4.0 percentage points. These results demonstrate the effectiveness and cross platform generalizability of the hybrid GUI and CLI paradigm for computer use agents.
|
| 210 |
From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection
2609.38010
|
cs.CVcs.AI
|
Mohamed Benkedadra, Aissa Saoudi, Maxime Gloesener, Sidi Ahmed Mahmoudi, Matei Mancas |
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic st...Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom $\Delta$-metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of $-50\%$ relative to real data, while CIA-only training shows a milder $-16.5\%$ degradation. Hybrid compositions significantly improve performance, with the 90\% real + 10\% Unity configuration achieving the best overall mAP@0.5 of $62.68\%$ ($+7.64\%$ over baseline), and the 90\% real + 10\% CIA configuration maximizing precision at $74.45\%$. Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.
|
| 211 |
Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation
2609.38019
|
cs.CV
|
Bangxun Tang |
We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, ye...We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, yet the mouth they render is an average mouth: the shape and texture of the lips, the arrangement of the teeth, and how much of them shows as the mouth opens are not that person's. The problem persists because nothing in current training or evaluation asks for the person's own mouth: perceptual losses accept any plausible mouth, face identity is carried mostly by the skin around it, and the released inference code of inpainting systems uses the unmasked target frame as the reference, which hides the gap. To address this, RGOR conditions every generated frame on frames from separate enrollment recordings of the same person and on HD patches of the mouth that bypass the VAE, and trains the generator against a paired judge that compares each rendered mouth with the person's reference and learns to reject a realistic mouth of someone else. We further build an evaluation protocol and use it to compare open-source and commercial lip-sync systems on held-out identities. Experiments show that RGOR achieves the best or second-best result on most metrics, and preserves the person's own lip and dental detail while keeping synchronization and the rest of the face intact.
|
| 212 |
EVO-WAM: Evolving World Action Models through Video-Action Verification
2609.38057
|
cs.CV
|
Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang |
Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source...Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately $2.5\times$ and $1.6\times$ their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.
|
| 213 |
RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA
2609.38072
|
cs.CV
|
Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan |
Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at infere...Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming pervious SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.
|
| 214 |
MUGEN: Interactive Panoramic World Exploration via Camera Control
2609.38077
|
cs.CV
|
Jiaming Tan, Zhen Li, Shuwei Shi, Minggui Teng, Siqi Yang |
Interactive panoramic video generation aims to synthesize immersive 360\textdegree{} videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: exist...Interactive panoramic video generation aims to synthesize immersive 360\textdegree{} videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: existing panoramic video datasets are often short, weakly annotated, or lack camera trajectories, while existing camera-controlled video generation models are designed for perspective videos and do not directly support panoramic geometry. In this paper, we introduce MUGEN and Wan360 to address these limitations. MUGEN is a large-scale real-world panoramic video dataset tailored to interactive 360-degree world exploration, comprising over 1,300 hours of at least 4K panoramic videos with rich semantic and geometric annotations. Built on MUGEN, we further present Wan360, a camera-controllable interactive panoramic video generation model. Panoramic videos are commonly represented by EquiRectangular Projection (ERP), which unfolds a spherical 360-degree view into a rectangular frame with cyclic longitude seams and pole distortions. To this end, Wan360 introduces three parameter-free ERP-aware components: periodic longitude RoPE for seam-consistent positional encoding, ERP-aware padding for reducing boundary artifacts, and random roll yaw for consistent learning. For camera control, Wan360 uses a panoramic Pl\"ucker embedding that represents camera motion with ERP rays rather than perspective pinhole rays. Experiments show that MUGEN serves as a data foundation for panoramic world exploration, and that Wan360 enables high-quality, temporally coherent, camera-controllable 360-degree video generation.
|
| 215 |
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
2609.38079
|
cs.CV
|
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han |
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generatio...Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.
|
| 216 |
VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents
2609.38086
|
cs.CV
|
Zheng Jiang, Houde Qian, Yiming Chen, Ling Li, Chaoyang Li |
Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, le...Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.
|
| 217 |
Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History
2609.38114
|
cs.CV
|
Weiqiang Wang, Zhuokun Chen, Yusheng Dai, Boying Li, Yi Zhang |
Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level...Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: https://anonymous.4open.science/w/self-aligned-forcing/.
|
| 218 |
GA-EIRFS: A Geometry-Augmented Repeat-Factor Sampling Method for Long-Tailed LiDAR 3D Object Detection
2609.38116
|
cs.CV
|
Taufiq Ahmed, Constantino \'Alvarez Casado, Daniel Herrera Castro, Sasan Sharifipour, Abhishek Kumar |
Long-tailed 3D object detection is treated as a class-frequency problem, but LiDAR supervision quality depends on object observability: similar frequencies can hide different geometric evidence. We introduce Geometry-Augmented Exponentially Weighted Instance-A...Long-tailed 3D object detection is treated as a class-frequency problem, but LiDAR supervision quality depends on object observability: similar frequencies can hide different geometric evidence. We introduce Geometry-Augmented Exponentially Weighted Instance-Aware Repeat Factor Sampling (GA-EIRFS), a detector-agnostic method that modulates a frequency-based repeat factor with a fixed geometry score combining point count, surface-normal entropy, and surface coverage. GA-EIRFS changes only frame-sampling probabilities, leaving the detector and inference unchanged. On nuScenes it improves mean average precision (mAP) and the nuScenes detection score (NDS) in four converged experiments with CenterPoint and PointPillars over two seeds; for CenterPoint at seed 666, mAP rises from 0.552 to 0.563 and bicycle AP from 0.306 to 0.359. Per-class gains correlate with the class sampling-weight increase (Spearman rho=0.70, p=0.025) but not with geometry score alone (rho=0.32, p=0.37), so geometry amplifies frequency-driven need. KITTI results vary across seeds, most for the rarest class. Code: https://github.com/Multimodal-Sensing-Lab/GA-EIRFS.
|
| 219 |
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents
2609.38119
|
cs.CV
|
Jinfa Huang, Jianming Xu, Jingyang Lin, Zhengyuan Yang, Jiebo Luo |
Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, ...Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent's context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).
|
| 220 |
HelixWorld: A Real-time Interactive Audio-Visual World Model
2609.38123
|
cs.CV
|
Lei Ke, Jiahao Pan, Zeyue Tian, Jiaming Wang, Haoyuan Huang |
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic d...World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.
|
| 221 |
CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer
2609.38136
|
cs.CV
|
Teng Zhou, Yunhao Chen |
Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and...Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at \href{https://github.com/0606zt/CLeaR}{https://github.com/0606zt/CLeaR}.
|
| 222 |
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
2609.38140
|
cs.CVcs.AI
|
Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang |
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward ...Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.
|
| 223 |
LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
2609.38146
|
cs.CV
|
Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao |
We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in control...We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.
|
| 224 |
FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation
2609.38152
|
cs.CV
|
Trong-Tung Nguyen, Jiahan Zhang, Anand Bhattad |
We introduce FracGen, a fracture-aware video generation model that produces plausible, controllable fracture dynamics from a single image of an intact object, conditioned on physics signals. To train FracGen, we build FracSim, a fracture-aware simulation frame...We introduce FracGen, a fracture-aware video generation model that produces plausible, controllable fracture dynamics from a single image of an intact object, conditioned on physics signals. To train FracGen, we build FracSim, a fracture-aware simulation framework that augments material point method (MPM) simulation with a continuum damage model, producing paired fracture videos and dense, pixel-aligned physical fields at no additional cost beyond standard rendering. FracGen leverages these maps in two ways: it is trained to jointly predict them alongside RGB video, encouraging the model to capture physical state rather than surface appearance; and it is supervised with physics-informed losses that encourage consistency among the predicted maps. As a result, FracGen captures distinct material-specific fracture behavior without expensive test-time simulation or per-scene tuning, while offering fine-grained control over where an object tears, how fast the crack propagates, and how much deformation precedes failure. We further introduce a benchmark for evaluating the physical plausibility of generated fracture video, and show through extensive experiments that FracGen outperforms existing video generation baselines in both physical and visual fidelity. Results are best viewed in our project website: https://fracgen.github.io/.
|
| 225 |
PowerSim: Differentiable Physics Simulation and Rendering with Power Diagrams
2609.38153
|
cs.CV
|
Trong-Tung Nguyen, Anand Bhattad |
We introduce PowerSim, a method to bring physically grounded, differentiable dynamics to PowerFoam's power diagram based 3D representation. PowerSim directly couples a pre-trained PowerFoam scene to the Material Point Method (MPM) by exploiting a natural align...We introduce PowerSim, a method to bring physically grounded, differentiable dynamics to PowerFoam's power diagram based 3D representation. PowerSim directly couples a pre-trained PowerFoam scene to the Material Point Method (MPM) by exploiting a natural alignment between the two: the geometric and appearance properties of each primitive correspond closely to the quantities MPM already tracks as an object deforms. Consequently, simulated motion can drive the scene's geometry and appearance directly, without an auxiliary representation in between. Built on this framework, we enable a range of applications on real and synthetic scenes: (1) simulating a static scene under user interaction, (2) recovering spatially varying material fields, (3) compositing primitives from independently captured scenes into a single simulation-ready scene and (4) ray-tracing reflections that update consistently as the object deforms. Our results suggest that PowerSim excels over previous frameworks for physically grounded dynamics, while unlocking unique advantages-such as secondary ray lighting effects on dynamic scenes. Results are best viewed on our project website: https://power-sim.github.io/.
|
| 226 |
LongLive-Plug: Once-for-All Distillation for Video Generation
2609.38154
|
cs.CV
|
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang |
Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically re...Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.
|
| 227 |
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
2609.38155
|
cs.CVcs.AI
|
Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia |
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, whil...Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.
|
| 228 |
DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses
2609.38156
|
cs.CV
|
Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Shaoteng Liu |
Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportu...Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA$^2$. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA$^2$ student performs better than the 25-step teacher and evaluated few-step distillers.
|
| 229 |
Rethinking Representations for World-Action Modeling
2609.38163
|
cs.CV
|
Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen |
World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction...World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.
|
| 230 |
Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data
2609.38165
|
cs.CV
|
Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini |
The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or ...The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for Segmentation (PAtteRNS), the first model to use self-attention mechanisms separately for each of the temporal, spectral, and spatial aspects of Sentinel-2 multispectral SITS data. To achieve fully-factorised attention in our proposed model, we introduce a novel parallel transformer architecture which significantly reduces the computational complexity of triple-factorised self-attention. We validate our architecture with an in-depth ablation study, and analyse the performance of our model against state-of-the-art crop segmentation models on multiple tile-size variants of the popular PASTIS and MTLCC datasets. Our findings show our model to outperform all others in the task of crop class segmentation, verified across multiple important segmentation metrics, with especially strong performance against compared models seen in the often under-reported parcel delineation quality, for which we use the Boundary IoU metric. We also find that flawed class groupings within datasets can have a significant negative impact on model performance, and report that alternate tile-size variants of crop segmentation datasets produce results incomparable to one-another, invalidating fair comparison between model performance when trained on different tile-sizes. Based on these findings, we suggest further work is required to standardise best practices when constructing SITS crop segmentation datasets, and to enable future dynamic-tile-sizing for ideal model performance.
|
| 231 |
Adversarial Training for Pixel Diffusion
2609.38170
|
cs.CV
|
Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Zhe Lin |
Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training cor...Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial post-training for pixel diffusion. Across two pixel backbones, the method jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality. We further investigate why it works. Frequency-band and power-law analyses show that the original models systematically underproduce natural-image high-frequency content, while adversarial post-training restores this missing spectral power. In contrast, perceptual loss also increases high-frequency content but sacrifices distribution fidelity and prompt alignment. Nearest-neighbor, recall, and matched no-GAN SFT controls further rule out memorization, mode dropping, and additional optimization as simple explanations. Finally, we examine the boundary of this effect. Under the tested latent diffusion configurations, the same procedure does not produce comparable joint gains and adds almost no decoded high-frequency power. These results identify direct output access to the image statistics being corrected as a key factor governing when adversarial post-training succeeds.
|
| 232 |
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
2609.38177
|
cs.CV
|
Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim |
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3...Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.
|
| 233 |
Point2Part: Unified 3D Partitioning from Point Prompts
2609.38180
|
cs.CV
|
Hao-Tang Tsui, Yu-Rou Tuan, Xiaoxuan Ma, Nicolas Ugrinovic, Takaaki Shiratori |
Existing 3D part decomposition methods do not necessarily partition the original shape into non-overlapping parts that collectively cover the entire shape, allowing overlaps or gaps that hinder downstream part-level applications. We instead formulate part deco...Existing 3D part decomposition methods do not necessarily partition the original shape into non-overlapping parts that collectively cover the entire shape, allowing overlaps or gaps that hinder downstream part-level applications. We instead formulate part decomposition as a joint partitioning of the entire shape, where the predicted parts are non-overlapping and jointly recover the entire shape. Our key insight is that part decomposition should consider all desired parts jointly, rather than modeling each part independently. To this end, we develop a promptable model for 3D part decomposition from images or meshes. Users can specify desired parts through 3D point prompts for controllable decomposition. Given one point prompt per desired part, our model produces the corresponding parts as a complete partition of the entire shape. We build on a pretrained 3D generation model and first obtain a shape latent from either an input image or mesh. We then introduce a prompt encoder that maps each 3D point prompt to a part token while attending to the shape latent. To decode the desired parts, we propose a novel part decoder jointly scoring the entire shape against all part tokens in a coarse-to-fine manner, assigning every position within the shape volume to exactly one part. We perform part decomposition in this shared shape latent space, enabling a unified model for image-to-part generation, mesh-to-part generation, and part segmentation. Our method outperforms existing works on all part-quality metrics across all three tasks, and improves compatibility among parts by an order of magnitude over previous SOTA methods. Code and models will be released.
|
| 234 |
PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units
2603.15106
|
cs.CV
|
Mark Deutel, Simon Geis, Axel Plinge |
Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typically requires DNN architectures to be specialized for each device separately. To avoid the huge manual effort, one can us...Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typically requires DNN architectures to be specialized for each device separately. To avoid the huge manual effort, one can use neural architecture search (NAS). However, many existing NAS methods are resource-intensive and time-consuming because they require the training of many different DNNs from scratch. Furthermore, they do not take the resource constraints of the target system into account. To address these shortcomings, we propose PrototypeNAS, a zero-shot NAS method to accelerate and automate the selection, compression, and specialization of DNNs to different target microcontroller units (MCUs). We propose a novel three-step search method that decouples DNN design and specialization from DNN training for a given target platform. First, we present a novel search space that not only cuts out smaller DNNs from a single large architecture, but instead combines the structural optimization of multiple architecture types, as well as optimization of their pruning and quantization configurations. Second, we explore the use of an ensemble of zero-shot proxies during optimization instead of a single one. Third, we propose the use of Hypervolume subset selection to distill DNN architectures from the Pareto front of the multi-objective optimization that represent the most meaningful tradeoffs between accuracy and FLOPs. We evaluate the effectiveness of PrototypeNAS on 12 different datasets in three different tasks: image classification, time series classification, and object detection. Our results demonstrate that PrototypeNAS is able to identify DNN models within minutes that are small enough to be deployed on off-the-shelf MCUs and still achieve accuracies comparable to the performance of large DNN models.
|
| 235 |
Toward a Culturally Adapted Chinese Language Agent: A Wizard-of-Oz Study of Nonverbal Behavior in Chinese-German Intercultural Interaction
2609.35150
|
cs.CVcs.MM
|
Siddhant Jain, Anna Lea Reinwarth, Dimitra Tsovaltzi, Rafael Math, Julia Renner |
Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose violation triggers subtle but meaningful nonverbal responses. For German learners of Mandarin Chinese, acquiring ...Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose violation triggers subtle but meaningful nonverbal responses. For German learners of Mandarin Chinese, acquiring this sensitivity is critical yet poorly supported by existing language-learning agents. We present a Wizard-of-Oz (WoZ) study design and supporting real-time system for collecting multimodal behavioral data from native Chinese speakers reacting to social norm violations by German learners. The system features a photorealistic MetaHuman avatar driven by Live Link face capture and MediaPipe upper-body tracking, a wizard console for real-time behavior selection, and synchronized multimodal logging across agent and learner streams. A layered annotation framework, based on psychological theory and covering non-observable socioemotional reactions, norm interpretation, verbal, and observable behavior thereof, and future supervision targets enables the corpus to support training of future automated cultural interpretation and behavior generation models. Four ecologically valid interaction scenarios, developed with cultural and pedagogical experts, provide the methodological and technical foundation for a culturally adapted conversational agent for Chinese language learning.
|
| 236 |
HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization
2609.35800
|
cs.CV
|
Nenad Banfic |
Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask. Image-se...Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask. Image-sensitivity and output-sensitivity scores select physical KV heads offline, with approximately 1/8 protected in the main experiments; their image keys and optionally values remain in bfloat16 (BF16), while the base quantizes unprotected image entries. Across eight VLMs, three base quantizers, and eight benchmarks (six discriminative and two generative), HeadGuard recovers a substantial fraction of lost accuracy on weaker quantizers, with the strongest gains for Qwen and InternVL. At 2 bits, the six-task discriminative mean over eight models rises from 0.436 to 0.580 on the weakest base; protection can also improve generated answers and caption fidelity to BF16 outputs. Mean accuracy gains persist across all three quantizers with both tested calibration datasets. Keys-only protection retains substantial recovery at lower modeled storage cost. Evaluated through simulated quantization, HeadGuard offers a composable way to improve low-bit VLM accuracy without replacing the underlying quantizer.
|
| 237 |
PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions
2609.35865
|
cs.CVcs.AI
|
Yida Lin |
Single-token typed-decision models answer a schema question by reading the logits of a few one-letter answer codes at a single position: they are fast and return a probability for every allowed answer, but they are trained with plain cross-entropy that ignores...Single-token typed-decision models answer a schema question by reading the logits of a few one-letter answer codes at a single position: they are fast and return a probability for every allowed answer, but they are trained with plain cross-entropy that ignores most of the structure in their training data. We study such a model whose data is curated as contrastive pairs---two contexts that differ in one edited fact that flips the answer---each carrying a machine-checked certificate that deleting the decisive sentence makes the fact unknown. We propose PACT, which turns this structure into four training terms that need no new annotation: a difference-in-differences margin over each pair that is invariant to any shared logit offset, a permutation-consistency term against answer-code position bias, an evidence-necessity term on certificate-verified ablated contexts, and an ordinal transport cost for rubric fields, plus a three-parameter contextual temperature. On a frozen 324-item holdout with three seeds, PACT matches the published recipe in accuracy ($84.6\%$ vs. $85.2\%$; McNemar $p \ge 0.50$ at every seed) while giving the lowest position bias of all runs (answer flips under relabelling $9.8\%$ vs. $13.8\%$) and the lowest ordinal error on rubric fields (MAE $0.232$ vs. $0.311$). Against a control with the same optimiser and schedule but cross-entropy only, PACT is significantly more accurate at two of three seeds, halves the seed-to-seed spread and lowers NLL by $26\%$. Seed-matched ablations and pre-specified falsification tests locate these gains precisely: no single term raises raw accuracy, and the method's value lies in robustness and stability rather than headline accuracy. Code, data splits, trained adapters, and all run records are available at https://github.com/BennyLinntu/PACT-Pairwise-Anchored-Calibrated-Tuning-for-Single-Token-Typed-Decisions.
|
| 238 |
The Decision Value of Perception Compute
2609.35910
|
cs.CV
|
Hoang Pham Cong, Ho Viet Duc Luong |
Adaptive perception spends extra computation on inputs where perception is expected to improve. When perception feeds a downstream decision system, a better perception output need not produce a better decision. We define the decision value of perception comput...Adaptive perception spends extra computation on inputs where perception is expected to improve. When perception feeds a downstream decision system, a better perception output need not produce a better decision. We define the decision value of perception compute as the change in downstream loss from escalating an input from a cheap to an expensive perception mode. Because this value can be negative, the allocation of perception compute should be judged against a budget-constrained decision oracle, with uniform full-fidelity inference as a baseline rather than an upper bound. We introduce DEEP (Decision Evaluation for Escalated Perception), a benchmark that scores pre-escalation allocators against this oracle under selection, latency and energy budgets, charging each allocator for its own computation. With deployed monocular geometry on KITTI and nuScenes, we find that 34--54% of the escalations that change downstream loss make it worse; harmful escalations also occur for the published PDM-Closed planner, evaluated open-loop on nuPlan with real detector outcomes. On nuScenes, perception-level gain frequently disagrees in sign with decision value. This mismatch has practical consequences: choosing among fixed deployable signals by missed-object perception gain rather than by decision value reduces realized test decision gain by 7.4% of the all-cheap loss on average. Learned allocators recover part of the oracle's value by finding beneficial escalations but select nearly as much harm as random, and once their own computation is charged at a 20% latency budget, only the lightweight routers, at about 3.5\% of a full detector pass, still beat random.
|
| 239 |
VehicleArena: A Realistic Urban Environment for Multi-Agent Driving
2609.35916
|
cs.CV
|
Jie Yang, Jiajun Chen, Jiazheng Zhou, Mianqiu Huang, Yining Zheng |
Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction ...Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent's driving decisions can reshape traffic flow, delays, risks, and subsequent observations for surrounding agents. The benchmark provides 112 evaluation tasks spanning single-agent and multi-agent driving. Across nine evaluated models, the highest arrival rates reach only 65.0% on single-agent tasks and 65.6% on multi-agent tasks, while strong passenger-request or cabin scores do not reliably translate into successful trip completion. Moreover, in matched multi-agent runs, every tested focal policy reduces the arrival rate of surrounding vehicles relative to the simulator's native traffic controller, revealing measurable externalities beyond the focal vehicle itself.
|
| 240 |
Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study
2609.36001
|
cs.CV
|
Rafael Garcia-Dias, Alexandre Triay Bagur, Chayanin Tangwiriyasakul, Virginia Fernandez, Parhom Esmaeili |
Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. Here we present FLIP (Federated Learning Interoperability Platform), an op...Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. Here we present FLIP (Federated Learning Interoperability Platform), an open-source, multi-application platform that makes FL training and evaluation repeatable. FLIP implements common FL workflows as a set of composable services: cohort queries against per-site structured databases, on-demand DICOM retrieval from institutional PACS, per-site project approval, and reusable FL job types. To demonstrate FLIP, we ran two distinct use cases, federated fine-tuning and federated evaluation, on synthetic chest X-ray cohorts across two client nodes based in the United Kingdom (UK) and Thailand. In FLIP, each institution independently approves its participation in each project and operates its own node under local IT security processes. This study makes an operational rather than an algorithmic claim. It does not compare federated with centralised training; for that question, we refer the reader to existing systematic reviews and meta-analyses. The central result is evidence that such platforms enable international FL collaboration and improve repeatability, auditability, and site-specific governance. We also present a comprehensive comparison of existing platforms to help researchers and operators choose the right platform for their use case.
|
| 241 |
On the spectral properties of generative denoiser Jacobians
2609.36210
|
cs.CV
|
Alexandros Graikos, Nebojsa Jojic, Dimitris Samaras |
Generative denoising models, such as diffusion and flow-matching, learn to sample from complex distributions by training a deep neural network denoiser to recover clean data from noise-corrupted samples. While such models are typically compared on the quality ...Generative denoising models, such as diffusion and flow-matching, learn to sample from complex distributions by training a deep neural network denoiser to recover clean data from noise-corrupted samples. While such models are typically compared on the quality of their synthesized samples, these metrics provide limited insight into how the underlying denoiser, which drives generation, differs. In this work, we propose to analyze the spectrum of the denoiser Jacobian as a tool to characterize these differences. Across pre-trained denoising models, we observe that better generative performance is associated with larger Jacobian eigenvalues. Motivated by this, we introduce a regularization scheme that controls the Jacobian spectrum by training the denoiser on perturbed inputs, with perturbations suppressing or amplifying Jacobian responses. On ImageNet, we test whether directly modifying the Jacobian spectral properties leads to improved generations. Our findings suggest that denoisers benefit from both strengthening responses along data-relevant principal eigen-directions and suppressing the noisy, data-irrelevant ones. This establishes the denoiser Jacobian as a useful tool for identifying differences between generative denoising models.
|
| 242 |
One-Step Next-Latent Prediction Is Not a World Model
2609.36227
|
cs.CV
|
Shitong Wang, Zhongang Cai, Yuzhou Hong |
Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that ...Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a conditional mean, and a mean is a kernel only in special cases. For a linear-Gaussian Markov latent, the mean transition and the innovation covariance are fixed by the one-step problem, and the open-loop squared error at horizon $K$ equals the trace of the sum of the pushed-forward innovation covariances. That error grows with $K$ after the one-step fit is exact. If the conditional mean is nonlinear, composing it is not the multi-step conditional mean. If the observation is a non-injective function of a Markov state, a memoryless one-step map does not determine future observations, while a short window can. An isotropy penalty is a function of the embedding marginal, so its partial derivative in the transition weights is zero. On a scalar autoregression with coefficient $0.9$, the one-step mean squared error is $0.998$ and the $16$-step open-loop error is $5.10$. On a hidden rotation, an eight-step window reaches $16$-step error $0.056$, while the current scalar alone reaches $0.778$. Raising the isotropy weight from $0.1$ to $10$ leaves eight-step latent error inside $[0.78,0.85]$ on three seeds.
|
| 243 |
PyroStack: A Multi-Band Spatio-Temporal Sub-Daily Dataset for Wildfires in the United States
2609.36315
|
cs.CVcs.AI
|
Arya Kondur, Giosue Migliorini, Cameron Schmitt, Francesco Immorlano, Tairan Wang |
Wildfires are an increasing hazard to ecosystems, air quality, and human systems, creating a growing need for datasets that support systematic development and evaluation of models for predicting fire spread across diverse landscapes. Effective prediction requi...Wildfires are an increasing hazard to ecosystems, air quality, and human systems, creating a growing need for datasets that support systematic development and evaluation of models for predicting fire spread across diverse landscapes. Effective prediction requires integrating meteorological conditions, fuels, vegetation, and topography at spatial and temporal resolutions suitable for both physical simulation and data-driven approaches. However, existing datasets often lack the resolution and coverage needed to capture these interacting controls. The PyroStack dataset addresses this gap by providing a harmonized, event-based collection of wildfire and environmental data across the contiguous United States and Alaska. It integrates satellite-derived fire observations with atmospheric reanalysis, vegetation, fuel characteristics, and topographic information into a unified framework spanning 6994 wildfires that occurred between 2012 and 2024 across a wide range of ecosystems and climate conditions. PyroStack offers spatial resolutions ranging from 30 m to 9 km and hourly temporal resolution, along with fire progression data at 12-hour intervals to support model initialization and evaluation. By combining broad spatial coverage with fine spatial and temporal detail, the dataset enables systematic analysis of wildfire dynamics and supports both physics-based and machine learning approaches, providing a foundation for benchmarking and improving fire spread models, with future extensions aimed at incorporating additional regions and fire suppression data streams to further advance wildfire prediction.
|
| 244 |
AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models
2609.36368
|
cs.CVcs.AI
|
Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi |
Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation...Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on training neural-based decoders. Both approaches struggle under limited supervision, while fine-tuning additionally requires access to model parameters, which is often unavailable for closed-source models. We introduce AdaKerNet, a novel learnable task-adaptive neural kernel decoder. AdaKerNet is fully agnostic to the parameters of the underlying MLLM and operates solely on its (frozen) rich representations obtained from the diverse available modalities. AdaKerNet relies on (i) a set of learnable, Lipschitz-controlled multimodal features derived from these MLLM representations; (ii) a reference kernel that provides a soft structural prior on those features; and (iii) a lightweight nonlinear neural predictor that adaptively deforms that structure. Learning the kernel representation and the neural predictor jointly within a unified optimization framework allows AdaKerNet to capture features and geometric relationships relevant to the downstream task. Numerical tests across four MLLMs: BLIP-2, LLaVA-1.5, Qwen2.5-VL, and Gemini Embedding 2, and multimodal inputs spanning text, audio, images, and tabular measurements demonstrate significant and consistent improvements over direct MLP, attention-, autoencoder- and kernel-based decoders, across a range of scarce-label budgets, with average error reduction of up to 41% across baselines. These results establish AdaKerNet as an effective approach for prediction from frozen multimodal representations in the scarce label regime. Additional structural ablations highlight the complementary contributions of AdaKerNet's components.
|
| 245 |
CAMEO: A Class-Activation-Mapped Equitable Overlay Framework for Fair and Robust Deep Learning-based Skin Condition Diagnosis
2609.36400
|
cs.CV
|
Youssef Attia, Debasmita Mukherjee |
Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues such as skin tone, device vignetting, and embedded rulers, rather than on lesion morphology. This undermines robu...Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues such as skin tone, device vignetting, and embedded rulers, rather than on lesion morphology. This undermines robustness and fairness across skin tones. This work asks whether Explainable AI (XAI), typically used only to audit a finished model, can instead be repurposed as an active training signal that corrects this shortcut without sacrificing diagnostic accuracy. We introduce CAMEO (Class Activation Mapped Equitable Overlay), a framework that improves skin-lesion classification by selecting stable model explanations and using them to separate lesions from their backgrounds. It then replaces the background with realistic synthetic skin while keeping the lesion unchanged. On HAM10000 and dark-skin ISIC images, CAMEO maintained accuracy while reducing background-driven errors by nearly four times. It also made the model's attention more consistent when backgrounds changed. Results across multiple tests show that reducing reliance on background information improves robustness, with Fitzpatrick-based backgrounds providing a realistic and interpretable approach. Results show that XAI-guided augmentation can make dermoscopic classifiers measurably more robust and fair at no cost to accuracy. They also clarify that it is the mechanism and not the specific tone palette that matters, and that the lasting contribution of XAI here lies in stability-screened, annotation-free lesion localisation rather than in the robustness number itself.
|
| 246 |
FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
2609.36416
|
cs.CVcs.AI
|
Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala |
Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds ...Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode while the rare bimanual effort that does label subtasks annotates only a fraction of its hours. We present FineART, a densely annotated bimanual manipulation dataset of 40,543 episodes, 1,718 hours, and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask, and show that mid-training it this way yields substantial gains. Specifically, success on a spatial disambiguation task increases from 32.0% to 100.0%, and step-by-step human subtask guidance lifts success on an unseen long-horizon task from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires one-tenth the data of baselines without mid-training and generalizes zero-shot to completely unseen tasks on the new hardware. We open-source the full dataset, model weights, and training code.
|
| 247 |
Losing the name before the box: measuring and repairing what narrow fine-tuning costs a detector outside its deployment vocabulary
2609.36426
|
cs.CV
|
Trung Minh Bui, Jongsul Moon, YoungOuk Kim, Jung-Hoon Hwang, Dongin Shin |
A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the r...A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give a longitudinal protocol: one pretrained checkpoint against its own fine-tuned descendants. It tracks held-out top-$K$ proposal coverage $C_\tau$: of categories pretraining covered and the vocabulary omits, the share of boxes a detector's top $K$ regions still cover. The quantity is the open-world proposal literature's; the longitudinal reading is not. $C_\tau$ falls while in-domain accuracy rises, on four architectures and three domains, by $5.12$ to $63.35$ points on boxes above $1024$ px$^2$. No in-domain number identifies the fall, and neither does detection average precision, which charges a missed and a misnamed box alike. On the one architecture scoring both, adaptation costs $87\%$ of the AP against a fifth of the coverage, and the naming goes first at all six depths of its freeze ladder, every run. What breaks is structured: three architectures sharing no pretraining run agree on which categories lose coverage, and those a model never learned do not lose any. A repair follows and needs no training: mixing a quarter of the pretrained state back, normalisation statistics included, raises coverage on every cell swept for at most $2.47$ points of in-domain accuracy. Seeing it costs one extra evaluation pass.
|
| 248 |
Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models
2609.36458
|
cs.CV
|
Abdullah All Tanvir, Xin Zhong |
Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions, raising a basic question: what local metric best captures semantically consequential variation? We propose ...Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions, raising a basic question: what local metric best captures semantically consequential variation? We propose Fisher-induced invariant representation geometry (Fisher-IRG), which measures local representation directions through their predictive sensitivity. Around each representation, we construct semantic-preserving and semantic-changing neighborhoods, aggregate their local Fisher information, and recover invariant directions through a contrastive generalized eigenvalue problem. Controlled displacement analyses first show that comparable Euclidean motion can have substantially different predictive consequences, supporting the need for a predictive geometry. Across language and vision models, Fisher-IRG yields stronger semantic-versus-nuisance predictive selectivity and generally more reproducible subspaces than covariance-based geometry, while recovering systematically distinct local directions. Representation interventions further localize semantic effects to the Fisher-derived subspace, and held-out separation and retrieval show that the recovered geometry generalizes beyond the discovery neighborhoods. These results support Fisher-IRG as a principled framework for characterizing local invariant representation geometry.
|
| 249 |
Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks
2609.36471
|
cs.CVcs.AI
|
Guoheng Sun, Chen Chen, Jin Wang, Ang Li, Teresa Lv |
World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize thi...World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, $3.62\times$ the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.
|
| 250 |
Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models
2609.36475
|
cs.CVcs.AI
|
Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li |
Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often ...Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here, we ask whether VLMs align with humans not only in choices, but also in where they look when making those choices. We study both VLMs and humans (N = 53), presenting them with the same stimuli, a pseudo-word and two images, and record participants' choices and eye movements, which we release. We find choice alignment in a few larger VLMs, but their saliency matches human gaze less closely than a center-bias baseline, a fixed Gaussian at the center of each image. Fine-tuning small VLMs on human choices brings their choice alignment to the level of a human majority-vote reference on unseen words and images, yet their attention still matches human gaze less closely than this baseline. Training model attention on human gaze raises attention-gaze correlation without improving choice alignment, and a single average gaze map per image position raises it by a similar amount. Matching human choices, or even human gaze patterns, is therefore not sufficient evidence of human-aligned cross-modal processing.
|
| 251 |
Distilling Privileged Control Barrier Functions into RGB-Only Safety Filters for Dynamic Visual Navigation
2609.36520
|
cs.CV
|
Seungyeon Yoo, Gawon Lee, Seungwoo Jung, Inkyu Jang, H. Jin Kim |
RGB-only end-to-end visual navigation policies remain vulnerable to collisions in real-world dynamic environments, motivating a dedicated safety layer. Existing visual Control Barrier Function (CBF) approaches seek to provide safety from RGB observations, but ...RGB-only end-to-end visual navigation policies remain vulnerable to collisions in real-world dynamic environments, motivating a dedicated safety layer. Existing visual Control Barrier Function (CBF) approaches seek to provide safety from RGB observations, but often rely on real-time rendering or explicit scene reconstruction and are primarily designed for static scenes, limiting their practicality for onboard deployment. We propose a teacher-student visual distillation framework that transfers the safety behavior of a privileged CBF teacher to an RGB-only student filter for dynamic environments. The student maps a short RGB history, robot velocity, and a nominal control action directly to a safe action, while the teacher uses ground-truth robot and obstacle states in a real-to-sim dynamic Gaussian Splatting environment. To reduce the teacher-student information gap, the teacher constructs safety constraints only from obstacles observable within the student's RGB history. It also accounts for obstacle-velocity uncertainty to improve robustness to motion variations, while action augmentation exposes the student to diverse safe and unsafe nominal actions to better capture the safety boundary. At deployment, the student requires only RGB observations and robot velocity, without explicit 3D reconstruction or online rendering. Experiments show that the proposed method outperforms visual CBF baselines and improves the safety of RGB-based navigation policies under dynamic obstacle motion. Project page: https://syeon-yoo.github.io/distill-cbf-site/.
|
| 252 |
PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation
2609.36638
|
cs.CV
|
Mingfeng Lin, Chengfei Cai, Lin Xu, Chengqian Ma, Yuxiang Wei |
Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts a...Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text--image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.
|
| 253 |
Quantum Fidelity Landscape-Guided Prior Calibration for Single-Circuit QGAN Image Generation
2609.36702
|
cs.CV
|
Xue Yang, Rigui Zhou, Dax Enshan Koh, Siong Thye Goh, Yitao Tang |
Quantum Generative Adversarial Networks (QGANs) have emerged as representative generative models in the Noisy Intermediate-Scale Quantum (NISQ) era and have attracted increasing attention in quantum machine learning. However, most existing QGAN methods rely on...Quantum Generative Adversarial Networks (QGANs) have emerged as representative generative models in the Noisy Intermediate-Scale Quantum (NISQ) era and have attracted increasing attention in quantum machine learning. However, most existing QGAN methods rely on patch-based decomposition strategies, which weaken the global consistency of generated images and increase quantum resource overhead. In this work, we investigate a simpler approach: pixel-level, end-to-end image generation using a single-quantum-circuit QGAN. By analyzing the structural matching relationship between the quantum prior and the target data distribution in Hilbert space, we provide a new theoretical perspective for understanding the training behavior of naive end-to-end QGANs. Specifically, we introduce the Quantum Fidelity Landscape (QFL), defined as the pairwise-fidelity structure induced by an ensemble of quantum states and preserved under shared unitary transformations of the quantum generation process. We show that, under a fixed Lipschitz readout, this invariant imposes a one-sided bound on decoded sample separation, motivating calibration of the prior-induced QFL before adversarial training. To validate this theoretical insight, we propose BasicQGAN, a QGAN framework incorporating quantum prior calibration. Before adversarial optimization, BasicQGAN aligns the prior-induced QFL with the data-induced QFL. Experimental results on small-scale grayscale image datasets show that BasicQGAN achieves stable and effective end-to-end pixel-level image generation while requiring fewer qubits and trainable parameters than representative patch-based quantum generators. Furthermore, experiments with different initial quantum-state ensembles show that QFL-calibrated ensembles achieve better generative performance.
|
| 254 |
Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients
2609.36770
|
cs.CVcs.AI
|
Aram Davtyan, Pablo Acuaviva, Sebastian Stapf, Paolo Favaro |
Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent ...Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using masked prediction of frozen DINOv3 features. We measure specialization by asking whether the best agent for an input aligns with its latent domain, and utilization by asking whether responsibility is distributed across agents. We progressively remove central control, ending with DISCO (DIStributed COllaboration) where each agent locally selects a helper, reads its internal state through a gradient free channel, and rewards its router only for the improvement that help provides. Specialization emerges and is useful. Randomly routed populations underperform a single generalist, while semantically routed populations outperform it, showing that specialization rather than population size drives the gain. Specialization persists without a central router, and gradient free communication lets nonexperts exploit emergent expertise. In DISCO, a random agent helped by the expert matches the solo generalist, while experts surpass it, including on data outside the specialization mixture. Local routers select the emergent expert for 98% of inputs. These effects persist across population size, model capacity, data imbalance, and finetuning seeds, providing measurable evidence for the dynamics needed by decentralized predictive pretraining.
|
| 255 |
DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients
2609.36779
|
cs.CV
|
Xiaotian Zhang, Yusheng Wang, Naoya Kagawa, Noritaka Takamura, Keiji Okuhara |
Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural n...Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations, enabling hand-eye calibration without fiducial markers in the calibration stage and providing interpretable optimization. While the state-of-the-art differentiable rendering methods achieve remarkable accuracy, the use of binary masks can result in the loss of internal profile details, reducing precision. Additionally, these methods can also suffer from unstable optimization and local minima. In this study, we propose a novel RGB-based differentiable rendering framework that provides richer geometric and appearance cues by incorporating color and mask geometric features, thereby improving calibration accuracy and optimization stability. Additionally, we propose a mask-guided image-to-image translation method to ensure explicit preservation of color and geometric consistency throughout the translation. Our approach is validated through both simulation and real-world experiments, with results demonstrating strong accuracy and robustness and clear improvements over existing differentiable rendering methods. Our method achieves a grasping success rate of 88.9% and insertion success rate of 57.4% on the UR5e real-world experiment, outperforming the state-of-the-art differentiable rendering hand-eye calibration method EasyHeC by 46.3 and 48.1 percentage points, respectively.
|
| 256 |
Rethinking Multimodal Fake News Detection in the Generative AI Era
2609.36850
|
cs.CVcs.MM
|
Wenbin Shen, Guoxuan Qin, Guangxu Yao, Baodong Wang, Yuanbo Rui |
Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimod...Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracity assessment and rarely characterizes how generativity differences affect the reliability of evidence. In contrast, AIGC detection primarily determines whether content is generated or modified by generative models, but it does not by itself establish whether the underlying news event is true. To bridge the separation between these tasks in data and evaluation, we construct Weibo26, a multimodal fake news detection dataset for generative-content scenarios. On this basis, we propose the Generativity-Aware Hierarchical Reasoning (GAHR) framework, which combines global judgment with local correction so that generativity information participates in news-veracity reasoning. Experiments on multiple existing fake news detection benchmarks and Weibo26 show that GAHR achieves competitive veracity-detection performance while effectively identifying generative content.
|
| 257 |
RAEGNet: Relation-Aware Evidence Graph Network for Harm-Aware Multimodal Fake News Detection
2609.36902
|
cs.CVcs.MM
|
Wenbin Shen, Guoxuan Qin, Guangxu Yao, Baodong Wang, Yuanbo Rui |
Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly foc...Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for differences in the degree of harm posed by different instances of fake news. To address these limitations, we design an Event-Level Evidence Retrieval Framework (ELERF) and propose a Relation-Aware Evidence Graph Network (RAEGNet). ELERF retrieves external evidence based on the complete event semantics of a news item. RAEGNet constructs a directed graph that incorporates news-evidence stance relations and evidence-evidence interaction relations, and introduces a conditional-harm branch to jointly model authenticity and potential harm. Experimental results demonstrate that RAEGNet outperforms multiple baseline methods across all evaluated metrics on Weibo-21, Fakeddit, and our self-constructed SSS dataset.
|
| 258 |
Track-and-Complete: Learning Humanoid Skills from a Single Failed Human Video
2609.36924
|
cs.CV
|
Sarmad Idrees, Jongeun Choi |
Learning humanoid skills from videos typically requires a successful human demonstration, which often demands custom data collection. Although failures have traditionally been treated only as negative examples in robot learning, they can still reveal a usable ...Learning humanoid skills from videos typically requires a successful human demonstration, which often demands custom data collection. Although failures have traditionally been treated only as negative examples in robot learning, they can still reveal a usable trajectory prefix before the task fails, as well as the intended outcome. To leverage this information from a failed-attempt video, we propose TRACC, a pipeline that imitates the useful portion of the motion trajectory and then completes the task based on the inferred task outcome. The usable motion prefix serves as prior knowledge until the failure occurs, after which the task-completion reward guides the policy to learn the intended task goal without requiring a successful task trajectory. We evaluate our method on six in-the-wild failed human tasks from the Oops! dataset. Our experimental results demonstrate the effectiveness of the proposed approach for learning from failed attempts when no successful demonstration is available. Thus, these findings establish failed human videos as a viable source of supervision for humanoid skill learning.
|
| 259 |
Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
2609.36965
|
cs.CV
|
Zexiao Wang, Zihao Zhang, Xudong Wang, Pan Wang, Ziyi Ye |
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their uti...System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev's average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at https://gulucaptain.github.io/Chinese-Jev/.
|
| 260 |
NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters
2609.37038
|
cs.CVcs.AI
|
Haoran Xu, Xingzhuo Guo, Yuchen Zhang, Jincheng Zhong, Jianmin Wang |
Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, ...Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instantiate this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill. Experiments on SEVIR and MRMS benchmarks show that NowcastDiT achieves state-of-the-art performance in both perceptual quality and meteorological skill. These results suggest that standard DiT can serve as an effective foundation for precipitation nowcasting.
|
| 261 |
Multi-Depth Temporal Fusion for Feedforward, Locally Trained Spiking Neural Networks
2609.37047
|
cs.CV
|
Aidin Attar, Eleonora Cicciarella, Michele Rossi |
We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convo...We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original framework combining residual-like connections with multi-depth feature aggregation and consensus. The full SNN pipeline features an early-vision front end, to convert raw visual data into sparse spike latencies, a four-layer convolutional backbone trained layerwise with unsupervised spike-timing-dependent plasticity (STDP), a deterministic Multi-Depth Temporal Fusion (MDTF) and a final classifier trained with reward-modulated spike-timing-dependent plasticity (R-STDP). Rather than replacing early features in deeper layers, the proposed MDTF preserves early temporal evidence, adding sparse residual events from intermediate layers, and incorporating deeper features only when they agree in time with earlier representations. The resulting architecture is experimentally validated across MNIST, Fashion-MNIST, CIFAR-10, and N-MNIST, delivering strong classification performance under a fully local learning regime. Selective multi-depth fusion significantly outperforms traditional STDP/R-STDP baselines on higher-variability visual tasks (achieving +18.2 pp on Fashion-MNIST and +29.2 pp on CIFAR-10). Furthermore, activity-budget analyses show that the network retains high accuracy even when removing a large fraction of late or weak spike events, confirming its high data efficiency and reduced event-processing requirements. The codebase is publicly available at github.com/aidinattar/multi-depth- temporal-fusion-snn.
|
| 262 |
V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving
2609.37098
|
cs.CV
|
Junwei You, Weizhe Tang, Can Wang, Yan Zhao, Jun Hua |
Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly expl...Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of the current scene, while the future consequences of prospective driving actions are rarely modeled explicitly. This limits the ability of the planner to anticipate how its decisions may interact with the evolving traffic environment. To address this issue, we propose V2X-WAM, a cooperative world action model that tightly couples cooperative scene understanding, action generation, and future-world reasoning. V2X-WAM constructs a reliability-aware spatiotemporal representation from vehicle- and infrastructure-side observations, while compressing infrastructure information into a compact quantized message for efficient communication. Based on the resulting cooperative representation, a multimodal planner generates prospective trajectories, which explicitly condition future occupancy and dynamic-flow prediction. The predicted world consequences are then fed back to refine the planned trajectory, forming a closed interaction between action and future-world evolution. Experiments on a large-scale real-world cooperative driving dataset demonstrate that V2X-WAM consistently improves planning accuracy and safety over representative end-to-end cooperative driving methods, while achieving stronger future-world prediction and substantially lower communication overhead. Ablation studies further validate the effectiveness of the proposed design.
|
| 263 |
Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interaction
2609.37148
|
cs.CV
|
Siddhant Jain, Dimitra Tsovaltzi |
Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone s...Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind of clean labeling that most machine learning pipelines are built around. We study this challenge through a case that is well grounded in psychological theory but rarely modeled computationally: self-compassion, the tendency to respond to one's own setbacks with patience rather than harsh self-criticism. We examine how it appears during structured reflective interviews in a technology-mediated training setting, where people naturally talk through socio-emotionally demanding situations. Since no existing dataset captures this kind of construct in this kind of setting, we collected and annotated 51 reflective dialog sessions using an independent, temporally overlapping annotation scheme grounded in established theory. We consolidate the underlying six-component psychological model into a three-class supervision space, balancing self-kindness and mindfulness against self-critical or overwhelmed states, and build a reproducible window-based pipeline that aligns video, audio, and text on a shared timeline. Unimodal models trained on each modality separately are compared against a simple probability-level fusion strategy, which yields modest but consistent gains over the best single modality. We close by discussing where each modality succeeds or struggles, what this suggests about how this kind of construct is actually expressed in reflective speech, and what would be needed to model it, and constructs like it, more effectively.
|
| 264 |
Scaling Full Conformal Image Classifiers
2609.37298
|
cs.CV
|
Julio Silva-Rodr\'iguez, Ender Konukoglu |
Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite ...Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally prohibitive at scale because it requires candidate-specific model refits at test time. We address this limitation by leveraging zero-shot vision-language models (VLMs) to guide scalable FCP in large label spaces. We introduce Targeted Full Conformal Prediction (T-FCP), which uses a lightweight inductive conformal predictor to prune unlikely labels and applies FCP only to the remaining candidates, reducing computation while retaining the formal guarantee of the combined conformal procedure. We further propose Stabilized Online LDA (SO-LDA), an efficient VLM adaptation solver based on rank-one inverse-covariance updates. Across multiple benchmarks, including ImageNet, T-FCP enables practical full-conformal image classification with modest test-time overhead, yielding efficient prediction sets and more stable empirical coverage than split conformal alternatives.
|
| 265 |
Encore: Few-Shot Agentic Discovery of Manipulation Strategies
2609.37359
|
cs.CVcs.AI
|
Yifan Kang, Zihan Wang, Zhiwen Fan, Bangya Liu |
Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only th...Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We introduce ENCORE, which gives the agent a few demonstrations as evidence to read rather than as training data. A deterministic builder distills each demonstration into a pack of multi-view keyframes, gripper events, frame strips, and the full trajectory. A coding agent studies the pack, writes a policy program against a fixed perception and action API, refines it iteratively over a few development rollouts, and freezes it before a sealed evaluation that never reveals the success signal. On LIBERO-PRO, the agent's first program already succeeds in half of the perturbed tasks with demonstrations and in one task without them, and the frozen programs outperform the strongest prior agentic system run with the same language model (96.3% against 89.3%). On RoboDojo tasks whose instructions leave the goal unstated, no program succeeds without demonstrations. ENCORE also runs on a real bimanual robot, learning cube handover and cup inversion from five demonstrations each.
|
| 266 |
Learning Social Navigation from Internet Videos in the Policy State Space
2609.37476
|
cs.CV
|
Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna, Diwen Liu |
Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monoc...Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning.
|
| 267 |
Hierarchical Compression of Vision-Language Model Benchmarks
2609.37515
|
cs.CVcs.AI
|
Hyunjong Ok, Seunggu Kang, Jaeho Lee |
Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fra...Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical framework operates in four stages: data cleaning to remove items answerable without the image and all-correct items, category representative selection to pick one benchmark per capability category, item pruning with Vision-Aware Variance (VAW), and category-count pruning. VAW combines inter-model variance with a vision-dependence score computed from multimodal embeddings alone, while encouraging coverage of diverse items within each benchmark. On models held out from item selection, it has the highest mean fidelity at the released 5% retention. The hierarchical design lets practitioners stop at any stage to match their compute budget; the released suite removes over 97% of items while preserving model rankings. Beyond compression, our analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.
|
| 268 |
RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation
2609.37530
|
cs.CV
|
Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou |
Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked de...Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth. Our analysis reveals that RAW-to-RGB processing materially shapes both action prediction and manipulation success, with different ISP dimensions exerting substantially different effects. Guided by these findings, we introduce RawVLA, a streaming neural ISP that adaptively renders RAW observations for frozen VLA policies while concentrating its capacity on the imaging factors relevant to embodied behavior. We further present RawVLA-Bench, a RAW-domain manipulation benchmark to expose image processing as an explicit evaluation variable across clean and adverse acquisition conditions. Experiments on RawVLA-Bench show that RawVLA preserves performance under standard conditions while substantially improving robustness under degraded imaging, establishing adaptive RAW processing as an effective interface between physical cameras and embodied policies.
|
| 269 |
When to Adapt: Multi-Signal Domain Shift Detection for Efficient Training-Free Adaptation in Open-Vocabulary Segmentation
2609.37602
|
cs.CV
|
Michele Antonazzi, Alejandra C. Hernandez, Jos\'e Araujo, Olov Andersson, Patric Jensfelt |
Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (...Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (VFMs) have improved open-vocabulary semantic segmentation, these models can still suffer from domain shift, which can significantly degrade performance if they are not adapted to the current environment. Training-free domain adaptation is a relevant paradigm for adaptation, consisting of adjusting the model online using lightweight adapters. Recent approaches apply this on a per-frame basis, which is impractical for deployments on resource-constrained robotic hardware. To tackle this, we propose a multi-signal domain shift detection method for training-free continual test-time adaptation (TF-CTTA) in open-vocabulary segmentation. Our method leverages temporal coherence across consecutive frames by monitoring and combining complementary aspects of domain shift (visual change, adapter mismatch, and semantic drift) to trigger adaptation only when needed. We validate our approach on a benchmark including indoor and outdoor environments and using real robotic data. We demonstrate that our approach maintains segmentation accuracy while substantially reducing adaptations, making training-free adaptation practical and feasible for long-term, real-world robotic deployments.
|
| 270 |
Procedural Core: A Compact Recurrent Initialization for Vision Transformers
2609.37631
|
cs.CV
|
Zachary Shinnick, Christian Intern\`o, Hemanth Saratchandran, Anton van den Hengel, Damien Teney |
Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure...Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic structure into a compact set of weights that can be reused across models. We train a minimal recurrent transformer on procedural data, then expand its weights to initialize transformers of arbitrary width and depth. The resulting initialization improves performance on image classification, self-supervised visual learning (DINO), and modeling natural language (FineWeb-Edu) and code (CodeParrot). For image classification, expanding a 1M-parameter core to initialize an 85M-parameter ViT-Base improves ImageNet top-1 accuracy by 2.2 pp over standard random initialization. Our analysis identifies recurrence as essential for learning compact weights that transfer across models. In ViTs, we localize a key benefit in the suppression of high-norm tokens that produces substantial improvements in zero-shot segmentation (ImageNet-S mAP 32.3 to 42.9), object localization (VOC07 CorLoc 9.9 to 18.4), and depth estimation (NYUv2 RMSE 1.104 to 0.998). This demonstrates that transformers need not start from a blank slate, and can be initialized with generic capabilities at low cost with no domain- or task-specific data.
|
| 271 |
MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators
2609.37670
|
cs.CVcs.AI
|
Haocheng Tang, Tianchi Xie, Xingqiao Lin |
MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_...MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_0$-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow's native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.
|
| 272 |
CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces
2609.37721
|
cs.CV
|
Sen Wang, Liu Liu, Xinjiang Wang, Zequn Chen, Haoyi Jiang |
Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-acti...Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. We further introduce semantic training strategies to improve transition learning and closed-loop conditioning. Without additional robot-action pretraining, CogWAM achieves 15.56 / 11.70 % Score/SR on RoboDojo and state-of-the-art performance on BiCoord, while real-world experiments demonstrate closed-loop dual-arm manipulation with 16.4 fewer Semantic State regenerations than step-wise updating.
|
| 273 |
Selective Channel Restoration for Backdoored Vision-Language Models
2609.37759
|
cs.CV
|
Shuming Liu, Zhifang Zhang, Suqin Yuan, Khin Mi Mi Aung, Zhuoyi Lin |
Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead duri...Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these limitations, we propose Perturb-Select-Restore (PSR), a post-training defense that performs sparse updates to the projection interface and introduces no additional computation during inference. We reveal that backdoored VLM projectors are substantially more sensitive to bounded perturbations than clean VLM projectors, a phenomenon we term projection fragility. Building on this finding, PSR identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values. Experiments across multiple tasks show that PSR reduces attack success rates to near zero while preserving clean-task performance.
|
| 274 |
It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
2609.37863
|
cs.CVcs.AI
|
Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem, Lotem Peled-Cohen |
Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English s...Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.
|
| 275 |
Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics
2609.37907
|
cs.CVcs.AI
|
Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio |
Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been propos...Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM's outcome and we analyse our models on per-key and balanced metrics such as $F_1^{macro}$. Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.
|
| 276 |
Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution
2609.37950
|
cs.CVcs.AI
|
Bingjun Luo, Jialin Guo, Siqi Li |
Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations f...Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the basis for self-improvement. We introduce Video-RSI, a framework for recursive self-improvement in which a video understanding agent uses its own language model to revise its harness. Through active video investigation, the model revisits the original training videos to test competing failure explanations with additional observations, grounding proposed changes in evidence beyond the existing trace. Cost-aware harness evolution turns these diagnoses into reusable revisions and determines which revisions to retain by considering both answer accuracy and visual cost. Across our evaluation settings on video understanding benchmarks, the evolved agent improves accuracy while processing fewer frames and achieves competitive accuracy-efficiency trade-offs against existing video understanding agents. These results demonstrate the potential for video understanding agents to improve their own evidence acquisition and use through harness evolution. Code is available at https://github.com/bingjunluo/Video-RSI .
|
| 277 |
PhysWAM: Physically Consistent World Action Model for Autonomous Driving
2609.37970
|
cs.CV
|
Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye |
World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for ...World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.
|
| 278 |
$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient
2609.37976
|
cs.CVcs.AI
|
Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang, Ge Liu |
LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the correspon...LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. $S^3$ establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.
|
| 279 |
Brain-SAD: A Brain-Inspired Safe Autonomous Driving Control Framework with Dynamic Fear-Oriented Constraint on Dual-Policy
2609.38016
|
cs.CVcs.AI
|
Huan Rong, Chao Yin, Anouar Imel, Yijie Xia, Tinghuai Ma |
Constrained Reinforcement Learning has recently gained increasing attention in the field of Safe Autonomous Driving, where the general mechanism is to maximize the expected reward while keeping the overall action risk bounded. In this way, the safety issues ar...Constrained Reinforcement Learning has recently gained increasing attention in the field of Safe Autonomous Driving, where the general mechanism is to maximize the expected reward while keeping the overall action risk bounded. In this way, the safety issues arising in AD can be mitigated through constrained actions. However, existing Constrained RL methods still lack dynamics on the imposed constraints. For instance, the action cost adopted by the existing Primal-Dual/soft-constrained methods is often defined as static state-to-cost mapping, and the safe-action projection in hard-constrained methods relies on the static projection with the fixed feasible region boundary estimated from offline demonstrations. The above drawback tightly couples the imposed constraints to the training scenarios, leaving the AD policy hard to handle different interaction scenarios, due to the improper state-level action-cost and the static projection boundary. Consequently, in this paper, we propose Brain-SAD, a brain-inspired safe autonomous driving control framework with dynamic fear-oriented constraints. By perceiving the current vehicle-interaction scene, Brain-SAD generates dynamic fear signal as fear reaction to online decide long-term policy for regular interaction or short-term policy for urgent-collision defense. In such two policy, the above fear-reaction will be constructed as the dynamic fear constraints, respectively reflecting the overall fear cost directly coupled with action-impact, and the dynamic fear boundary of the feasible region derived from different risky neighbors, both of which will in turn serve for the online policy optimization. Experimental results show that Brain-SAD outperforms existing methods, achieving higher success rate in shorter task-completion and collision-recovery time, and exhibits stronger reliability across continuous intersections of fluctuating complexity.
|
| 280 |
doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving
2609.38028
|
cs.CVcs.AI
|
Parthib Roy, Yash Tandon, Marcus Blennemann, Giovanni Tapia Lopez, Angel Martinez-Sanchez |
Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as dri...Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.
|
| 281 |
Pow3R-SLAM: Real-Time RGB-D SLAM with 3D Reconstruction Priors
2609.38054
|
cs.CV
|
Christopher Kolios, Ishaan Mehta, Sasa Janjic, Yeganeh Bahoo, Sajad Saeedi |
We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incor...We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incorporate depth as a prior on the network's prediction, rather than as geometry to fuse. Where traditional RGB-D SLAM systems struggle with sparsity in the depth images, Pow3R utilizes the available depth to give a better-conditioned pointmap, while inferring the depths in empty regions from the two-view photometric, depth, and intrinsic data. Evaluated against MASt3R-SLAM following its protocol on 24 sequences from TUM, 7-Scenes, and Replica, Pow3R-SLAM runs 1.6x faster in wall time, has 15% lower mean trajectory error, a 3.1x lower unscaled error, and produces denser maps, with a 30% lower Chamfer distance. We also introduce a hybrid variant that runs 2.1x faster than MASt3R-SLAM at 25.3 frames per second (FPS), while maintaining improved tracking and mapping accuracy. Against ORB-SLAM3 in RGB-D mode, Pow3R-SLAM is more accurate on TUM, 7-Scenes, and ETH3D-SLAM, and completes every TUM sequence. While Pow3R-SLAM can struggle on a small set of self-similar scenes, its overall performance shows that adding depth as a prior for two-view 3D reconstruction SLAM can be beneficial. A project webpage is available at: https://ChrisKolios.github.io/Pow3R-SLAM , and code will be made open-source upon acceptance.
|
| 282 |
WorldLine: Action-Driven Visual Simulation for Robotic Manipulation
2609.38059
|
cs.CV
|
Shenghe Zheng, Wenbo Li, Jiyao Zhang, Bin Xia, Haoyang Huang |
Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often f...Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at \href{https://zhengsh123.github.io/WorldLine/}{project page}.
|
| 283 |
From Routing Signals to Selective Review: Visual regrounding in MoE VLMs
2609.38111
|
cs.CV
|
Hongzhu Guo, Mohsen Fayyaz, Nanyun Peng |
Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-...Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, hidden states, or uncertainty measures. We present the first framework to leverage internal routing decisions in Mixture-of-Experts (MoE) VLMs to detect target absence before generation and guide selective correction. We extract target-token routing probabilities from Qwen3-VL-30B-A3B-Instruct and Gemma-4-26B-A4B-it, train a separate L2-regularized linear detector for each model, and use its predictions to selectively invoke a target-aware review prompt. Using routing alone, the Qwen and Gemma detectors achieve ROC-AUCs of 0.9988 and 0.9956 on GQA-Inpaint and retain 0.8095 and 0.7781 on the external OBER dataset, respectively. The resulting routing-gated policy improves end-to-end accuracy on GQA-Inpaint and OBER by +22.25% and +12.17% for Qwen, and by +13.42% and +1.39% for Gemma, without modifying model weights. Further analysis shows that the signal is localized to the target-object token, emerges in early MoE layers, and is distributed across partially substitutable experts. Although cross-dataset threshold shifts require recalibration, false-positive review causes limited harm overall, suggesting that intervention risk can be controlled through joint selection of the detector threshold and review prompt. Overall, we show that routing probabilities alone preserve actionable information about visual perception, allowing computation already produced by an MoE VLM to support low-cost detection and selective visual regrounding.
|
| 284 |
Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
2609.38172
|
cs.CV
|
Zihan Wang, Zhen Wu, Pieter Abbeel, Rocky Duan, Jitendra Malik |
Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body...Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.
|
| 285 |
Hybrid Feature Learning for Handwriting Verification
1812.02621
|
cs.CV
|
Seyed Mohammad Abuzar Hashemi, Mihir Chauhan, Jun Chu, Sargur Srihari |
We propose an effective Hybrid Deep Learning (HDL) architecture for the task of determining the probability that a questioned handwritten word has been written by a known writer. HDL is an amalgamation of Auto-Learned Features (ALF) and Human-Engineered Featur...We propose an effective Hybrid Deep Learning (HDL) architecture for the task of determining the probability that a questioned handwritten word has been written by a known writer. HDL is an amalgamation of Auto-Learned Features (ALF) and Human-Engineered Features (HEF). To extract auto-learned features we use two methods: First, Two Channel Convolutional Neural Network (TC-CNN); Second, Two Channel Autoencoder (TC-AE). Furthermore, human-engineered features are extracted by using two methods: First, Gradient Structural Concavity (GSC); Second, Scale Invariant Feature Transform (SIFT). Experiments are performed by complementing one of the HEF methods with one ALF method on 150000 pairs of samples of the word "AND" cropped from handwritten notes written by 1500 writers. Our results indicate that HDL architecture with AE-GSC achieves 99.7% accuracy on seen writer dataset and 92.16% accuracy on shuffled writer dataset which out performs CEDAR-FOX, as for unseen writer dataset, AE-SIFT performs comparable to this sophisticated handwriting comparison tool.
|
| 286 |
AdvMT: Adversarial Motion Transformer for Long-term Human Motion Prediction
2401.05018
|
cs.CV
|
Sarmad Idrees, Seokman Sohn, Jongeun Choi |
Human motion prediction is a crucial capability for advanced robotic systems that interact with humans. In facilities with dynamic human-robot collaboration settings, robots must anticipate human movements to ensure safety, prevent collisions, and optimize coo...Human motion prediction is a crucial capability for advanced robotic systems that interact with humans. In facilities with dynamic human-robot collaboration settings, robots must anticipate human movements to ensure safety, prevent collisions, and optimize cooperative tasks. Traditionally, motion forecasting is treated as a sequential modeling problem using historical pose data, but achieving long-term accuracy and physical realism remains challenging. We present Adversarial Motion Transformer (AdvMT), a novel approach that integrates a Transformer-based motion encoder with a temporal continuity discriminator to address these challenges. The Transformer captures rich spatio-temporal dependencies across human joints, while adversarial training with a continuity discriminator enforces smooth, natural motion trajectories that adhere to biomechanical constraints. Our training scheme includes a bone-length consistency term and adversarial loss to reduce common artifacts like pose freezing or unnatural transitions. In experiments on the Human3.6M motion dataset, AdvMT achieves state-of-the-art long-horizon prediction accuracy while also delivering robust short-term predictions. These improvements strengthen the prediction foundation for physical AI in manufacturing and human-robot collaboration, where anticipating human motion is a prerequisite for safe and efficient robot coordination.
|
| 287 |
Segment Anything for Dendrites from Electron Microscopy
2411.02562
|
cs.CV
|
Zewen Zhuo, Ilya Belevich, Ville Leinonen, Eija Jokitalo, Tarja Malm |
Segmentation of cellular structures in electron microscopy (EM) images is fundamental to analyzing the morphology of neurons and glial cells in the healthy and diseased brain tissue. Current neuronal segmentation applications are based on convolutional neural ...Segmentation of cellular structures in electron microscopy (EM) images is fundamental to analyzing the morphology of neurons and glial cells in the healthy and diseased brain tissue. Current neuronal segmentation applications are based on convolutional neural networks (CNNs) and do not effectively capture global relationships within images. Here, we present DendriteSAM, a vision foundation model based on Segment Anything, for interactive and automatic segmentation of dendrites in EM images. The model is trained on high-resolution EM data from healthy rat hippocampus and is tested on diseased rat and human data. Our evaluation results demonstrate better mask quality compared to the original and other fine-tuned models, leveraging the features learned during training. This study introduces the first implementation of vision foundation models in dendrite segmentation, paving the path for computer-assisted diagnosis of neuronal anomalies.
|
| 288 |
TeD-Loc: Text Distillation for Weakly Supervised Object Localization
2501.12632
|
cs.CV
|
Shakeeb Murtaza, Soufiane Belharbi, Alexis Guichemerre, Marco Pedersoli, Eric Granger |
Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, without requiring explicit bounding-box annotations. Given their reliance on classification objectives, traditional WSOL ...Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, without requiring explicit bounding-box annotations. Given their reliance on classification objectives, traditional WSOL methods, like class activation mapping, tend to focus on the most discriminative object regions, often missing the full spatial extent. Although vision-language models like CLIP encode rich semantic priors, their global text and class-token embeddings are not explicitly aligned with local patch embeddings, limiting patch-level localization. Recent methods such as GenPrompt address this limitation, but at the cost of increased complexity, as they rely on conditional denoising and elaborate prompt-learning strategies. In this paper, we propose Text Distillation for Localization (TeD-Loc), which distills knowledge from CLIP text embeddings to patch embeddings through contrastive alignment, thereby enabling patch-level foreground/background localization. A localization-guided classification module is also introduced, which uses localization scores to aggregate foreground patch embeddings for joint classification and localization within a single model. In addition, a QR-based orthogonalization of class text embeddings is applied before distillation to improve discrimination for semantically similar classes. Extensive experiments show that TeD-Loc improves Top-1 Loc by ~5% on CUB and ILSVRC, and PxAP by ~31% on histopathology benchmarks, while achieving more efficient inference than GenPrompt.
|
| 289 |
From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identification
2503.12232
|
cs.CV
|
Yan Jiang, Hao Yu, Xu Cheng, Haoyu Chen, Zhaodong Sun |
Aiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distr...Aiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distributed across multiple devices/entities, raising privacy and ownership concerns that make existing centralized training impractical for VI-ReID. To tackle these challenges, we propose L2RW, a benchmark that brings VI-ReID closer to real-world applications. The rationale of L2RW is that integrating decentralized training into VI-ReID can address privacy concerns in scenarios with limited data-sharing regulation. Specifically, we design protocols and corresponding algorithms for different privacy sensitivity levels. In our new benchmark, we ensure the model training is done in the conditions that: 1) data from each camera remains completely isolated, or 2) different data entities (e.g., data controllers of a certain region) can selectively share the data. In this way, we simulate scenarios with strict privacy constraints which is closer to real-world conditions. Intensive experiments with various server-side federated algorithms are conducted, showing the feasibility of decentralized VI-ReID training. Notably, when evaluated in unseen domains (i.e., new data entities), our L2RW, trained with isolated data (privacy-preserved), achieves performance comparable to SOTAs trained with shared data (privacy-unrestricted). We hope this work offers a novel research entry for deploying VI-ReID that fits real-world scenarios and can benefit the community.
|
| 290 |
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
2506.01004
|
cs.CVcs.AI
|
Tong Zhang, Victor Escorcia, Juan C Leon Alcazar, Bernard Ghanem |
Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybrid while preserving the source video's motion and layout. We propose MoCA-Video, a training-free framework that steers ...Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybrid while preserving the source video's motion and layout. We propose MoCA-Video, a training-free framework that steers a frozen video-diffusion denoising trajectory through concept-localized reference injection. At selected low-noise steps, MoCA-Video uses concept attention to localize the target object and injects the reference latent into the localized region, where object structure has formed but appearance remains editable. A momentum-based correction carries the injected prediction across frames to encourage coherent concept integration through the sequence. We further introduce CASS, a CLIP-based metric that measures the output's directional alignment shift toward the reference and away from the source prompt. Using the denoiser's internal attention avoids an external localization model; in our A100 FP16 setup, MoCA-Video takes 3.2 seconds per output frame, excluding preprocessing. Across the evaluated baselines, MoCA-Video achieves the highest CASS, rel-CASS, and ImageReward, while LPIPS-T and FVD expose separate temporal-coherence and video-quality trade-offs.
|
| 291 |
Moving Object Detection from Moving Camera Using Focus of Expansion Likelihood and Segmentation
2507.13628
|
cs.CV
|
Masahiro Ogawa, Qi An, Atsushi Yamashita |
Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understanding in robotics. Existing approaches often rely primarily on optical flow, which struggle to detect moving object...Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understanding in robotics. Existing approaches often rely primarily on optical flow, which struggle to detect moving objects in complex, structured scenes involving camera motion. To address this limitation, we propose Focus of Expansion Likelihood and Segmentation (FoELS), a method based on the core idea of integrating both optical flow and texture information. FoELS computes the focus of expansion (FoE) from optical flow and derives an initial motion likelihood from the outliers of the FoE computation. This likelihood is then fused with a segmentation-based prior to estimate the final moving probability. The method effectively handles challenges including complex structured scenes, rotational camera motion, and parallel motion. Comprehensive evaluations on the DAVIS 2016 and FBMS-59 datasets, along with real-world traffic videos including parallel, cross-direction, opposite-direction, and crowded scenes, demonstrate its effectiveness and state-of-the-art performance.
|
| 292 |
Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
2508.13009
|
cs.CV
|
Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang |
Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynamics and interactive behaviors. However, existing interactive world models depend on bidirectional attention and le...Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynamics and interactive behaviors. However, existing interactive world models depend on bidirectional attention and lengthy inference steps, severely limiting real-time performance. Consequently, they are hard to simulate real-world dynamics, where outcomes must update instantaneously based on historical context and current actions. To address this, we present Matrix-Game 2.0, an interactive world model generates long videos on-the-fly via few-step auto-regressive diffusion. Our framework consists of three key components: (1) A scalable data production pipeline for Unreal Engine and GTA5 environments to effectively produce massive amounts (about 1200 hours) of video data with diverse interaction annotations; (2) An action injection module that enables frame-level mouse and keyboard inputs as interactive conditions; (3) A few-step distillation based on the casual architecture for real-time and streaming video generation. Matrix Game 2.0 can generate high-quality minute-level videos across diverse scenes at an ultra-fast speed of 25 FPS. We open-source our model weights and codebase to advance research in interactive world modeling.
|
| 293 |
Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models
2508.13744
|
cs.CVcs.AI
|
Yeji Park, Minyoung Lee, Sanghyuk Chun, Junsuk Choe |
Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While this degradation has been observed in prior work, its nature remains poorly underst...Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While this degradation has been observed in prior work, its nature remains poorly understood. We empirically observe visual elements from different images become entangled in the model's representations and responses. We refer to this phenomenon as cross-image information leakage. To address this issue, we propose FOCUS, a training-free and architecture-agnostic method. FOCUS masks all but one image with random noise, guiding the model to focus on the single clean image. This process is applied across the target images to obtain logits under partially masked contexts. These logits are aggregated and then refined using a noise-only reference input, which suppresses the leakage and yields more accurate outputs. FOCUS consistently improves performance on diverse multi-image benchmarks. We further show that FOCUS generalizes to video understanding, extending its applicability beyond static multi-image inputs. This demonstrates that FOCUS offers a general solution for enhancing multi-image reasoning without additional training or architectural modifications.
|
| 294 |
Achieving detailed medial temporal lobe segmentation with upsampled isotropic training from implicit neural representation
2508.17171
|
cs.CV
|
Yue Li, Pulkit Khandelwal, Rohit Jena, Long Xie, Michael Duong |
Imaging biomarkers in magnetic resonance imaging (MRI) are important tools for diagnosing, tracking and treating Alzheimer's disease (AD). Neurofibrillary tau pathology in AD is closely linked to neurodegeneration and generally follows a pattern of spread in t...Imaging biomarkers in magnetic resonance imaging (MRI) are important tools for diagnosing, tracking and treating Alzheimer's disease (AD). Neurofibrillary tau pathology in AD is closely linked to neurodegeneration and generally follows a pattern of spread in the brain, with early stages involving subregions of the medial temporal lobe (MTL). Accurate segmentation of MTL subregions is needed to extract granular biomarkers of AD progression. MTL subregions are often imaged using T2-weighted (T2w) MRI scans that are highly anisotropic due to constraints of MRI physics and image acquisition, making it difficult to reliably model MTL subregions geometrically and extract morphological measures, such as thickness. In this study, we propose a segmentation framework for MTL subregions in isotropic space, in which an implicit neural representation is used to construct the isotropic training atlas from the anisotropic low-resolution T2w data, with T1w MRI as an auxiliary modality to support the INR and segmentation. In an independent test set, the morphological measures extracted using this isotropic model showed stronger effect sizes than those from models trained on anisotropic data in distinguishing participants with mild cognitive impairment (MCI) from cognitively unimpaired individuals. In the test-retest analysis, the morphological measures extracted using the isotropic model showed greater stability than those from the anisotropic segmentation. This study demonstrates improved reliability of MRI-derived MTL subregion biomarkers without additional atlas annotation effort, which may more accurately quantify and track the relationship between AD pathology and brain atrophy for monitoring disease progression.
|
| 295 |
Geometrically Constrained and Token-Based Probabilistic Spatial Transformers
2509.11218
|
cs.CVcs.AI
|
Johann Schmidt, Tom Siegl, Martin Becker, Sebastian Stober |
Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful consideration is required for reliable use in high stakes settings. A model should stay robust under such transformations, expos...Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful consideration is required for reliable use in high stakes settings. A model should stay robust under such transformations, expose why a correction was applied, and signal when its input is ambiguous. While geometrically equivariant architectures provide a mathematically grounded solution, they often limit model flexibility through strict symmetry constraints and incur significant computational overhead. Spatial Transformer Networks (STNs) offer a data-driven, flexible alternative for learning pseudo-equivariances to affine transformations. However, STNs have historically been restricted to convolutional architectures and suffer from training instability. To address this, we introduce a novel STN framework. It leverages the global modeling capabilities of transformers to regress the affine transformation acting on the input. For this, we decompose affine transformations into interpretable primitives, regressed under adaptable geometric constraints, thereby preventing the training instability typically caused by degenerate transformations. By sharing weights between the localization network and the classification backbone, the framework requires minimal computational overhead. Extensive experiments on challenging insect biodiversity and medical imaging benchmarks demonstrate that our approach achieves superior predictive performance under diverse spatial transformations while maintaining high efficiency. Code is available at https://github.com/johSchm/TokenSTN.
|
| 296 |
Hybrid Approach for Enhancing Lesion Segmentation in Fundus Images
2509.25549
|
cs.CVcs.AI
|
Mohammadmahdi Eshragh, Emad A. Mohammed, Behrouz Far, Ezekiel Weis, Carol L Shields |
Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical to improving survival rates, but misdiagnosis or delayed diagnosis can lead to poor outcomes. Despite advancements in AI...Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical to improving survival rates, but misdiagnosis or delayed diagnosis can lead to poor outcomes. Despite advancements in AI-based image analysis, diagnosing choroidal nevi in colour fundus images remains challenging, particularly for clinicians without specialized expertise. Existing datasets often suffer from low resolution and inconsistent labelling, limiting the effectiveness of segmentation models. This paper addresses the challenge of achieving precise segmentation of fundus lesions, a critical step toward developing robust diagnostic tools. While deep learning models like U-Net have demonstrated effectiveness, their accuracy heavily depends on the quality and quantity of annotated data. Previous mathematical/clustering segmentation methods, though accurate, required extensive human input, making them impractical for medical applications. This paper proposes a novel approach that combines mathematical/clustering segmentation models with insights from U-Net, leveraging the strengths of both methods. This hybrid model improves accuracy, reduces the need for large-scale training data, and achieves significant performance gains on high-resolution fundus images. The proposed model achieves a Dice coefficient of 89.7% and an IoU of 80.01% on 1024*1024 fundus images, outperforming the Attention U-Net model, which achieved 51.3% and 34.2%, respectively. It also demonstrated better generalizability on external datasets. This work forms a part of a broader effort to develop a decision support system for choroidal nevus diagnosis, with potential applications in automated lesion annotation to enhance the speed and accuracy of diagnosis and monitoring.
|
| 297 |
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
2510.09741
|
cs.CV
|
Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra, Jeonghwan Kim, Madhav Kanda |
Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce AttWarp, a lightweight method that allocates more resolution to query-relevant con...Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce AttWarp, a lightweight method that allocates more resolution to query-relevant content while compressing less informative areas, all while preserving global context. At test time, the approach uses an MLLM's cross-modal attention to perform rectilinear warping of the input image, reallocating spatial resolution toward regions the model deems important, without changing model weights or architecture. This attention-guided warping preserves all original image information but redistributes it non-uniformly, so small objects and subtle relationships become easier for the same model to read while the global layout remains intact. Across five benchmarks (TextVQA, GQA, DocVQA, POPE, MMMU) and four MLLMs (LLaVA, Qwen-VL, InternVL, and InstructBLIP), AttWarp consistently improves accuracy, strengthens compositional reasoning, and reduces hallucinations, outperforming four competitive baselines that manipulate raw images at test time. Together, these results show that attention-guided warping prioritizes information relevant to the query while preserving context, and that the same MLLMs perform better when given such warped inputs.
|
| 298 |
Reverberation: Learning the Latencies Before Forecasting Trajectories
2511.11164
|
cs.CV
|
Conghao Wong, Ziqian Zou, Beihao Xia, Xinge You |
Bridging the past to the future, connecting agents both spatially and temporally, lies at the core of the trajectory prediction task. Despite great efforts, it remains challenging to explicitly learn and predict latencies, i.e., response intervals or temporal ...Bridging the past to the future, connecting agents both spatially and temporally, lies at the core of the trajectory prediction task. Despite great efforts, it remains challenging to explicitly learn and predict latencies, i.e., response intervals or temporal delays with which agents respond to various trajectory-changing events and adjust their future paths, whether on their own or interactively. Different agents may exhibit distinct latency preferences for noticing, processing, and reacting to a specific trajectory-changing event. The lack of consideration of such latencies may undermine the temporal continuity of forecasting systems, leading to implausible or unintended trajectories. Inspired by reverberation in acoustics, we propose a new reverberation transform and the corresponding Reverberation (Rev for short) trajectory prediction model, which predicts both individual latency preferences and their stochastic variations accordingly, by using two explicit and learnable reverberation kernels, enabling latency-conditioned and controllable trajectory prediction under both non-interactive and social latencies. Experiments on multiple datasets, whether of pedestrians or vehicles, demonstrate that Rev achieves competitive accuracy while revealing interpretable latency dynamics across agents and scenarios. Qualitative analyses further verify the properties of the reverberation transform, highlighting its potential as a general latency modeling approach.
|
| 299 |
Think, Then Look: Active Spatial Reasoning for House-Scale 3D Scene Understanding
2512.03284
|
cs.CV
|
Hongpei Zheng, Shijie Li, Lin Qian, Zhenghao Li, Qijun Yang |
Spatial reasoning in large-scale 3D environments remains challenging for current vision--language models, which are typically constrained to room-scale scenarios. We formalize Active House-Scale Spatial Reasoning (AHSR), a new paradigm in which a model reasons...Spatial reasoning in large-scale 3D environments remains challenging for current vision--language models, which are typically constrained to room-scale scenarios. We formalize Active House-Scale Spatial Reasoning (AHSR), a new paradigm in which a model reasons over a pre-built house-scale 3D map via virtual spatial tool invocations to answer spatial questions, without exhaustive scene-wide processing. To support AHSR research, we introduce H$^2$U3D (Holistic House Understanding in 3D), the first benchmark targeting house-scale 3D scene understanding, featuring environments with an average aggregate floor area of 250.8 m$^2$ and up to three floors, together with hierarchical coarse-to-fine visual representations. Building on H$^2$U3D, we propose SpatialReasoner, an AHSR framework trained via supervised fine-tuning with self-correction, followed by reinforcement learning with a task-aware adaptive exploration reward. SpatialReasoner achieves state-of-the-art performance on H$^2$U3D with 64.9% overall accuracy, outperforming strong baselines including GPT-5.4 and Gemini-3.5-Flash, and generalizes effectively to MT-HM3D and HM-EQA. These results demonstrate the clear advantage of active map-directed exploration over passive scene-wide processing in house-scale 3D understanding.
|
| 300 |
Event-based Scene Synthesis via Inter-Frame Residual Alignment
2512.17323
|
cs.CV
|
Jiyun Kong, Jun-Hyuk Kim, Jong-Seok Lee |
Event-based scene synthesis reconstructs target RGB frames from sparse image observations and asynchronous event streams, encompassing both video frame prediction and interpolation. Existing event-based synthesis methods commonly estimate optical flow to warp ...Event-based scene synthesis reconstructs target RGB frames from sparse image observations and asynchronous event streams, encompassing both video frame prediction and interpolation. Existing event-based synthesis methods commonly estimate optical flow to warp the observed frames toward the target time, but are vulnerable to inaccurate flow under large motion and occlusion and often rely on flow supervision or pretrained estimators. In this work, we propose EvFRA, an Event-based scene synthesis framework based on inter-Frame Residual Alignment. We identify a structural correspondence between event measurements and frame-to-frame scene changes, and exploit this correspondence for target frame synthesis. Our training pipeline consists of two stages: 1) an Event-to-Residual Alignment Variational Autoencoder (ER-VAE) aligns the event frame captured between the anchor and target frames with the corresponding inter-frame residual, and 2) a ControlNet-conditioned diffusion model is fine-tuned to denoise the residual latent using event data. Our method outperforms state-of-the-art methods by up to 2.61 dB and 1.85 dB in PSNR for frame prediction and interpolation, respectively, with consistent SSIM improvements. Code is available at https://github.com/jiyun-kong/EvFRA.
|
| 301 |
Training-Free Global Geometric Association for 4D LiDAR Panoptic Segmentation
2512.18991
|
cs.CVcs.AI
|
Gyeongrok Oh, Youngdong Jang, Jonghyun Choi, Suk-Ju Kang, Guang Lin |
Dominant paradigms for 4D LiDAR panoptic segmentation are usually required to train deep neural networks with large superimposed point clouds or design dedicated modules for instance association. However, these approaches perform redundant point processing and...Dominant paradigms for 4D LiDAR panoptic segmentation are usually required to train deep neural networks with large superimposed point clouds or design dedicated modules for instance association. However, these approaches perform redundant point processing and consequently become computationally expensive, yet still overlook the rich geometric priors inherently provided by raw point clouds. To this end, we introduce \textsc{Geo-4D}, a simple yet effective training-free framework that unifies spatial and temporal reasoning, enabling holistic LiDAR perception over long time horizons. Specifically, we propose a global geometric association strategy that establishes consistent instance correspondences by estimating an optimal transformation between instance-level point sets. To mitigate instability caused by structural inconsistencies in point cloud observations, we propose a global geometry-aware soft matching mechanism that enforces spatially coherent point-wise correspondences grounded in the spatial distribution of instance point sets. Furthermore, our carefully designed pipeline, which considers three instance types-static, dynamic, and missing-offers computational efficiency and occlusion-aware matching. Our extensive experiments across both SemanticKITTI and nuScenes demonstrate that our method consistently outperforms state-of-the-art approaches, even without additional training or extra point cloud inputs.
|
| 302 |
EVolSplat4D: Efficient Volume-based Gaussian Splatting for 4D Urban Scene Synthesis
2601.15951
|
cs.CV
|
Sheng Miao, Sijin Li, Pan Wang, Dongfeng Bai, Bingbing Liu |
Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle to balance reconstruction time with quality. While state-of-the-art neural radiance fields and 3D Gaussian Splatti...Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle to balance reconstruction time with quality. While state-of-the-art neural radiance fields and 3D Gaussian Splatting approaches achieve photorealism, they often rely on time-consuming per-scene optimization. Conversely, emerging feed-forward methods frequently adopt per-pixel Gaussian representations, which lead to 3D inconsistencies when aggregating multi-view predictions in complex, dynamic environments. We propose EvolSplat4D, a feed-forward framework that moves beyond existing per-pixel paradigms by unifying volume-based and pixel-based Gaussian prediction across three specialized branches. For close-range static regions, we predict consistent geometry of 3D Gaussians over multiple frames directly from a 3D feature volume, complemented by a semantically-enhanced image-based rendering module for predicting their appearance. For dynamic actors, we utilize object-centric canonical spaces and a motion-adjusted rendering module to aggregate temporal features, ensuring stable 4D reconstruction despite noisy motion priors. Far-Field scenery is handled by an efficient per-pixel Gaussian branch to ensure full-scene coverage. Experimental results on the KITTI-360, KITTI, Waymo, and PandaSet datasets show that EvolSplat4D reconstructs both static and dynamic environments with superior accuracy and consistency, outperforming both per-scene optimization and state-of-the-art feed-forward baselines.
|
| 303 |
Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Science
2602.03915
|
cs.CVcs.AI
|
Levi Lingsch, Georgios Kissas, Johannes Jakubik, Siddhartha Mishra |
Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned, generated, and generalized to new tasks. While foundational for image and video generation, the a...Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned, generated, and generalized to new tasks. While foundational for image and video generation, the application of tokens to physical simulation remains nascent. Because existing tokenizers are designed for the perceptual requirements of natural images, they struggle with scientific data, which exhibits large dynamic ranges and requires exact preservation of physical and spectral properties. In this work, we investigate the performance of a suite of image tokenizers across metrics designed to measure PDE fidelity. Observing that these baselines struggle to simultaneously capture fine geometric details and precise physical magnitudes, we propose Phaedra, a novel tokenizer inspired by classical shape-gain quantization and the paradigm of basis functions coupled with continuous coefficients. Phaedra acts as a highly effective nonlinear compression algorithm, massively reducing dataset footprints while maintaining physical fidelity. We demonstrate that Phaedra consistently improves reconstruction across diverse 2D gridded PDE solutions, generalizes robustly to unseen PDE types and real-world Earth observation data, and is competitive with continuous models in downstream proof-of-concept operator learning and masked autoencoding tasks.
|
| 304 |
From Concept Erasure to Style Purification: Contrastive Eigenbases for Artist Style Protection
2602.08059
|
cs.CVcs.AI
|
Tong Zhang, Ru Zhang, Jianyi Liu |
Text-to-image diffusion models can reproduce specific artists visual styles at extremely low cost, raising copyright and deployment safety concerns about unauthorized style mimicry. Existing model-side protection methods generally follow ordinary concept erasu...Text-to-image diffusion models can reproduce specific artists visual styles at extremely low cost, raising copyright and deployment safety concerns about unauthorized style mimicry. Existing model-side protection methods generally follow ordinary concept erasure, emphasizing aggressive deletion or redirection of target styles. However, our causal intervention analysis shows that the central issue is not insufficient erasure strength, but a mismatch between artist styles and this paradigm: unlike ordinary object concepts, artist styles do not form compact, localized editable semantic units. Consequently, sparse editing and fixed retain lists struggle to suppress target styles while preserving generation utility. We therefore reformulate artist style protection as style purification, suppressing target style expression during inference while preserving the requested content and visual structure. We propose CAPE (Contrastive Artist Style Purification with Eigenbases), a training-free framework against artist style mimicry. CAPE constructs contrastive triplets around the target request and formulates style direction estimation as a generalized eigenvalue problem, capturing style-related directions that remain stable across content variations and are less affected by shared content. During inference, CAPE employs the Adaptive Suppression Controller to assign suppression strengths to different tokens based on Q, K, and V responses, and performs target style suppression on the K and V paths of self-attention. Experimental results show that CAPE effectively weakens target artist characteristics, including brushstrokes, textures, and local color processing, while better preserving major semantic entities, scene composition, and visual structures.
|
| 305 |
ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
2602.09839
|
cs.CV
|
Yijie Lin, Guofeng Ding, Haochen Zhou, Haobin Li, Mouxing Yang |
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal r...Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal retrieval from two complementary perspectives: (i) knowledge domains (five domains with 17 subtypes), which characterize the content and expertise retrieval relies on, and (ii) reasoning skills (six categories), which characterize the type of inference over multimodal evidence required to identify the correct candidate. Specifically, ARK evaluates retrieval with both unimodal and multimodal queries and candidates, covering 16 heterogeneous visual data types. To avoid shortcut matching during evaluation, most queries are paired with targeted hard negatives that require multi-step reasoning. We evaluate 25 representative text-based and multimodal retrievers and observe a pronounced gap between knowledge- and reasoning-intensive retrieval, with fine-grained visual and spatial reasoning as persistent bottlenecks. We further show that enhancements such as re-ranking, rewriting, and agentic retrieval yield consistent gains, but substantial headroom remains.
|
| 306 |
COMiT: Learning Structured Visual Tokens through Sequential Communication
2602.20731
|
cs.CVcs.AI
|
Aram Davtyan, Yusuf Sahin, Yasaman Haghighi, Sebastian Stapf, Pablo Acuaviva |
Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a com...Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations. COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the existing token sequence. After several iterations, the resulting message conditions a flow-matching decoder that reconstructs the complete image. The encoder and decoder are implemented within a single transformer and trained end-to-end using flow-matching reconstruction and semantic representation-alignment objectives. COMiT substantially improves compositional generalization and relational reasoning over prior methods. Our experiments show that, while semantic alignment helps ground the representation, attentive sequential tokenization is critical for inducing more interpretable, object-centric token structures.
|
| 307 |
Evaluating Generative Models via One-Dimensional Code Distributions
2603.08064
|
cs.CV
|
Zexi Jia, Pengcheng Luo, Yijia Zhong, Jinchao Zhang, Jie Zhou |
Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality...Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the space of discrete visual tokens, where modern 1D image tokenizers compactly encode both semantic and perceptual information and quality manifests as predictable token statistics. We introduce Codebook Histogram Distance (CHD), a training-free distribution metric in token space, and Code Mixture Model Score (CMMS), a no-reference quality metric learned from synthetic degradations of token sequences. To stress-test metrics under broad distribution shifts, we further propose VisForm, a benchmark of 210K images spanning 62 visual forms and 12 generative models with expert annotations. Across AGIQA, HPDv2/3, and VisForm, our token-based metrics achieve state-of-the-art correlation with human judgments. We will release all code and datasets to facilitate future research, with the code publicly available at https://github.com/zexiJia/1d-Distance.
|
| 308 |
HARMONI: Aligning Human and Scene Priors for Multi-View 4D Reconstruction
2603.12789
|
cs.CV
|
Sangmin Kim, Minhyuk Hwang, Geonho Cha, Dongyoon Wee, Jaesik Park |
Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independently trained human and scene priors often produces misalignment in scale and depth. We observe that the two prio...Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independently trained human and scene priors often produces misalignment in scale and depth. We observe that the two priors have complementary strengths. The scene prior provides consistent depth but approximate scale, while the human prior provides fixed body scale but less reliable depth. Motivated by this observation, we present HARMONI, a feed-forward framework that reconstructs cameras, scene, and humans with identities from monocular or multi-view video without test-time optimization. We introduce bidirectional anchoring, in which scene depth guides human placement, while human keypoints calibrate the scene scale to match the human body. While most existing approaches target monocular inputs and multi-view methods rely on optimization or re-identification, our approach naturally extends to multiple views. Building on bidirectional anchoring, we introduce a multi-view fusion module that merges per-view estimates, while reducing the influence of unreliable views and visually similar individuals. Experiments show that our framework outperforms previous human-scene methods in global motion and multi-view pose estimation by up to 28% and 64%, while running 28 times faster than optimization-based approaches. Project page: https://nstar1125.github.io/harmoni.
|
| 309 |
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
2603.14733
|
cs.CV
|
Yue Zhang, Liqiang Jing, Jia Li, Yapeng Tian, Xinya Du |
Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single input and perform direc...Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single input and perform direct inference, which introduces training-inference mismatch, information loss from frame compression, and a lack of explicit cross-video coordination. Meanwhile, current multi-video benchmarks primarily emphasize event-level comparison, leaving identity-level matching, fine-grained discrimination, and structured multi-step reasoning underexplored. To address these gaps, we introduce MVX-Bench, a Multi-Video Cross-Dimension Benchmark that reformulates 11 classical computer vision tasks into a unified multi-video question-answering framework, comprising 1,442 questions over 4,255 videos from diverse real-world datasets. We further propose SAMA, a Skill-Augmented Agentic Framework for Multi-Video Understanding, which integrates visual tools, task-specific skills, and a conflict-aware verification mechanism to enable iterative and structured reasoning. Experimental results show that SAMA outperforms strong open-source baselines and GPT on MVX-Bench, and ablations validate the effectiveness of skill design and conflict resolution.
|
| 310 |
FusionBERT: Multi-View Image--3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
2604.02583
|
cs.CV
|
Wei Li, Yufan Ren, Hanqing Jiang, Jianhui Ding, Zhen Peng |
We propose FusionBERT, a novel multi-view visual fusion framework for image--3D multimodal retrieval. Existing image--3D representation learning methods predominantly focus on feature alignment of a single object image and its 3D model, limiting their applicab...We propose FusionBERT, a novel multi-view visual fusion framework for image--3D multimodal retrieval. Existing image--3D representation learning methods predominantly focus on feature alignment of a single object image and its 3D model, limiting their applicability in realistic scenarios where an object is typically observed and captured from multiple viewpoints. Although multi-view observations naturally provide complementary geometric and appearance cues, existing multimodal large models rarely explore how to effectively fuse such multi-view visual information for better cross-modal retrieval. To address this limitation, we introduce a multi-view image--3D retrieval framework named FusionBERT, which innovatively utilizes a cross-attention-based multi-view visual aggregator to adaptively integrate features from multi-view images of an object. The proposed multi-view visual encoder fuses inter-view complementary relationships and selectively emphasizes informative visual cues across multiple views to get a more robustly fused visual feature for better 3D model matching. Furthermore, FusionBERT proposes a normal-aware 3D model encoder that can further enhance the 3D geometric feature of an object model by jointly encoding point normals and 3D positions, enabling a more robust representation learning for textureless or color-degraded 3D models. Extensive image--3D retrieval experiments on both synthetic 3D models and real-world industrial mechanical objects demonstrate that FusionBERT achieves significantly higher retrieval accuracy than SOTA multimodal large models under both single-view and multi-view settings, establishing a strong baseline for multi-view multimodal retrieval.
|
| 311 |
Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs
2604.03114
|
cs.CVcs.AI
|
Zhangyun Tan, Zeliang Zhang, Jiani Liu, Susan Liang, Yolo Y. Tang |
Vision-language models (VLMs) may need to forget visual concepts after deployment because of privacy, copyright, licensing, safety, or policy changes. Conventional machine unlearning modifies model parameters, which may be costly or inaccessible for API-only m...Vision-language models (VLMs) may need to forget visual concepts after deployment because of privacy, copyright, licensing, safety, or policy changes. Conventional machine unlearning modifies model parameters, which may be costly or inaccessible for API-only models. Prompt-based suppression offers a training-free alternative, but does it make a concept inaccessible or merely change the model's answer? We investigate this question in off-the-shelf VLMs. Our visually grounded, multi-probe evaluation first verifies that a model recognizes each concept from the image, then tests its recoverability through multiple-choice, short-answer, and indirect queries. Across objects, scenes, and identities, prompt suppression reduces short-answer recall for some concepts while leaving multiple-choice and indirect performance largely unchanged. Explicitly listing the concepts to suppress often increases short-answer recall, suggesting that the list itself cues the answer. Beyond prompting, decoding constraints, representation editing, and parameter updates can suppress particular responses while the same concept remains detectable through other queries. A model may stop naming a visual concept yet still identify or use it when queried differently.
|
| 312 |
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
2604.08995
|
cs.CV
|
Zile Wang, Zexiang Liu, Jiaxing Li, Kaichen Huang, Baixin Xu |
With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle to simultaneously achieve memory-enabled long-term temporal consistency and high-...With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle to simultaneously achieve memory-enabled long-term temporal consistency and high-resolution real-time generation, limiting their applicability in real-world scenarios. To address this, we present Matrix-Game 3.0, a memory-augmented interactive world model designed for 720p real-time longform video generation. Building upon Matrix-Game 2.0, we introduce systematic improvements across data, model, and inference. First, we develop an upgraded industrial-scale infinite data engine that integrates Unreal Engine-based synthetic data, large-scale automated collection from AAA games, and real-world video augmentation to produce high-quality Video-Pose-Action-Prompt quadruplet data at scale. Second, we propose a training framework for long-horizon consistency: by modeling prediction residuals and re-injecting imperfect generated frames during training, the base model learns self-correction; meanwhile, camera-aware memory retrieval and injection enable the base model to achieve long horizon spatiotemporal consistency. Third, we design a multi-segment autoregressive distillation strategy based on Distribution Matching Distillation (DMD), combined with model quantization and VAE decoder pruning, to achieve efficient real-time inference. Experimental results show that Matrix-Game 3.0 achieves up to 40 FPS real-time generation at 720p resolution with a 5B model, while maintaining stable memory consistency over minute-long sequences. Scaling up to a 2x14B model further improves generation quality, dynamics, and generalization. Our approach provides a practical pathway toward industrial-scale deployable world models.
|
| 313 |
ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models
2604.12251
|
cs.CV
|
Xinliang Wang, Yifeng Shi, Zhenyu Wu |
3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constraints. Current generative restoration approaches are often limited by insufficient temporal coherence, a lac...3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constraints. Current generative restoration approaches are often limited by insufficient temporal coherence, a lack of explicit spatial constraints, and a lack of large-scale training data, resulting in multi-view inconsistencies, erroneous geometric hallucinations, and limited generalization to diverse real-world artifact distributions. In this paper, we present ArtifactWorld, a framework that resolves 3DGS artifact repair through systematic data expansion and a homogeneous dual-model paradigm. To address the data bottleneck, we establish a fine-grained phenomenological taxonomy of 3DGS artifacts and construct a comprehensive training set of 107.5K diverse paired video clips to enhance model robustness. Architecturally, we unify the restoration process within a video diffusion backbone, utilizing an isomorphic predictor to localize structural defects via an artifact heatmap. This heatmap then guides the restoration through an Artifact-Aware Triplet Fusion mechanism, enabling precise, intensity-guided spatio-temporal repair within native self-attention. Extensive experiments demonstrate that ArtifactWorld achieves state-of-the-art performance in sparse novel view synthesis and robust 3D reconstruction. Code available at: https://github.com/fyting/ArtifactWorld.
|
| 314 |
Robust Promptable Video Object Segmentation
2605.12006
|
cs.CV
|
Sohyun Lee, Yeho Gwon, Lukas Hoyer, Konrad Schindler, Christos Sakaridis |
The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in safety-critical domains. This paper offers the first comprehensive study on robust PVOS (RobustPVOS). We fir...The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in safety-critical domains. This paper offers the first comprehensive study on robust PVOS (RobustPVOS). We first construct a new, comprehensive benchmark with two real-world evaluation datasets of 351 video clips and more than 2,500 object masks under real-world adverse conditions. At the same time, we generate synthetic training data by applying diverse and temporally varying corruptions to existing VOS datasets. Moreover, we present a new RobustPVOS method, dubbed Memory-object-conditioned Gated-rank Adaptation (MoGA). The key to successfully performing RobustPVOS is two-fold: effectively handling object-specific degradation and ensuring temporal consistency in predictions. MoGA leverages object-specific representations maintained in memory across frames to condition the robustification process, which allows the model to handle each tracked object differently in a temporally consistent way. Extensive experiments on our benchmark validate MoGA's efficacy, showing consistent and significant improvements across diverse corruption types on both synthetic and real-world datasets, establishing a strong baseline for future RobustPVOS research. Our benchmark is publicly available at https://sohyun-l.github.io/RobustPVOS_project_page/.
|
| 315 |
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
2605.12237
|
cs.CV
|
Shuo Ni, Tong Wang, Jing Zhang, Di Wang, He Chen |
Vision-Language Models (VLMs) are increasingly used to analyze ultra-high-resolution (UHR) Earth observation imagery, yet they face a severe scale mismatch between broad scene context and micro-scale targets. We refer to this phenomenon as a "resolution illusi...Vision-Language Models (VLMs) are increasingly used to analyze ultra-high-resolution (UHR) Earth observation imagery, yet they face a severe scale mismatch between broad scene context and micro-scale targets. We refer to this phenomenon as a "resolution illusion": higher input resolution provides access to more visual detail, but does not necessarily translate into reliable perception of task-relevant micro-evidence. To benchmark this challenge, we introduce UHR-Micro, a benchmark comprising 11,072 instructions grounded in 1,212 UHR images, designed to evaluate VLMs on micro-scale evidence in native-resolution Earth observation imagery. UHR-Micro spans diverse target scales, task families, and visual conditions, with each sample having an objectively verifiable target. Experiments with representative high-resolution VLMs show substantial failures in localizing and interpreting task-relevant evidence, despite access to high-resolution inputs. Further analysis shows that model scaling alone is insufficient, while localized evidence substantially improves performance, pointing to evidence access as a major bottleneck. Motivated by this finding, we propose Micro-evidence Active Perception (MAP), which constructs a spatially grounded evidence state from localized observations and supplements it when the current evidence is insufficient. Across two backbone VLMs, MAP improves UHR-Micro performance by 8.65 percentage points on average. UHR-Micro and MAP provide a framework for diagnosing and improving high-resolution reasoning in Earth observation VLMs. Datasets and source code were released at https://github.com/MiliLab/UHR-Micro.
|
| 316 |
FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery
2605.14854
|
cs.CVcs.AI
|
Patrick Kwon, Chen Chen |
Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root structure are often relatively well constrai...Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root structure are often relatively well constrained, whereas distal articulations such as the arms and legs are more uncertain. Building on this observation, we propose FactorizedHMR, a two-stage framework that treats these two regimes differently. A deterministic regression module first recovers a stable torso-root anchor, and a probabilistic flow-matching module then completes the remaining non-torso articulation. To make this completion reliable, we combine a composite target representation with geometry-aware supervision and feature-aware classifier-free guidance, preserving the torso-root anchor while improving single-reference recovery of ambiguity-prone articulation. We also introduce a synthetic data pipeline that provides the paired image-camera-motion supervision under diverse viewpoints. Across camera-space and world-space benchmarks, FactorizedHMR remains competitive with strong baselines, with the clearest gains in occlusion-heavy recovery and drift-sensitive world-space metrics.
|
| 317 |
Observation-Aligned Mask Priors for Learning Physical Fields from Authentic Occlusions
2605.16818
|
cs.CVcs.AI
|
Chiyuan Ma, Zihan Zhou, Tianshu Yu |
Learning physical fields directly from incomplete observations is challenging because authentic occlusions are structured, sample-dependent, and often missing not at random, whereas existing methods typically rely on heuristic masking rules or predefined mask ...Learning physical fields directly from incomplete observations is challenging because authentic occlusions are structured, sample-dependent, and often missing not at random, whereas existing methods typically rely on heuristic masking rules or predefined mask distributions. We propose Observation-Aligned Mask Priors, a framework that learns the distribution of authentic observation masks and uses it to construct context-query partitions for training from incomplete data. Specifically, we pretrain a Bayesian Flow Network (BFN) on binary observation masks to capture real occlusion topologies, then guide BFN sampling with a globally normalized cross-entropy objective to generate sample-specific masks aligned with each sparse observation. The intersection between the guided mask and the observed mask defines the context, and the remaining observed entries become query targets for a diffusion-based reconstruction model. We show that this intersection-based partitioning gives every valid observed dimension a strictly positive probability of being queried, preventing zero-query dead zones and local generative collapse. Experiments on three real-world oceanographic datasets with authentic satellite occlusions, across resolutions up to 256$\times$256, show consistent improvements over strong diffusion baselines in MSE and PSNR. These results demonstrate that learning mask priors from authentic occlusions is an effective alternative to heuristic masking for learning from incomplete physical observations without access to fully observed fields.
|
| 318 |
Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction
2605.16981
|
cs.CV
|
Kejun Ren, Lei Jin, Tianxin Huang, Lianming Xu, Li Wang |
Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TTT3R-style per-token gates across five benchmarks and discover a structural bottleneck: the gate is intrinsically ...Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TTT3R-style per-token gates across five benchmarks and discover a structural bottleneck: the gate is intrinsically bounded in magnitude (median $0.31$; never exceeding $0.6$) and nearly frame-invariant, yielding an effective memory horizon of only $\sim$3 frames per state token, which serves as the structural origin of long-sequence drift. We trace this to a missing axis: existing inference-time methods modulate updates only at the per-token, intra-frame level, while the orthogonal frame-level question of \emph{how strongly each frame should contribute to the state} has been treated as content-independent. We close this gap with a scalar frame-level gate $\alpha_t \in (0, 1]$ derived in closed form from frame-to-frame changes of internal features---a graded write weight, inspired by classical Simultaneous Localization and Mapping (SLAM) keyframe selection, that never discards a frame and requires no parameters, no training, and no extra forward pass. Across six benchmarks spanning camera pose, video depth, and 3D reconstruction at sequence lengths up to $4,661$ frames, our gate cuts ATE by $51\%$ on long TUM-RGBD pose sequences, reduces AbsRel by $13.0\%$ on Bonn video depth, and on KITTI long-sequence pose estimation surpasses both LongStream and Keyframe-VO in average ATE, while retaining strictly constant memory at zero training cost.
|
| 319 |
Leveraging Latent Visual Reasoning in Silence
2605.18641
|
cs.CV
|
Dongyao Zhu, Zhen Wang, Xi Xiao, Han Jiang, Saeed Vahidian |
Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of these latent tokens at inference remains ambiguous. We show that replacing latent ...Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of these latent tokens at inference remains ambiguous. We show that replacing latent tokens with random noise or removing them completely causes little performance degradation across spatial reasoning benchmarks. Reinforcement learning further diminishes the latent generation behavior after post-training. These observations raise a central question: \textit{Is latent visual reasoning still meaningful}? We argue that its value should be measured by how effectively latent tokens guide learning, rather than whether they persist as an inference-time format. Our analysis shows that latent reasoning is unevenly favorable across question types, yet hard task-level routing for applying latent generation is brittle. Motivated by these findings, we propose an attention-based reward that encourages generated latent tokens to interact with later text tokens during RL. This reward promotes latent utilization when the latent mode is activated while preserving the flexibility to use pure-text reasoning. Experiments show that our method improves performance across perception and visual reasoning benchmarks, even when latent tokens are rarely generated after post-training. Our results highlight that, without explicit expression at inference, latent visual reasoning can shape better visual grounding and more accurate textual reasoning \textbf{in silence}. Code \& model: \href{https://github.com/ddydyd32/silent-lvr/tree/master}{\faGithub}.
|
| 320 |
Vision Harnessing Agent for Open Ad-hoc Segmentation
2605.19410
|
cs.CV
|
Zilin Wang, Stella X. Yu |
Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc concepts, where the grounding may not exist as one learned mask and must often be constructed from image eviden...Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc concepts, where the grounding may not exist as one learned mask and must often be constructed from image evidence through parts, relations, exclusions, and collections. We propose a Vision-guided Ad-hoc Segmentation Agent (VASA), the first vision harnessing agent for open ad-hoc segmentation. VASA is training-free and couples a VLM agent, a segmentation foundation model, and a visual harness that maintains a working mask to make visual progress persistent, inspectable, and editable. Rather than revising text prompts alone, it plans visual operations, invokes segmentation tools, inspects results, edits the mask, and recovers from errors. We construct PARS, a new benchmark that turns part-level labels into open ad-hoc concepts through long-form definition queries. We show that VASA is consistently effective across six VLMs with varying capabilities. Using Qwen3-VL 32B Thinking as the VLM, VASA outperforms various baselines on PARS, surpassing SAM3 Agent by 13.5%-25.3%. On RefCOCOm, VASA improves over SAM3 Agent by 4.8%-8.8% and over other agentic baselines by more. VASA also remains competitive with SAM3 Agent on ReasonSeg for common, named concepts. These results validate VASA's agentic visual construction for open ad-hoc segmentation.
|
| 321 |
Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models
2605.20624
|
cs.CVcs.AI
|
Taesung Kwon, Jonghyun Park, Hyungjin Chung, Jong Chul Ye |
Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high initial latency caused by holistic video restoration, and low throughput resulting from multiple VAE passes to...Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high initial latency caused by holistic video restoration, and low throughput resulting from multiple VAE passes to enforce measurement consistency in pixel space. To overcome these limitations, we propose Autoregressive Video Inverse problem Solver (AVIS). The AVIS framework leverages autoregressive video diffusion models to restore videos in a streaming manner, naturally eliminating latency bottlenecks. Specifically, AVIS initializes reverse diffusion with a measurement-consistent estimate, reducing the required sampling steps. Compared to leading non-autoregressive solvers, AVIS drastically reduces initial latency from 114s to 4s and increases throughput from 0.71 to 1.18 FPS while achieving superior restoration quality. We further introduce a highly accelerated variant, dubbed AVIS Flash, that enforces measurement consistency solely on the first chunk. AVIS Flash substantially boosts throughput to 5.91 FPS on a single RTX 4090 GPU while maintaining competitive performance and achieving a favorable efficiency-performance trade-off, paving the way toward real-time deployment.
|
| 322 |
Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion
2605.21907
|
cs.CV
|
Gang Dai, Yining Huang, Yiming Xia, Guohao Chen, Shuaicheng Niu |
Test-Time Scaling (TTS) paradigm offers a promising perspective for enhancing the generation performance of diffusion models. However, current solutions largely restrict their search to predefined noise candidates or suffer from inflexible exploration across t...Test-Time Scaling (TTS) paradigm offers a promising perspective for enhancing the generation performance of diffusion models. However, current solutions largely restrict their search to predefined noise candidates or suffer from inflexible exploration across the denoising trajectory. To bridge this gap, we propose RTS, a novel Reward-guided Trajectory Scaling method to fully unlock the generative potential of diffusion models. Unlike existing methods, RTS facilitates the synthesis of refined, high-fidelity images via two core innovations: 1) a coarse-to-fine noise optimization mechanism that exploits historical search experience to actively steer the exploration toward high-reward regions and 2) a unified sparse test-time scaling framework featuring PCA-driven curvature analysis, which eliminates temporal redundancy by flexiblely allocating compute to a sparse set of key timesteps that represent critical shifts in the denoising direction. Extensive experiments across SD v3, FLUX, and Qwen-Image architectures demonstrate that RTS outperforms baselines, improving the GenEval score by 20.7%, 15.6%, and 12.2%, respectively. Notably, empirical findings indicate that these key points primarily cluster in the mid-stage of the trajectory, distinct from the structure-sensitive early phases and the late attribute refinement phases.
|
| 323 |
Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
2605.25333
|
cs.CV
|
Tianshuo Xu, Yichen Xie, Depu Meng, Chensheng Peng, Quentin Herau |
Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity problem: pretrained video diffusion transformers already possess KV-cache mechani...Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity problem: pretrained video diffusion transformers already possess KV-cache mechanisms capable of non-local retrieval, but they are rarely trained to use them as dynamic memory. We introduce ReMind, a framework eliciting dynamic memory behavior via memory-oriented data, event-aware training, and cache adaptation. Organized around a taxonomy of 100+ dynamic events, we build a camera-annotated training mixture combining VLM-filtered real videos, generated hard dynamics, synthetic camera loops, and memory-interruption augmentations. Each clip is converted into a frame graph with protected anchors, degraded intervals, and explicit temporal gaps. A node-structured curriculum -- including node-drop, noisy memory, frontier continuation, and reference-cache training -- forces the model to retrieve relevant past states across interruptions rather than relying solely on local continuity. PM-RoPE, an elegant camera-phase RoPE extension, unlocks spatiotemporal retrieval at a single-attention cost while preserving pretrained pathways. ReMind achieves the best overall scores on STEVO-Bench and recovery tasks. Furthermore, general image-to-video evaluations confirm this curriculum avoids catastrophic forgetting. We have released our code, data, and models on our project page \href{https://remind-applied.github.io/}{https://remind-applied.github.io/}.
|
| 324 |
Generalizing Geometry-Guided Mamba as a Plug-and-Play Context Module for CNN-based Semantic Segmentation
2606.08866
|
cs.CV
|
Hsin-Jui Pan, Sheng-Wei Chan, Chun-Po Shen, Jen-Shiun Chiang |
CNN-based semantic segmentation networks usually rely on context heads such as ASPP, PPM, or attention modules to enlarge the receptive field. These heads are effective but may introduce heavy computation, memory cost, or boundary leakage. This paper revisits ...CNN-based semantic segmentation networks usually rely on context heads such as ASPP, PPM, or attention modules to enlarge the receptive field. These heads are effective but may introduce heavy computation, memory cost, or boundary leakage. This paper revisits Directional Geometric Mamba (G-Mamba) from DGM-Net and studies it as a plug-and-play context aggregation module rather than a completely new segmentation architecture. The key idea is to inject geometric guidance into the selective scan process, allowing long-range feature propagation to be modulated by boundary and centripetal-flow cues. We replace the original context heads of six representative CNN segmentation models, including DeepLabV3+, DANet, CCNet, PSPNet, PSANet, and OCRNet, while keeping the ResNet-101 backbone unchanged. On CCNet, we additionally compare serial and parallel combinations of criss-cross attention and the G-Mamba block, with the parallel head performing best. Results on Cityscapes show consistent mIoU gains with only moderate extra GFLOPs at $1024\times1024$ resolution, suggesting that geometry-guided SSM modules can serve as practical alternatives or enhancements to conventional CNN context heads.
|
| 325 |
Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning
2606.10196
|
cs.CVcs.AI
|
Ghodsiyeh Rostami, Po-Han Chen, Mahdi S. Hosseini |
Parameter-efficient fine-tuning often selects trainable parameters before adaptation using architectural heuristics, without accounting for their varying importance during training. We introduce \textbf{FisherAdapTune}, which progressively selects parameter gr...Parameter-efficient fine-tuning often selects trainable parameters before adaptation using architectural heuristics, without accounting for their varying importance during training. We introduce \textbf{FisherAdapTune}, which progressively selects parameter groups based on temporal drift in their Fisher information. Under a local Gaussian approximation, we bound the divergence between the fine-tuned posterior and pretrained prior by accumulated Fisher-weighted update costs, motivating curvature-aware selection. FisherAdapTune measures Jensen-Shannon distance between successive Fisher-value distributions and uses an adaptive threshold to freeze stabilized groups. Across VTAB-1k classification tasks, it achieves the highest macro Top-1 accuracy among the compared methods with a smaller average trainable set than full fine-tuning. Across four segmentation backbones, it maintains competitive in-distribution performance and improves zero-shot transfer in several settings. Its selections reveal architecture-dependent patterns where input, output, and normalization parameters can remain trainable as attention and MLP groups freeze at different rates. The results support Fisher structural drift as a task-dependent signal for allocating updates during adaptation. We release our \href{https://github.com/AtlasAnalyticsLab/FisherAdapTune}{code} publicly to enable further application of our proposed approach.
|
| 326 |
JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space
2606.13345
|
cs.CV
|
Xinnan Zhu, Ruijie Xu, Jiayu Ying, Daoguo Dong, Jiachen Xu |
Existing 3D scene editing methods typically rely on per-scene optimization over explicit 3D representations or cascaded edit-and-reconstruct pipelines, resulting in high test-time cost, limited 3D awareness, and structural inconsistencies. To couple appearance...Existing 3D scene editing methods typically rely on per-scene optimization over explicit 3D representations or cascaded edit-and-reconstruct pipelines, resulting in high test-time cost, limited 3D awareness, and structural inconsistencies. To couple appearance synthesis with geometry prediction, we adapt a pretrained unified RGB-geometry latent space to feed-forward scene editing. Given a source video and an edited reference image, JointEdit3D performs asymmetric latent inpainting: it observes only the edited RGB reference latent and jointly generates the remaining RGB latents and the entire geometry latent along the source trajectory. JointEdit3D introduces a dedicated SceneAnchor Branch to inject source-scene structure without forcing direct copying, and adopts edit/background-aware losses to balance edited-region fidelity with unedited-content preservation. To address the lack of paired resources for standardized 3D scene editing evaluation, we introduce SceneEdit3D-15K, a dataset with 15K paired editing samples and renderer-provided 3D annotations, together with SceneEdit3D-Bench, a curated 100-sample benchmark. Experiments show that JointEdit3D improves edited-region quality and 3D structural completeness over prior baselines while maintaining competitive background preservation.
|
| 327 |
ActiveSAM: Fast and Accurate Open-Vocabulary Semantic Segmentation with Frozen SAM 3
2606.16996
|
cs.CVcs.AI
|
Tran Dinh Tien, Zhiqiang Shen |
Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset voc...Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes. We introduce ActiveSAM, a training-free inference framework that turns SAM 3 into an active-vocabulary segmenter. ActiveSAM first canonicalizes and expands class prompts, then uses evidence-proportional grounding to estimate an image-conditioned active set from a low-resolution presence preview. Only retained prompts receive full-resolution mask prediction, using bucketed prompt multiplexing with the frozen SAM 3 decoder. The preview stage uses only class-presence evidence and skips unnecessary segmentation-head computation. To resolve overlapping concept responses, exclusive concept decoding compares each pixel's joint score vector with class signatures estimated once per vocabulary from unlabeled images. ActiveSAM requires no weight updates, no oracle class-presence labels and no per-dataset hyperparameter tuning. Across eight OVSS benchmarks, ActiveSAM improves the speed-accuracy tradeoff of training-free open-vocabulary semantic segmentation, outperforming the current state-of-the-art SegEarth-OV3 by +2.1 mIoU on average while running much faster, with 7.3-12.2x speedups on large-vocabulary datasets. ActiveSAM also achieves the highest accuracy under image corruptions that simulate real-world distribution shift, making it well-suited for deployment in noisy-input domains such as autonomous driving and embodied AI. Code is available at https://github.com/VILA-Lab/ActiveSAM
|
| 328 |
Do Gaussian Scenes Contain Enough Structure for Intrinsic Segmentation?
2606.18623
|
cs.CV
|
Mohamed Rayan Barhdadi, Hasan Yazar, Erchin Serpedin, Mehmet Tuncel, Hasan Kurban |
Gaussian segmentation is usually posed as transferring object knowledge from 2D foundation models into a 3D representation. This leaves a fundamental question unanswered: how much object structure is already encoded by a trained gaussian scene? We investigate ...Gaussian segmentation is usually posed as transferring object knowledge from 2D foundation models into a 3D representation. This leaves a fundamental question unanswered: how much object structure is already encoded by a trained gaussian scene? We investigate this question with GS-IntSeg, an intrinsic, mask-free, and training-free method that constructs partitions using only: gaussian geometry, opacity, spherical-harmonic radiance, and deformation trajectories. On the dynamic Neu3D and HyperNeRF datasets, GS-IntSeg obtains a mean of 0.677 mIoU across multi-view and monocular scenes without masks, external features, or segmentation training. This intrinsic formulation also enables GS-IntSeg to require approximately 2.5 minutes per HyperNeRF scene on a consumer RTX 5080 to construct its partitions, over 10x faster than SAM-based methods that require mask generation, feature rendering, and other stages. These results suggest that gaussians alone can approach mask-supervised performance in gaussian scenes segmentation. While a gap remains between intrinsic and foundation model based methods, to our knowledge, GS-IntSeg is the first mask-free approach to gaussian segmentation, motivating a re-evaluation of the assumption that gaussian scene segmentation must be based on external masks and pointing toward an alternative faster, more generalizable segmentation approach.
|
| 329 |
EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
2606.20092
|
cs.CV
|
Ganlin Yang, Zhangzheng Tu, Yuqiang Yang, Sitong Mao, Junyi Dong |
Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historic...Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historical context, they either suffer from severe information bottlenecks, incur high latency via decoupled dual systems, or rely on unselective buffers that accumulate massive visual redundancies. To address these limitations, we introduce EventVLA, an end-to-end framework founded on the concept of sparse visual evidence memory that comprises two core components: foundational visual anchors to retain initial and short-term contexts, and a dynamic Keyframe Evidence Memory (KEM) module. Specifically, KEM directly predicts future keyframe probabilities from the VLA's latent embeddings to autonomously capture and store sparse, task-critical visual events. This foresight-driven mechanism empowers the policy to dynamically evaluate the future causal utility of current observations, preserving transient visual evidence before it becomes unobservable. Furthermore, we propose RoboTwin-MeM, a diagnostic benchmark specifically designed to evaluate non-Markovian manipulation tasks with interactive visual evidence. Extensive evaluations show that across 17 memory-requiring simulation tasks and 4 real-world bimanual tasks, EventVLA achieves an average success rate improvement of +40% over state-of-the-art memory-augmented VLAs.
|
| 330 |
DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests
2606.20223
|
cs.CV
|
Hugo Magaldi, Theau d'Audiffret, Etienne Francois Akomo-Okoue, Bala Amarasekaran, Naomi Anderson |
Camera-trap monitoring in African tropical forests increasingly extends beyond closed-canopy interiors to riverbanks, clearings, and park edges. Among available open tools for African forest camera-trap classification, DeepForestVision is the only one providin...Camera-trap monitoring in African tropical forests increasingly extends beyond closed-canopy interiors to riverbanks, clearings, and park edges. Among available open tools for African forest camera-trap classification, DeepForestVision is the only one providing a matched offline workflow for both photographs and videos, and previous work showed that it outperformed other available baselines on a comparable benchmark. However, it was designed for closed-canopy, ground-level forest interiors and uses a 35-class prediction space that becomes too coarse when deployments encounter arboreal primates, birds, semi-aquatic taxa, or human-associated confounders such as livestock. We present DeepForestVisionV2, an ecology-driven expansion from 35 to 64 prediction classes (61 animal classes plus human, vehicle, and blank) designed to address three recurrent deployment gradients: vertical stratification, scene openness, and anthropogenic interfaces. DeepForestVisionV2 retains the same offline workflow and is trained on 1,535,010 photographs and 243,354 videos from multi-country African tropical-forest projects. Evaluation combines a cross-country cropped-photo validation set, used to assess robustness across sites and camera-trap settings, with three held-out Uganda video benchmarks spanning the targeted gradients. On the validation set, DeepForestVisionV2 reaches 0.86 accuracy, 0.82 macro-F1, and 0.81 balanced accuracy. On the deployment benchmarks, it preserves or improves baseline accuracy despite its harder classification task, while increasing the number of identified taxa from 22 to 29 in forest-interior videos and from 4 to 9 at riverbanks. In the park-edge use case, it raises accuracy from 0.62 to 0.86 and reduces false alarms from 11 to 0. These results show that DeepForestVisionV2 materially improves field utility while preserving robustness across sites, habitats, and camera-trap settings.
|
| 331 |
Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection
2606.20752
|
cs.CV
|
Ziba Parsons, Ang Li |
Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent studies have revealed its vulnerability to backdoor attacks. Existing attacks typically require white-box acces...Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent studies have revealed its vulnerability to backdoor attacks. Existing attacks typically require white-box access or label modification and focus on geometric attacks such as object disappearance or bounding-box manipulation. In this paper, we present Mirage, a black-box and clean-label backdoor attack against deep neural network-based LiDAR 3DOD. Mirage injects a small number of label-consistent poisoning samples into the training set, causing the model to learn a malicious association between a trigger pattern and an attacker-chosen target class while preserving normal training semantics. As a result, the compromised model behaves normally on benign inputs yet systematically misclassifies triggered objects as the target class during deployment. We evaluate Mirage on multiple state-of-the-art LiDAR 3DOD models and benchmark datasets. Experimental results show that Mirage achieves a 73% misclassification success rate with a poisoning rate of only 0.5%, while maintaining detection performance close to that of benign models.
|
| 332 |
UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation
2607.12896
|
cs.CV
|
Yunzhou Li, Jiesi Hu, Yanwu Yang, Hanyang Peng, Chenfei Ye |
Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fragmented by prompt paradigms and spatial dimensions. Visual in-context learning, interactive segmentation, and la...Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fragmented by prompt paradigms and spatial dimensions. Visual in-context learning, interactive segmentation, and language-guided segmentation are typically handled by paradigm-specific models, while 2D and 3D images are also modeled separately. Such isolation prevents heterogeneous annotations and data from being jointly absorbed by a single scalable model and limits cross-paradigm knowledge transfer. To address this bottleneck, we propose UniMedSeg, a Transformer-centric universal segmentation framework that maps visual examples, geometric interactions, language instructions, and 2D/3D images into a shared sequence space, enabling heterogeneous medical supervision to be jointly learned through a unified in-context interface without prompt- or dimension-specific branches. To overcome the long-sequence memory bottleneck caused by visual contexts, we introduce Decoupled Split Attention, which reduces attention complexity to linear while preserving hardware-friendly computation and focused context-target interaction. Extensively trained and evaluated on a large corpus curated from 27 public datasets, UniMedSeg achieves state-of-the-art performance across visual in-context, interactive, and language-guided segmentation without task-specific fine-tuning, demonstrating strong generalization on diverse held-out tasks. The code and model weights are publicly available at https://github.com/Lii1228/UniMedSeg
|
| 333 |
Importance-Aware OBS Pruning for Diffusion Models
2607.20048
|
cs.CV
|
Ba-Thinh Lam, Srijan Das, Hieu Le |
We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps -- derived from conditioning signals or ...We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps -- derived from conditioning signals or model attention -- into the pruning objective. This produces parameter rankings aligned with perceptual relevance rather than uniform reconstruction error. On MS-COCO dataset, our proposed approach consistently retains subject fidelity and structural correctness at high compression ratios where conventional pruning causes visible degradation. These results demonstrate that content-aware objectives are key to perceptually faithful compression of generative models.
|
| 334 |
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification
2607.28225
|
cs.CV
|
Haoqing Wang, Xingrun Xing, Ziheng Li, Jianyuan Guo, Yehui Tang |
Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multi-modal reasoning. However, recent s...Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multi-modal reasoning. However, recent studies have revealed that such models often use tools unfaithfully. Many process images are irrelevant to the question (e.g., the crops miss the queried target), yet the tool call still receives full credit and the model still answers correctly. Such decorative or misaligned tool calls waste computation and reveal that the model does not faithfully use the evidence it retrieves. This may stem from two limitations of prevailing methods: the tool reward fails to distinguish useful from useless calls, and tool feedback carries no signal of usefulness. To this end, we introduce FaithEyes, a multi-agent self-judging framework. Concretely, we use a VLM to judge whether each process image helps answer the question. The judgement is injected into the reasoning context as part of the tool observation to help subsequent reasoning, and meanwhile is used to scale the tool reward by the helpful-tool ratio to suppress reward hacking. To keep judgement available at evaluation, we further design a multi-agent framework where the model itself serves as a subagent to judge the tool calls from the main agent, eliminating any dependence on external models at inference. Training via a two-stage SFT + RL pipeline on adapted open-source data, FaithEyes attains competitive or superior accuracy across visual perception and reasoning benchmarks, while substantially improving tool faithfulness and reducing inference cost. The homepage is at https://github.com/Mosi-AI/FaithEyes.
|
| 335 |
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
2607.28595
|
cs.CV
|
Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He |
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks. We rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness and Tool Effect. Mode Ad...The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks. We rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness and Tool Effect. Mode Adaptiveness characterizes whether an MLLM recognizes when tools are necessary and invokes them accordingly, avoiding unnecessary computational overhead while improving performance on problems requiring tool assistance. Tool Effect characterizes whether tools extend the model's capabilities on problems unsolvable through tool-free reasoning without introducing errors on problems it can already solve. Our analysis quantifies these properties and reveals that existing models exhibit limited Mode Adaptiveness, while tool-use gains on hard examples are largely offset by harm on easy ones. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model trained with supervised fine-tuning (SFT) and reinforcement learning (RL). Its RL stage combines Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion. Necessity-Aware Adaptive Reward encourages tool-free solutions when they succeed while preserving full reward for successful tool use when tool-free rollouts fail. Hint-Guided Capability Expansion uses verified, answer-free expert hints to recover learning signals from all-wrong rollout groups, aiming to extend tool-use capability on the hardest problems. Across 13 benchmarks, Beacon achieves the highest average score among the evaluated open-source models and ranks first on 11 benchmarks. On five diagnostic benchmarks, it improves the average tool-available accuracy over its tool-free accuracy by 1.96 points and achieves the largest tool-gain minus tool-harm score (+3.14 points). These results show Beacon's advanced performance, Mode Adaptiveness, and the net benefit of tool use.
|
| 336 |
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
2608.07468
|
cs.CV
|
Zongchuang Zhao, Xin Zhou, Tianyang Xu, Zhengyang Sun, Kaixuan Zhou |
In autonomous driving, World-Action Models (WAMs) have improved end-to-end planning by transferring video dynamics priors to action prediction, but many still couple planning with future-video generation at inference, incurring substantial computational overhe...In autonomous driving, World-Action Models (WAMs) have improved end-to-end planning by transferring video dynamics priors to action prediction, but many still couple planning with future-video generation at inference, incurring substantial computational overhead. We present SimWAM, a simple yet effective WAM that leverages future-video prediction solely as a training-time supervision signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without future-frame generation at inference. This design supports multiple pretrained video backbones and independent action-expert scaling within a shared attention interface, while preserving the joint learning objective. Moreover, we apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Experiments show that SimWAM achieves $91.9$ PDMS on NAVSIM with a favorable trade-off between accuracy and latency among world-model-based planners, while transferring zero-shot to nuScenes. It also achieves competitive planning accuracy on WOD-E2E and PhysicalAI-Autonomous-Vehicles. These results position SimWAM as a plain yet solid baseline for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/.
|
| 337 |
Gated Spatial Redundancy Projection for Pathology Transformer Attentions
2608.08374
|
cs.CV
|
Zhiyuan Yang, Jiahao Cheng, Vincent Quoc-Huy Trinh, Mahdi S. Hosseini |
Transformer models are increasingly used for whole-slide image analysis in computational pathology. Yet, WSIs differ fundamentally from natural images: neighbouring patches often contain highly similar tissue type, stain, texture, and cellular composition. We ...Transformer models are increasingly used for whole-slide image analysis in computational pathology. Yet, WSIs differ fundamentally from natural images: neighbouring patches often contain highly similar tissue type, stain, texture, and cellular composition. We identify this local spatial redundancy as a pathology-specific failure mode of self-attention, where dominant neighbourhood features can be repeatedly mixed into patch-tokens and weaken subtle diagnostic or prognostic deviations. We propose Gated Spatial Redundancy Projection (Gated SRP), a lightweight drop-in correction module for self-attention layers. For each patch token and attention head, Gated SRP estimates a local redundancy axis from neighbouring value vectors, projects the attention output onto this axis, and applies a learned signed gate to correct the redundancy-aligned component geometrically. Across five TCGA survival cohorts, Gated SRP obtains the highest mean C-index among the compared attention variants in all cohorts, with an average improvement over the base attention, while adding only +0.02% parameters. Across five slide-level classification datasets, it improves the base attention on 12 of 16 reported metrics and achieves the best AUC on three datasets. Code is publicly available at https://github.com/AtlasAnalyticsLab/GatedSRP.
|
| 338 |
Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning
2608.19669
|
cs.CV
|
Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale |
Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further ref...Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. In this paper, we identify two key limitations of this framework, one in each stage. First, the SFT stage typically relies on an off-the-shelf vision encoder to encode the helper image, yielding suboptimal latent representations that may not be well aligned with the downstream reasoning task. Second, existing RL methods treat the latent component only through deterministic regularization, which constrains policy drift but does not create alternative latent trajectories for exploration. To address these limitations, we propose Scaffolding Minds. Our approach learns a dedicated scaffolding encoder that provides an optimized target in latent space, and learns both the mean and variance of the RL sampler. We further show that these two improvements are complementary, together yielding substantial gains over strong baselines. Empirically, our method improves over the strongest latent reasoning baseline by +9.5 points on FrozenLake spatial planning, with the gain widening to +19 points on the 32x32 grids, and by +5.6 points on average across nine visual-centric reasoning benchmarks.
|
| 339 |
Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
2609.04802
|
cs.CVcs.AI
|
Tianyidan Xie, Shenyi Wang, Qiang Tang, Mingjie Wang, Zhicheng Qiu |
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (cl...Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (clip-level video-language embeddings), keep it only as raw coordinates (geometric SLAM), or organise it around immediate task context (agent working memories). None of them gives the agent a per-object timeline whose state transitions are themselves queryable in language. Our key contribution is \textbf{Linguistic Trajectory Encoding} (LTE), which compresses dynamic object motion histories via a hybrid representation combining natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location, while representing motion with geometric waypoints and linguistic descriptions to preserve accuracy. To evaluate these capabilities across extended time horizons, we construct the \textbf{Spatial Memory Benchmark} (SMB) from EgoLife multi-day recordings, targeting capabilities absent in existing benchmarks: semantic trajectory retrieval and long-horizon object retrieval. On SMB, the LTE-based system achieves $45.3\%$ success in semantic trajectory retrieval and $48.7\%$ in long-horizon object retrieval, outperforming structured-memory and VLM baselines (best prior: $31.9\%$ and $34.4\%$). LTE achieves trajectory compression by factors of $8.7\times$ to $26.1\times$ with sub-second query latency on $24$\,h video. On Ego4D natural-language queries, the system reaches $28.75\%$ / $55.10\%$ R@1/R@5, $+15.80$ / $+31.30$ pts over EgoVLPv2.
|
| 340 |
SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection
2609.04886
|
cs.CVcs.AI
|
Yongchun Lin, Xinliang Zhang, Yun Zou, Zhixuan Xiao, Liang Lei |
Changes in sensor height and viewpoint alter object-level point distributions, making cross-platform LiDAR unsupervised domain adaptation (UDA) difficult. Self-training uses labeled source scans and unlabeled target scans, yet a retained prediction may provide...Changes in sensor height and viewpoint alter object-level point distributions, making cross-platform LiDAR unsupervised domain adaptation (UDA) difficult. Self-training uses labeled source scans and unlabeled target scans, yet a retained prediction may provide a useful target location while enclosing sparse foreground returns, background clutter, or points inconsistent with the predicted box. We refer to this mismatch as box-point inconsistency. We introduce SimFuse3D, which preserves the target placement and repairs the associated pseudo object using measured geometry from labeled source scans. Object Memory retrieves a similar labeled source instance. Target Simulation places the retrieved source geometry at the target location, aligns its points with the target viewing geometry, and filters the aligned crop to approximate the target observation. Confidence-Guided Multi-Stage Localization Reweighting (CMLR) maps each target pseudo-object confidence score to a bounded weight shared by RPN localization and R-CNN box regression. All components operate only during adaptation, leaving the detector architecture and inference graph unchanged. Across six cross-platform transfers, SimFuse3D consistently outperforms Pi3DET-Net and achieves the best performance among the compared adaptation methods on nearly all metrics. On nuScenes-to-KITTI, it ranks first among the compared adaptation methods with both evaluated detectors.
|
| 341 |
CST-WM: A Causally Structured World Model for Embodied Visual Tracking
2609.06302
|
cs.CV
|
Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang |
Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning over future target evidence wit...Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning over future target evidence with an action-conditioned world model. In logged tracking data, however, the behavior policy's actions are correlated with where the target is, so a generic predictor can learn a shortcut: it writes the current action directly into its prediction of target evidence, instead of letting the action affect that evidence only by moving the robot and changing what it observes. We call this failure causal hallucination; the resulting rollouts look plausible but rank candidate actions for the wrong reason. We propose CST-WM, a causally structured world model whose state is split into target-evidence, robot, and observation branches. Its transition removes the same-step edge from action to target evidence but keeps the path through robot motion and the resulting views, so candidate actions are still distinguished by their predicted ego-motion. With rollout-based model-predictive control, a single model handles both steady following and re-acquisition after target loss. On EVT-Bench and Habitat 3.0, covering standard tracking, target-loss recovery, and cross-dataset transfer, CST-WM improves following, distance-range control, safety, and re-acquisition over reactive trackers and world-model baselines, and removing the action mask causes the largest drop in re-acquisition among our ablations. Offline, CST-WM has lower multi-step rollout error, and its ranking of candidate actions agrees better with the simulator's. On a Unitree Go2 quadruped, CST-WM succeeds in 20 of 30 real-world trials under occlusion, distractor crossing, and fast motion, against 14 for TrackVLA.
|
| 342 |
AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow
2609.10723
|
cs.CVcs.AI
|
Junran Wang, Zehao Jin, Tianyu Luan, Ruixuan Deng, Tian Qiu |
Text-to-image diffusion transformers (DiTs) are powerful generators, yet direct prompting provides limited control interface for style intensity and can fail to suppress unwanted concepts. To enable these controls, we introduce AcFlow, an inference-time contro...Text-to-image diffusion transformers (DiTs) are powerful generators, yet direct prompting provides limited control interface for style intensity and can fail to suppress unwanted concepts. To enable these controls, we introduce AcFlow, an inference-time controller that transports intermediate layer image-token activations through a learned concept-conditioned velocity field while keeping the base DiT frozen. A textual concept description specifies the desired intervention, while the integration horizon provides a continuous control parameter. The field produces token-varying, activation-dependent updates. With parameters shared across concepts within each task family, one field covers over 15,000 style descriptions or over 1,000 suppression concepts, and generalizes to concepts unseen during training without per-concept fitting. On style control, AcFlow achieves the best style--content trade-off among the evaluated baselines in the high-style-alignment regime. At a fixed operating point, AcFlow attains style--content alignment of 0.5365/0.2860, compared with 0.4397/0.2684 for the baseline with the highest style alignment. On concept suppression, AcFlow reduces the fraction of images showing the concept from 95.3%/82.1% to 41.6%/40.5% on held-in/held-out concepts, including cases where deleting them from the prompt fails to remove them. Our analyses support the learned velocity field as an adaptive control mechanism, with update directions varying across tokens and depending on their activation states. Our code is available at https://github.com/Nove1yst/AcFlow.
|
| 343 |
BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable Registration
2609.11472
|
cs.CV
|
Qianliang Wu, Haobo Jiang, Guangwei Gao, Shuo Chen, Weiping Ding |
Reliable matching between partially observed, deforming point clouds requires global context and fine geometric detail. Coarse candidate selection can exclude correct fine-level correspondences. We present \paper, a unified conditional transport framework with...Reliable matching between partially observed, deforming point clouds requires global context and fine geometric detail. Coarse candidate selection can exclude correct fine-level correspondences. We present \paper, a unified conditional transport framework with the matching matrix itself as the evolving state. Coarse diffusion establishes global matching hypotheses; hierarchy-preserving lifting expands them into a structured high-resolution source. Geometry-conditioned ODE and SDE bridges continue refinement in the complete fine-level candidate space, allowing coarse errors to be corrected. The deterministic endpoint-parameterized conditional flow matching (CFM) design improves matching through iteration, while the SDE forward drift enables effective few-step refinement. Experiments on 4DMatch and 4DLoMatch demonstrate competitive matching and non-rigid registration, with transfer to CAPE and DeepDeform without target-domain training.
|
| 344 |
PhysPlan: Grounded Physical State Reasoning and Graph-Guided Optimization for Physically Plausible Video Generation
2609.13006
|
cs.CV
|
Minh-Loi Nguyen, Xuan-Vu Le, Trung-Nghia Le, Tam V. Nguyen, Minh-Triet Tran |
Video diffusion models (VDMs) synthesize photorealistic content, yet they often fail to follow the course that a physical phenomenon should take within a given scene. Recent training-free methods let a vision-language model (VLM) plan the phenomenon and guide ...Video diffusion models (VDMs) synthesize photorealistic content, yet they often fail to follow the course that a physical phenomenon should take within a given scene. Recent training-free methods let a vision-language model (VLM) plan the phenomenon and guide a frozen VDM toward the plan; however, such plans are derived from the prompt and consumed as whole keyframes or trajectories, which leaves unspecified where the consequences land in the observed scene and turns incidental visual details into optimization targets. We observe that a phenomenon specified in words unfolds as sparse, local changes to the physical state of the observed scene. Building on this observation, we present PhysPlan, a training-free image-to-video framework that represents a phenomenon as a grounded state graph and uses this graph to decide what, where, and when the guidance constrains. Grounded Physical State Reasoning decomposes the phenomenon into physical deltas, each stating which objects change, to what state, and by which physical rule, and translates each delta into graph edits, verified by deterministic checks, that leave all other objects unchanged. Graph-Guided Test-Time Optimization renders a keyframe for each state, measures the denoised estimates only along the properties selected by the edits, and concentrates the update on the edited objects. On PhyGenBench and Physics-IQ, PhysPlan raises its base model from 0.52 to 0.77 and from 27.1 to 38.2, surpassing the strongest prior I2V method (0.60 and 34.6), and lowers FVD by over 20%. Project page: https://physplan.github.io
|
| 345 |
Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs
2609.18210
|
cs.CV
|
Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang |
UAV-borne imaging has advanced from megapixel to gigapixel sensors, shifting aerial perception from recognizing individual targets to understanding entire dynamic scenes. We characterize this demand as Wide-area Spatio-temporal Scene Understanding (WSTU), whic...UAV-borne imaging has advanced from megapixel to gigapixel sensors, shifting aerial perception from recognizing individual targets to understanding entire dynamic scenes. We characterize this demand as Wide-area Spatio-temporal Scene Understanding (WSTU), which requires wide-area coverage, per-target resolution, and temporal continuity at once, a combination existing datasets lack. To fill this gap, we introduce an ultra-High-resolution (12768x9564) Airborne Remote-sensing Dataset (HARD) annotated at three levels for object detection, multi-object tracking, and scene-level visual question answering. Ultra-high-resolution imagery raises per-frame processing time to seconds. At that scale latency can no longer be ignored in evaluation. Thus, we propose a latency-aware metric for multi-object tracking called streaming-HOTA (s-HOTA). Extensive baseline experiments show how ultra-high-resolution processing reshapes each task. For detection, the end-to-end pipeline affects accuracy and speed as much as the detector itself does. For tracking, high latency charges the association axis far more unevenly than the detection axis, and association is where pipelines diverge. As a result, the pipeline that performs best offline can lose its lead under s-HOTA. For VQA, vision-language models remain weak at cross-frame identity binding and cannot transfer their single-frame gains to it. Together these findings show that the baselines we evaluate fall short of WSTU. HARD provides the data and the systematic baselines to advance it.
|
| 346 |
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
2609.18860
|
cs.CVcs.AI
|
Girish A. Koushik, Diptesh Kanojia, Helen Treharne |
When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-con...When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages $0.740$ versus $0.432$ for native macro-F1, residual reconstruction reaches $0.486$, and Gemma improves from $0.532$ to $0.714$. These gains measure how accessible the label is to a supervised readout; they do not show that the model's native generation already applies such a decision rule. Under the evaluated scales, Qwen silent-feature ablation is $24-63$ times more probe-sensitive, whereas routed-feature patching on literal yes/no tasks is $16-140$ times more output-sensitive. Native-only threshold calibration explains much, but not all of the gap: on five tasks with matched probe scores, it recovers $69.8$\% of the raw native-to-probe difference, while direct routing adds $0.094$ mean macro-F1 beyond calibrated native scoring. Joint gold-label, probe-KL, and pairwise LoRA supervision improves dedicated FHM prediction, but a gold-only adapter performs better on the shared seven-task mean. A case study of Gemma-3-12B on the Facebook Hateful Memes dataset finds a distributed rank-32 image-prompt interaction, reaching $0.756$ versus $0.685$ native macro-F1. Robustness controls show that the signal is not explained solely by accompanying OCR and depends on paired visual evidence, and that it extends beyond English. In many of the errors we study, the evidence is represented but does not reach the answer; therefore, routing is a common bottleneck in harmful meme classification.
|
| 347 |
FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants
2609.20769
|
cs.CV
|
Tianao Li, Xinhui Qian, Emma Alexander |
Flow matching has emerged as the state-of-the-art generative model and has been used for plug-and-play (PnP) priors to solve inverse problems in computational imaging. However, existing flow-based inverse solvers assume linear forward models and/or make simpli...Flow matching has emerged as the state-of-the-art generative model and has been used for plug-and-play (PnP) priors to solve inverse problems in computational imaging. However, existing flow-based inverse solvers assume linear forward models and/or make simplifying approximations in posterior sampling. To circumvent these problems, we introduce FlowSGS, a flow-based posterior sampling method using Split Gibbs Sampling (SGS) to decompose the posterior into a likelihood step and a prior step. Specifically, we sample from the likelihood step using Langevin dynamics and leverage the Stochastic Interpolants (SI) framework to integrate a pretrained flow model into the prior step. We provide a form for the prior step that uses SI's reverse-time SDE, and show connections to previous PnP methods. Moreover, with the aid of the flow prior's straight probability paths and a novel timestep correction technique for the reverse-time SDE, FlowSGS requires fewer network evaluations in its prior step than plug-and-play diffusion samplers. Our experiments show state-of-the-art performance on a range of inverse problems. For the first time, we provide an experiment on a nonlinear inverse problem (Fourier phase retrieval) for flow-based inverse solvers.
|
| 348 |
4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors
2609.21176
|
cs.CV
|
Haitao Huang, Shenghao Zhao, Boyuan Tian, Shin-Fang Chng, Songlin Yang |
This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cann...This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.
|
| 349 |
Rethinking Vision Architectures with Gated Linear Attention and KAN
2609.22506
|
cs.CV
|
Ali Mehizel, Oussama Khaldi |
Vision Transformers devote most of their parameters to MLPs for channel mixing, but still rely on quadratic multi-head self-attention for token interactions. While linear attention fixes the complexity problem, bringing it down to O(N), it is usually just pair...Vision Transformers devote most of their parameters to MLPs for channel mixing, but still rely on quadratic multi-head self-attention for token interactions. While linear attention fixes the complexity problem, bringing it down to O(N), it is usually just paired with the same fixed-activation MLP as before. Kolmogorov-Arnold Networks take a different approach, placing learnable univariate functions on the edges instead. However, existing vision KANs either retain standard attention or remove attention entirely, so the two ideas have not been effectively combined. We introduce LKAT (Linear Kolmogorov-Arnold Transformer) to close this gap: an isotropic ViT-style encoder that couples chunk-wise Gated Linear Attention with a two-layer KAN feed-forward block, backed by an I/O-aware fused RBF-KAN kernel to make radial-basis grid functions efficient in practice. Under a shared DeiT-style training recipe, LKAT-B outperforms ViT-B/16, ViT-5-B, and Mixer-B/16 on ImageNet-100, while Tiny, Small, and Base variants scale consistently on CIFAR-10/100. ImageNet-100 pretraining also transfers effectively to CIFAR fine-tuning, suggesting that gated linear attention and KAN-based radial basis functions provide complementary inductive biases for mid-scale visual representation learning. Code: https://github.com/mehizelali/linear-kan-transformer
|
| 350 |
HOIBlender: Blending Lightweight Detection with Vision-Language Priors for Efficient Human-Object Interaction Detection
2609.23431
|
cs.CV
|
Junwen Chen, Keiji Yanai |
Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under severe long-tail supervision. Recent methods improve accuracy with stronger detectors and vision-language priors,...Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under severe long-tail supervision. Recent methods improve accuracy with stronger detectors and vision-language priors, but many still stack heavy transformer encoders, intricate denoising schedules, or post-hoc semantic calibration on top of the detector. We present \textbf{HOIBlender}, an efficient HOI detector named after its core design principle: blending detector-grounded visual tokens, spatial subject-object reasoning, and BLIP-2 semantic priors inside one lightweight decoding pipeline. HOIBlender builds on an RF-DETR/LW-DETR-style foundation with a DINOv2 backbone and selects top-$K$ image-conditioned tokens directly from the multi-scale projector as subject and object candidates, removing the dedicated encoder stage retained by prior HOI methods. A dual-stage decoder first stabilizes human-object geometry and then performs verb and HOI classification through progressive BLIP-2 prior fusion, with classifier weights initialized from BLIP-2 text embeddings for long-tail categories. Grouped-query training further enriches optimization without increasing inference cost. Across three model scales (Nano, Small, 2XL), HOIBlender consistently outperforms SOV-STG-VLA and Hybrid-SOV-VLA on HICO-DET, reaching $44.49$ Default Full mAP in only $9$ training epochs while maintaining competitive latency and parameter budgets. These results show that lightweight detection, structured spatial-semantic decoding, and deeply integrated vision-language priors can be blended into a single efficient HOI pipeline.
|
| 351 |
VideoGen-Agent: Reinforcing Video Generation Agents
2609.24997
|
cs.CV
|
Binxu Li, Haoyi Duan, Yuhui Zhang, Yaohui Zhang, Zihao Lin |
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. ...Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to use external tools for video generation. The agent coordinates augmentation, generation, and verification tools through multi-turn interactions, using the prompt and intermediate observations to guide its decisions. We train a shared policy on a category-balanced dataset spanning six tasks. Supervised fine-tuning on teacher-generated trajectories establishes tool-use behavior, which is then refined through reinforcement learning. A category-aware hybrid reward evaluates tool-call validity, task-appropriate tool use, and generated video quality. We further introduce VABench, a held-out benchmark of 600 prompts covering procedural knowledge, single- and multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6. Upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons. These results support learning tool use across video-generation tasks and show that the trained agent can benefit from subsequent advances in generation tools. Project page: https://andyca111.github.io/VideoGen_Agent/
|
| 352 |
Sex Estimation from Footwear Outsole Impressions Using CNN Transfer Learning and Interpretable Image Statistics
2609.25386
|
cs.CV
|
Jinyi Niu, Ziyi Song, Weining Shen |
Footwear outsole impressions are a common form of forensic pattern evidence, yet quantitative methods for estimating wearer attributes from these images remain relatively underdeveloped. We investigate binary sex estimation from footwear outsole impressions by...Footwear outsole impressions are a common form of forensic pattern evidence, yet quantitative methods for estimating wearer attributes from these images remain relatively underdeveloped. We investigate binary sex estimation from footwear outsole impressions by comparing convolutional neural network (CNN) transfer learning with traditional feature-based classification. Using a publicly available outsole-impression dataset, we adopt a shoe-level training and test partition that keeps replicate scans of the same physical shoe together to reduce data leakage. We evaluate pretrained CNNs through end-to-end fine-tuning, frozen feature extraction followed by support vector machine classification, and hybrid feature fusion incorporating handcrafted, geometric, and metadata-derived descriptors. Fine-tuned CNNs achieve the strongest overall predictive performance and substantially outperform traditional classifiers trained on the manually specified descriptors alone, while frozen-feature approaches offer a less computationally demanding alternative. Exploratory analysis of low-dimensional CNN representations reveals associations with frequency threshold ratio, image contrast, and wavelet-based summaries, providing a connection between learned representations and measurable properties of outsole impressions. These findings suggest that CNN transfer learning captures discriminative information beyond the descriptors considered and offers a promising approach to footwear-based forensic screening. Further validation on independently collected and casework-like impressions is needed before operational use.
|
| 353 |
KwaiMind Technical Report
2609.26375
|
cs.CV
|
Boheng Zhang, Fan Yang, Jia Sun, Junlong Wu, Wenwu Ou |
Commercial image editing requires product identity preservation, accurate text rendering, and user appeal alongside general editing quality. We present KwaiMind, an image editing system combining general capabilities with e-commerce specialization. An agent-ba...Commercial image editing requires product identity preservation, accurate text rendering, and user appeal alongside general editing quality. We present KwaiMind, an image editing system combining general capabilities with e-commerce specialization. An agent-based data engine maintains approximately 1.8 million high-quality editing pairs. Built on a multimodal diffusion transformer, KwaiMind undergoes continued pre-training and supervised fine-tuning, followed by preference optimization and online reinforcement learning. A general-purpose vision-language judge and specialized rewards for click-through rate (CTR), text rendering, and product consistency guide specialized policies, which are consolidated through on-policy distillation. We introduce Ecom-Bench, covering 11 commercial editing tasks with task-specific visual evaluation and CTR-based ranking. KwaiMind achieves the strongest overall scores among evaluated open-source editors on ImgEdit, GEdit, both language splits of REDEdit, and Ecom-Bench visual quality, and the highest aggregate CTR ranking score among compared systems. Offline, CTR-guided optimization increases the proportion of generated images whose predicted CTR exceeds that of the original product image from 12.16% to 37.41%. In an online A/B experiment, CTR-based selection of product main images yields an approximately 2.44% relative increase in actual CTR. These results demonstrate the value of domain-specific data and reward-driven alignment for commercial image editing.
|
| 354 |
RoadOcc Learns When to Persist, Transport, or Refresh Memory for Roadside Occupancy Prediction
2609.27677
|
cs.CV
|
Xiaokai Bai, Lei Yang, Songkai Wang, Lianqing Zheng, Si-Yuan Cao |
Fixed roadside cameras repeatedly observe a stable scene overlaid by sparse moving traffic. Temporal memory can recover weak observations, but reusing moving evidence at stale locations can corrupt occupancy predictions. Motion compensation addresses displacem...Fixed roadside cameras repeatedly observe a stable scene overlaid by sparse moving traffic. Temporal memory can recover weak observations, but reusing moving evidence at stale locations can corrupt occupancy predictions. Motion compensation addresses displacement, while reliance on the resulting history remains a separate learning problem. We introduce RoadOcc, which learns soft routing among fixed-coordinate history (\emph{Persist}), velocity-addressed history (\emph{Transport}), and current evidence (\emph{Refresh}). Motion state and class-consistent historical support supervise these source choices. Dynamic-aware cross-attention (DCA) updates candidate locations, multi-scale voxel velocity estimation (VVE) constructs transport addresses from multi-scale current--history correspondence, and velocity-guided dynamic sparse fusion (VDSF) combines routed evidence under fixed sparse-token budgets. On InfraOcc, RoadOcc reaches 65.29 mIoU and 32.37 dynamic mIoU, gains of 4.44 and 4.71 over STCOcc. Controlled address experiments show that VVE raises dynamic mIoU by 0.87 over fixed-coordinate reading. Across three seeds, supervised P/T/R adds 1.40 dynamic points over motion-corrected retrieval, while removing Refresh costs 0.32 points. Results from two transfer models, Occ3D-nuScenes, and longer intervals provide additional support. Code will be released.
|
| 355 |
TaskIR: Task-Driven Image Restoration via Degradation Adaptation and Task Feedback
2609.31170
|
cs.CV
|
Yanjie Tu, Qingsen Yan, Axi Niu, Wenxuan Cai, Tao Hu |
Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Diff...Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Different degradations impose distinct restoration demands, and insufficient restoration may leave residual degradations and artifacts that impair object boundaries and semantic cues, thereby compromising downstream task performance. To address these challenges, we propose TaskIR, a two-stage task-driven unified image restoration framework that integrates degradation-adaptive restoration with task feedback refinement. In Stage I, a Degradation Representation Module (DRM) extracts degradation representations, enabling a Degradation-Guided Transformer Block (DGTB) to dynamically modulate feature transformations for adaptive restoration. In Stage II, a Task-to-Restoration Feedback Generation module (TRFG) transforms heterogeneous task features into restoration feedback by modeling task-representation discrepancies associated with the current restoration. Subsequently, a Selective Task Feedback Refinement module (STFR) assesses feedback relevance and selectively refines intermediate restoration features to mitigate interference with well-restored content. Extensive experiments demonstrate that TaskIR achieves competitive restoration quality and downstream task performance across diverse degradations and tasks.
|
| 356 |
Temporal-Attention Head Specialization During Video Diffusion Training
2609.31654
|
cs.CVcs.AI
|
Taewoo Ha, Shafayat Mowla Anik, Dae Yeol Lee, Byeong Kil Lee, Jeeho Ryoo |
Video diffusion transformers depend on temporal attention to coordinate information across frames, yet nearly everything known about this mechanism comes from analyzing trained models, so when and where temporal-attention structure forms during training remain...Video diffusion transformers depend on temporal attention to coordinate information across frames, yet nearly everything known about this mechanism comes from analyzing trained models, so when and where temporal-attention structure forms during training remains poorly characterized. Population averages can also hide it, since a few specializing heads and a diffusing majority cancel in the mean. We therefore conduct a checkpoint-resolved census of every temporal-attention head across nine Open-Sora STDiT training runs spanning three model scales (306M to 1.03B parameters), scoring each head with an entropy-normalized measure of cross-frame attention concentration (CFAC) under a preregistered change-point and effect-size selection rule. The census reveals the sparse picture that averages obscure. Aggregate CFAC is flat or decreasing in every run, while a small minority of heads, roughly 4--13% in full-grid runs, develops pronounced concentration. Across seeds, the reproducible signal is positional but block-level. Selected heads repeatedly arise in the first temporal block, whereas individual head coordinates do not reproduce once block membership is accounted for. Among the analyzed 760M selected heads, attention maps converge to a small repertoire of local frame-routing motifs, self-frame diagonals and adjacent-frame bands, even when the responsible coordinates differ across runs. Correlation and ablation analyses do not establish a causal link to generated video quality, and we bound our claims accordingly. Beyond this STDiT family, the study contributes a transferable methodology. Checkpoint-resolved, per-head analysis under fixed selection rules can expose sparse temporal organization in other factorized video diffusion transformers and, with adapted routing metrics, in joint spatio-temporal architectures.
|
| 357 |
TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
2609.32013
|
cs.CVcs.AI
|
Quinlan Sykora, Sourav Biswas, Christopher Diehl, Andrew Cunningham, Thomas Gilles |
We present TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to prior work, TriO utilizes three distinct sensor modalities (camera, LiDAR, and RADAR) as both inputs and sources of self-s...We present TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to prior work, TriO utilizes three distinct sensor modalities (camera, LiDAR, and RADAR) as both inputs and sources of self-supervision, eliminating the need for additional human annotations. Thanks to its novel supervision, the model is able to segment any occupancy from the drivable surface, overcoming the limitations of existing open-set methods in handling long-tail objects. TriO achieves state-of-the-art results in multiple 3D and 4D tasks, including occupancy, flow, and LiDAR prediction, as well as zero-shot road obstacle segmentation across multiple datasets such as Argoverse 2, and Spotting the Unexpected.
|
| 358 |
EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding
2609.32352
|
cs.CVcs.AI
|
Gujie Shao, Zixun Xie, Xuechun Xing, Ruixiang Wang, Ziyun Lan |
Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or spec...Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluate whether VLMs can move beyond disease recognition toward comparative reasoning and fine-grained spatial grounding. We introduce EyeVQA, a unified visual question answering benchmark for comprehensive evaluation of ophthalmic VLMs. EyeVQA is constructed from 21 available ophthalmic datasets and contains 20,000 question-answer pairs spanning six disease groups and seven question types: Single-Choice, Multi-Select, Variable-Select, True-False, Ranking, Point Location, and Bounding Box. Gold answers are deterministically derived from source-provided diagnoses, severity grades, clinical findings, segmentation masks, bounding boxes, and anatomical landmarks, enabling reproducible evaluation without relying on model-generated annotations. Notably, 44.5% of the questions require reasoning across multiple images, extending evaluation beyond conventional single-image medical VQA. We benchmark fourteen representative general-purpose, scientific, and medically specialized VLMs under a unified zero-shot protocol. The best-performing model only achieves an overall score of 62.8, while substantial gaps remain in spatial grounding and cross-task generalization. These results highlight the limitations of current VLMs in comprehensive ophthalmic visual understanding and establish EyeVQA as a diagnostic benchmark for developing more reliable and spatially grounded ophthalmic multimodal models. The project page is available at https://github.com/PKUTHM/EyeVQA.
|
| 359 |
Harnessing Coupled Stream Completion For Human-Object Interaction Modeling
2609.32551
|
cs.CVcs.AI
|
Dawei Guan, Di Yang, Jiangtao Wang |
Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories & rotations, and hand articulation to remain coordinated. These components differ in scale and dynamics, but must agree on contact, relative pose, and timi...Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories & rotations, and hand articulation to remain coordinated. These components differ in scale and dynamics, but must agree on contact, relative pose, and timing. A shared representation may limit the distinct structure of each stream, while independent generation prevents each stream from responding to changes in the others. Latent supervision alone also does not directly constrain contact after decoding. We propose TRACE, a continuous latent framework that keeps stream states separate and couples their updates. TRACE encodes body, object, and hand motion into separate latents and predicts each stream velocity from the complete current interaction state. Geometric losses on decoded motion further constrain contact and object-relative motion over time. The same model supports completion of any single absent stream from the other two. Frozen flow features also serve as input to a language model for HOI understanding. Experiments on InterAct, OMOMO, and BEHAVE show that joint completion training improves generation and that frozen flow features improve understanding over raw-motion encoding. On InterAct, TRACE achieves the highest contact precision, recall, and F1 among the compared methods.
|
| 360 |
Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity
2609.32876
|
cs.CVcs.CLcs.LGcs.AI
|
Yishu Zhang, Yun Li, Daiwei Zhang |
State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as patholog...State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.
|
| 361 |
PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation
2609.33236
|
cs.CV
|
Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao |
Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeate...Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional history. Encoding history through text can compress visual evidence and introduce semantic bias, whereas retaining visual history expands multimodal context. We introduce PARSEE-VAD, a two-module framework that separates semantic evidence acquisition from score-state evolution. Proposition-Aware Reasoning (PAR) extracts structured propositional evidence from the current causal window and conditionally activates more specific queries when coarse evidence warrants further refinement. By sharing a reusable causal visual prefix across queries, PAR reduces redundant computation through selective execution. Streaming Evidence Escalation (SEE) maps the acquired proposition evidence into a compact score-domain event state through current evidence escalation, then propagates only the resulting bounded state across decisions to support temporal continuity. Experiments on four benchmarks demonstrate strong training-free online performance while selective routing reduces specialist computation and score-state propagation remains sparse. These results support a current-first principle for streaming multimodal inference: resolve present semantics first, then use compact historical state only to repair residual continuity gaps.
|
| 362 |
PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers
2609.33384
|
cs.CV
|
Yutong Wang, Xingtong Ge, Enhuai Liu, Yunke Wang, Tianfan Xue |
Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method th...Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with activation geometry to guide offline calibration. Isolated block--step interventions estimate propagation risk, which prioritizes sensitive trajectory states during row-radius selection. With these radii fixed, response-subspace correction uses neighboring-code edits to reduce residual components along dominant activation directions. Both stages preserve the original 4-bit weight representation. Controlled interventions show that short-horizon propagated error predicts final latent error more reliably than immediate block-output error, supporting calibration beyond local reconstruction objectives. Evaluations on Wan models, Self Forcing, and MiniMax-H3 demonstrate improvements in key consistency and dense-reference metrics while remaining competitive on other attributes across model scales and generation paradigms.
|
| 363 |
SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation
2609.33399
|
cs.CV
|
Jiali Chen, Zhengteng Lin, Zuqi Wang, Shirong Lin, Xi Yu |
In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have ena...In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.
|
| 364 |
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
2609.33419
|
cs.CV
|
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen |
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify w...Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture-objective study at roughly 170M ~ 190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
|
| 365 |
SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera
2609.33462
|
cs.CV
|
Shriram Damodaran, Soumyaratna Debnath, Cheston Tan, Lin Wang |
Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use of Multi-modal Large Language Models (MLLMs) for omnidirectional spatial reasoning. However, most MLLMs are train...Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use of Multi-modal Large Language Models (MLLMs) for omnidirectional spatial reasoning. However, most MLLMs are trained on conventional 2D perspective images and struggle with the severe distortions and wrap-around discontinuities induced by spherical geometry. Enabling them to generalize to non-Euclidean 3D spaces without retraining therefore remains challenging. We propose SphMind, a training-free, plug-and-play framework that decouples semantic perception from geometric reasoning. Rather than requiring MLLMs to learn spherical geometry internally, SphMind preserves their semantic capabilities while handling geometry externally. We introduce a Spherical Harmonics-based Spatial Graph (SHSG) that models spatial relationships through equivariant transformations on the sphere, together with Inference-Time Geometric Grounding (IGG), a model-agnostic closed-loop optimization process that aligns MLLM representations with spherical geometric constraints during inference. Experiments on three benchmarks show that SphMind achieves over 21.4% average improvement in directional reasoning on MP3D and Stanford2D-3D, outperforms prompt-engineering baselines by 8.7% on the real-world ODI-Bench, and improves rotational invariance by 5.9% under panorama rotations, without additional training or dataset-specific tuning. In-the-wild evaluations further show that SphMind resolves directional reasoning queries that baseline vision-language models fail to answer correctly.
|
| 366 |
PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization
2609.33558
|
cs.CV
|
Liang Peng, Shizhuo Mu, Bohan Tan, Wenyuan Wang, Chen Zhao |
3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotat...3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search frames. The benchmark baseline predicts cuboids after feature modeling, leaving their geometry unused for subsequent feature refinement. We investigate whether complete intermediate cuboids can improve query and proposal representations before final decoding. We introduce Progressive Geometric Learning for 3DVQL (PGL-3D), a predict--select--refine--re-predict framework that uses intermediate cuboids to guide the aggregation of search evidence and update query and proposal representations. A shared head first predicts a complete cuboid for every proposal. Query--Tube--Memory (QTM) then selects reference observations by combining proposal association, cuboid quality, frame response, and target absence, since association confidence alone establishes neither target presence nor geometric accuracy. The center, size, and orientation of each selected cuboid define soft pooling weights over query-conditioned proposal features. The pooled memory updates the query and proposal representations, and the head re-predicts from the updated features. A training-only objective, ST-D9O, supervises cuboid geometry at every stage by adding boundary, signed-distance, and soft-overlap terms to parameter regression. PGL-3D achieves a mean stAP of $0.270 \pm 0.004$ on 3DVQL, compared with $0.044$ reported for LaF. Ablations support the benefits of geometry-guided feature updates, while stage-wise analyses show improved cuboid accuracy. Replacing the geometry objective in our PROT3D reproduction with ST-D9O improves mAO on GSOT3D from $21.63\%$ to $25.78\%$. Our code and models will be released.
|
| 367 |
Natural Image Autoencoder-Based fMRI Representations for Trait and State Prediction
2609.34167
|
cs.CV
|
Juhyeon Park, Yeonwoo Kim, Peter Yongho Kim, Yansen Wang, Mingqing Xiao |
Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, wh...Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a frozen Deep Compression AutoEncoder (DCAE) pre-trained exclusively on natural images and pairs them with a task specific readout. For trait prediction, FReD summarizes frame-wise representations by their temporal mean and log-standard deviation and applies linear probing, with late fusion across two normalization schemes. For state prediction, it represents each frame as a single token and models temporal dependencies with a shallow Transformer. Across four resting-state datasets spanning six trait-prediction targets, linear probes on frozen DCAE features generally outperform those on fMRI foundation model representations and remain competitive with fully fine-tuned fMRI foundation models. On three task-fMRI state-prediction tasks, a temporal readout on DCAE features performs comparably to the strongest foundation models evaluated. A Gaussian injection analysis further shows that localized signal changes are recovered more accurately from the frozen DCAE features than from the evaluated foundation-model representations. Together, these results show that strong performance on current fMRI benchmarks is possible without fMRI-specific representation pre-training, making frozen natural-image features as a useful baseline for assessing its added value.
|
| 368 |
CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving
2609.34387
|
cs.CV
|
Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen |
Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth,...Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for deciding how to reason in time-critical situations. We therefore propose CAR-VLA, a unified driving VLA model that jointly considers scene complexity and dynamic risk to guide reasoning depth, urgency, and focus. CAR-VLA maps four complexity--risk categories to three reasoning modes: \textit{Fast Intuition} for direct trajectory generation in simple low-risk scenes, \textit{Slow Thinking} for deliberate reasoning in complex low-risk scenes, and \textit{Reflex Response} for compact, hazard-focused reasoning in high-risk scenes regardless of complexity. Rather than merely shortening deliberation, Reflex Response centers reasoning on the most critical hazard and the immediate safe response. We train CAR-VLA through progressive supervised learning that links scene assessment, reasoning-mode selection, and trajectory generation, followed by reasoning-augmented reinforcement learning to improve driving quality and reasoning behavior. Experiments on NAVSIM v1(91.1 PDMS), NAVSIM v2(90.3 EPDMS), and Navhard(35.0 EPDMS) demonstrate competitive driving performance. Qualitative comparisons on navtest and in-house high-risk scenarios further illustrate risk-aware reasoning and hazard-responsive trajectory generation. The code for this paper will be released publicly at: https://github.com/chenxl124578/CAR-VLA.git
|
| 369 |
GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior
2609.34579
|
cs.CVcs.AI
|
Yajiao Xiong, Youyu Luan, Xiaoyu Zhou, Yongtao Wang |
Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of fore...Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-enhanced novel view synthesis via a disentangled 3D prior. Specifically, GenNVS models foreground objects and the background with 3D Gaussian Splatting and aligns them through a coarse-to-fine geometric optimization process to form a unified 3D scene. This scene conditions a video diffusion model through the proposed Dual-Stream Masking mechanism, which guides synthesis by jointly exploiting rendered validity masks and geometry-aware warping. Experimental results show that GenNVS performs favorably against recent methods in both visual quality and geometric accuracy, while naturally supporting flexible scene editing.
|
| 370 |
Counterfactual Attention Policy Distillation for Temporal Video Grounding
2609.34581
|
cs.CV
|
Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang |
Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we...Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distillation (OPD) and propose a new training regime for MLLMs termed Counterfactual Attention Policy Distillation (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.
|
| 371 |
DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding
2609.34651
|
cs.CV
|
Jiantao Lin, Yingjie Xu, Mingzhi Sheng, Yangkai Wei, Hao Chen |
Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D inf...Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid positions rather than to the underlying surface geometry. This mismatch limits coherence across seams and disconnected UV islands. We propose DirectUV, an image-conditioned UV texture diffusion framework that operates in the latent UV space of a pretrained image VAE, in which a Diffusion Transformer denoises the UV latent given a single input image and a coarse UV map. At its core, Surface-Aware Positional Encoding (SAPE) replaces the standard 2D-grid positional encoding with encodings derived from per-token 3D surface coordinates obtained via UV-to-surface correspondence. As positional encodings define the distance metric used by attention, SAPE enables tokens to interact according to 3D positional proximity derived from surface correspondence rather than UV-grid distance, restoring coherence across seams and disconnected islands. A multi-level extension further assigns different attention heads to progressively finer subdivisions of the same latent UV patch, allowing the model to reason about surface structure at multiple granularities. Experiments show that DirectUV produces sharper and more globally consistent textures than other baselines, with the largest improvements in occluded and view-unseen regions where projection-based methods leave gaps or stretched textures.
|
| 372 |
CapField-OPD: Learning Continuous Capability Fields via Joint-Anchored Multi-Teacher On-Policy Distillation for Flow Models
2609.34658
|
cs.CV
|
Pengyang Ling, Jiazi Bu, Yujie Zhou, Yibin Wang, Zeqiang Lai |
Reward-specialized post-training produces strong experts for flow-based generative models, while multi-teacher on-policy distillation (OPD) consolidates their capabilities into a single student. Existing methods, however, route each prompt to a single teacher ...Reward-specialized post-training produces strong experts for flow-based generative models, while multi-teacher on-policy distillation (OPD) consolidates their capabilities into a single student. Existing methods, however, route each prompt to a single teacher according to its semantic category, implicitly binding the desired capability to prompt content. This coupling makes capability invocation vulnerable to prompt perturbations and prevents users from explicitly adjusting the strength of the desired capability at inference time. In this work, we introduce CapField-OPD, an OPD framework that integrates multiple teachers into a continuous capability field through explicit capability coordinates. We use teacher models as anchors to construct this field, with the coordinates determining how their outputs are combined. Each capability configuration thus receives a unique supervision target, and capability control no longer depends on prompt semantics. Since the training anchors may not be optimal at inference time, we further profile the learned field on a small calibration set. The coordinate with the highest mean reward serves as the recommended default, while coordinates that are frequently optimal offer a promising candidate set for test-time scaling. Extensive experiments on compositional generation, text rendering, and visual aesthetics demonstrate that CapField-OPD consolidates multiple specialized teachers into a single student while preserving or surpassing their performance, reliably invokes the desired capabilities under semantics-preserving prompt variations, and supports continuous capability control and coordinate-based test-time scaling.
|
| 373 |
From Pixel Generation to Topological Inference: Structural Dual Super-Resolution for Trustworthy Cross-Physical-Domain Trabecular Morphology Learning
2609.34716
|
cs.CV
|
Fan Zhang, Yi Zhang, Ling Wang |
Clinical CT and UHRCT cannot resolve individual trabeculae, whereas synchrotron radiation microCT (SRuCT) provides high-resolution references but is not applicable for in vivo imaging. The two domains differ by a 32x resolution gap, are only coarsely paired, a...Clinical CT and UHRCT cannot resolve individual trabeculae, whereas synchrotron radiation microCT (SRuCT) provides high-resolution references but is not applicable for in vivo imaging. The two domains differ by a 32x resolution gap, are only coarsely paired, and exhibit severe physical differences including partial volume effects, noise, and artifacts. Existing super-resolution networks and pretrained-prior methods (GLEAN/StyleGAN2, Stable SR/LDM) underperform because they target pixel generation---diverse details and SSIM/PSNR---and do not model these physical differences. Pixel generation for a 32x resolution gap is intrinsically ill-posed. We propose a paradigm shift from pixel generation to topological inference: deterministically predicting invariant microstructures from macro-scale low-resolution inputs, evaluated by morphological parameters. The core of our 2D morphology learning lies in training on 2D slices while evaluating on 3D morphological parameters, ensuring that the learned representations capture true three-dimensional trabecular topology rather than 2D pixel statistics. We realize this paradigm via structural dual super-resolution, coupling forward physical degradation (micro-to-macro) with inverse structural inference (macro-to-micro) through structural duality constraints. The method is an end-to-end, few-shot, compact structural dual network (SDN), comprising a bidirectional modeling network, a multi-scale structural consistency discriminator, and four structural duality constraints. On the test set, SDN achieves morphological parameters largely consistent with SRuCT across six metrics, with SSIM reaching 0.8. Trained on 3.2 um SSRF data, the model generalizes well to 3.25 um isotropic BSRF data from an independent source, validating cross-source generalization and confirming trustworthy structural inference rather than pixel generation.
|
| 374 |
LVMT: Video Mask Transformer for Long-term Video Segmentation
2609.34895
|
cs.CV
|
Narges Norouzi, Niccol\`{o} Cavagnero, Idil Esen Zulfikar, Bastian Leibe, Gijs Dubbelman |
Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object i...Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradients. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time. Second, to allow training on long videos, we introduce Truncated Query Propagation (TQP), a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients. The resulting model is called the Long-term Video Mask Transformer (LVMT). Extensive experiments on six benchmarks show that LVMT sets a new state of the art across a range of video segmentation tasks, while retaining the speed of the highly efficient model it is based on, making it 10X faster than the prior state of the art. Code: https://www.tue-mps.org/lvmt
|
| 375 |
What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling
2609.34981
|
cs.CV
|
Renping Zhou, Zanlin Ni, Zihao Fan, Guohao Fu, Zeyu Liu |
World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along wit...World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirely for acceleration. We find that latent WAMs, despite matching explicit ones on in-distribution tasks, fail to retain the generalization benefits that originally motivated WAMs. To demonstrate this, we evaluate generalization along three axes: environmental perturbation, data efficiency, and task generalization. Controlled comparisons with a matched backbone, training data, and budget reveal consistent degradation across all three axes when the action expert no longer conditions on future representations. Further analysis shows that the gap arises almost entirely from the first denoising step: the benefit comes from preparing the future, not generating it. We therefore propose Simple-WAM, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the training-time noise schedule to this inference behavior. Across simulation and real-world tasks, Simple-WAM achieves the best of both worlds, leading explicit WAMs in generalization performance with efficiency comparable to Latent WAMs. Project Page: https://zrporz.github.io/Simple-WAM-Web/
|
| 376 |
Beyond Selection: Token Parameterization for Extreme Visual Token Compression
2609.35232
|
cs.CVcs.LG
|
Rui Zhong, Yu Li, Zheyu Yan, Cheng Zhuo |
Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training comple...Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under $23\times$--$64\times$ compression and remains competitive at $144\times$, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%--86.7% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to approximately 36% end-to-end speedup and using $16.6\times$/$78.8\times$ lower compressor latency/FLOPs.
|
| 377 |
From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
2609.35491
|
cs.CVcs.AI
|
Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li |
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminat...Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nystr\"om--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.
|
| 378 |
GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
2609.35734
|
cs.CV
|
Kerui Ren, Tao Lu, Linning Xu, Changjian Jiang, Mu Huang |
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene struc...Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.
|
| 379 |
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
2501.18157
|
cs.CVcs.SDeess.AScs.MM
|
Joanna Hong, Sanjeel Parekh, Honglie Chen, Jacob Donley, Ke Tan |
Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come with several constrai...Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come with several constraints such as increased sensory requirements, computational cost, and modality synchronization, to mention a few. These challenges constrain the direct uses of these multimodal solutions in real-world applications. In this work, we develop approaches where the learning happens with all available modalities but the deployment or inference is done with just one or reduced modalities. To do so, we propose a Multimodal Training and Unimodal Deployment (MUTUD) framework which includes a Temporally Aligned Modality feature Estimation (TAME) module that can estimate information from missing modality using modalities present during inference. This innovative approach facilitates the integration of information across different modalities, enhancing the overall inference process by leveraging the strengths of each modality to compensate for the absence of certain modalities during inference. We apply MUTUD to various audiovisual speech tasks and show that it can reduce the performance gap between the multimodal and corresponding unimodal models to a considerable extent. MUTUD can achieve this while reducing the model size and compute compared to multimodal models, in some cases by almost 80%.
|
| 380 |
Gondola: Grounded Vision Language Planning for Robotic Manipulation
2506.11261
|
cs.CVcs.AI
|
Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid |
Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretability and weak robustness in complex, lo...Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretability and weak robustness in complex, long-horizon tasks. To address these challenges, we employ a modular manipulation framework that separates high-level planning from low-level control. At its core is Gondola, a grounded vision-language planning model that generates structured plans with explicit pixel-level object grounding before action execution. Given multi-view observations and planning history, Gondola predicts the next-step plan as interleaved textual instructions and multi-view segmentation masks corresponding to target objects and goal locations. To train Gondola, we construct synthetic datasets that provide explicit supervision for short-horizon grounded planning, multi-view referring expression, and long-horizon compositional reasoning. By coupling grounded plan generation with a 3D-based execution policy, our framework achieves state-of-the-art performance on the challenging GemBench benchmark. The system further demonstrates promising transfer to real robots. Ablation studies confirm that pixel-level grounding and the proposed planning-oriented supervision are critical for effective high-level reasoning. Project webpage: https://cshizhe.github.io/projects/robot_gondola.html
|
| 381 |
CPATTA: Conformal Supervision Allocation For Active Test-Time Adaptation
2509.25692
|
cs.CVcs.AI
|
Tingyu Shi, Fan Lyu, Haihua Zhu, Dadi Wang, Shaoliang Peng |
Active Test-Time Adaptation (ATTA) improves model robustness under domain shift by selectively querying human annotations at deployment, but existing methods use heuristic uncertainty measures and suffer from low data selection efficiency, wasting human annota...Active Test-Time Adaptation (ATTA) improves model robustness under domain shift by selectively querying human annotations at deployment, but existing methods use heuristic uncertainty measures and suffer from low data selection efficiency, wasting human annotation budget. We propose Conformal Prediction Active TTA (CPATTA), which first brings principled, conformal uncertainty with coverage-aware online calibration into ATTA. CPATTA employs smoothed conformal scores with a top-$K$ certainty measure, an online weight-update algorithm driven by pseudo coverage, a domain-shift detector that adapts human supervision, and a staged update scheme that balances human-labeled and model-labeled data. Extensive experiments demonstrate that CPATTA consistently outperforms the state-of-the-art ATTA methods by around 5% in accuracy.
|
| 382 |
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
2601.01762
|
cs.CV
|
Yanhao Wu, Haoyang Zhang, Fei He, Rui Wu, Yanhu Shan |
Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. While state-of-the-art (SOTA) methods use parallel planning architectures, they fail to explicitly couple speed decisio...Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. While state-of-the-art (SOTA) methods use parallel planning architectures, they fail to explicitly couple speed decisions with agent behavior along the driving path, leading to suboptimal coordination. To address this, we propose a cascaded framework that transforms longitudinal planning from an independent prediction task into a path-conditioned reasoning process. On the model side, we introduce an anchor-based regression design that conditions longitudinal prediction on the lateral drive path, and reformulate longitudinal planning as 1D displacement prediction along the path. This reduces geometric uncertainty and sharpens the model's focus on interaction-driven dynamics. On the data side, we introduce a planning-oriented data augmentation strategy that simulates rare safety-critical events by programmatically inserting agents and relabeling longitudinal targets to enforce collision avoidance. Evaluated on the challenging Bench2Drive benchmark, our method achieves SOTA performance with a driving score of 89.07 and a success rate of 73.18%, demonstrating significantly improved coordination and safety. Further evaluation on Fail2Drive confirms strong generalization to rare edge cases where parallel formulations typically fail. Project page:https://yanhaowu.github.io/AlignDrive/.
|
| 383 |
Beyond Pixels: A Vector-to-Graph Framework for Reliable Schematic Auditing
2602.11678
|
cs.CVcs.AI
|
Chengwei Ma, Zhen Tian, Zhou Zhou, Zhixian Xu, Xiaowei Zhu |
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state-of-the-art models fail to capture topology and symbolic logic in engineering schematic...Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state-of-the-art models fail to capture topology and symbolic logic in engineering schematics, as their pixel-driven paradigm discards the explicit vector-defined relations needed for reasoning. To overcome this, we propose a Vector-to-Graph (V2G) pipeline that converts CAD diagrams into property graphs where nodes represent components and edges encode connectivity, making structural dependencies explicit and machine-auditable. On a diagnostic benchmark of electrical compliance checks, V2G yields large accuracy gains across all error categories, while leading MLLMs remain near chance level. These results highlight the systemic inadequacy of pixel-based methods and demonstrate that structure-aware representations provide a reliable path toward practical deployment of multimodal AI in engineering domains. To facilitate further research, we release our benchmark and implementation at https://github.com/gm-embodied/V2G-Audit.
|
| 384 |
Formalizing the Sampling Design Space of Diffusion-Based Generative Models via Adaptive Solvers and Wasserstein-Bounded Timesteps
2602.12624
|
cs.CV
|
Sangwoo Jo, Sungjoon Choi |
Diffusion-based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by high sampling costs. While prior work focuses on training objectives or individual solvers, the broader sampling d...Diffusion-based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by high sampling costs. While prior work focuses on training objectives or individual solvers, the broader sampling design problem, specifically solver selection and scheduling, remains largely governed by static heuristics. We propose SDM, a principled, training-free sampling framework that adapts both the numerical solver and the timestep schedule to the intrinsic properties of the diffusion trajectory. By analyzing the PF-ODE dynamics, we show that velocity variation is small in high-noise stages and increases near the data manifold, identifying intervals where solver order is most consequential. In parallel, we introduce an offline-calibrated adaptive scheduling method that explicitly controls the local Wasserstein discretization error and projects the calibrated trajectory to a prescribed NFE budget. We further extend the formulation to a mixed-transition Wasserstein error bound, providing a unified error-propagation view of adaptive scheduling and solver selection within the overall SDM framework. Across standard benchmarks, with extensions to modern ODE samplers, high-resolution synthesis, and text-to-image generation, SDM achieves improved sample quality compared to baseline methods, attaining an FID of 1.93 on CIFAR-10, 2.41 on FFHQ, and 1.98 on AFHQv2, with a reduced number of function evaluations compared to existing samplers. Our code is available at https://github.com/aiimaginglab/sdm.
|
| 385 |
Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
2602.15382
|
cs.CV
|
Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He |
Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We ...Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We introduce the Vision Wormhole, which repurposes the visual input interface of Vision-Language Models (VLMs) for continuous communication between frozen heterogeneous agents. A Universal Visual Codec encodes each sender's latent rollout into a fixed-size message, maps it through a shared reference space, and decodes received messages into the receiver's image-token span. Per-model codecs and affine reference maps form a hub-and-spoke architecture with $O(N)$ components for $N$ models. Each model learns its codec independently through self-distillation on anchor texts, and shared-anchor alignment enables reuse across communication partners. Across four VLM families, six team configurations, and nine reasoning benchmarks, Vision Wormhole improves accuracy by 6.0 percentage points on average over text-mediated MAS and achieves a 1.69$\times$ geometric-mean speedup in batch-normalized end-to-end runtime.
|
| 386 |
Learning to Recorrupt: Noise Distribution Agnostic Self-Supervised Image Denoising
2603.25869
|
cs.CV
|
Brayan Monroy, Jorge Bacca, Juli\'an Tachella |
Self-supervised image denoising methods have traditionally relied on architectural constraints, pseudo-pair constructions, or specialized loss functions to avoid the trivial identity mapping. Among these, approaches such as Noisier2Noise or R2R create training...Self-supervised image denoising methods have traditionally relied on architectural constraints, pseudo-pair constructions, or specialized loss functions to avoid the trivial identity mapping. Among these, approaches such as Noisier2Noise or R2R create training pairs by adding synthetic noise to noisy images. While effective, these recorruption-based approaches require precise knowledge of the noise distribution, which is often unavailable. We present Learning to Recorrupt (L2R), a self-supervised framework that does not require exact knowledge of the noise distribution. Our method introduces a learnable recorruptor jointly optimized with the denoiser through a min--max saddle-point objective. The proposed method achieves state-of-the-art performance among methods without prior knowledge of the noise distribution across unconventional and heavy-tailed noise distributions, such as log-gamma and Laplace, as well as spatially correlated noise, while obtaining competitive results on real-world denoising benchmarks characterized by signal-dependent noise.
|
| 387 |
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
2604.15086
|
cs.CVcs.SDcs.MM
|
Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang, Lipan Zhang |
Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability under visual-text c...Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability under visual-text conflict and imprecise stylistic control due to entangled temporal and timbre information in reference audio. Moreover, the lack of standardized benchmarks limits systematic evaluation. We propose ControlFoley, a unified multimodal V2A framework that enables precise control over video, text, and reference audio. We introduce a joint visual encoding paradigm that integrates CLIP with a spatio-temporal audio-visual encoder to improve alignment and textual controllability. We further propose temporal-timbre decoupling to suppress redundant temporal cues while preserving discriminative timbre features. In addition, we design a modality-robust training scheme with unified multimodal representation alignment (REPA) and random modality dropout. We also present VGGSound-TVC, a benchmark for evaluating textual controllability under varying degrees of visual-text conflict. Extensive experiments demonstrate state-of-the-art performance across multiple V2A tasks, including text-guided, text-controlled, and audio-controlled generation. ControlFoley achieves superior controllability under cross-modal conflict while maintaining strong synchronization and audio quality, and shows competitive or better performance compared to an industrial V2A system. Code, models, datasets, and demos are available at: https://github.com/xiaomi-research/controlfoley.
|
| 388 |
NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
2605.08144
|
cs.CVcs.AI
|
Haokai Zhao, Da Xing, Hanqun Cao, Tinson Xu, Xinyu Xiang |
Training a diffusion model involves two sources of randomness for each data sample: the timestep and the Gaussian noise realization. The timestep has been studied extensively through scheduling and weighting, whereas the impact of the noise realization at a gi...Training a diffusion model involves two sources of randomness for each data sample: the timestep and the Gaussian noise realization. The timestep has been studied extensively through scheduling and weighting, whereas the impact of the noise realization at a given timestep is still underexplored. In this work, we examine whether different noise instances are equally informative. We introduce NoiseRater, a network that scores an individual noise instance conditioned on the data sample and timestep. The rater is learned through bilevel optimization, where its scores reweight the diffusion loss in the inner loop, and it is updated to reduce validation loss after the inner-loop updates. Using the trained rater to select training noise, we observe three properties of training noise. First, noise realizations at the same timestep are not equally useful: the rater's top-scored noise improves performance over i.i.d.\ sampling, while its bottom-scored noise degrades it. Second, this utility is contextual, depending jointly on the image, the class, and the timestep. Third, noise selection is complementary to timestep-level design, retaining most of its gain when combined with existing scheduling and weighting schemes. These findings establish instance-level noise valuation as a new axis for understanding and improving diffusion training. Code is available at https://github.com/JoeZhao527/Noise-Rater.
|
| 389 |
SCAMP: Sparse-anchor Control is One Small Projection
2605.14716
|
cs.CV
|
Pengcheng Fang, Tengjiao Sun, Xiaoyu Zhan, Yanwen Guo, Hansung Kim |
Authoring with a text-to-motion generator needs sparse anchors: chosen joints, at chosen frames, at given positions. Meeting them currently costs a conditioning branch trained for the task, or hundreds of per-clip optimisation steps in the architecture's nativ...Authoring with a text-to-motion generator needs sparse anchors: chosen joints, at chosen frames, at given positions. Meeting them currently costs a conditioning branch trained for the task, or hundreds of per-clip optimisation steps in the architecture's native variables. In any generator that decodes a continuous state through a frozen differentiable decoder, the anchors ask for little and leave most of the state free: a few hundred numbers against a state of tens of thousands. Every control method is a choice among the states that satisfy them, and the choices differ along the directions the anchors cannot see and the motion can. SCAMP makes the choice that moves none of them: damped Gauss-Newton in the space of the anchors, through the frozen decoder alone, training-free, with one dimensionless damping constant. Every increment is a combination of the rows of the anchors' Jacobian, so the correction is orthogonal to everything the anchors never see, and the system solved is the size of the request rather than of the state. Applied unchanged to seven published generators spanning diffusion, token and latent designs, it matches or exceeds in anchor error every released control method it is measured against, and closes the anchors on hosts that ship none. Confined to those rows, a correction can only take the shapes the decoder admits, so what it costs belongs to the decoder, and holding the solver fixed makes that cost measurable: it divides by decoder family, windowed decoders staying within a small multiple of the unconstrained generator's foot skating where analytic recoveries multiply it several times over. Built to that criterion, our own generator reaches 0.083 m anchor error at FID 0.102 in 0.50 s per clip. The decoder's temporal support is a design criterion for controllable motion generation.
|
| 390 |
AdvantageFlow: Regularized Advantage-Weighted RL in Flow Models
2605.26013
|
cs.CVcs.AI
|
Branislav Kveton, Anup Rao, Subhojyoti Mukherjee, Krishna Kumar Singh, Viet Dac Lai |
We present AdvantageFlow, a forward-process reinforcement learning (RL) algorithm for rectified flow models. The algorithm minimizes an advantage-weighted prediction loss, which maximizes reward, regularized by the rollout policy, which convexifies the objecti...We present AdvantageFlow, a forward-process reinforcement learning (RL) algorithm for rectified flow models. The algorithm minimizes an advantage-weighted prediction loss, which maximizes reward, regularized by the rollout policy, which convexifies the objective and makes its optimization stable. Our objective can be viewed as fitting a local reward-improving target distribution. The rollout regularization arises as a variance reduction step. We evaluate AdvantageFlow empirically on text-to-image generation with Stable Diffusion 3.5 Medium and FLUX.1, and compare it to both forward- and reverse-process RL algorithms.
|
| 391 |
EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control
2606.08495
|
cs.CV
|
Haoyang Ge, Peng Ren, Yukun Shi, Cong Huang, Kun Li |
Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic interfaces, but neither offers a scalabl...Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic interfaces, but neither offers a scalable and interactive prior for broad full-body behavior. We introduce EgoPriMo (Egocentric Motion Prior for Humanoid Robots), a unified framework that learns such priors from egocentric human demonstrations. Given egocentric observations and a text prompt, EgoPriMo reconstructs, generates, and forecasts SMPL-based full-body motion. Language is used as a high-level control signal rather than a complete motion specification. At the core of EgoPriMo is a Triple-stream DiT that jointly models body dynamics, egocentric visual context, and text; task-conditioning masks route different tasks and missing-modality data through the same checkpoint. Experiments on Nymeria and EgoExo4D show that one checkpoint improves egocentric motion generation over UniEgoMotion while supporting reconstruction and forecasting; the generated SMPL motions can also be executed by a Unitree humanoid controller. These results indicate a practical path from scalable egocentric observations to generalizable and interactive humanoid motion priors.
|
| 392 |
Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport
2606.21030
|
cs.CV
|
Yinhuan Huang, Hao Cao, Pu chen, Wenqi Guo, Jungong Han |
Diffusion-based image compression has achieved strong perceptual quality at ultra-low bitrates. However, existing codecs are often tied to specific backbones and specialized components, making diverse, rapidly evolving generative models difficult to reuse. Thi...Diffusion-based image compression has achieved strong perceptual quality at ultra-low bitrates. However, existing codecs are often tied to specific backbones and specialized components, making diverse, rapidly evolving generative models difficult to reuse. This raises a natural question: Can modern generative foundation models be connected to image compression through a simple and extensible interface? Two insights guide our design: stronger generative priors make a simpler codec interface viable, and generation and compression can be intrinsically linked through latent transport. We therefore propose Gen2-IC with two stages: (1) Latent Compression maps clean image latents to entropy-constrained latents; and (2) Latent Transport refines them with one near-terminal update based on the pretrained model. Gen2-IC requires neither auxiliary conditioning signals nor task-specific backbone modifications. With lightweight adaptation and no distillation, it supports fast encoding and one-step decoding across multiple bitrates. We validate Gen2-IC on SD-2.1, SANA-1.5, FLUX.1-dev, and Qwen-Image-2512, spanning U-Net and Transformer architectures as well as diffusion and flow-matching formulations. With stronger priors, Gen2-IC delivers gains below 0.05 bpp: the Qwen variant leads diffusion-based generative codecs in reconstruction fidelity (PSNR), perceptual similarity (LPIPS and DISTS), and recognizer-based semantic fidelity (OCR CER/WER and face-ROI similarity) across four benchmarks.
|
| 393 |
NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics
2606.31764
|
cs.CV
|
Jingye Qiu, Shizhe Zhou |
Differentiable rendering of planar rational splines remains largely underexplored, despite their widespread use in vector graphics and design. Existing differentiable vector renderers primarily focus on B\'ezier curves and rely on analytic rasterization, which...Differentiable rendering of planar rational splines remains largely underexplored, despite their widespread use in vector graphics and design. Existing differentiable vector renderers primarily focus on B\'ezier curves and rely on analytic rasterization, which can suffer from gradient instability and limited flexibility. We propose NURBS Splatting, a unified framework that represents planar rational curves as continuous Gaussian fields. By sampling Gaussians along the curve parameter domain and inside closed regions, rendering is reformulated as a smooth accumulation process with stable gradients. Our method naturally supports long splines, rational weights, non-uniform knots, and closed-region filling. We demonstrate its effectiveness in calligraphy reconstruction, vectorization frameworks, and long-spline image abstraction, showing improved stability and reconstruction quality over existing approaches.
|
| 394 |
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
2608.09853
|
cs.CV
|
Dongchi Huang, Hongyin Zhang, Bohan Hou, Siteng Huang, Zhian Su |
General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal ...General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's $\tau_a$ of 0.704 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. As a zero-shot reward model, RynnValue serves a range of downstream applications. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline; used for data filtering, it improves multi-task behavior cloning success from 35.0% to 42.5%; and applied as inference-time value guidance, it lifts a frozen policy's success from 67.5% to 80.0%. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.
|
| 395 |
Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory
2608.16889
|
cs.CVcs.AI
|
Bingxin Xu, Yuzhang Shang, Emilio Ferrara |
Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) and world-action models (WAMs) increasingly master individual skills, yet the chain still fails: errors compound beyond the policy's ability...Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) and world-action models (WAMs) increasingly master individual skills, yet the chain still fails: errors compound beyond the policy's ability to correct, and one subtask silently constrains the next. A promising pathway freezes the VLA and puts an LLM coding agent in charge: it plans in language, moves in free space with analytic primitives, invokes the VLA only for contact-rich segments, and writes adaptation into language memory. Yet applied to long horizons, this recipe breaks twice. (1) Its competence comes from whole-task exploration at test time, whose cost is exponential in the number of stages: if one stage needs T episodes, a K-stage task needs on the order of T^K, and a failure does not reveal which stage caused it. (2) It has no representation of transitions: the VLA primitive carries an exit but no entry condition, and a subtask can succeed in a form its successor cannot use. We present BATON to address both failures. Against (1), BATON makes the subtask the unit of exploration: each subtask is explored in the cheap short-horizon regime and its solution stored in memory; a long-horizon trajectory is then composed from these solutions rather than discovered whole. Exploration cost becomes linear (KT), and each failure is attributed to one stage. Against (2), BATON equips exploration with a transition-aware memory. Within a subtask, a verifier agent governs the invocation transition: the VLA is invoked only after the wrist view confirms the scene is ready. Across subtasks, a handoff transition restores an entry state disturbed by the predecessor's residue, and a lookahead transition selects the strategy whose outcome the successor can inherit. On the RoboMemArena benchmark, BATON improves task success by 37.7% and cumulative success by 29.7% over the SoTA.
|
| 396 |
Unified-protocol voxel-level pulmonary embolism annotations for three public CT angiography datasets
2608.24486
|
cs.CV
|
Qihang Sun, Zhongxiao Liu, Bailiang Jian, Shenman Qiu, Jingyuan Wang |
Reliable clot-volume quantification and subsequent risk assessment in pulmonary embolism depend on precise segmentation of emboli on computed tomography pulmonary angiography. Deep learning models for this task must be trained on accurate voxel-level labels. T...Reliable clot-volume quantification and subsequent risk assessment in pulmonary embolism depend on precise segmentation of emboli on computed tomography pulmonary angiography. Deep learning models for this task must be trained on accurate voxel-level labels. The three public datasets that provide such labels were annotated under different protocols, and some of their studies contain unlabeled emboli or labels that are discontinuous across slices. This Data Descriptor presents voxel-level pulmonary embolism annotations for 149 of the 166 studies in these datasets. A primary rater drew all annotations under a single protocol. A thoracic radiologist with more than 20 years of experience reviewed and revised them. Three raters at three different centers independently annotated a subset of 15 studies. The subset was selected by source dataset and embolus location. Technical validation quantifies volumetric agreement with the source annotations, changes in within-mask attenuation, and inter-rater agreement on the subset. The dataset is intended to allow segmentation models to be developed and compared under a common reference standard.
|
| 397 |
ClusterAttention: A training-free speedup of bidirectional attention
2608.26965
|
cs.CV
|
Kasper Nordenram, Amelie Dittmann |
We introduce ClusterAttention, a general training-free speedup of bidirectional attention at large token counts. We point out two common assumptions in contemporary training-free methods; attention sparsity, and context that can be leveraged, such as structure...We introduce ClusterAttention, a general training-free speedup of bidirectional attention at large token counts. We point out two common assumptions in contemporary training-free methods; attention sparsity, and context that can be leveraged, such as structure in the input or multiple similar forward passes, and show when they fail. Our proposed method utilizes a fast attention-aware recursive clustering method, and compensation of excluded clusters through their mean. The clustering method gives power-of-two cluster sizes, allowing block-sparse attention to match dense attention in GPU throughput. On TabPFN-3 arXiv:2605.13986, a model where none of the assumptions hold, ClusterAttention is to our knowledge the first method to provide a substantial speedup over the default attention, while consistently keeping over 99\% of its accuracy. On the largest dataset from the TALENT benchmark suite, it makes processing of the training dataset close to 8x faster at nearly 11x attention speedup. ClusterAttention is also competitive with domain-specific methods, while avoiding any of the domain-specific engineering. On video-generation with Wan 2.1-T2V-14B arXiv:2503.20314 it produces output closer to dense attention at a larger speedup (1.8x vs 1.4x) than SVOO arXiv:2603.18636, a leading method in this domain, with both evaluated without offline calibration.
|
| 398 |
Hydra: A Navigation World Action Model with Discrete Latent Planning and Continuous Flow-Matching Execution
2608.28995
|
cs.CV
|
Mohammad Nazeri, Alexandyr Card, Samira Huber, Anuj Pokhrel, Yujun Wang |
World models let robots imagine possible futures, but exploiting this capability for real-time planning is bottlenecked by a representation misalignment: generative models and planners operate on decoupled manifolds, requiring computationally expensive decodin...World models let robots imagine possible futures, but exploiting this capability for real-time planning is bottlenecked by a representation misalignment: generative models and planners operate on decoupled manifolds, requiring computationally expensive decoding of every candidate back to the high-dimensional observation space for evaluation. In this paper, we present Hydra, a discrete World Action Model that tackles this by establishing a unified latent manifold over visual states, physical poses, and control actions. By compressing this manifold through modality-specific Vector-Quantized bottlenecks, Hydra yields discrete vocabularies of kinodynamic intents and visual states. This enables Discrete Latent Planning (DLP), where candidates are sampled directly from the shared manifold and ranked by a Kinematic-Perceptual Cost within the discrete latent space. To bridge discrete planning with the continuous commands required for physical actuation, Hydra pairs DLP with conditional Flow Matching to map selected intents to smooth execution trajectories. Evaluated on two physical robotic platforms, Hydra outperforms state-of-the-art navigation world models in goal-directed planning, while matching or exceeding the closed-loop execution capabilities of leading reactive navigation policies.
|
| 399 |
Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
2609.00355
|
cs.CVcs.AI
|
Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo |
Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can il...Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target's already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree in one pass and commits exactly its greedy output. In one production engine at a fixed round budget, GLANCE decodes up to 3.05 times faster than autoregressive decoding and outpaces the production EAGLE3-VL head on average and by about 11% on grounded tasks. An entropy law explains when drafting pays, predicting the longest accepted blocks on grounded tasks, where the target's next-token entropy is lowest. Our code is available at https://github.com/js-lee-AI/GLANCE.
|
| 400 |
MeshSplatBench: A Unified Benchmark for Triangle- and Mesh-Based Neural Rendering
2609.01306
|
cs.CV
|
Kaixuan Zhang, Minxian Li, Mingwu Ren, Xiatian Zhu |
Triangle- and mesh-based neural rendering aims to bridge neural scene representations and existing graphics engines (\textit{e.g.}, Unity and Blender) by leveraging triangle primitives compatible with standard rasterization hardware. However, existing methods ...Triangle- and mesh-based neural rendering aims to bridge neural scene representations and existing graphics engines (\textit{e.g.}, Unity and Blender) by leveraging triangle primitives compatible with standard rasterization hardware. However, existing methods are developed and evaluated under inconsistent settings, with limited comparison and little investigation into practical graphics engine deployment. This gap significantly hinders the understanding of their real-world usability. To address this issue, we introduce MeshSplatBench, the first benchmark for systematic evaluation of triangle- and mesh-based neural rendering from native rendering to graphics engine deployment. We propose a hierarchical deployment protocol with two options: (1) Standard deployment, using a conventional opaque mesh pipeline with vertex colors and hardware Z-buffering; and (2) Dedicated deployment, incorporating method-specific engine implementations to preserve appearance and compositing properties (e.g., alpha blending). For mesh splatting, we further introduce a structural audit to evaluate the topological and geometric integrity of exported surfaces for downstream graphics applications. Extensive evaluations reveal three key findings: (1) graphics engine deployment introduces noticeable quality degradation across methods, while mesh splatting approaches achieve relatively better robustness under standard deployment; (2) dedicated deployment can preserve most rendering fidelity at the cost of approximately 6-30$\times$ slowdown; and (3) explicit connectivity and shared vertex indexing in current mesh splatting methods remain insufficient to guarantee manifoldness or global connectivity. Our benchmark demonstrates that rasterizability alone does not imply graphics readiness and highlights the importance of evaluating practical engine compatibility. The benchmark and source code will be publicly released.
|
| 401 |
From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key-Value Extraction in Noisy Document Settings
2609.17538
|
cs.CV
|
Zahra Anvari |
Large language models (LLMs) have demonstrated strong capabilities in document key-value pair (KVP) extraction, yet controlled evaluations of their robustness to optical character recognition (OCR) output remain limited. This leaves an important gap in underst...Large language models (LLMs) have demonstrated strong capabilities in document key-value pair (KVP) extraction, yet controlled evaluations of their robustness to optical character recognition (OCR) output remain limited. This leaves an important gap in understanding their reliability in real-world OCR-to-LLM pipelines. Unlike end-to-end Vision-Language Models (VLMs), which jointly perform visual perception and semantic extraction, modular pipelines allow these stages and their errors to be isolated and audited. We introduce a controlled benchmark that distinguishes downstream LLM extraction behavior from upstream OCR degradation. It evaluates 136 experimental configurations and 17,688 document-level inferences generated with deterministic decoding across five instruction-tuned open-weight LLMs (2B-8B parameters), three datasets, and four text-quality conditions. The evaluation combines a full zero-shot comparison, targeted one- to three-shot experiments, and a sensitivity analysis of 40 configurations across 20 frozen demonstration sets. By separating Key Recall (annotated-field recovery) from Exact Match and Value F1 (exact and partial value recovery, respectively), we test whether OCR degradation affects field identification and value reproduction differently across models. Our findings challenge three practical assumptions: (1) clean-text performance reliably predicts real-world robustness, (2) model rankings remain consistent across annotation-derived Gold and OCR-derived text, and (3) additional few-shot demonstrations monotonically improve extraction accuracy. The observed model-ranking reversals and unstable few-shot gains expose important reliability risks under noisy document conditions. We release the benchmarking framework, dataset splits, and evaluation scripts to support reproducible research.
|
| 402 |
Are Coreset Selection Methods Worth Their Cost?
2609.22894
|
cs.CVcs.AI
|
Yangze Liu, Zhongyi Han |
Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behi...Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.
|
| 403 |
Video-to-Music Generation for Gameplay Videos
2609.31810
|
cs.CVcs.AIcs.SDcs.MM
|
Felipe Marra, Lucas N. Ferreira |
Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are r...Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are rendered graphics, music is mostly synthetic audio, and soundtracks loop across entire levels rather than following on-screen events. We introduce a new dataset of 217.6 hours of Super Nintendo (SNES) gameplay video paired with 485 hours of clean soundtracks, free of sound effects and voice-overs, matched to gameplay audio via audio fingerprinting. With this dataset, we train a simple encoder-decoder transformer that passes video features directly to a MusicGen decoder, comparing different encoding strategies: textual descriptions (T5), independent frames (ViT), or spatiotemporal patches (ViViT). Each encoder is tested both frozen and fine-tuned, while the decoder is always fine-tuned. Frozen encoders match or outperform their fine-tuned counterparts on every metric, and the frozen ViViT achieves the best overall results. We compare this model with state-of-the-art baselines using both objective metrics and a listening study (N = 96). Despite having up to 18% fewer parameters, our model outperforms all baselines on objective metrics, surpasses GVMGen in the listening study, and performs comparably to OSSL.
|
| 404 |
Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision
2609.34768
|
cs.CVcs.AI
|
Shuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin |
Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete actio...Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete action classification. We present a cross-modal teacher-student framework that lifts commercial radar to full-body, per-frame, metric 3D mesh reconstruction with per-joint uncertainty. Three innovations: (1) a mesh-foundation-model teacher - SAM 3D Body produces whole-body MHR ground truth (70 joints, 18,439 mesh vertices) from a single RGB frame with zero training, slashing annotation cost by orders of magnitude; (2) StudentPoseFormer - set encoding with masked attention pooling, a temporal Transformer, and a CVAE multi-hypothesis head that outputs both the pose mean and per-joint variance, honestly reporting where the radar cannot see; and (3) a multi-stage ground-truth quality pipeline (confidence gating, depth validation, temporal smoothing, bone-length consistency, bad-frame rejection) plus systematic information-lever ablations. On the public MM-Fi benchmark (same TI IWR6843 sensor, cross-subject), our full configuration reaches 7.45 cm 12-joint MPJPE, with ablations proving the causal value of point accumulation (k = 3, -0.34 cm), Doppler (-0.85 cm; -2 cm at the wrist on fast actions), and velocity loss (-0.27 cm). On our own synchronized radar + RGB-D corpus with block-level held-out splits, the pipeline achieves 21.47 cm end-to-end (per-joint hierarchy from 4.8 cm at the hip to 34.7 cm at the wrist - matching physical information limits), could be improved to 15 cm with ~30k diverse samples, and a scaling law shows sample diversity, not volume, is the binding constraint. Deployment inference is radar-only - no camera, no image.
|
| 405 |
Don't Throw Away the Tail: Action Upcycling for Policy Acceleration
2609.34911
|
cs.CVcs.LGcs.AI
|
Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin |
Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A ...Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose Action Upcycling, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find that discarded actions stay close to their replanned versions as long as the action velocity remains smooth. Action Upcycling therefore extends the execution horizon up to the point where the velocity begins to fluctuate. Extensive experiments on simulated and real-world manipulation tasks show that Action Upcycling reduces policy calls by 1.2-1.7x with no loss in success rate, across multiple Vision-Language-Action Models (VLAs) and even a World Action Model (WAM). It applies to any chunked policy at negligible cost and is orthogonal to other policy acceleration methods such as few-step sampling and streaming action decoding, opening a new axis for policy acceleration.
|
| 406 |
$\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
2609.35288
|
cs.CVcs.LG
|
Berker Demirel, Cl\'ementine Domin\'e, Valentino Maiorca, Marco Fumero, Marco Mondelli |
Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use t...Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and potentially limit downstream transfer. To address this, we introduce SACReg, a spectral anti-collapse regularizer motivated by an analysis of $\lambda$-balance, which captures the relative scale of weight matrices across layers. In a two-layer linear network, we show that (i) $\lambda$-balance prevents collapse, and (ii) our regularizer applied to the backbone induces $\lambda$-balance. In the nonlinear case, this regularizer leads to anti-collapse as well and, in realistic architectures on ImageNet100, it empirically increases the representations' ranks. We apply SACReg to JEPA and propose $\lambda$-JEPA, which improves over LeJEPA and VISReg on ImageNet-1k classification and in average linear-probe transfer performance across eight downstream image datasets. On video self-supervised learning, $\lambda$-JEPA improves over LeVJEPA and V-JEPA 2 on the Something-Something-v2 and Kinetics-400 benchmarks. Code is available at https://github.com/berkerdemirel/lambda-jepa.
|
| cs.LG 1314 papers | ||||
| 1027 |
Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase
2609.31630
|
cs.LG
|
Zhang Yanhai |
Replay-based continual learning almost always consolidates in a dedicated offline phase or by interleaving replayed samples with the input stream, whereas brains also consolidate during wakefulness through local sleep, brief use-dependent off-periods of indivi...Replay-based continual learning almost always consolidates in a dedicated offline phase or by interleaving replayed samples with the input stream, whereas brains also consolidate during wakefulness through local sleep, brief use-dependent off-periods of individual circuits. We ask whether a network trained by local, biologically constrained rules can consolidate with no offline phase at all. An isolation rule confines replay updates to hidden synapses invisible to the current input under k-winner-take-all dynamics, with optimiser state advanced only inside the mask; a refractory rotation rule makes units that have just fired sit out the next competition, widening the consolidable set; a homeostatic pressure and a relative-novelty gate decide when replay bursts fire and when rotation runs. This inverts the usual direction of non-interfering continual learning: the hidden computation on the current input is held invariant (exactly on the proven channels, and for all but 0.3% of waking samples per update elsewhere) while past memories are written into the degrees of freedom the current batch leaves unused. On class-incremental split-MNIST the system reaches 91.6+-0.3% with no offline phase, at or above the best offline-night schedule on two held-out splits, tied with DER++ and above experience replay, ER-ACE, A-GEM and unmasked local replay; in a single pass it leads DER++ (91.8% against 90.1%) while the night falls to 76.9%. The advantage is largest at small buffers and gives way to the backpropagation references at large ones; on split CIFAR-10 the system leads offline rehearsal and experience replay but trails ER-ACE and DER++. Rotation carries most of the gain; isolation adds the invariance guarantee. The mechanism is not tied to the local rule: under the same schedule a backpropagation network with k-WTA hidden layers gains from rotation, and isolation is again free on top of it.
|
| 1028 |
OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit
2609.31631
|
cs.LG
|
Dezhi Li, Lujun Li, Qiyuan Zhu, Hao Gu, Bei Liu |
Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencie...Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE{\dag}, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33$\times$ faster search and 1.55$\times$ inference speedup. Codes will be available after acceptance.
|
| 1029 |
EEGAgentBench: Benchmarking LLM Agents on Short- and Long-Horizon EEG Analysis
2609.31632
|
cs.LG
|
Huyu Wu, Weining Weng, Yuchen Liu, Yiqiang Chen, Yang Gu |
Electroencephalography (EEG) analysis is evolving from short-segment classification toward long-horizon interpretation that demands iterative evidence accumulation, multi-step reasoning, and coordinated use of specialized signal-processing tools. Although larg...Electroencephalography (EEG) analysis is evolving from short-segment classification toward long-horizon interpretation that demands iterative evidence accumulation, multi-step reasoning, and coordinated use of specialized signal-processing tools. Although large language models (LLMs) have recently shown promise as autonomous agents for EEG analysis, existing EEG agentic evaluations remain fragmented, covering limited tasks over narrow temporal horizons with inconsistent protocols, and providing no comprehensive assessment of agents' reasoning, tool-use, and workflow construction capabilities. To address this gap, we propose \textbf{EEGAgentBench}, a unified benchmark for systematically evaluating LLM agents on short- and long-horizon EEG analysis. EEGAgentBench spans six representative EEG applications ranging from knowledge question answering to sleep staging. It encompasses signal durations from 2 seconds to nearly 23 hours, with prediction targets ranging from class labels to event intervals and epoch-level sequences. This design supports unified evaluation across knowledge reasoning, short-horizon interpretation, long-horizon event detection, and sequential understanding. The benchmark further provides 10 deterministic EEG analysis tools that expose only task-relevant signal measurements. Agents must therefore select tools autonomously, accumulate evidence iteratively, and construct multi-step workflows. For evaluation, we benchmark 29 frontier LLMs from 15 model families. Results demonstrate that EEGAgentBench effectively distinguishes agent capabilities beyond model scale and inference cost, while revealing substantial limitations of current LLM agents in long-horizon EEG analysis, particularly in sustained evidence accumulation and multi-step reasoning.
|
| 1030 |
Enhancing generalization in endwall film cooling prediction: Incorporating the superposition principle into transformer-based neural operators
2609.31633
|
cs.LG
|
Qineng Wang, Liming Song, Tianyuan Liu, Zhendong Guo |
In this study, a physics-enhanced neural operator framework is proposed to enhance the generalization prediction ability of the cooling layout of a turbine endwall with variable number of film holes. Specifically, inspired by the film cooling superposition pri...In this study, a physics-enhanced neural operator framework is proposed to enhance the generalization prediction ability of the cooling layout of a turbine endwall with variable number of film holes. Specifically, inspired by the film cooling superposition principle, we propose a film cooling prediction model, namely superposition-based deep neural operator (SDNO), that divides the endwall temperature field prediction into two stages. In the first stage, the cooling layout of a turbine endwall is divided into several sub-parts with randomly assigned film holes, and a Transformer-based neural operator network, namely Calculate Net, is designed to predict the temperature field of each sub-part. Then, in the second stage, another neural operator network, i.e., Super Net, is trained to combine the temperature fields predicted by Calculate Net for each sub-part and obtain the superposed temperature field of the full cooling layout. Additionally, instead of directly taking the film cooling contours as pixel plots, a signed distance function (SDF) which is sensitive to the variable locations of cooling holes, is designed to encode the location information of cooling holes. Furthermore, the proposed endwall film cooling prediction model is trained with the samples that changing the number of film holes from 1-5 with variable locations. Then, the trained prediction shows excellent generalization prediction ability, which can accurately predict the film effectiveness of the cooling layout with 10-20 film cooling holes that are unseen in the training samples. The proposed SDNO also improves prediction accuracy relative to the fully supervised baseline. With the above, the effectiveness of our proposed prediction model has been well demonstrated.
|
| 1031 |
Symmetry-quotient Flatness and Generalization
2609.31634
|
cs.LG
|
Taiki Miyagawa |
This paper develops a theorem-level pipeline in symmetry-quotient settings: quotient linear stability implies quotient flatness, quotient flatness implies input smoothness, and input smoothness yields generalization under local covering assumptions. Flatness i...This paper develops a theorem-level pipeline in symmetry-quotient settings: quotient linear stability implies quotient flatness, quotient flatness implies input smoothness, and input smoothness yields generalization under local covering assumptions. Flatness is often associated with generalization, and Stochastic Gradient Descent (SGD) is frequently viewed as implicitly biased toward flat solutions. However, standard flatness measures are typically defined in the raw parameter space and are therefore not invariant under function-preserving symmetries such as positive rescaling. We develop a symmetry-aware theory of quotient flatness, quotient linear stability, input smoothness, and generalization on quotient spaces of neural-network parameters. For square loss and models equipped with function-preserving group actions, we define quotient flatness as the trace of the Hessian of the empirical loss on the regular quotient manifold. We show that quotient flatness controls input smoothness through a quotient-space analogue of the flatness-to-smoothness argument. We also prove that one-step mean-square quotient linear stability of the linearized SGD dynamics implies an explicit quotient-flatness bound in terms of the batch size and learning rate, and extend this analysis to higher-order tensor moments. Finally, under local covering and boundedness assumptions, we derive population generalization bounds in terms of quotient flatness and, consequently, in terms of quotient linear stability.
|
| 1032 |
What Next-Event Accuracy Cannot See: Closed-Loop Evaluation of Emergency Department Trajectory Simulators
2609.31635
|
cs.LG
|
Zhen Xuen Brandon Low |
Clinical trajectory models are usually evaluated by next-event accuracy on observed histories. Simulation is different: models must condition on their own generated events, allowing errors to compound. Although this problem is well known in sequence modelling,...Clinical trajectory models are usually evaluated by next-event accuracy on observed histories. Simulation is different: models must condition on their own generated events, allowing errors to compound. Although this problem is well known in sequence modelling, it has not been systematically quantified for clinical trajectory simulators. We developed EDSim-Bench to evaluate this failure mode using 425,028 MIMIC-IV-ED stays, with external replication on MC-MED, and release the evaluation protocol and scoring code. Starting from held-out visit prefixes, models generate the remainder of each visit and are evaluated on termination, event composition, timing, conditional fidelity, and occupancy forecasting, with a train-only order-3 n-gram as a reference baseline. Despite next-event accuracies within 0.001, three neural architectures behaved very differently under rollout. Across seeds, one Transformer recipe ranged from 0.43 to 0.96 in termination score and from 4.2- to 137-fold the divergence of the n-gram; no prefix-trained neural model approached the n-gram on termination or event composition. Inference-time interventions improved termination but did not jointly recover composition and timing. Supervising every eligible sequence position rather than only the final prefix position was associated with one to two orders of magnitude lower divergence across Transformer, GRU, and LSTM models, with the pattern persisting under model scaling, temporal shift, and external-site evaluation. Nevertheless, even the best model generated visits approximately half as long as observed, and model rankings reversed on occupancy forecasting, a downstream quantity relevant to bed management. These results show that next-event accuracy is insufficient to evaluate clinical trajectory simulators and motivate closed-loop evaluation across seeds, rollout criteria, and downstream tasks.
|
| 1033 |
Grounding Vision-Language Models in Driving Semantics: A Multi-Dataset Predicate Framework for Explainable Reasoning
2609.31636
|
cs.LG
|
Mohamed Chouai, Fazli Faruk Okumus, Stefan Kugele |
Vision-language models are increasingly used for driving-scene understanding, yet the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation. This paper introduces a deterministic multi-dataset pred...Vision-language models are increasingly used for driving-scene understanding, yet the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation. This paper introduces a deterministic multi-dataset predicate framework that derives driving-scene semantics from measurable geometric, kinematic, temporal, map, and traffic-control evidence. Dataset-specific interfaces are used only to recover the required scene information, while predicate definitions remain unchanged across nuPlan and nuScenes and are materialised in a common Predicate Knowledge Graph. Quantitative semantic validation against manually annotated predicate relations on 200 scenarios from each dataset yields macro F1 scores of 0.94 on nuPlan and 0.93 on nuScenes, with an average cross-dataset difference of 0.02 across the shared predicates. The Predicate KG is further evaluated using a frozen LLaVA-OneVision-7B model on the nine NuPlanQA subtasks. Predicate grounding achieves the highest accuracy among the evaluated visual-input conditions in seven of nine NuPlanQA subtasks, including Traffic Light (53.2% to 71.5%), Situation Assessment (76.2% to 86.1%), and Action Recommendation (82.9% to 89.0%). Weather/Lighting remains essentially unchanged (89.4% vs. 88.8%), consistent with the absence of corresponding predicates, while Predicate KG only input outperforms metadata-only input in eight of nine subtasks. The results show that deterministic predicates provide a consistent and traceable semantic representation and, under oracle grounding, can reduce visual dependence for reasoning tasks covered by the predicate vocabulary.
|
| 1034 |
FIDAL: Diversity-Aware Federated Active Learning Under Real-World Distribution Shifts
2609.31637
|
cs.LG
|
David Due\~nas Gaviria, Shadi Albarqouni |
Federated learning enables collaborative model training across institutions without centralizing data, yet high annotation costs, domain shifts, and class imbalance remain major obstacles, especially when irrelevant out-of-distribution (OOD) samples dilute the...Federated learning enables collaborative model training across institutions without centralizing data, yet high annotation costs, domain shifts, and class imbalance remain major obstacles, especially when irrelevant out-of-distribution (OOD) samples dilute the labeled data. Existing active learning methods target uncertainty or diversity within in-distribution (ID) data and overlook unknown samples in federated clinical settings. We propose FIDAL, an open-set federated active learning framework that combines calibrated global-local evidential uncertainty, support-set diversity weighting, and adaptive OOD rejection. The rejection gate thresholds a foundation-model Gaussian-coverage signal per client and per round with Otsu's criterion, so that highly informative ID samples are queried while irrelevant outliers are excluded without any hand-tuned threshold. Evaluated on three multi-center medical imaging benchmarks (dermatology, histopathology, and mammography with organically occurring artifacts) in realistic open-set scenarios, FIDAL outperforms detector-based open-set methods by up to about 12 percentage points of balanced accuracy and is the only method on the accuracy-ID purity Pareto front of all three benchmarks. At an equal query budget it spends at least 1.3 times fewer annotations on OOD samples than every accuracy-matched baseline, saving an estimated 7-29 hours of expert reading on the mammography benchmark. By labeling only a fraction of the data pool, it matches or exceeds fully supervised performance across modalities. These results highlight the value of integrating uncertainty, diversity, and OOD rejection in open-set federated active learning for medicine.
|
| 1035 |
Energy-aware frugal Bayesian optimization
2609.31638
|
cs.LG
|
Gaston Plat, Paul Saves, Nathalie Bartoli, Thierry Lefebvre, Joseph Morlier |
Modern design optimization frameworks aim first and foremost for models with the most accurate predictions without balancing computational overhead. It remains a reason why scaled architecture and multidisciplinary design optimization problems are difficult to...Modern design optimization frameworks aim first and foremost for models with the most accurate predictions without balancing computational overhead. It remains a reason why scaled architecture and multidisciplinary design optimization problems are difficult to address, even with sample-efficient Bayesian optimizers. In this paper, a metric quantifying the computational energy footprint is introduced within a Bayesian optimization framework to guide the parameter setting of a model towards configurations that balance both performance and frugality. The computer experiments highlighted existing tradeoffs between optimum convergence and the underlying energy footprint, and sometimes resulted in both a better-found optimum and lower energy consumption.
|
| 1036 |
When Does Domain Adaptation Help on Physical Vibration Sensors? A Held-Out-Bearing Study of Neural-Operator and Convolutional Models
2609.31639
|
cs.LG
|
Kumbha Nagaswetha, Rabi Pathak |
Diagnosing rolling-element bearing faults from vibration is a canonical physical-sensing task and a widely used benchmark for domain adaptation under operating-condition shift. Accuracies above 99 percent are commonly reported, but under evaluation splits that...Diagnosing rolling-element bearing faults from vibration is a canonical physical-sensing task and a widely used benchmark for domain adaptation under operating-condition shift. Accuracies above 99 percent are commonly reported, but under evaluation splits that place the same physical bearing in both training and test. We revisit the task under a held-out-bearing protocol, assigning every bearing unit entirely to either the training or the test set, and find that source-only transfer is far weaker than such numbers suggest: on a change of shaft speed it reaches only $0.36$, against a target-supervised ceiling of 0.97. We then study what governs transfer. Treating computed order tracking, a shaft-angle resampling that places fault frequencies at fixed shaft orders independent of running speed, as a controlled change of representation, we find that a Fourier Neural Operator raises source-only transfer from $0.36$ to $0.61$ on the speed shift, where the fault peaks move, while a convolutional network of matched feature dimension stays near chance in both representations. The representation also decides whether unsupervised alignment can work: with the same normalized RBF-MMD loss and no target labels, the operator reaches 0.71 in the frequency domain but 0.95 in the order domain, within 0.02 of the target-supervised ceiling and above $0.86$ on every held-out bearing fold. Once the representation is right, a small label budget adds little. These results indicate that, for this task, the input representation rather than the alignment method decides whether adaptation helps. A second dataset, whose held-out units are fault diameters rather than bearings, shows that the same protocol exposes failures that even a target-supervised model cannot avoid.
|
| 1037 |
Measure Learning at Steady State: A BIRD-SQL Formula 1 Case Study
2609.31640
|
cs.LG
|
Manoj Bajaj |
Continual Learning Bench scores learning as short-horizon gain versus a reset baseline and finds naive full-context ICL strongest among the memories it tested. We treat ICL as one learning system and score it on a longer shared-world schedule. Steady-state lea...Continual Learning Bench scores learning as short-horizon gain versus a reset baseline and finds naive full-context ICL strongest among the memories it tested. We treat ICL as one learning system and score it on a longer shared-world schedule. Steady-state learning is the gap versus baseline on a pre-set late window (last 40 of 174 BIRD-SQL formula-1 questions). We split the score into exploration efficiency (SQL probes), task reward (hits), and delivery cost (API dollars and context size). On gpt-5.6-luna, late probes fall from 4.6-5.6 to 0.95 while hits rise only modestly and ICL context grows to about 95k tokens with cost roughly doubling. Short-horizon gain understates the late probe saving and misses the cost inversion, so we find that unbounded ICL is a poor candidate for the learning mechanism.
|
| 1038 |
Product-Aware Deterministic Rounding for Quantized Matrix Multiplication
2609.31641
|
cs.LG
|
Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, Keita Teranishi, Sudip Seal |
Scalar rounding decisions interact through matrix multiplication. We study deterministic product-aware rounding after scales, clipping bounds, and grids are fixed, with each active scalar choosing between adjacent levels. For dynamic activation rounding, null-...Scalar rounding decisions interact through matrix multiplication. We study deterministic product-aware rounding after scales, clipping bounds, and grids are fixed, with each active scalar choosing between adjacent levels. For dynamic activation rounding, null-space reduction preserves the relaxed product while leaving at most $r$ fractional decisions, where $r$ is the rank of the active gap-weighted weight block. Conditional-expectation completion gives a deterministic polynomial-time algorithm with squared product error at most $ \mathrm{OPT}_{\mathrm{dyn}}+r\nu_{\max}^2/4$, where $\mathrm{OPT}_{\mathrm{dyn}}$ is the best admissible error and $\nu_{\max}$ is the largest row norm of that block. For reusable static weights, the exact expected product-loss metric is the uncentered input second moment with fixed output bias; free bias recalibration yields the centered covariance. Exact optimization is NP-hard even at rank one. In balanced blocks with $K=1024$ and $r=16$, conditional- expectation completion attains a dither-normalized median error of $0.010$, compared with $0.899$ for round-to-nearest. Clipping-aware initialization reduces median normalized error by a factor of $43.4$ at ten-percent clipping. Held-out Digits experiments show that retaining the input mean or correcting the output bias improves median product error over round-to-nearest in all four tested bit-width and calibration-size settings.
|
| 1039 |
Information Design Against Gaming and Learning Adversaries
2609.31643
|
cs.LG
|
Madhava Gaikwad |
A principal who deploys a binary classifier with an abstention option must decide which queries the mechanism abstains on. The right choice depends on the adversary. A gaming adversary already knows the classifier and tries to manipulate features across the bo...A principal who deploys a binary classifier with an abstention option must decide which queries the mechanism abstains on. The right choice depends on the adversary. A gaming adversary already knows the classifier and tries to manipulate features across the boundary, so the principal does best by abstaining on queries close to that boundary. The same boundary-localizing rule is the worst possible choice against a learning adversary who does not know the classifier: each abstention now tells the adversary that the boundary is nearby, which is enough to drive a binary search. We analyze this tension. The two natural defenses, abstaining at a fixed rate and abstaining near the boundary, are Blackwell-incomparable: neither can be simulated by post-processing the other's responses. The number of queries needed to reconstruct the boundary to error $\eps$ is $\tilde\Theta(d/\eps)$ under the first defense and $\Theta(d \log(1/\eps))$ under the second, where $d$ is the VC dimension of the classifier family and $\tilde\Theta$ suppresses factors polylogarithmic in $d$ and $1/\eps$. The first rate is a worst case over query distributions; no reconstruction algorithm can close the gap at the distributions that attain it. We characterize the Pareto frontier between the two defense objectives, and confirm both rates on seven binary-classification tasks spanning tabular, image, and language-model-feature inputs: label-plus-counterfactual access extracts the boundary with up to $200\times$ fewer queries than a published label-only baseline.
|
| 1040 |
MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning
2609.31644
|
cs.LG
|
Sourabh Kulkarni, Ksheeraj Sai Vepuri, Basar Demir, Jason Bohrer, Emily Shen |
Reinforcement learning (RL) is widely used in language-model post-training to maximize rewards assigned to individual model outputs, such as scores from binary verifiers or reward models trained on human feedback. However, applications such as synthetic-data g...Reinforcement learning (RL) is widely used in language-model post-training to maximize rewards assigned to individual model outputs, such as scores from binary verifiers or reward models trained on human feedback. However, applications such as synthetic-data generation, fairness-related constraint satisfaction, and policy exploration require controlling the distribution of outputs across model generations rather than only maximizing expected reward. We propose a general RL-based framework for \textit{Distribution Matching} allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution. Empirically, we demonstrate that dominant post-training recipes such as Group Relative Policy Optimization (GRPO) reduce output diversity by concentrating policy probability towards a single mode. Entropy regularization and sampling temperature can improve the spread of the distribution but have constrained effectiveness, limited to apply only in token space and toward uniform distributions. We show that prior work in this area is a specific case of Distribution Matching involving the $L_2$ divergence. We then propose reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justification. Finally, we demonstrate the effectiveness of our approach on a set of experiments involving mathematical reasoning and programming.
|
| 1041 |
STAR: Adaptive Spatial-Temporal Normalization for Unified Microservice Incident Management
2609.31645
|
cs.LG
|
Xinhua Miao, Linyu Zhu, Bowei Yang, Zhengong Cai |
Automated incident management in large-scale microservice systems relies on learning robust representations from multimodal observability data, including metrics, logs, and traces. Although recent self-supervised frameworks enable unified modeling for anomaly ...Automated incident management in large-scale microservice systems relies on learning robust representations from multimodal observability data, including metrics, logs, and traces. Although recent self-supervised frameworks enable unified modeling for anomaly detection (AD), failure triage (FT), and root cause localization (RCL), they often struggle with non-stationary temporal dynamics and heterogeneous service dependency structures. In this paper, we propose STAR, a Spatial-Temporal Adaptive Representation learning framework that explicitly addresses these challenges through adaptive normalizations. STAR introduces two tightly coupled mechanisms: Temporal Adaptive Normalization (TAN), which dynamically normalizes multivariate time series using multi-scale temporal context, and Spatial Adaptive Normalization (SAN), which performs structure-aware normalization over service dependency graphs. Unlike prior methods that treat normalization as static or task-agnostic, STAR formulates it as a learnable, context-conditioned transformation aligned with the intrinsic properties of microservice systems. The resulting adaptive representations are integrated into a unified self-supervised framework, enabling end-to-end unsupervised support for AD, FT, and RCL tasks. Extensive experiments on two real-world microservice benchmarks demonstrate that STAR consistently outperforms all state-of-the-art baselines, yielding significant and stable improvements across all three tasks. Our results highlight adaptive normalization as a principled and effective mechanism for robust multimodal representation learning in complex software systems.
|
| 1042 |
Typed Temporal Interaction Features for Simulation-Backed Forecasting of Open-Source Game Release Incidents
2609.31647
|
cs.LG
|
Shayma Alkobaisi, Anas Ali |
Open-source video-game quality depends on inter-actions among code, assets, configuration, tests, contributors, and issue workflows, yet conventional defect predictors usually flatten or omit these relations. We investigate release-level forecasting of a quali...Open-source video-game quality depends on inter-actions among code, assets, configuration, tests, contributors, and issue workflows, yet conventional defect predictors usually flatten or omit these relations. We investigate release-level forecasting of a quality incident within thirty days using GAMEQUALGRAPH-Pilot, a typed temporal feature pipeline with calibrated risk estimates and effort-aware ranking. Because the accessible OS-SGameBench materials do not provide manually audited release dates and outbreak labels, the executed evaluation is explicitly simulation-backed rather than an empirical claim about real games. Five seeded worlds each contain 120 projects and 24 releases, with project-disjoint validation and future cross-project testing. The pilot obtains an AUPRC of 0.520, AUROC of 0.673, Brier score of 0.207, and 29.68% effort-aware recall at a twenty-percent testing budget. Its closest local comparator, Static-Hetero-Reimpl, reaches 0.522 AUPRC; the -0.002 difference is not statistically significant after Holm correction. Inference requires 0.023 milliseconds per release in the measured environment. Ablations and controlled missingness, drift, engine, project-size, alert-threshold, and attribution analyses expose where typed interactions help and where they fail. Results support the reproducibility of the proposed protocol, not deployment effectiveness. Real OSSGameBench release reconstruction, stratified label audits, and official graph-model comparisons remain mandatory before journal submission or operational use in practice. This boundary protects research integrity and supports credible evaluation.
|
| 1043 |
Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management
2609.31648
|
cs.LG
|
Lyes Saad Saoud, Oualid Doukhi, Ehsan Reihani, Saeed Sepasi, Deok Jin Lee |
Vision-Language-Action (VLA) models are studied mainly in robotics, where visual observations and language instructions are mapped to physical actions. This paper introduces Energy Vision-Language-Action (EVLA), a controlled multimodal benchmark for intent-con...Vision-Language-Action (VLA) models are studied mainly in robotics, where visual observations and language instructions are mapped to physical actions. This paper introduces Energy Vision-Language-Action (EVLA), a controlled multimodal benchmark for intent-conditioned residential energy management. EVLA frames battery scheduling as a multimodal trajectory-prediction problem in which an RGB energy-field representation, a numerical operating state, and a natural-language objective are mapped to a 16-step battery-action trajectory generated by a finite-horizon sampling-based reference generator. Source windows are derived from public residential electrical-load data, while electricity price, battery state of charge, indoor temperature, and time of day are generated benchmark metadata. A hidden operating regime is encoded only through energy-field texture, enabling paired visual changes while the explicit numerical state is fixed. Crossing 439,203 retained base windows with three hidden regimes and five language objectives yields 6,588,045 multimodal instances. An initial study evaluates 36 configurations over three training seeds using fixed subsets of 5,000 training, 500 validation, and 500 test instances. In the MobileNet-family comparison, removing processed language increases trajectory mean-squared error from 0.3856 +/- 0.0039 to 0.8628 +/- 0.0001, whereas removing vision yields 0.3843 +/- 0.0013, comparable to the full model. The results show strong asymmetry in modality use: the processed-language pathway is strongly associated with prediction quality, while the current RGB pathway provides no aggregate error advantage. These results characterize the fixed pilot subset and executed protocol rather than full-benchmark training. EVLA provides a controlled setting for studying how semantic intent and latent context influence residential energy-action prediction.
|
| 1044 |
From Phase Transition to Systemic Failure: A Decoupled Analytics Framework for GNN Robustness
2609.31656
|
cs.LG
|
Shuai Yan, Dan Peng, Jie Li, Ke Wang |
Data quality is a major bottleneck for the reliable deployment of graph neural networks (GNNs) in real-world graph mining tasks. Among various sources of degradation, label noise and feature distribution shift (hereafter referred to as distribution shift) are ...Data quality is a major bottleneck for the reliable deployment of graph neural networks (GNNs) in real-world graph mining tasks. Among various sources of degradation, label noise and feature distribution shift (hereafter referred to as distribution shift) are two common yet fundamentally different challenges. To study their effects under controlled conditions, this paper constructs a synthetic homophilic graph regression benchmark in which the two factors can be manipulated separately. A total of 41 configurations and 410 runs are conducted to evaluate the behavior of representative GNN models under varying noise and shift conditions. The results show two distinct patterns. First, under additive label corruption, performance remains relatively stable over a broad range of noise settings and begins to deteriorate sharply only after an observed transition region around the 50 percent noise ratio. Second, under extreme feature distribution shift, all tested models suffer substantial degradation, with test MSE increasing by 48 times to 316 times and correlation dropping by 73 percent to 89 percent. These findings suggest that, in the present controlled setting, GNNs are considerably more tolerant to moderate label perturbation than to severe distribution mismatch. The study provides a controlled empirical baseline for understanding how data quality affects GNN-based graph mining systems and offers practical implications for deployment-oriented monitoring and model maintenance.
|
| 1045 |
Cross-Material Support Transfer for Core-Loss Prediction Under Waveform Covariate Shift
2609.31659
|
cs.LG
|
Cong Yao, Chunye Gong |
Power magnetic materials are characterized on the sinusoidal and triangular waveforms that excitation hardware conveniently produces, whereas deployed converters expose cores to trapezoidal, PWM-shaped flux trajectories, so loss models must predict exactly whe...Power magnetic materials are characterized on the sinusoidal and triangular waveforms that excitation hardware conveniently produces, whereas deployed converters expose cores to trapezoidal, PWM-shaped flux trajectories, so loss models must predict exactly where their training data are thinnest. The final test of the MagNet Challenge embeds a deliberately extreme instance of this characterization-deployment mismatch: for material D, trapezoids form 16.4% of the test set but only 1.4% of the training set. The 95th-percentile relative error, hereafter p95, of the best submission, built on sequential transfer learning, stalled at 15.9%, the worst among the five materials. This paper shows that the obstacle is missing information under covariate shift rather than class imbalance, and that the missing support can be borrowed from sibling materials instead of being extrapolated. Controlled experiments first refute the imbalance reading: four standard remedies fail, and raising the trapezoidal share to the test-set level degrades accuracy further. The proposed material-identity support transfer, MIST, then trains one 2784-parameter predictor jointly on all five challenge materials. Material identity enters through feature-wise linear modulation, or FiLM, the scarce material's true-label loss is reweighted, and material D receives no fine-tuning, so that the bias of its trapezoid-free training set is never re-installed. MIST lowers the five-seed material-D p95 from 20.39+/-2.03% to 12.38+/-0.92% and the trapezoidal-class p95 from 37.4+/-8.8% to 15.16+/-1.69%, surpassing the best submission with one-sixth of its parameters and no fine-tuning stage; removing material identity at matched capacity inflates the error by an order of magnitude. These results argue that scarce materials should be characterized jointly with their siblings.
|
| 1046 |
Beyond the Graph: An Adaptive Meta-Learner Fuses Explainability, Weather, and Dynamics for Robust Bus ETA Prediction
2609.31667
|
cs.LG
|
Pratham Payra, Jagadish |
Accurate bus Estimated Time of Arrival (ETA) prediction is vital for urban mobility, passenger satisfaction, and transit efficiency, yet existing models falter against nonlinear spatiotemporal dynamics, data sparsity, and factors such as weather. This paper pr...Accurate bus Estimated Time of Arrival (ETA) prediction is vital for urban mobility, passenger satisfaction, and transit efficiency, yet existing models falter against nonlinear spatiotemporal dynamics, data sparsity, and factors such as weather. This paper proposes HYB(nm), an adaptive hybrid ensemble framework that dynamically fuses five complementary models - a historical baseline (MST-AV), periodical temporal pattern analysis (GDRN-DFT), Koopman Neural Operators for nonlinear dynamics (KOOP-NET), weather-integrated feature-engineered neural networks (FENN), and real-time graph convolutional networks (MGCN) - via a meta-learner attuned to real-time context. Evaluated on GPS and weather data from three Kolkata bus routes comprising more than 4,000 trips, the framework leverages the individual strengths of its components (for example, the low-latency explainability of MST-AV, the weather resilience of FENN, and the network-dynamics capture of MGCN) to deliver the superior robustness of HYB(2), state-of-the-art accuracy rivalling leading graph neural networks, and balanced trade-offs in stability and efficiency across prediction horizons and operating conditions. The extensible HYB(k) architecture equips transit agencies with flexible tools, ranging from economical single models to tailored high-fidelity hybrids, advancing predictive, equitable urban transport.
|
| 1047 |
NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization
2609.31669
|
cs.LG
|
Gautam Kishore |
We present NanoForecast v0.5, a 6.5M-parameter forecaster that competes with models 31x its size (TimesFM, 200M parameters) after training pipeline fixes and no architecture change. Retraining the v0.3 architecture with corrected loss-scope handling, tensor sh...We present NanoForecast v0.5, a 6.5M-parameter forecaster that competes with models 31x its size (TimesFM, 200M parameters) after training pipeline fixes and no architecture change. Retraining the v0.3 architecture with corrected loss-scope handling, tensor shape alignment, and wider augmentation coverage cuts overall Mean Absolute Scaled Error by 43.8% under one fixed protocol (MASE 3.030 to 1.704) on the same data and compute budget. NanoForecast v0.5 beats TimesFM on all three ETT datasets (MASE 0.676/1.110/0.287 vs. 0.705/1.360/0.545) and on exchange rate (4.317 vs. 4.383); TimesFM keeps a clear lead on the high-cardinality electricity and traffic sets. Against PatchTST (15M+ parameters, official configuration), v0.5 wins all three ETT sets. Training takes about 12 hours on a single cloud GPU (NVIDIA T4, Google Colab) and inference needs no GPU (measurements in this paper are on an Apple M4 CPU). We release all code, pretrained checkpoints, and evaluation framework under Apache 2.0 at https://github.com/eulogik/NanoForecast
|
| 1048 |
Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions
2609.31675
|
cs.LG
|
Sagar Deb, Devam Shah, Ashwanth Krishnan |
We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents recover causal graph structure from observations and budget-constrained hard interventions. ACDB pairs a linear-Gaussian world generator wi...We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents recover causal graph structure from observations and budget-constrained hard interventions. ACDB pairs a linear-Gaussian world generator with a fixed observe-intervene-submit API and a three-layer scoring contract that separates skeleton recovery, DAG recovery, and intervention efficiency. On the current six-level ladder, PC with a greedy active orientation heuristic is the strongest non-oracle method (directed F1 42.7%, SHD 4.79), ahead of Claude Sonnet 4.6 raw active (31.7%, 7.25) and GPT-5.4 raw active (22.9%, 9.27). The most informative diagnostic is the precision-recall decomposition: PC under-commits with high precision, LLMs over-commit with lower precision, and statistical-tool access often increases abstention rather than useful intervention. A structure-blind random DAG baseline reaches 23.6% directed F1 on this dense v0 ladder; a density probe lowers this floor to 16.9%, motivating the v1 calibration pass. The current results should therefore be read as a benchmark audit and calibration report, not as evidence that current LLMs solve active causal discovery.
|
| 1049 |
Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?
2609.31680
|
cs.LG
|
Yutong Feng, Bowen Liao, See Kiong Ng, Yuxuan Liang |
Joint-embedding predictive architectures (JEPA) have emerged as a promising self-supervised pretraining paradigm for time series, learning representations by predicting target embeddings in latent space rather than reconstructing raw signals. Yet evidence on t...Joint-embedding predictive architectures (JEPA) have emerged as a promising self-supervised pretraining paradigm for time series, learning representations by predicting target embeddings in latent space rather than reconstructing raw signals. Yet evidence on their benefits remains mixed, and most studies test only a single backbone or a narrow set of architectures, leaving unclear whether JEPA pretraining is a reliable improvement or one that depends heavily on the downstream model. We address this gap through a large scale evaluation of one JEPA instantiation across nine backbones and eleven benchmarks spanning temporal and spatio-temporal forecasting, the most extensive cross architecture assessment of JEPA for time series to date. We find that the benefit of this instantiation varies sharply across backbones, producing consistent gains for some architectures and consistent degradation for others, even on the same dataset. This pattern holds across both task families, indicating the variability is a general property of this instantiation rather than a dataset specific artifact worth accounting for when choosing a backbone in practice.
|
| 1050 |
Autonomous Research Project Management as an Agent Skill: A Case Study in Exact Spectral Spatial Regression
2609.31683
|
cs.LG
|
Alexander Chen (University of New South Wales), Jeffrey Meng (University of New South Wales), Bram Hoex (University of New South Wales, GreenDynamics), Tong Xie (University of New South Wales |
This work presents an end-to-end demonstration of autonomous machine learning research conducted by an agent skill on consumer hardware. The demonstration evaluates an FFT-based Kernel Ridge Regression (KRR) solver for regular spatial grids using 2005 monthly ...This work presents an end-to-end demonstration of autonomous machine learning research conducted by an agent skill on consumer hardware. The demonstration evaluates an FFT-based Kernel Ridge Regression (KRR) solver for regular spatial grids using 2005 monthly NOAA Kaplan SST v2 anomaly fields on a $36 \times 72$ grid. This was autonomously executed by DeepSeek V4 Flash, orchestrated by our agent skill suite within DeepSeek Harness (DSH). Experiments were executed on CPU-only hardware (Apple M2 Pro; 78.7 s solver time, 1.57 GB peak RSS). Long-horizon state was decoupled into a file-based epic- and issue-tracking substrate. Across 74 sub-agent sessions, the agent demonstrated closed-loop scientific resilience: routing two failed hypothesis review gates back to literature retrieval, patching bootstrap indexing bugs, and executing with only four discrete human steering events. Finally, we reflect on autonomous research governance, arguing that scientific credibility requires inspectable state, falsifiable review gates, and transparent reporting of negative results, urging the machine learning community to favour agent-accessible structured formats over static PDF manuscripts.
|
| 1051 |
What does FFN compression change downstream? Same-state causal restoration in diffusion language models
2609.31685
|
cs.LG
|
Shaurya Omar |
Diffusion language models (DLMs) enable flexible, parallel generation, but their iterative denoising remains computationally expensive, motivating increasingly aggressive compression. Existing compression objectives largely measure how well compressed computat...Diffusion language models (DLMs) enable flexible, parallel generation, but their iterative denoising remains computationally expensive, motivating increasingly aggressive compression. Existing compression objectives largely measure how well compressed computation approximates the original locally, but local error does not reveal which removed computations actually matter to the downstream denoising trajectory. We introduce Same-State Causal Restoration (SSR), which restores the original FFN on the exact current input reached by the compressed model and measures how the resulting trajectory changes. To our knowledge, this is the first direct measurement of the same-current-input closed-loop effect of removed FFN computation in DLM compression. Across LLaDA-8B-Instruct and Dream-v0-Instruct-7B, compressed-side state ranks this downstream effect substantially better than local NMSE at fixed denoising phase, while controlled interventions show that correction structure matters beyond magnitude. Using task-label-free calibration, SSR freezes a single restoration window for held-out inference. Under aggressive LLaDA compression, restoring only four transitions recovers 89.9% of the lost accuracy while retaining an estimated 36.8% whole-model MAC saving and outperforming an equal-budget local-error baseline. Dream further shows that restoring dense behavior and repairing the final task are distinct outcomes.
|
| 1052 |
3-D Emissions Mapping and Social Cost Estimation for US Domestic Aviation at West Coast Hubs
2609.31686
|
cs.LG
|
Hesam Shafiei Nia, Don MacKenzie |
Existing aviation emissions inventories lack accurate trajectory data for high-resolution social cost and health impact assessment. This paper develops a 3-D emissions map by reconstructing flight trajectories for US west coast hubs to estimate regional enviro...Existing aviation emissions inventories lack accurate trajectory data for high-resolution social cost and health impact assessment. This paper develops a 3-D emissions map by reconstructing flight trajectories for US west coast hubs to estimate regional environmental and near-airport health impacts. A physics-informed autoencoder (AE) is applied to ADS-B trajectory records for January 2025 covering US west coast hubs. The encoder combines a Convolutional Neural Network (CNN), a Bi-GRU, and a 3-D CNN with skip connection; the decoder is a Temporal Convolutional Network (TCN). It is benchmarked against a baseline-AE and cubic spline interpolation. Emissions are mapped via EUROCONTROL Base of Aircraft Data (BADA) performance tables and ICAO Engine Emissions Databank (EEDB) emission indices, with altitude corrections via Boeing Fuel Flow Method 2 (BFFM2). Social costs are quantified for all flight phases, with health impacts assessed for Landing and Takeoff cycles within 50 km of each hub. The proposed AE model outperforms both a TCN-AE and cubic spline interpolation across 5% to 50% missing rates. Monetizing the emissions inventory shows NOx produces a small net cooling effect in direct climate forcing, while accounting for 99.7% of monetized air-quality and health cost despite being under 0.4% of CO2 by mass, making it the dominant health-cost driver. To our knowledge, this is among the first studies combining AE-based trajectory reconstruction with separate spatial-temporal feature encoding and altitude-based emissions modeling to produce a regional aviation emissions inventory for air quality, climate impact and population exposure. The resulting emissions map and social cost estimates provide quantitative context for environmental impact assessment and near-airport health policy evaluation for US domestic aviation.
|
| 1053 |
Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerical Non-Determinism?
2609.31796
|
cs.LG
|
Alizishaan Khatri |
Activation probes are increasingly used to monitor LLMs in deployment. A probe is typically trained under one inference configuration, then used under whatever batch size and numerical precision the serving stack uses. Because common GPU kernels are not batch-...Activation probes are increasingly used to monitor LLMs in deployment. A probe is typically trained under one inference configuration, then used under whatever batch size and numerical precision the serving stack uses. Because common GPU kernels are not batch-invariant and floating-point formats round differently, the activations seen at deployment are not the ones the probe was trained on. We measure what that costs for Llama-3.1-8B, Qwen3-8B and Gemma-3-4B across batch sizes 4, 8 and 16 and float32, bfloat16 and float16, training 768 probes on one configuration, evaluating each on every other, and comparing verdicts example by example. Probes are stable, but aggregate accuracy is the wrong instrument for showing it: it understates how many verdicts change by a factor of two to nine. At the prompt, accuracy never moves by more than 0.47 percentage points across 1,392 transfers and only 0.076% of verdicts change; under float32 with only the batch size varied, none of 201,960 verdicts change. During decoding the flip rate rises to 2.8%, but rows whose realised tokens matched flip in only 0.12-0.15% of cases, while rows whose tokens diverged flip in 12.9%: the cause is the text, not the arithmetic. A bfloat16 batch-size change flips the first generated token for 2.1% of rows and leaves 25% on different tokens by token 20. Flips are symmetric, Cohen's kappa stays above 0.94, and AUROC moves by at most 0.05 points. Underneath, activations move about as much as the format's rounding: a bfloat16 batch-size change perturbs them by a median relative L2 of 1e-2, roughly 8x the float16 figure. Probes absorb this; the model's own next-token argmax does not. Robustness evaluations of activation monitors should report per-example agreement rather than aggregate accuracy, separate representational noise from input change, and state the serving configuration.
|
| 1054 |
seq2cause: One Autoregressive Backbone, Four Causal Discovery Tasks in Event Sequences
2609.31801
|
cs.LG
|
Hugo Math |
Complex systems such as vehicles, patients, or genomes emit discrete event sequences whose operative question is causal, not predictive: which events cause which other events, and which cause higher-level outcomes such as failures or diseases? This question de...Complex systems such as vehicles, patients, or genomes emit discrete event sequences whose operative question is causal, not predictive: which events cause which other events, and which cause higher-level outcomes such as failures or diseases? This question decomposes along two axes -- dependency type (event $\to$ event vs.\ event $\to$ outcome) and causal scope (single sequence vs.\ population) -- yielding four structurally distinct regimes with different identifiability conditions. No existing method addresses more than one, because all assume multi-stream structure with low vocabulary, and none scales beyond a few hundred event types. We present \textsc{Seq2Cause}, a unified framework that resolves all four regimes through a single shared primitive: a pretrained autoregressive model repurposed as an amortized conditional independence testing engine requiring no task-specific retraining. We establish a prediction--causality duality: the model's excess cross-entropy simultaneously bounds causal identification error across all four regimes, so that every improvement in next-token prediction tightens causal guarantees for free. On nonlinear SCMs (vocabularies up to $8{,}000$ types) and real-world vehicle diagnostic logs ($29$K event types, $474$ failure outcomes), \textsc{Seq2Cause} is the first method to populate all four regimes at scale with a single frozen backbone. Existing methods are either inapplicable, inaccurate, or computationally intractable in this setting.
|
| 1055 |
Medium-Term Multi-Resolution Electric Load Forecasting using Economic Data and Foundation Model
2609.31806
|
cs.LG
|
Lindas Eloi, Goude Yannig, Ciais Philippe |
Accurate medium-term, from a few months to a few years, electricity load forecasts are crucial for informed decision-making in power plant maintenance scheduling, load dispatch and price settlement. Being comprised between Long-Term Load Forecasting (LTLF) whi...Accurate medium-term, from a few months to a few years, electricity load forecasts are crucial for informed decision-making in power plant maintenance scheduling, load dispatch and price settlement. Being comprised between Long-Term Load Forecasting (LTLF) which uses mostly economic projections and appliances development scenarios, and Short-Term Load Forecasting (STLF) driven by weather, calendar and autoregressive patterns, Medium-Term Load Forecasting (MTLF) requires both extrapolation capabilities and variability modeling. Yet, it remains unclear if MTLF can benefit from economic indicators, and especially at which forecast horizon and resolution. To address these challenges we investigated the impact of socioeconomic data on predictions issued 1 month and up to 48 months in advance for France at monthly and daily resolution using a tabular Foundation Model (FM). A dataset covering 20 years of observations of electricity load, weather variables and economic features such as consumer price and production indices, electric vehicle counts or employment is created for the study. To avoid noisy data, we used a new feature selection pipeline, creating ensemble of expert models with diverse feature subsets, to demonstrate that selected economic covariates improve forecast skill by 20% over 2015-2025. This enhancement is steady across lead times and resolutions limiting the Mean Absolute Percentage Error to 4% for monthly granularity and 5% for daily granularity. Explainability of the models is investigated through feature and context importance. Results showed that the FM is limited in the context it leverages pointing towards potential computational savings with a reduced context, while feature importance of economic predictors grows with the forecast horizon. This suggests that including economic data in MTLF could bridge the gap with LTLF leading to seamless forecasts.
|
| 1056 |
Relational Compression: A Framework for Relational Fidelity in Constrained Representations
2609.31816
|
cs.LG
|
Yaniv Shulman |
What should a compressed representation preserve when the information of interest lies in relationships among elements rather than in the elements themselves? We formulate relational compression in the classical source-description-reconstruction sense, but wit...What should a compressed representation preserve when the information of interest lies in relationships among elements rather than in the elements themselves? We formulate relational compression in the classical source-description-reconstruction sense, but with relational structure itself as the fidelity-bearing content. Each instance specifies the source relation, retained description, reconstructed or evaluated relation, fidelity criterion, and constrained resource. We use this interface to situate selected methods from graph summarization, spectral sparsification, similarity-preserving representation, and relational distillation within a common formulation while keeping their different reconstruction and resource assumptions explicit. We develop finite-codeword collision as one concrete realization. Same-codeword probability yields a relational geometry linking pair-specific alignment and separation to aggregate R\'enyi-2 occupancy and the spherical geometry of categorical assignments, with exact objective correspondences to squared-Euclidean centroid reconstruction and normalized graph association and cut. Graph and image studies illustrate complementary routes within the finite-codeword family: graph- and teacher-defined relational requirements act directly on equality or collision, while reconstruction acts through a joint decoder. Together, these results illustrate how distinct relational requirements can be formulated and tested within a common constrained-representation framework.
|
| 1057 |
Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problems
2609.31848
|
cs.LG
|
Joanna Marks, Gabriel Rioux, Riccardo Passeggeri |
The Gromov-Wasserstein (GW) distance measures the discrepancy between metric measure (mm) spaces and identifies optimal alignments between them based solely on their intrinsic structure. Since it identifies isomorphic mm spaces, it provides a natural notion of...The Gromov-Wasserstein (GW) distance measures the discrepancy between metric measure (mm) spaces and identifies optimal alignments between them based solely on their intrinsic structure. Since it identifies isomorphic mm spaces, it provides a natural notion of distance for heterogeneous datasets which may admit isomorphic representations. In order to accelerate computation of GW distances, many practitioners employ entropic regularization to obtain an Entropic GW (EGW) problem. The most popular EGW solver is the Mirror Descent (MD) algorithm, which reduces EGW computations to an iterative process where an entropic optimal transport (EOT) problem is solved at each iteration. Despite its widespread use, the convergence of MD for this problem has only been established for restricted classes of costs. On the other hand, a recently proposed dual gradient method is available for general costs, but requires a choice of step size which depends on the regularization parameter. To address these two issues, we introduce Averaged Mirror Descent (AMD), which averages consecutive MD steps, and prove its convergence for arbitrary costs. Then, we establish that the dual gradient method with a fixed step size also converges for arbitrary costs at the cost of a more complicated iteration. In both cases, we also account for inexact iterations which are inescapable in practice. We compare the empirical performance of these methods across various settings and, in particular, show that AMD and the dual gradient method both converge on an example where classical MD fails.
|
| 1058 |
Deep Reinforcement Learning for Equity Trading: Benchmarking Actor-Critic Methods with Forward Retraining
2609.31870
|
cs.LG
|
Bicheng Wang, Xinyi Zhang |
Consistently profitable trading is difficult because equity markets are noisy, non-stationary, and only partially predictable from historical data. We benchmark five deep reinforcement learning (DRL) actor-critic methods: A2C, PPO, DDPG, TD3, and SAC, that lea...Consistently profitable trading is difficult because equity markets are noisy, non-stationary, and only partially predictable from historical data. We benchmark five deep reinforcement learning (DRL) actor-critic methods: A2C, PPO, DDPG, TD3, and SAC, that learn trading actions end-to-end from market states, and compare them with a supervised price-forecasting baseline. Using daily data for 20 large-capitalization S&P 500 stocks from 2000 to 2020, enriched with trend-following technical indicators and log min-max scaling, we train on 2000-2018 and backtest on 2019-2020. Each agent is evaluated both when trained once and under forward retraining, in which it is retrained on all data available before each successive test window. DDPG achieves the highest annual return (55.5%), Sharpe ratio (1.38), and alpha (0.22), but also the highest market beta (1.24). TD3 and SAC offer a better risk-return balance, with Sharpe ratios of 1.37 and 1.33 and maximum drawdowns of about 25%. Forward retraining improves A2C, PPO, and SAC, leaves TD3 essentially unchanged, and reduces DDPG's annual return from 55.5% to 29.8%, consistent with TD3's greater robustness to hyperparameters. The forecasting baseline has the smallest maximum drawdown (9.6%) and the lowest beta (0.31), underscoring a trade-off between the higher returns of end-to-end DRL and the lower risk of forecast-driven strategies.
|
| 1059 |
TemporalGraphLLM: Temporal Graph Neural Networks with Large Language Models for Dynamic Text-Attributed Graphs
2609.31881
|
cs.LG
|
Moran Beladev, Or Eitan, Gilad Katz, Lior Rokach |
Dynamic text-attributed graphs (DTAGs), where nodes, edges, and textual attributes evolve over time, are crucial in applications such as social networks, citation graphs, and knowledge graphs. However, existing approaches struggle to jointly model the temporal...Dynamic text-attributed graphs (DTAGs), where nodes, edges, and textual attributes evolve over time, are crucial in applications such as social networks, citation graphs, and knowledge graphs. However, existing approaches struggle to jointly model the temporal evolution of graph structures and the semantic richness of textual attributes. While Temporal Graph Neural Networks (TGNNs) capture evolving node relationships, they often lack contextual text reasoning. Conversely, Large Language Models (LLMs) excel in textual understanding but struggle with structured graph reasoning in temporal settings. To bridge this gap, we propose TemporalGraphLLM, a novel framework that can integrate any temporal GNN with an LLM for enhanced reasoning in DTAGs. Our approach fine-tunes LLMs using graph-time-aware instruction tuning and novel temporal GNNs injection to replace dedicated added tokens with graph embeddings. TemporalGraphLLM effectively leverages pretrained TGNNs within an LLM framework to achieve state-of-the-art performance on edge classification, link prediction, and edge-based text generation tasks. Extensive evaluation on real-world dynamic graph datasets demonstrates state-of-the-art performance. Our findings highlight the synergistic potential of LLMs and TGNNs, opening new directions for learning on evolving graphs.
|
| 1060 |
DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance
2609.31882
|
cs.LG
|
Zhengyi Guo, Jiayuan Sheng, Wenpin Tang |
Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into ...Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction $\nabla\log h$ under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified $h$-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.
|
| 1061 |
FARE: Deep Reinforcement Learning For Fair Exposure Constrained Uncertainty Aware Financial Content Personalization
2609.31890
|
cs.LG
|
Arundeep Chinta, Lucas Vinh Tran, Jay Katukuri |
Content personalization systems in financial services must ensure fair exposure across diverse offerings-a requirement driven by contractual obligations and the need to prevent "rich-get-richer" dynamics where content with high click-through rate (CTR) dominat...Content personalization systems in financial services must ensure fair exposure across diverse offerings-a requirement driven by contractual obligations and the need to prevent "rich-get-richer" dynamics where content with high click-through rate (CTR) dominates while other relevant products receive minimal visibility. Share of Voice (SOV) constraints, which guarantee each content category a target fraction of top-position exposure, address this by promoting product diversity and balanced user discovery. While re-ranking layers atop CTR models are common in practice, we propose two key novelties: (1) framing SOV-constrained ranking as a deep reinforcement learning problem analogous to constrained trade execution in algorithmic finance, and (2) explicitly incorporating CTR prediction uncertainty into the agent's state space and policy design-enabling larger ranking adjustments for high-uncertainty predictions where deviation from CTR-optimal ordering is less costly. We introduce FARE (Fair Ranking Executor), a modular uncertainty-aware execution layer that translates any black-box CTR model's predictions into SOV-fair rankings without retraining the underlying model. Our uncertainty-weighted proportional control policy (FARE-PC) and learned neural policies (FARE-ES, FARE-PPO) demonstrate that uncertainty-aware approaches can substantially reduce SOV deviation from fairness targets while minimizing engagement loss, with gradient-free evolution strategies outperforming policy gradient methods on synthetic data and the ordering reversing on KuaiRand-Pure.
|
| 1062 |
CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense
2609.31893
|
cs.LG
|
Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan, Hyunwoo Oh, SungHeon Jeong |
Deep reinforcement learning has become a prominent approach to autonomous cyber defense. Existing methods are predominantly model-free and consequently require extensive environment interaction. World models provide an alternative by learning predictive dynami...Deep reinforcement learning has become a prominent approach to autonomous cyber defense. Existing methods are predominantly model-free and consequently require extensive environment interaction. World models provide an alternative by learning predictive dynamics and optimizing policies through imagined trajectories, yielding substantial gains in sample efficiency in robotics and embodied control. Extending this paradigm to cybersecurity raises a fundamental question: what should constitute the "world" in a cyber world model? We introduce CyberWorld, a Dreamer-style world modeling framework that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of the defended network. Across all four scoreable CyberWheel attack strategies, the graph-based CyberWorld variant exceeds a strategy-agnostic control after 3.6k-15.8k environment steps, compared with millions of steps required by model-free PPO. Across representation choices, graph structure provides greater robustness under topology-dependent attacks, while simpler representations remain competitive in overall performance. Among successful runs, the number of episodes required to reach the control remains approximately constant as network size increases from 15 to 100 hosts. These results establish learned cyber dynamics as a sample-efficient and scalable basis for autonomous defense, and identify world representation as a central design axis for robustness and scalability.
|
| 1063 |
Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training
2609.31900
|
cs.LG
|
Emre Can Acikgoz, Yang Li, Zeyu Leo Liu, Srijan Bansal, Dilek Hakkani-T\"ur |
Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that...Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a stage that improves the current model can make the next stage less effective. Through controlled experiments with Qwen3 models on math and science reasoning, we first characterize OPD across nine student-teacher pairs spanning 2x to 53x parameter ratios and show that OPD effectiveness depends on student-teacher compatibility rather than teacher scale alone. The surrounding stages of OPD reshape this compatibility in three ways: (1) A brief SFT warm-up improves subsequent OPD, while an RLVR-strengthened student regresses under distillation from the same teacher. (2) Adapting the teacher with RLVR raises downstream OPD accuracy in proportion to the capability it adds. Following these two interventions, we find that combining teacher adaptation and student warm-up alone raise average OPD accuracy from 29.2\% to 43.8\% (50\% relative improvement) after the same number of distillation steps, with additional preparatory training. (3) At comparable accuracy, OPD leaves a stronger initialization for downstream RLVR than SFT, with a gap that widens as RL compute scales. Our results suggest that each post-training stage should be chosen not only for the capability it adds, but for the learning interface it creates for the next stage.
|
| 1064 |
ROTE: Benchmarking Neural Memorization on Complexity-Controlled Symbolic Sequences
2609.31918
|
cs.LG
|
Xinye Chen, Stefan G\"uttel, Mohammad Mozaffari |
We introduce ROTE (RollOut Testing of Exact memorization), a benchmarking protocol for evaluating symbolic memorization of neural architectures. We study memorization and the extension of symbolic rules in neural sequence models by using sequences whose comple...We introduce ROTE (RollOut Testing of Exact memorization), a benchmarking protocol for evaluating symbolic memorization of neural architectures. We study memorization and the extension of symbolic rules in neural sequence models by using sequences whose complexity is regulated by Lempel--Ziv--Welch (LZW) compression. Under ROTE, each architecture is trained as the same finite-context conditional predictor and is evaluated using teacher-forced one-step prediction as well as closed-loop rollout on the withheld symbols. Following a shared prediction-and-rollout evaluation routine, the benchmark evaluates gated recurrent, minimal recurrent, attention-based, and hybrid recurrent-attention models with their native computational characteristics preserved. Beyond standard predictive metrics, the benchmark reports normalized string distances, training time, memory usage, and parameter count across an LZW-complexity sweep. The study establishes a connection between the complexity of algorithmic sequences and the memorization capacity of neural architectures, revealing the trade-offs involving memorization quality, rollout stability, and computational expense. Our software and reproducible experimental code can be obtained from https://github.com/nla-group/rote.
|
| 1065 |
Resource-Aware Federated Mixture-of-Experts with Adaptive Pruning for Onboard Learning in LEO Satellite Constellations
2609.31932
|
cs.LG
|
Mohamed Shaaban, Mohamed Elmahallawy, Marius Bernahrndt, Tobias Hecking |
Low-Earth-orbit (LEO) satellites are increasingly expected to perform onboard learning for applications such as disaster response and environmental monitoring. However, conventional federated learning (FL) is ill-suited to onboard satellite learning, as it ass...Low-Earth-orbit (LEO) satellites are increasingly expected to perform onboard learning for applications such as disaster response and environmental monitoring. However, conventional federated learning (FL) is ill-suited to onboard satellite learning, as it assumes computational, memory, and communication resources beyond the capabilities of resource-constrained LEO platforms, often necessitating the transmission of raw imagery to ground stations. We present COSMIC-FL, a resource-aware FL framework for efficient onboard learning in LEO satellite constellations. COSMIC-FL introduces two complementary Mixture-of-Experts (MoE) architectures: a Sliced design that shares backbone representations while activating task-specific channel subsets, and a Modular design that employs lightweight gating to route inputs to physically separated expert networks. A semantic class-to-expert mapping enables each satellite to train, update, and communicate only the expert paths relevant to its local data. To further improve efficiency, COSMIC-FL integrates staged optimization with three structured pruning strategies: server-side pruning, client-side fixed-ratio pruning with mean-vote aggregation, and adaptive client-side per-layer pruning based on aggregated importance and a MAD-based gap criterion. Combined with semantic expert routing, these techniques jointly adapt computation and model sparsity to both data semantics and layer importance, yielding a favourable accuracy--efficiency trade-off for heterogeneous space platforms. Experiments on six image classification benchmarks under highly non-i.i.d. settings show that COSMIC-FL maintains competitive accuracy while reducing communication, computation, and energy consumption by up to 80% over SOTA FL methods. We further validate COSMIC-FL on an NVIDIA Jetson AGX Orin, confirming its efficiency gains under realistic embedded deployment constraints.
|
| 1066 |
Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders
2609.31938
|
cs.LG
|
Jiaming Zhang, Wu Yang, Shuai Tao, Wulong Liu |
Generative world models can provide visual rollouts for embodied planning, yet their feasibility on edge devices depends not only on the learned model but also on how the execution runtime represents its operations. We introduce a cache-aware lowering that exp...Generative world models can provide visual rollouts for embodied planning, yet their feasibility on edge devices depends not only on the learned model but also on how the execution runtime represents its operations. We introduce a cache-aware lowering that expresses supported causal Conv3D calls as batched spatial Conv2D operations while preserving pretrained weights, temporal-cache semantics, convolution parameters, bias placement, and output layout. Across the complete Cosmos3-Edge image-to-video pipeline on a 64-GB NVIDIA Jetson AGX Orin, the proposed route accelerates VAE decoding by approximately $7\times$ and reduces complete-generation latency by more than $2\times$, while repeated decoder evaluations maintain complete fast-path coverage without fallbacks. The unchanged lowering also improves Cosmos3-Nano and transfers to LingBot-World's architecturally distinct Wan2.1 VAE. A clean-device comparison against fully specialized TensorRT shows that TensorRT provides a further $1.36\times$ steady-state improvement, but requires substantially greater per-module and per-runtime-state AOT specialization. Same-latent BF16 and FP32 evaluations characterize the finite-precision differences introduced by the alternative execution order. Together, these results position cache-aware lowering as a lightweight runtime optimization that recovers most of the available decoder acceleration without modifying the learned models themselves.
|
| 1067 |
Simple Extensions of Single-Objective Acquisition Functions and Hedge Strategies for Multi-Objective Bayesian Optimization
2609.31940
|
cs.LG
|
Haris Moazam Sheikh |
Multi-objective Bayesian optimization (MOBO) is commonly approached through specialized acquisition functions or scalarization schemes designed to explicitly account for trade-offs among non-preferential objectives. In this work, we show that such complexity m...Multi-objective Bayesian optimization (MOBO) is commonly approached through specialized acquisition functions or scalarization schemes designed to explicitly account for trade-offs among non-preferential objectives. In this work, we show that such complexity might be unnecessary. We propose a framework that extends standard single-objective acquisition functions directly to the multi-objective setting through a hypervolume-based transformation. We further extend hedge strategies for acquisition functions, which are typically used only in single-objective optimization, to the multi-objective regime. Our approach requires minimal modification to existing Bayesian optimization pipelines and avoids the need for bespoke multi-objective formulations. We demonstrate how a broad class of commonly used single-objective acquisition functions and hedge strategies can be adapted in a principled manner to handle multiple objectives, while preserving their intuitive interpretation and computational efficiency. Empirically, we evaluate the proposed methods across a range of synthetic and real-world multi-objective benchmarks. Despite their simplicity, our extensions consistently match or outperform more complex state-of-the-art MOBO methods in terms of optimization performance and sample efficiency. These results suggest that effective multi-objective Bayesian optimization can be achieved by reusing and carefully extending well-established single-objective acquisition strategies, offering a simpler and more flexible alternative to existing approaches.
|
| 1068 |
On-Policy Attention Linearization
2609.31947
|
cs.LG
|
Arian Raje, Anupam Nayak, Anthony Fei, Akaash Parthasarathy, Mohamed Abdelfattah |
Hybrid transformer architectures that replace most softmax attention layers with linear attention offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained f...Hybrid transformer architectures that replace most softmax attention layers with linear attention offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained full-attention transformers. However, these distilled models often collapse on long-context retrieval and reasoning tasks, particularly when operating in thinking mode, where the efficiency gains of hybrid architectures matter most. Since linear attention layers must compress context into a fixed-size state, their errors compound over long sequences. As off-policy distillation never teaches the student model to recover from this drift, tasks that necessitate longer sequence lengths become especially challenging. We introduce On-Policy Attention Linearization (OPAL) in which the hybrid attention student samples its own long-context trajectories and receives dense supervision from the frozen full-attention teacher. Applying OPAL to Qwen3-4B and MiMo-7B-RL-0530, we recover $87$--$94\%$ of full-attention performance on commonsense reasoning, $100\%$ on needle-in-a-haystack (NIAH) retrieval, and $83$--$93\%$ on mathematical reasoning with only 3B training tokens. We achieve these results without supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR). Compared with the strongest prior linearization method, which recovers $68\%$ of its teacher's retrieval performance and $21.6\%$ absolute average mathematical reasoning accuracy, OPAL fully recovers retrieval and achieves $67.6$--$72.2\%$ on math reasoning.
|
| 1069 |
Model-Agnostic Online Certificate-Driven Calibration for Time Series Forecasting Under Distribution Shift
2609.31960
|
cs.LG
|
Chenfeng Huang, Zixuan Ma, George Michailidis |
Time series out-of-distribution generalization requires forecasters to remain reliable when deployment dynamics differ from training conditions due to covariate shift, concept shift, and temporal dependence. Probably Approximately Correct Bayesian domain adapt...Time series out-of-distribution generalization requires forecasters to remain reliable when deployment dynamics differ from training conditions due to covariate shift, concept shift, and temporal dependence. Probably Approximately Correct Bayesian domain adaptation provides computable certificates by decomposing target risk into a source risk term, a source-to-target mismatch term, and a complexity term, but standard analyses rely on independent sampling and distributional stability, assumptions that are violated in time series by serial dependence and nonstationary shift. We propose a model-agnostic online martingale Probably Approximately Correct Bayesian framework that yields finite-sample certificates under temporal dependence and distribution shift. The certificate replaces independent-sample concentration with martingale concentration that adapts to loss scale and predictable variation. We use the certificate as a surrogate regularizer for online calibration by training a gated residual Bayesian head on top of a fixed forecasting backbone, producing a corrective update that reverts to the backbone prediction when the gate is closed. Online calibration combines a source risk anchor, a posterior-shift penalty, and a time-adaptive mismatch term computed from target windows observed before forecasting. It follows a predict-then-update protocol in which outcomes become available only after forecasting and are used to update subsequent predictions. Experiments across convolutional, attention-based, and large language model-based forecasters show improved stability and accuracy under covariate and concept shift.
|
| 1070 |
Model Casting and Low-Parameter Gating: Towards More Sparsely Activated FFNs
2609.31975
|
cs.LG
|
Maria Lomeli, Antoine Groudiev, Matthijs Douze, Lo\"ic Cabannes, Pierre Emmanuel Mazar\'e |
This paper introduces model casting, a mid-training recipe that drastically sparsifies the activations within the Feed-Forward Network (FFN) layer. With this strategy, at inference time, we first compute the output of the gating matrix and, thanks to its high ...This paper introduces model casting, a mid-training recipe that drastically sparsifies the activations within the Feed-Forward Network (FFN) layer. With this strategy, at inference time, we first compute the output of the gating matrix and, thanks to its high sparsity, we avoid computations with the two other matrices, reducing the FLOP count by up to 3x. While this theoretical speedup is an upper bound, model casting translates into significant speedups both on CPU and GPU. We then introduce LoPA Gating, a new FFN design that increases the maximum theoretical speedup. It is a low-FLOPs parameterization of the gating matrix that overcomes the 3x cap by allocating fewer FLOPs and parameters to the gating matrix, compared to the two other FFN matrices that are sparsely activated. We consider two cases: (i) we cast a pre-trained model with a sparsity inducing activation; (ii) we train with LoPA from scratch. In all settings, we significantly outperform existing pruning solutions and regular RELU-fication. For instance, at matched quality, we achieve a 3.2x FLOP speedup with LoPA Casting, against 1.6x at best for competing methods top-p and TEAL. Using dedicated kernels, we achieve an actual 3.31x speed-up on GPU at 90% sparsity, past the 3x ceiling of standard gating; RELU-fication, meanwhile, plateaus below 80% sparsity.
|
| 1071 |
Understanding the Subspace Stabilization of the Hessian and Gradient Covariance Matrix
2609.31983
|
cs.LG
|
Fangshuo Liao, Anastasios Kyrillidis |
The phenomenon of the top subspace stabilization of the Hessian matrix is an surprising and critical aspect in study of the second-order information of neural network training. Prior work argues that the top subspace of the Hessian stabilizes by measuring the ...The phenomenon of the top subspace stabilization of the Hessian matrix is an surprising and critical aspect in study of the second-order information of neural network training. Prior work argues that the top subspace of the Hessian stabilizes by measuring the overlap between the top subspaces of the step-wise Hessian, and explains this stabilization with diminishing parameter change in the late phase of training. In this paper, we define a new instability metric for the subspace evolution, and use it to detect subspace stabilization that is independent of the magnitude of parameter change. In the meantime, we observe that the gradient covariance matrix has a similar property of its top subspace to the Hessian. By using a between-class and within-class decomposition of the gradient covariance matrix, we identify an explicit form that gives a near-perfect approximation of the top-$(C-1)$ subspace of the Hessian and the gradient covariance matrix. In the gradient flow set-up, we show that the slow evolution of the idenfied approximation is due to the separation between the outlier and the bulk eigenvalues of the Hessian matrix, thus providing an explanation to the phenomenon of the top subspace stabilization of the Hessian matrix.
|
| 1072 |
Lagrangian and Hamiltonian Neural Networks With a Dissipative System
2609.31988
|
cs.LG
|
V. Rayamajhi, J. Singal |
We investigate the applicability of Lagrangian and Hamiltonian Neural Network models to a dissipative system that has explicit time dependence in its Lagrangian, Hamiltonian, and total energy. To do so we consider these neural network models for simulated syst...We investigate the applicability of Lagrangian and Hamiltonian Neural Network models to a dissipative system that has explicit time dependence in its Lagrangian, Hamiltonian, and total energy. To do so we consider these neural network models for simulated systems of a harmonic one-dimensional, one-component oscillator with damping, as well as without damping for comparison. We find that both the Lagrangian and Hamiltonian approaches are able to predict the empirical physical behavior of the damped oscillator systems and to effectively ``learn'' to varying degrees the underlying Lagrangians and Hamiltonians, as has previously been shown to be the case with undamped oscillator systems. These investigations elucidate important properties of Lagrangian and Hamiltonian mechanics, including properties that are not manifest when considering systems without explicit time dependence.
|
| 1073 |
Mechanistic Interpretability Reveals Shared Causal Subspaces in Brain-to-Speech Decoders
2609.31992
|
cs.LG
|
Maryam Maghsoudi, Ayushi Mishra, Sanghamitra Dutta |
Decoding covert speech, such as mimed or imagined, from brain activity is harder than decoding vocalized speech. Cross-modal transfer, where information from one speech form helps decode another, is a promising remedy; yet how a decoder internally represents a...Decoding covert speech, such as mimed or imagined, from brain activity is harder than decoding vocalized speech. Cross-modal transfer, where information from one speech form helps decode another, is a promising remedy; yet how a decoder internally represents and processes brain activity from different speech forms remains unclear. In this work, we ask: which internal neurons of a decoder carry cross-modal information, and are these neurons shared across different speech forms? To answer these questions, we leverage mechanistic interpretability, using recordings of the same sentences in vocalized, mimed, and imagined input pairs for activation patching. We insert the decoder's internal activity for a sentence in one condition into its processing of the same sentence in another and measure the change in decoding accuracy. We find that no single neuron drives this benefit; instead, it arises from small groups of neurons, with vocalized speech as the most useful source. These groups are largely condition-specific in the early stage of the decoder but overlap in the later stage. These findings point toward more data-efficient covert speech decoders through training objectives that encourage shared later-stage representations learned mainly from vocalized data.
|
| 1074 |
Can Circuit Alignment Predict OOD Generalization?
2609.31996
|
cs.LG
|
Ayan Banerjee, Abhra Chaudhuri, Josep Llados, Umapada Pal, Anjan Dutta |
Can out-of-distribution (OOD) generalization be predicted from a trained model's weights alone, without any target-domain data? Existing representational similarity metrics (CKA, SVCCA, RSA) compare activations rather than forecast generalization. We show they...Can out-of-distribution (OOD) generalization be predicted from a trained model's weights alone, without any target-domain data? Existing representational similarity metrics (CKA, SVCCA, RSA) compare activations rather than forecast generalization. We show they are provably insensitive to structural rerouting in the computational graph, the very change distribution shift induces. We close this gap with the Circuit Alignment Score (CAS), which compares class-specific circuits across domains via graph kernels, decomposed into same-class coherence and cross-class confusion. Casting CAS as a Lebesgue integral over the domain distribution, we prove its Monte Carlo estimate recovers the ground-truth ranking of learners by OOD accuracy, with pairwise inversion error vanishing at rate $O(1/M)$, where $M$ is the number of sampled domains. Across $48$ learners on PACS, CAS attains $0.88$ rank correlation with OOD accuracy, versus $0.58$ (CKA), $0.23$ (SVCCA), and $0.14$ (RSA), with similar trends on other benchmarks and even against data-dependent methods, making it the first provably consistent predictor of distributional robustness requiring neither target-domain data nor labels. The code is available at: https://github.com/ayanban011/ACE
|
| 1075 |
Integrating Language Models into Listened and Imagined Speech Decoding from MEG
2609.31997
|
cs.LG
|
Maryam Maghsoudi, Sai Samrat Kankanala, Shihab A. Shamma, Sriram Ganapathy |
Decoding imagined speech is an important goal for brain-computer interfaces but remains challenging due to weak neural responses, low signal-to-noise ratio, and limited imagined-speech datasets. Language models provide strong contextual cues for text predictio...Decoding imagined speech is an important goal for brain-computer interfaces but remains challenging due to weak neural responses, low signal-to-noise ratio, and limited imagined-speech datasets. Language models provide strong contextual cues for text prediction, but how much they can help neural decoding and whether their contribution differs for decoding perceived and imagined speech remains unclear. To investigate this, we use a paired listened-imagined MEG dataset and incorporate language-model information at two stages. First, we train a contrastive neural decoder that aligns MEG representations with acoustic and contextual language representations, improving cross-subject word decoding for both listened and imagined speech. Second, at inference, we introduce a neural-constrained beam-search framework that combines neural evidence with language-model next-word probabilities. We find that imagined-speech decoding benefits more from the language model than listened-speech decoding. For Imagined speech, the best-performing balance between neural and language-model evidence shifts toward the language model, and the gain over neural-only decoding is larger. Together, these results suggest that language priors are most useful when neural evidence is weaker, making them particularly valuable for imagined-speech BCIs.
|
| 1076 |
VC Dimension and Expressivity of Real-Valued Transformers
2609.31999
|
cs.LG
|
Gavin Dooley, Andy Yang, Yijia Jessica Zhu, David Chiang, Peter Cholak |
Whereas previous results on abilities and limitations of transformers have restricted the definition of transformers in various ways, here we study softmax-attention, multi-layer transformers operating on real values, with very few additional assumptions. Appl...Whereas previous results on abilities and limitations of transformers have restricted the definition of transformers in various ways, here we study softmax-attention, multi-layer transformers operating on real values, with very few additional assumptions. Applying results from real geometry, we obtain upper bounds on the VC dimension and split VC dimension of such transformers ($O(n^4)$ and $O(n^6)$, respectively, where $n$ is the input length). Conversely, we also construct specific transformers witnessing lower bounds on these quantities ($\Omega(n)$ in each case). These results have some notable consequences. For example, within the class of symmetric (permutation-invariant) functions, we show that transformers can uniformly express all functions over an alphabet of one symbol and non-uniformly express all functions over an alphabet of two symbols, but cannot (even non-uniformly) express some functions over an alphabet of six symbols. We also prove limitations on how many bits of a real number a transformer can access.
|
| 1077 |
Is invariance all you need for algorithmic fairness? Removing demographic information can create new bias
2609.32004
|
cs.LG
|
Aditya Parikh, Eike Petersen, Stella Frank, Enzo Ferrante, Melanie Ganz |
Encoded demographic information in internal model representations is a commonly assumed risk factor for algorithmic bias, with demographic representation invariance often being touted as the ideal state. However, while demographic shortcut learning is a genuin...Encoded demographic information in internal model representations is a commonly assumed risk factor for algorithmic bias, with demographic representation invariance often being touted as the ideal state. However, while demographic shortcut learning is a genuine threat, some degree of encoding is necessary when demographics correlate with target labels. Here, we show, mathematically and empirically, that enforcing demographic invariance can actually hamper bias mitigation and even create new biases. We distinguish marginal from class-conditional representation invariance, and show that they imply the standard group fairness notions of demographic parity and equalized odds, respectively. We evaluate the effects on predictive performance and fairness of enforcing both invariance types, both theoretically and empirically across five tabular and two chest X-ray imaging datasets. Our findings support our mathematical argument that demographic representation invariance is neither desirable nor sufficient for fairness.
|
| 1078 |
Human Activity Recognition via Ultra-Wideband Data: A Framework for Dimensionality Reduction, Pattern Discovery, and Predictive Modeling
2609.32008
|
cs.LG
|
Nahid Sahel Gozin, Reza Sedaghat, Prathap Siddavaatam |
Recent advances in sensor technology have enabled more effective human activity recognition (HAR), particularly in real-time systems with limited computational resources. However, Ultra-Wideband (UWB) radar data remain challenging due to high dimensionality, n...Recent advances in sensor technology have enabled more effective human activity recognition (HAR), particularly in real-time systems with limited computational resources. However, Ultra-Wideband (UWB) radar data remain challenging due to high dimensionality, noise, complexity, and nonlinear characteristics. This research proposes a framework to efficiently reduce data size, uncover significant patterns, and classify six activity types (Standff, Liedown, Noactivity, Sit, Stand, and Walk) from UWB signals with high accuracy. Two novel dimensionality reduction techniques are introduced in this paper. The first, Clustered Polynomial Expansion with Incremental PCA (CPE-IPCA), combines clustering and polynomial feature expansion with Incremental PCA, preserving 100% of the variance in only 50 components. The second, Post-PCA Standardization Approach (PPSA), standardizes data after PCA and retains 99.1% of the variance in 80 components, achieving superior compression and computational efficiency compared to conventional nonlinear methods. Frequent patterns are identified using Apriori and FP-Growth, which are then classified with Random Forest and a Vector Space Model (VSM). The framework achieves 100% accuracy with Random Forest on CPE-IPCA and 99% on PPSA, while VSM attains 100% precision, recall, and F1 on PPSA and near-perfect performance on CPE-IPCA (precision 1.00, recall 0.98-1.00, F1 0.99-1.00), demonstrating a fast, interpretable, and robust HAR system suitable for healthcare, assisted living, and smart environments.
|
| 1079 |
Representation Learning for Exact Preimages
2609.32018
|
cs.LG
|
Konstantin Hess, Stefan Feuerriegel |
Modern neural predictors can model highly nonlinear maps, but many scientific and engineering tasks require reasoning in the opposite direction: given a performance or safety level, the goal is to characterize the preimage, that is, the complete set of inputs ...Modern neural predictors can model highly nonlinear maps, but many scientific and engineering tasks require reasoning in the opposite direction: given a performance or safety level, the goal is to characterize the preimage, that is, the complete set of inputs which meet the desired target level and optimize over that set. For expressive neural predictors, however, such preimages typically have no explicit representation and are expensive to recover or optimize over. This creates a fundamental three-way challenge between expressive forward prediction, accurate preimage approximation, and tractable optimization over the preimage for downstream tasks. We introduce TRIO (tractable representations for preimage learning and inverse optimization), a framework for learning representations that make these objectives compatible by construction. Our key contribution is a preimage factorization: the forward model remains expressive through nonlinear radial transformations (including neural networks), while, under inversion, each transformation reduces to a single scalar radius, which yields simple geometric level sets. This yields an explicit geometric representation that is reusable for downstream optimization over the preimage, and, for linear objectives, we show that this admits a closed-form global solution. We finally prove a universal approximation theorem which shows that TRIO can approximate any continuous forward map and its entire family of potentially disconnected, nonconvex preimages arbitrarily well. Hence, TRIO combines expressive forward modeling, exact preimage recovery, and tractable global downstream optimization over preimages by design.
|
| 1080 |
Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation
2609.32048
|
cs.LG
|
Debamita Ghosh, George K. Atia, Yue Wang |
Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for ad...Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for addressing such uncertainty, yet existing methods often rely on restrictive assumptions or scale poorly to large state and joint action spaces. We study online learning in general-sum DRMGs with general function approximation and $\phi$-divergence uncertainty sets. We propose RoMEX-$\phi$, a model-free framework that integrates equilibrium-based exploration with dual fitted learning. Through a functional dual representation of the robust multi-agent Bellman operator, RoMEX-$\phi$ enables tractable worst-case value estimation from nominal interaction data using a centered empirical robust discrepancy. We introduce the robust Multi-Agent Decoupling Coefficient (robust MADC) to characterize the intrinsic exploration complexity arising from strategic interactions and adversarial transition uncertainty. We establish sublinear robust regret guarantees governed by the robust MADC rather than explicitly by the state and joint action space sizes, replacing tabular dependence with intrinsic function-class complexity. Numerical experiments on a scalable general-sum DRMG under total variation uncertainty show that RoMEX-$\phi$ is substantially more resilient to transition shifts than its non-robust counterpart while remaining competitive with an exact tabular robust baseline. Our results provide a scalable framework for distributionally robust multi-agent reinforcement learning with general function approximation.
|
| 1081 |
Graph Forward Distribution Matching for Molecular Inverse Design
2609.32056
|
cs.LG
|
Yihan Zhu, Yuhan Liu, Brett Savoie, Tengfei Luo, Meng Jiang |
Achieving precise control over multiple properties without sacrificing chemical validity remains a central challenge in molecular inverse design. Existing reinforcement learning (RL) methods fine-tune graph diffusion models by treating **reverse** sampling as ...Achieving precise control over multiple properties without sacrificing chemical validity remains a central challenge in molecular inverse design. Existing reinforcement learning (RL) methods fine-tune graph diffusion models by treating **reverse** sampling as a sequential policy, using a single terminal reward to optimize hundreds of coupled decisions. They often suffer from instability, validity collapse, and limited property gains. We introduce GraphFDM (Graph Forward Distribution Matching), a new online RL paradigm for graph diffusion that performs optimization through the **forward** process. GraphFDM uses valid generations to define a reward-tilted target distribution jointly optimized over graph size and molecular structure for each property condition, incorporating reinforcement signals into supervised learning without storing reverse trajectories. We derive the unique optimal target, prove a condition-wise improvement guarantee, and show that the fixed graph-size prior of standard graph diffusion leaves an irreducible matching gap. In multi-conditional polymer and small-molecule generation, GraphFDM achieves the lowest MAE on every target property, with reductions of up to 53.0\% relative to the strongest baselines and chemical validity above 0.99. It further generalizes to out-of-distribution property combinations.
|
| 1082 |
Depth Laws for the Precision Floor of Trained Neural Networks: Amplification, Residual Scaling, and a Quantization-Aware Training Paradox
2609.32060
|
cs.LG
|
Ahmad S. Tarawneh |
How many bits does a network need before its accuracy collapses, and how does this grow with depth? We study the precision floor, the perturbation level or bit-width at which accuracy falls halfway to chance, in MLPs, CNNs, Vision Transformers and nine pretrai...How many bits does a network need before its accuracy collapses, and how does this grow with depth? We study the precision floor, the perturbation level or bit-width at which accuracy falls halfway to chance, in MLPs, CNNs, Vision Transformers and nine pretrained language models, under post-training quantization (PTQ) and quantization- or noise-aware training (QAT). (i) A first-order theory sets the floor through one full-precision quantity, the predictive amplification $G$: $\eta_c=\Lambda/G$, and $G^2$ grows linearly in depth at a rate proportional to the squared residual branch scale. (ii) The predicted equality $\alpha_{PTQ}=\rho$ of depth exponents holds within 95% intervals in twelve of thirteen trained architectures and in GPT-2 from 12 to 48 layers, with $\Lambda=1.45\pm14\%$ across trained architectures. (iii) Residual branches scaled by $1/\sqrt{D}$ and pre-normalisation remove the depth penalty, and each quantizer turns noise into bits at a rate fixed by its step rule, giving $b_c=(\alpha/\gamma)\log_2 D+C$. (iv) A QAT paradox: noise-aware training roughly doubles the tolerable noise of shallow networks, but the gain decays with depth, so the depth law steepens ($\alpha_{QAT}/\alpha_{PTQ}=1.45$-$1.47$ on two datasets, ten seeds each). Decision margins, cross-layer error cancellation and heavy tails do not set the floor.
|
| 1083 |
When Should a Human Take Back Control? Optimal Delegation under Turbulent AI Risk
2609.32083
|
cs.LG
|
Haoze Yan, Julien Roze, Ved Upadhyay, Unal Tatar, Thibaut Mastrolia |
Deploying AI systems requires deciding when to delegate tasks and when humans should intervene to monitor and mitigate risk induced by AI operations. These decisions are challenging when failures cluster: a hallucination or harmful output can trigger further e...Deploying AI systems requires deciding when to delegate tasks and when humans should intervene to monitor and mitigate risk induced by AI operations. These decisions are challenging when failures cluster: a hallucination or harmful output can trigger further errors, creating periods of elevated risk. We introduce a continuous-time framework for learning adaptive human oversight under such turbulent AI risk. Existing oversight and delegation formulations condition on history but do not model incident clustering, or its suppression by supervision effort, jointly with the delegation decision and this study fixes this gap. The self-exciting dynamics capture how risk events increase the likelihood of subsequent events, making their timing and history central to decision-making. We formulate a stochastic control problem combining human actions, monitoring effort, and switching between human-AI-assisted operation and full AI delegation, balancing operational rewards against oversight costs, and cascading AI-failures and induced uncertainty. Human participation is an endogenous component of risk management: the policy determines both when oversight is needed and how much effort to allocate. We study a relaxed switching formulation and propose Hawkes-PPO, a policy-gradient method that uses a bank of exponential filters of observed incident times. In a synthetic environment it attains a higher risk-adjusted objective than either fixed regime and approaches an approximate full-information oracle. We illustrate our results with numerical simulations by examining how cascade risks influence intervention and delegation, connecting reinforcement learning with adaptive human oversight of AI systems. In particular, we illustrate the benefit of our switching strategy and Hawkes-PPO algorithm to monitor the project efficiently along time, reducing turbulent risks occurrences and costs.
|
| 1084 |
Emergent One-Third Scaling Law as Attention Tries to Concentrate
2609.32100
|
cs.LG
|
Yizhou Liu, Sara Kangaslahti, Jeff Gore |
The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a s...The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distributions. What happens with multiple softmax functions, as in LLMs, is unclear. Here, we show through toy models that any softmax learning peaked distributions, regardless of its position in the model, can develop logit magnitudes that grow in a power law with exponent $1/3$, becoming a training bottleneck whose loss contribution decays as a power law with the same exponent $1/3$. The overall loss therefore obeys $1/3$ scaling whenever at least one softmax learns peaked distributions. We confirm that many softmax functions in LLMs learn peaked distributions and that LLM loss scaling matches this $1/3$ prediction. Moreover, logit growth dynamics reveal that attention heads, rather than the language modeling head, are the bottleneck likely driving the $1/3$ loss scaling in LLMs. Attention trying to concentrate on specific information, which is the heart of Transformers, may therefore also be the heart of the neural scaling law of training.
|
| 1085 |
LLM Unlearning Evaluation with TRIAGE
2609.32103
|
cs.LG
|
Danial Ataee, Peter Triantafillou |
Large language models can memorize private or harmful information, motivating machine unlearning methods that remove targeted knowledge while preserving other capabilities. However, existing evaluations rely primarily on behavioral benchmarks, which assess \em...Large language models can memorize private or harmful information, motivating machine unlearning methods that remove targeted knowledge while preserving other capabilities. However, existing evaluations rely primarily on behavioral benchmarks, which assess \emph{whether} a model appears to forget but provide limited insight into \emph{how} unlearning changes the model or affects related knowledge. We introduce \textit{TRIAGE} (\textit{Tripartite Representation-internal Introspection for Adjacency Gap Evaluation}), a benchmark-agnostic evaluation framework for characterizing these changes. TRIAGE uses diagonal approximations of the Fisher information and Hessian to measure changes in parameter sensitivity and local curvature, and utilizes a Forget / \emph{Adjacent-Retain} / \emph{Generic-Retain} partition to quantify an \emph{adjacency gap} in semantically related knowledge. Based on the magnitude and distribution of these changes, TRIAGE further classifies each algorithm's update as \emph{no-op}, \emph{partially localized}, \emph{collateral dominant}, or \emph{globally destructive}. Across 12 unlearning methods, four language models, and the WMDP, TOFU, and MUSE benchmarks, we find that methods with similar behavioral forgetting can produce substantially different internal changes and patterns of collateral damage. These signatures also vary across models and benchmarks, indicating that the effects of unlearning are not determined solely by the unlearning algorithm. TRIAGE can be applied alongside existing unlearning benchmarks to complement behavioral evaluation with a model-internal view of how unlearning reshapes the model's parameter space and affects retained knowledge.
|
| 1086 |
An Attention-Driven Heterogeneous GNN Model for Credit Card Fraud Detection
2609.32106
|
cs.LG
|
Kathiresan Jayabalan, Sethuraman Radhakrishnan |
The global transition to a cashless economy has placed credit cards as the key element of digital transactions, acclaimed for their easy use, speed, and acceptance in most places. However, the growing dependence on this payment method has led to an escalation ...The global transition to a cashless economy has placed credit cards as the key element of digital transactions, acclaimed for their easy use, speed, and acceptance in most places. However, the growing dependence on this payment method has led to an escalation of the risks associated with credit card (CC) fraud. Detecting this type of fraud is a difficult task because the patterns are constantly changing, there is a data imbalance, and it is necessary to identify the legitimate transactions and the fraud ones at the same time. This study addresses this challenge by proposing a credit card fraud detection (CCFD) framework using a data balancing technique and a deep learning (DL) model. The proposed fraud detection model is trained and evaluated by collecting the dataset called Credit Card Fraud Detection from the Kaggle repository. As the dataset is highly imbalanced, we utilized the Synthetic Minority Oversampling Technique (SMOTE)-Tomek technique to balance the dataset. Further, the balanced dataset is classified using the Heterogeneous Graph Neural Network (HGNN) model. The HGNN model represent various transactions using a heterogeneous graph architecture and by using an attention based message passing technique, it managed to consider the complex relationships, time factors, and user behavior. The integration of SMOTE-Tomek in the model further boosted its capacity to identify fraudulent transactions, while lowering the rate of false positives. The HGNN model attained a 99.97% accuracy, a 99.48% F1-score, a 99.15% precision, and a 98.97% recall. The findings indicates that this model is effective and can be applied to real-world CCFD scenarios.
|
| 1087 |
SAMBAR: Selective Anchoring via Method of Multipliers for Balanced Knowledge Acquisition and Retention in Vision-Language-Action Models
2609.32108
|
cs.LG
|
Aayushi Shrivastava, Xunlan Zhou, Hongrui Zhao, Ziyu Chen, Negar Mehr |
Vision-Language-Action (VLA) models leverage large-scale pretraining to ultimately achieve generalist manipulation. Deployed VLA policies must support continual learning to acquire new tasks over time. Teaching a VLA a new task generally requires finetuning it...Vision-Language-Action (VLA) models leverage large-scale pretraining to ultimately achieve generalist manipulation. Deployed VLA policies must support continual learning to acquire new tasks over time. Teaching a VLA a new task generally requires finetuning it on demonstrations of that task. However, naively finetuning on downstream tasks causes the policy to forget earlier tasks and degrades generalist capabilities. This failure is known as catastrophic forgetting. Most continual learning methods counter it by replaying data from earlier tasks. However, the old task demonstrations are not always readily available. In this paper, we introduce SAMBAR, a continual learning algorithm that prevents catastrophic forgetting during VLA finetuning without requiring access to the demonstrations of any previously learned task. We propose to cast continual learning as a constrained optimization problem and solve it with the method of multipliers. In our approach, the method of multipliers drives the policy to learn the new task without the model parameters drifting far away from their previous values. In contrast to a standard regularization penalty, the method of multipliers raises the penalty as the constraint violation accumulates by using a dual variable. We also selectively anchor the parameters critical to previous tasks to preserve past knowledge, leaving other parameters free for new task acquisition. The combination of dual variable and selective anchoring, therefore, balances knowledge acquisition with knowledge retention. We evaluate our method, SAMBAR, on the LIBERO simulation benchmark and on hardware. When sequentially finetuning on a VLA, every replay-free baseline we compare against completely forgets the first task it learned, whereas SAMBAR retains every task it has learned.
|
| 1088 |
How Reusable Are Benchmarks with Richer Feedback?
2609.32109
|
cs.LG
|
Youssef Allouah, John Duchi |
We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex com...We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the number of criteria, reaching the $\Theta(\sqrt{k})$ cost of answering $k$ adaptive statistical queries with only $O(\log k)$ criteria, at fixed accuracy and confidence. In attacks on multi-task large language model benchmarks with five to ten criteria, feedback restricted to nondominated task profiles produces large reused-to-held-out score gaps and frequent false winners. These results challenge a prominent explanation for prior observed reliable benchmark reuse---that developers mainly respond to convincing improvements over the current best---in rich-feedback settings, while leaving open how often ordinary model development encounters this vulnerability.
|
| 1089 |
Handwritten Digit Leakage from Smartphone Motion Sensors Across Unseen Users and Phone Models
2609.32117
|
cs.LG
|
Constantino \'Alvarez Casado, Erkka Rantahalvari, Matteo Pedone, Matti Matilainen, Manuel Lage Ca\~nellas |
Smartphone motion sensors support interactive applications, but their readings may also reveal touchscreen input beyond their intended use. Assuming known drawing intervals, we study whether handwritten digits remain predictable across users and devices, as a ...Smartphone motion sensors support interactive applications, but their readings may also reveal touchscreen input beyond their intended use. Assuming known drawing intervals, we study whether handwritten digits remain predictable across users and devices, as a 10-class problem on 19,628 HuMIdb recordings from 481 participants. We compare handcrafted features with classical machine learning algorithms, MiniRocket kernels, and a compact sensor patch transformer on accelerometer, linear acceleration, gyroscope, and gravity signals. The transformer achieves 57.74\% accuracy and 82.64\% top-3 accuracy on 75 unseen participants, and 58.77$\pm$0.95\% over 3 seeds for unseen participants on 9 unseen phone models. Low motion recordings remain informative, accuracy is not monotonic in motion level, and the tested contrastive pretraining, augmentation, and derived signals give no consistent gains. Digits are thus predictable beyond familiar users and phone models under assumed segmentation, while acquisition-order shortcuts limit conclusions about practical privacy exposure. Code available at: https://github.com/Arritmic/motion-digit-leakage.
|
| 1090 |
What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models
2609.32124
|
cs.LG
|
Leonel Aguilar |
When adapting pre-trained models through fine-tuning, freezing weights alone might not preserve performance, as updates elsewhere can change the inputs to the frozen core, ultimately affecting overall performance. We first analyse the case where a selected fro...When adapting pre-trained models through fine-tuning, freezing weights alone might not preserve performance, as updates elsewhere can change the inputs to the frozen core, ultimately affecting overall performance. We first analyse the case where a selected frozen core can be isolated and propose removal-value, a capacity-based score that approximates HOPE's removal cost averaged over removal orders. We show that in VGG-8, cutting paths from trainable neurons into a frozen core makes selection using this score useful: $70\%$ frozen preserves $5.22\pm0.51$ percentage points more old-task accuracy than DEFT at similar new-task accuracy. In transformers, shared residual streams leave paths into frozen neurons open. For this case, we derive drift-value, a forward-only proxy for the output disturbance from updating each weight entry under a local update model. In language models, at 40 epochs, this policy exceeds adapted Wanda and RIA freezing scores in settings with substantial retention loss, while its differences from Fisher remain unresolved. After 160 epochs on Qwen2.5-1.5B, it retains $0.0433\pm0.0102$ more than static Fisher. In DINOv3 vision-transformer adaptation to point clouds, drift-value retains $0.440$ image accuracy versus $0.187$ for a random mask of the same count. These results motivate Guarded Freezing: select by removal-value when incoming paths are cut, and by drift-value when they remain.
|
| 1091 |
Latency-Aware Client Assignment for Parallel Split Learning With Global Sampling
2609.32132
|
cs.LG
|
Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink |
In cross-silo split learning, Parallel Split Learning with Global Sampling forms representative pooled batches when class distributions differ across clients, but ignores client delay when several clients can supply the same class. We introduce Latency Budgete...In cross-silo split learning, Parallel Split Learning with Global Sampling forms representative pooled batches when class distributions differ across clients, but ignores client delay when several clients can supply the same class. We introduce Latency Budgeted Parallel Split Learning with Global Sampling, which separates each pooled batch's integer class target from the choice of clients that supply its examples. The flow variant formulates this assignment as an integral network-flow problem and minimizes modeled client-side completion time for the current target. The fast variant uses a greedy next-completion rule to reduce schedule-construction cost. Both preserve the target stream and use every local example once per epoch. A planning rule selects between the variants while accounting for the cost of constructing both candidate schedules. On CIFAR-10, the flow variant reduces modeled training time by 6.75%, with a 0.30 percentage-point decrease in final accuracy. On Tiny ImageNet with 20 candidate classes per client, the fast variant reduces modeled time by 16.87% and reaches all four validation targets earlier than the latency-unaware baseline. Across 405 schedule comparisons, the planning rule stays within 2% of the lower realized cost in 96.54% of cases. In our evaluation, latency-aware provider assignment reduces modeled training time without changing the prescribed class targets, while the preferred variant depends on whether assignment savings outweigh schedule-construction overhead.
|
| 1092 |
REALM: Regime-Switching, Explainable, and Activation-Induced Linear Models
2609.32141
|
cs.LG
|
Xiaoran Cheng, Sen Na, Jia Li |
Deep ReLU networks are piecewise-affine mappings that partition the input space into cells, each characterized by a distinct activation pattern. This structure motivates fitting a local linear model within each cell to preserve predictive accuracy while improv...Deep ReLU networks are piecewise-affine mappings that partition the input space into cells, each characterized by a distinct activation pattern. This structure motivates fitting a local linear model within each cell to preserve predictive accuracy while improving interpretability. The challenge is to identify regimes that are stable, data-adaptive, and easy to explain. We propose REALM, a mixture of linear models whose regimes are induced by neural activation patterns. Because the number of activation cells in a deep neural network (DNN) can grow rapidly with depth, we first distill a deep teacher into a wide, shallow student network (WSSN), then binarize and cluster its hidden-layer activations to define the regimes and fit a linear model within each regime. Since the regimes are discovered from internal structure, the router does not carry the predictive burden. To make regime assignment interpretable, we train a multiclass logistic regression, the explanatory gate, to reproduce the regime assignments. The two-level structure is interpretable at both stages in terms of raw tabular or learned convolutional features: the gate identifies features that determine regime assignments, while the linear models identify features that drive predictions within each regime. We analyze an idealized setting that illustrates a trade-off between partition complexity and stability: as the number of regimes grows, finer partitions can improve approximation but may reduce regime-assignment stability. Experiments on tabular and image datasets show that REALM achieves competitive predictive performance relative to other DNN-guided mixture surrogates and inherently interpretable models while producing stable regime-level explanations.
|
| 1093 |
Spectral Reversal: Counteracting Singular Value Bias for Graph Prompting
2609.32143
|
cs.LG
|
Hanxu Yang, Yuhuan Zhao, Xiaodong He, Zhao Kang |
Pre-training Graph Neural Networks (GNNs) via self-supervised learning has become a dominant paradigm, yet efficiently adapting frozen encoders remains a challenge. Graph prompting offers a parameter-efficient alternative to fine-tuning, but existing methods l...Pre-training Graph Neural Networks (GNNs) via self-supervised learning has become a dominant paradigm, yet efficiently adapting frozen encoders remains a challenge. Graph prompting offers a parameter-efficient alternative to fine-tuning, but existing methods largely treat pre-trained models as opaque feature extractors, ignoring their internal spectral structure. In this work, we identify a systematic phenomenon in pre-trained GNNs, which we term spectral bias: optimization during pre-training disproportionately aligns representations with directions associated with large singular values, leaving low-energy directions under-explored. We show that these underutilized directions can encode complementary information that is beneficial for downstream adaptation, especially under distribution shift. To leverage this insight, we propose Spectral Reverse Prompt (SRP), a prompting framework that rebalances the spectral contributions of frozen GNN encoders. SRP applies a learnable soft-thresholding mask in the spectral domain to down-weight dominant directions while amplifying weaker ones. In addition, SRP incorporates a null-space augmentation module that captures variation in directions with minimal activation under the frozen encoder. Extensive experiments across multiple benchmarks demonstrate that SRP achieves state-of-the-art performance with minimal additional parameters, highlighting that reweighting spectral components is a principled and effective strategy for parameter-efficient graph adaptation.
|
| 1094 |
Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions
2609.32146
|
cs.LG
|
Arjun Narayanan, Per-Olof Persson |
A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bou...A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vertex irregularity of any all-quadrilateral mesh of a given domain purely based on its topology and corner angles. We train a reinforcement learning agent to build decompositions that reach this bound, which we call par. It acts directly on the mesh's half-edge data structure through local edits, with a policy network whose convolutions follow the mesh's own connectivity, so it applies unchanged to domains larger than any seen in training. The reward targets the floor directly, and it is sparse: random play reaches it on no domain with more than eight sides. We overcome this exploration barrier via behaviour cloning on optimal meshes that are trivial to construct, walked backward into demonstrations, before training it with PPO. On 96 held-out domains the agent produces an all-quadrilateral mesh on every one, a usable one on 95.7 on average, and a provably optimal one on 90; Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none, and even at three to fourteen times the elements never produces a more regular mesh. On 64 domains twice the training size the agent completes all, is usable on 62, and keeps a median excess over par below one against Gmsh's 39 at the same element count.
|
| 1095 |
CAFE: Counterfactual Prediction via Fast Posterior Estimation
2609.32167
|
cs.LG
|
Xinyan Han, Xiaoyu Lin, Hao Zou, Xingxuan Zhang, Bo Li |
Counterfactual prediction estimates an individual's outcome under an alternative intervention given their factual observations. Such outcomes are generally not identifiable from observational data without additional assumptions. Even within the class of fully ...Counterfactual prediction estimates an individual's outcome under an alternative intervention given their factual observations. Such outcomes are generally not identifiable from observational data without additional assumptions. Even within the class of fully observed additive noise models (ANMs), different causal graphs can generate the same observational distribution yet imply different individual counterfactual outcomes. Predictions based on a single estimated graph ignore this structural uncertainty. We therefore target a Bayesian counterfactual posterior predictive distribution that combines predictions from plausible SCMs. We introduce CAFE (\textbf{C}ounterf\textbf{A}ctual Prediction via \textbf{F}ast Posterior \textbf{E}stimation), an amortized inference framework that directly approximates the Bayesian counterfactual posterior predictive distribution. We pretrain a transformer-based model on synthetic counterfactual tasks generated from a diverse prior over ANMs. Given an observational dataset, an individual's factual observations, and an intervention, CAFE approximates the corresponding posterior predictive distribution in a single forward pass. Experiments show that CAFE accurately predicts individual counterfactual outcomes in identifiable settings and approximates the posterior predictive distribution when structural uncertainty induced by observationally indistinguishable causal graphs exists. Strong performance in realistic manufacturing and viticulture settings further demonstrates its empirical robustness beyond the assumptions of the training prior.
|
| 1096 |
Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles
2609.32170
|
cs.LG
|
Yongyi Guo, Zifan Xu, Ziping Xu, Kelly W. Zhang |
The performance of online reinforcement learning depends critically on design choices, especially those that affect exploration. These choices are often selected by fitting a simulator to offline data, evaluating candidate algorithms in that simulator, and dep...The performance of online reinforcement learning depends critically on design choices, especially those that affect exploration. These choices are often selected by fitting a simulator to offline data, evaluating candidate algorithms in that simulator, and deploying the best-performing one. The simplest Plug-In selection rule simply selects the best performing algorithm on the fitted simulator, making evaluations unreliable when the offline data used to fit the simulator are limited. We investigate Uncertainty-Aware selection, which forms an ensemble of simulators---for example, obtained by bootstrap resampling---and selects the online algorithm with the best average performance across the ensemble. While ensemble-based approaches have been used to mitigate distribution shift and facilitate sim-to-real transfer, we formally show that this approach can mitigate the effects of limited data when fitting the simulator and theoretically has significant regret gains compared to Plug-In selection in multi-armed bandits. We also empirically investigate the Uncertainty-Aware selection approach in deep RL experiments on robotic control tasks that involve selecting reward-shaping hyperparameters, and show that it leads to more reliable selection and improved online performance.
|
| 1097 |
Efficient Support Recovery of Mixtures of Sparse Linear Classifiers with Less Measurements
2609.32176
|
cs.LG
|
Xiaxin Li, Arya Mazumdar |
The support recovery problem in mixture of linear classifiers intends to identify which features actually matter when data is generated by a mixture of several linear decision rules. In particular, the aim is to recover the support (nonzero coordinates) of $l$...The support recovery problem in mixture of linear classifiers intends to identify which features actually matter when data is generated by a mixture of several linear decision rules. In particular, the aim is to recover the support (nonzero coordinates) of $l$ unknown $k$-sparse vectors from sign measurements. Each measurement is generated by selecting one of the $l$ vectors uniformly at random, and returning the sign of its inner product with a chosen measurement vector. In this paper, we propose adaptive and non-adaptive schemes that significantly improve upon prior results by reducing the number of measurements and achieving sublinear decoding time simultaneously. In particular, our adaptive constructions substantially reduce measurements compared to existing approaches, while also lowering decoding complexity from super-quadratic to sublinear in the ambient dimension. We further provide a non-adaptive scheme that improves previous measurement bounds while maintaining efficient decoding. Overall, our approach yields a more efficient trade-off between sample complexity and decoding time for support recovery in mixture models than previously known methods.
|
| 1098 |
Analytic-Walk Rotary Positional Encodings for Graphs
2609.32178
|
cs.LG
|
Jiaqing Xie, Yuxin Wang, Xipeng Qiu |
Rotary position encodings make attention sensitive to relative position, but extending them to graphs requires choosing how graph structure enters the rotation. Previous works assign each node a rotation from spectral coordinates, so the rotary factor between ...Rotary position encodings make attention sensitive to relative position, but extending them to graphs requires choosing how graph structure enters the rotation. Previous works assign each node a rotation from spectral coordinates, so the rotary factor between two nodes depends only on their endpoints and cannot distinguish the routes connecting them. We introduce \textit{Analytic-Walk Rotary Positional Encodings} (AW-RoPE), which place the rotations on edges and sum the transported features over all walks, so contributions along different routes can reinforce or cancel. An exact variant evaluates the complete sum by a differentiable linear solve, and a sparse variant truncates it at a finite depth. We prove forward and parameter-derivative truncation bounds at fixed inputs and parameters. Both variants act on projected queries and keys, and the sparse recurrence also augments message-passing networks. Across five synthetic tasks both variants reduce nRMSE by $15$--$58\%$ relative to the strongest baseline, and on real superpixel, peptide and OGB benchmarks the sparse recurrence attains the best mean on every dataset with Performer kernels and on twelve of thirteen datasets with GIN. Analysis shows that AW-RoPE can distinguish routes whose only cue is how two endpoints are connected, while node-wise rotary encodings cannot.
|
| 1099 |
Rethinking Cross-Channel Importance in Time-Series Forecasting
2609.32187
|
cs.LG
|
Yong-Hoon Choi, Kwang-Hyun Park, Youngjin Cho |
Cross-channel modeling is central to multivariate time-series forecasting, yet channels that are statistically related, predictively useful, and actually used by a trained forecaster are often treated as if they defined the same notion of importance. We show t...Cross-channel modeling is central to multivariate time-series forecasting, yet channels that are statistically related, predictively useful, and actually used by a trained forecaster are often treated as if they defined the same notion of importance. We show that they need not coincide. Cross-channel dependency structures change substantially across future offsets, and horizon-adaptive source selection improves a controlled Ridge predictor in 21 of 32 dataset--prediction-length conditions, with a mean gain of $5.16\%$. This selected-set signal also transfers to a matched nonlinear predictor. Yet imposing the same horizon-specific source logic on iTransformer yields only 11 of 20 wins and a mean gain of $0.208\%$, with little alignment between controlled and neural gains. Functional interventions further show that strong forecasters use cross-channel information, while their source-reliance rankings agree little with controlled utility or with one another across iTransformer, TimesNet, and a cross-channel TimeMixer. As a constructive consequence, bounded post-hoc support improves a frozen channel-independent forecaster in 12 of 16 dataset--horizon conditions, with a positive aggregate bootstrap interval. Cross-channel importance should therefore be interpreted relative to the forecasting mechanism and question that define it: related $\neq$ useful $\neq$ used.
|
| 1100 |
DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation
2609.32215
|
cs.LG
|
Dwipam Katariya, Thomas Caputo, Akshat Shreemali, Juan Manuel Origgi, Nikita Seleznev |
Transformers have redefined sequential recommendation by effectively modeling dynamic user behaviors and long-range dependencies. However, they remain inherently inefficient: standard architectures operate at a fixed rate, allocating comparable computation to ...Transformers have redefined sequential recommendation by effectively modeling dynamic user behaviors and long-range dependencies. However, they remain inherently inefficient: standard architectures operate at a fixed rate, allocating comparable computation to every item in a user's history regardless of its information content. This leads to prohibitive computational overhead on long sequences and increased sensitivity to behavioral noise. To address this, practitioners often resort to lossy sequence compression, staged modeling, or truncation. This limits the model's ability to leverage the full context of long histories during inference. Inspired by the recent success of Byte Latent Transformer, we propose DP-Rec, a dynamic latent patching architecture for recommendation. DP-Rec shifts from item-level modeling to patch-level modeling by segmenting interaction sequences using contrastive entropy surprise to identify informative behavioral boundaries. A lightweight patch encoder compresses these temporally contextualized segments into a reduced set of dynamic latent behavior vectors, which are then processed by a larger latent transformer and decoded for next-item prediction. Extensive experiments show that, under constrained computational budgets, DP-Rec scales effectively to long sequences and achieves a superior efficiency-accuracy trade-off over both non-compressed and fixed-size compression baselines.
|
| 1101 |
CompassPlay: Rewarding the Proposer for Where It Moves the Solver
2609.32228
|
cs.LG
|
Sophia Xiao Pu, Ximeng Sun, Jiang Liu, Jialian Wu, Emad Barsoum |
In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the pro...In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the proposer through gradient alignment. The reward favors tasks whose solver loss gradients align with those of reference tasks representing the target capabilities. It draws on a first-order approximation to learning progress and scores each eligible task without additional solver training. Our experiments show gains in performance and training efficiency. In coding self-play with Qwen2.5-Coder-7B, a small external reference set guides task generation. CompassPlay improves average accuracy over AZR's difficulty reward by 1.5 percentage points on in-domain coding and 2.7 on out-of-domain mathematics. In Lean4 theorem proving, CompassPlay matches the difficulty baseline's 150-iteration cumulative coverage with 40\% fewer GPU-hours.
|
| 1102 |
"Where Can I Trust You?": Boundary-Aware Evaluation of Surrogate Fidelity
2609.32230
|
cs.LG
|
Jackson Eshbaugh |
Surrogate models are commonly evaluated by how often they agree with their teacher model over an evaluation set. Local variation in this agreement is well known, but its structure and consequences are less clear. We ask whether disagreement is systematically c...Surrogate models are commonly evaluated by how often they agree with their teacher model over an evaluation set. Local variation in this agreement is well known, but its structure and consequences are less clear. We ask whether disagreement is systematically concentrated near the teacher's decision boundary and whether retaining that structure provides information beyond a single global score. Across several datasets and surrogate model classes, we find substantially lower fidelity near teacher decision boundaries under two different methods of identifying near-boundary examples. Moreover, conditioning agreement on confidence-defined regions improves prediction of teacher--surrogate agreement when evaluation-set composition changes, relative to the global score alone. Yet surrogates that agree equally well with the teacher both globally and near the decision boundary can respond very differently to changes selected using the surrogate itself. Finally, we show that independently trained deep teachers can agree on most predictions while identifying different examples as lying near their decision boundaries, complicating the use of those boundaries as stable reference regions for evaluating surrogates. Together, these results show that surrogate fidelity depends not only on how often a surrogate agrees with its teacher, but also on where that agreement holds and, for deep models, how stable the teacher's decision boundary is across training runs.
|
| 1103 |
Arithmetic Simplicity in Stochastic Gradient Methods
2609.32240
|
cs.LG
|
Bin Fu, Pengfei Gu, Jose Nunez, Fabian Vazquez |
A gradient descent method is arithmetically simple if the operations are limited to $+,-, \times$, and division $x/2^t$ with integer $t$. An arthmetically simple gradient method is easy to implement in chip design. We show how to transform AdaGrad, Adam, and A...A gradient descent method is arithmetically simple if the operations are limited to $+,-, \times$, and division $x/2^t$ with integer $t$. An arthmetically simple gradient method is easy to implement in chip design. We show how to transform AdaGrad, Adam, and AdamW into arithmetically simple. AdamW is based on the recursion $x_{t+1}=(1-\lambda\eta)x_t-\frac{\eta }{s}m_t$ and Adam is the special case of AdamW with $\lambda=0$. We transform them into a static case with $s=S(T)$, where $T$ is the number of iterations, and $S(T)$ is a fixed function. The convergence analysis is given for a static Adam, which is also arithmetically simple.
|
| 1104 |
Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes
2609.32248
|
cs.LG
|
Hongfu Gao, Songxin Zhang, Zejian Xie, Bingyi Jing, Zhou Wang |
Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks. However, variability in evaluation outcomes across runs may produce unsupported claims of model superiority on the ben...Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks. However, variability in evaluation outcomes across runs may produce unsupported claims of model superiority on the benchmark, a risk compounded by leaderboard updates. In this paper, we propose BB-EDGE (Benchmark-Weighted and Block-Factorized e-processes for Directed Graph Evaluation), a principled framework that represents an LLM leaderboard as a directed graph whose edges certify pairwise mean-performance advantages, with anytime-valid family-wise error rate (FWER) control. Concretely, for each direction, BB-EDGE constructs an empirical-Bernstein e-process by factorizing evidence over protocol-defined blocks and assigning stakes proportional to the corresponding block weights, then applies direct e-Holm across these $e$-processes to certify directional advantages as edges. Theoretically, we characterize weight-proportional linear stakes under heterogeneous benchmark-average nulls and prove anytime FWER control under arbitrary within-block and cross-pair dependence. BB-EDGE further supports anytime-valid Top-$k$ certification and simultaneous rank intervals. Extensive experiments on synthetic data and four real-world benchmarks demonstrate that BB-EDGE maintains anytime FWER control while achieving high efficiency.
|
| 1105 |
Representation Editing for Multimodal Test-Time Adaptation
2609.32263
|
cs.LG
|
Longfei Huang, Xiangyu Wu, Yang Yang |
Multimodal test-time adaptation (TTA) aims to adapt a pretrained multimodal model online to distribution shift across modalities using unlabeled test data, showing broad potential in real-world applications. However, existing methods primarily focus on adjusti...Multimodal test-time adaptation (TTA) aims to adapt a pretrained multimodal model online to distribution shift across modalities using unlabeled test data, showing broad potential in real-world applications. However, existing methods primarily focus on adjusting fused features to bridge the source-target gap, lacking explicit control over intermediate representation misalignment, which is a key driver of performance drop under distribution shift. In this work, we tackle this challenge from the perspective of representation engineering. Unlike previous TTA methods that update fusion weights in place, we propose FourIer Representation Editor (FIRE), a novel multimodal TTA approach that directly edits semantically rich intermediate representations. Specifically, we first adopt representation editors into each intermediate layer of the unimodal encoders, enabling layer-wise calibration of unimodal representations. To further enhance the diversity and stability of the low-rank editing subspaces, each representation editor performs frequency domain mixing via the fast Fourier transform to construct structured bases. Moreover, we introduce multi-level adaptation objectives to optimize these editors, jointly promoting cross-modal semantic alignment, source-target statistical alignment, and asymmetric prediction consistency. In this way, FIRE yields aligned unimodal representations for fusion and further improves prediction reliability. Extensive experiments on two widely used multimodal benchmarks under various corruption types demonstrate the superiority of FIRE over existing multimodal TTA methods.
|
| 1106 |
When Can Old Evaluations Certify a New Model? Label-Efficient Release Decisions under Evaluator Drift
2609.32267
|
cs.LG
|
Joyanta Jyoti Mondal, Mridul Banik, Md. Shifatul Ahsan Apurba, Md Masud Al Mahmud |
Releasing a model update requires certifying that its current-population risk stays below a threshold. Trusted labels are expensive, while a cheap evaluator, such as an LLM judge, scores every example. Reusing evaluator errors from earlier audits is tempting, ...Releasing a model update requires certifying that its current-population risk stays below a threshold. Trusted labels are expensive, while a cheap evaluator, such as an LLM judge, scores every example. Reusing evaluator errors from earlier audits is tempting, but when may such evidence replace current labels? It depends on the status of history. If the errors can change invisibly, no label-free test detects the change, and every valid, useful certifier must keep buying labels at a rate we characterize; if a bound on the change is assumed, label-free certification is valid at an explicit error cost. For the middle ground, where history is informative but untrusted, we propose \emph{portfolio vigilance}, a sequential certifier mixing a betting expert guided by history with one that learns only from current labels; history affects only how it bets, so validity holds for any history. The contribution is not prior-informed betting or expert mixtures, but separating history that may enter validity from history that may only guide label collection. In a canonical model, accurate history shortens decisions but never raises the evidence growth rate; stale history can destroy it. On held-out CIFAR-10N and DICES-990 data, portfolio vigilance needs 0.465 (95\% CI $[0.327,0.575]$) and 0.740 ($[0.618,0.877]$) times the labels of a matched prediction-powered monitor, with no observed false certification, and fewer labels on all six external blocks. Under corrupted advice it stays within 8.0\% of its better component, while trusting history alone costs up to 1.66 times as much. In post-confirmatory repeated-judge experiments on DICES-990 and ToxicChat, changing a fixed LLM judge's rubric moves its scores beyond run-to-run variation; the portfolio then needs 0.790 ($[0.667,0.909]$) and 0.631 ($[0.520,0.770]$) times the labels of the matched monitor, and fewer than trusting history alone.
|
| 1107 |
Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement
2609.32268
|
cs.LG
|
Hitoshi Iyatomi |
Standard autoencoder (AE) inference uses a single encoder-decoder pass, though the latent may not be optimal for each sample under a fixed decoder. We ask whether a trained AE can reveal information for improving its own reconstruction. Repeated application of...Standard autoencoder (AE) inference uses a single encoder-decoder pass, though the latent may not be optimal for each sample under a fixed decoder. We ask whether a trained AE can reveal information for improving its own reconstruction. Repeated application of a frozen AE to its reconstruction produces transient image- and latent-space trajectories, termed Self-Reconstruction Dynamics (SRD). Although this degrades fidelity in the AEs studied here, SRD contains sample-specific information for correcting the reconstruction. We propose SRD-guided Reconstruction Refinement (SRD-RR), which predicts a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. We also introduce MSE-recov, an MSE recovery ratio relative to an empirical decoder-optimized reference. Across six datasets, SRD-RR recovers 38.6% of the empirically recoverable MSE gap with one transition and 45.3% with two. A two-transition variant trained without direct access to original images, using an SRD-derived pseudo-target, achieves 40.7% recovery and a 1.74 dB average PSNR gain. Removing trajectory information reduces the gain, while cross-sample trajectory assignment causes severe degradation, confirming strong sample specificity. Nonlinear SRD-conditioned refinement consistently outperforms fixed and trained linear latent correction. On a pretrained DINOv2-based representation autoencoder (RAE) with substantially different latent dynamics, SRD conditioning again improves a matched trajectory-free predictor. However, pixel-MSE latent refinement reveals a strong mismatch between pixel fidelity and perceptual quality, while the SRD-derived pseudo-target mitigates this degradation. Overall, SRD is a useful sample-specific refinement signal, while the objective determines how it translates into pixel and perceptual quality.
|
| 1108 |
Certification Frontiers for Gaussian LoRA: Independent Priors, Posterior Risk, and Prediction-Preserving Balancing
2609.32271
|
cs.LG
|
Joyanta Jyoti Mondal, Ibne Farabi Shihab |
Post-hoc Bayesian fine-tuning places Gaussians around trained low-rank adapters, yet a calibrated posterior does not by itself yield a useful generalization certificate. Such a posterior admits an informative PAC-Bayes certificate only when both the loss of it...Post-hoc Bayesian fine-tuning places Gaussians around trained low-rank adapters, yet a calibrated posterior does not by itself yield a useful generalization certificate. Such a posterior admits an informative PAC-Bayes certificate only when both the loss of its sampled predictors and its KL divergence from an admissible prior are small. In this research, we characterize this certification frontier for Gaussian LoRA posteriors and separate three interventions: changing the prior, changing the stochastic predictor, and changing only how its complexity is counted. First, an exact isotropic KL envelope eliminates the prior scale and yields a width threshold that excludes posterior widths before sampling, while a zero-KL floor identifies targets that no complexity reduction can reach at a measured risk bound. Second, we minimize KL in closed form over the full $GL(r)$ symmetry of the low-rank factors, leaving every sampled adapter product unchanged, and derive the noncentral objective required when the prior center is trained on an independent split. On a small-pool RoBERTa audit of 567 configurations, the recorded 64-draw summaries imply a certificate floor of $0.7298$ even with zero KL and Chernoff accounting, so reducing complexity alone cannot certify these posteriors at the recorded budgets. For a stable posterior in a controlled Gaussian-factor task, Chernoff accounting certifies risk below $0.1$ on 20 of 20 datasets with 1024 draws, whereas Hoeffding certifies none. On deliberately deformed synthetic rank-four factors, matrix balancing reduces KL by $29.3\%$ on average beyond scalar balancing without changing any prediction. Numerical split-prior scenarios make the remaining risk and complexity budgets explicit.
|
| 1109 |
Learning Through Game: Skewed Transfer of Tabular Knowledge to Strengthen Image Model
2609.32272
|
cs.LG
|
Longfei Huang, Shangdong Yang, Yang Yang |
Multimodal tabular-image learning is gaining growing attention, yet it faces challenges due to tabular data unavailable at test time. A practical solution involves transferring tabular knowledge to images during training to enhance the performance of image mod...Multimodal tabular-image learning is gaining growing attention, yet it faces challenges due to tabular data unavailable at test time. A practical solution involves transferring tabular knowledge to images during training to enhance the performance of image models at inference. However, the overlooked yet important challenges lie in the modality imbalance between images and tables, as well as their asymmetric modality relationship in cross-modal transfer, which limits the auxiliary role of tabular data. To address these issues, we propose Skewed Knowledge Transfer (SKT), which asymmetrically transfers tabular knowledge to improve the image model by adaptive integration of modality gradients in a shared parameter space. Specifically, we first introduce a multimodal shared head, which allows the model to benefit from cross-modal structure without adding additional parameters. We then design a two-step Nash Bargaining strategy to effectively leverage tabular gradients. In the first step, SKT seeks a point of modality balance and uses preference awareness in the second step to steer combined gradients toward image-beneficial directions. Furthermore, we theoretically analyze the Pareto improvement and convergence of SKT. To this end, tabular knowledge is explicitly transferred to enhance image models. Empirical experiments on widely used tabular-image datasets reveal that SKT consistently improves image unimodal performance by using tabular data as auxiliary information.
|
| 1110 |
Topology-Adaptive Hyperbolic Graph Attention Networks Guided by the Hyperbolic Sombor Index
2609.32275
|
cs.LG
|
Haifang Cao, Boan Tao, Xiyuan Gao, Timing Li, Yu Wang |
Hyperbolic geometry has emerged as a principled space for representing hierarchical graphs. However, existing hyperbolic graph neural networks typically rely on shared curvature configurations and feature-driven attention, failing to explicitly exploit local h...Hyperbolic geometry has emerged as a principled space for representing hierarchical graphs. However, existing hyperbolic graph neural networks typically rely on shared curvature configurations and feature-driven attention, failing to explicitly exploit local hierarchical topological patterns. To bridge this gap, we introduce the Hyperbolic Sombor Index (HSO) as a lightweight structural prior for capturing hierarchy-indicative degree stratification. Building on this, we propose \textbf{HSO-GAT}, a topology-adaptive hyperbolic graph attention network that unifies geometric adaptation and message propagation. Specifically, it comprises two complementary modules: HSO-Guided Local Curvature Adaptation, which performs adaptive node-wise geometric scaling from aggregated node-level HSO signals, and HSO-Gated Hyperbolic Graph Attention, which enables structure-aware message passing through feature-conditioned gating. Theoretically, we establish the monotonic sensitivity of edge-level HSO to degree imbalance and analyze the validity and radial scaling properties of node-adaptive hyperbolic mappings. Extensive experiments on eight benchmark datasets demonstrate that HSO-GAT consistently achieves state-of-the-art performance in both node classification and link prediction tasks.
|
| 1111 |
HyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation Modeling
2609.32276
|
cs.LG
|
Peiyu Zhang (University of Southern California), Heng Ping (University of Southern California), Nikos Kanakaris (Amazon Web Services), Yucheng Zhao (University of Tennessee, Knoxville) |
Multi-label classification (MLC) requires predicting multiple relevant labels for each instance, where a central challenge is modeling complex label dependencies arising from co-occurrence patterns. Existing approaches are limited in capturing high-order label...Multi-label classification (MLC) requires predicting multiple relevant labels for each instance, where a central challenge is modeling complex label dependencies arising from co-occurrence patterns. Existing approaches are limited in capturing high-order label correlations, relying on implicit learning through contrastive objectives or pairwise attention mechanisms without structural guidance. We propose HyperLabel, an encoder-decoder framework that explicitly models label dependencies through hypergraph neural networks. Our contributions are twofold: (i) We construct a label hypergraph where sample-defined hyperedges naturally encode multi-way co-occurrence patterns, providing explicit structural prior knowledge that captures relationships beyond pairwise interactions. (ii) We propose a unified cross-modal learning approach where HGNN+ performs bidirectional message passing to integrate feature information with label structure, and a shared cross-attention decoder processes both modalities through complementary learning objectives. Extensive experiments on seven benchmark datasets demonstrate that HyperLabel achieves state-of-the-art performance, with particularly significant improvements on macro-F1 scores (+10.3% on Delicious, +8.2% on Bibtex), validating that explicit hypergraph structure effectively captures complex label relationships. The code is available at https://github.com/iZHpy/Multi-label_hypergraph .
|
| 1112 |
SIMANF: Sample Free Learning of Unnormalized Distributions via Simulated Annealing in Normalizing Flows
2609.32279
|
cs.LG
|
Vikas Kanaujia |
Efficiently learning and sampling from high dimensional, multimodal unnormalized distributions without target samples remains a challenging problem. Although normalizing flows can generate samples efficiently, training based on the reverse KL divergence using ...Efficiently learning and sampling from high dimensional, multimodal unnormalized distributions without target samples remains a challenging problem. Although normalizing flows can generate samples efficiently, training based on the reverse KL divergence using only the unnormalized target density may suffer from mode collapse. We introduce SIMANF, a sample free framework that integrates simulated annealing with normalizing flows. SIMANF progressively transforms the target distribution from a smooth initial form to the original target distribution and trains the flow sequentially across these stages. By transferring the learned representation between stages, the method promotes mode coverage while progressively capturing finer features of the target distribution. Following annealing, a final refinement stage combines the reverse KL divergence with an importance weighted forward KL objective using samples generated by the flow. SIMANF requires no target samples during training and uses only the unnormalized density. We demonstrate its effectiveness on Many-Well distributions and high dimensional Scalar Phi4 lattice field theory distribution.
|
| 1113 |
A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents
2609.32288
|
cs.LG
|
Indranil Halder, Rastri Dey, Cengiz Pehlevan |
Pre-training data poisoning of large language models is usually studied using targeted backdoors and their survival through safety post-training, which leaves open a more basic question: how does a model's clean data performance degrade as the poison rate $\va...Pre-training data poisoning of large language models is usually studied using targeted backdoors and their survival through safety post-training, which leaves open a more basic question: how does a model's clean data performance degrade as the poison rate $\varepsilon$ grows? Motivated by our controlled pre-training runs of OLMo-style models, in which the relative clean data validation perplexity increase $\Delta$ between poisoned and clean models matched in architecture, token budget, and optimization schedule is well fit by a power law $\Delta\approx C\varepsilon^{a}$ with a non-integer exponent, we ask what such a law requires theoretically. We first prove an analyticity barrier: whenever the contaminated objective depends analytically on $\varepsilon$ around a nondegenerate clean data optimum, $\Delta$ is generically quadratic in $\varepsilon$, so a generic non-integer exponent is a signature of genuinely singular structure. We then supply that structure in solvable truncated ridge regression with heavy-tailed covariates, controlled by $q_\star$, and a label-shift poisoning. Our central result is that the excess risk scaling exponent depending on the order of limits: in the higher dimensional proportional regime it is $\epsilon^{q_\star/(q_\star+2)}$, whereas taking the ample-data limit first gives $\epsilon^{2-2/q_\star}$, and the limits do not commute. We confirm this prediction through several numerical simulations. Finally, we argue that finite training time plays the role of truncation on the curvature spectrum in local LLM pre-training, deriving the observed scaling law under heavy tailed inverse curvature spectrum as a modeling hypothesis.
|
| 1114 |
A Journey to the Edge of Stability
2609.32290
|
cs.LG
|
Jaerin Lee, Kyoung Mu Lee |
It has recently been found that deep learning often occurs at the "edge of stability (EoS)," where the maximum Hessian eigenvalue of the model is stabilized at a value reciprocal to the learning rate. However, what happens before we reach that regime? We fix a...It has recently been found that deep learning often occurs at the "edge of stability (EoS)," where the maximum Hessian eigenvalue of the model is stabilized at a value reciprocal to the learning rate. However, what happens before we reach that regime? We fix a deep learning problem and vary first order optimization methods with dense learning rate sweeps. We then track the characterizing quantities of a learning trajectory: the loss, the sharpness, and the alignment between consecutive gradients. To our surprise, if we scale the learning rate by the dc gain of the optimizer, these traces from the sweeps from different optimizers almost perfectly overlap across a large range of learning rates. The dc-normalized optimizers have another role that only becomes apparent in high learning rates: they select when the sharpness value detaches from this universal curve and enters the edge of stability. Upon this discovery, we specify three distinct regimes with respect to the dc-adjusted learning rate: the low-LR regime where the trajectory is nearly insensitive to the optimizer, the high-LR regime, where the optimizer governs the sharpness according to the EoS reciprocal rule, and the in-between mid-LR regime where so-called progressive sharpening originates independently of the optimizer. This distinguishes the role of the optimizer, the learning rate, and the model in shaping the learning progress.
|
| 1115 |
FUND: Density Flow for Sampling Unnormalised Distributions
2609.32296
|
cs.LG
|
Vikas Kanaujia, Vipul Arora |
Efficient sampling from Boltzmann distributions is central to modelling complex physical systems. Markov Chain Monte Carlo (MCMC) methods suffer from critical slowing down, high autocorrelation, and poor mode-mixing, limiting their scalability. Recent advances...Efficient sampling from Boltzmann distributions is central to modelling complex physical systems. Markov Chain Monte Carlo (MCMC) methods suffer from critical slowing down, high autocorrelation, and poor mode-mixing, limiting their scalability. Recent advances, like Boltzmann Generators, offer a promising alternative but remain constrained by costly MCMC-based training, inefficient sampling, and poor ergodicity. We introduce an algorithm for learning Boltzmann distributions that does not require any true samples for training. Our approach draws inspiration from flow matching but departs fundamentally from sample-trajectory matching to distribution-trajectory matching. The algorithm iteratively reshapes the target distribution, using model generated samples to guide learning and ensure comprehensive mode coverage. We validate our method on standard benchmarks, including a 2D Gaussian mixture, Many-Well distributions, and high-dimensional scalar $\phi^4$ theory. The proposed approach not only improves sampling performance and accuracy over traditional MCMC and flow-based baselines but also establishes a new method for sample-free learning of complex physical distributions.
|
| 1116 |
Refresh or Realize? Compute Allocation in Drifting Models
2609.32298
|
cs.LG
|
Sipeng Chen, Xu Zheng, Shibo Li |
Drifting Models train a one-step generator by recomputing a finite-sample drift field at every iteration and taking an optimizer step toward the drifted target. The field says how generated samples should move, but the step is taken in parameters shared by all...Drifting Models train a one-step generator by recomputing a finite-sample drift field at every iteration and taking an optimizer step toward the drifted target. The field says how generated samples should move, but the step is taken in parameters shared by all samples, so the motion the network actually makes need not match the motion it was given. This leaves a basic training question open: should extra compute go into fitting the current target more closely, or into recomputing the field? We study it on ImageNet 256x256. Holding the target fixed for k optimizer steps and measuring the realized displacement, we find that deeper fitting does bring the network closer to the frozen target, and that the number of steps needed before it makes any net progress drops from about sixteen early in training to one later on. When the extra steps come for free, k=2 also lowers FID. Once they are paid for, the result flips: at approximately matched measured wall-clock, spending the budget on fresh fields gives lower FID than deeper fitting, on both training seeds. The target itself shows why a fresh field is worth so much. Redrawing the finite support rotates its direction far more than a parameter update does (cosine ~0.3-0.6 against ~0.95), and a correction that is optimal in field space is not reliably better in FID than a parameter-free one. For Drifting, fitting each target well and spending compute well are different goals.
|
| 1117 |
When Does Synergy Help Active Feature Acquisition? A PID-Based Study
2609.32301
|
cs.LG
|
Jie Li, Maruf A. Dhali, Hjalmar R. Bouma |
Active feature acquisition (AFA) sequentially selects informative features under budget constraints. However, existing policies rarely distinguish whether information contributed by interacting features is redundant, unique, or synergistic. We introduce SynAFA...Active feature acquisition (AFA) sequentially selects informative features under budget constraints. However, existing policies rarely distinguish whether information contributed by interacting features is redundant, unique, or synergistic. We introduce SynAFA, a state-dependent AFA policy that combines pairwise joint information and conditional information, with Partial Information Decomposition (PID) characterizing its information structure. Across five tabular datasets and MNIST-loop, performance is heterogeneous. SynAFA shows its strongest gains at low budgets on PhysioNet, where synergistic and redundant feature pairs are supported by permutation tests, but its advantage diminishes as budgets increase and does not depend on pair proposals. SynAFA performs significantly worse than nearly all baselines on MiniBooNE, and than CAE on MNIST-loop. Controlled synthetic experiments further show that, within budget, SynAFA's advantage rises as joint information becomes more synergy-dominated, including when total joint information is held approximately constant. Further analyses show that the budget-dependent erosion is not resolved by non-greedy local search, which improves a diagnostic set-level objective but leaves predictive performance no better, often significantly worse, and that improvements in this objective are weakly aligned with the fixed classifier's predictive utility. These findings characterize when pairwise synergy can benefit AFA while exposing a persistent challenge in translating local information into effective acquisition objectives.
|
| 1118 |
On the Capability and Limitation of Hard Prompt
2609.32302
|
cs.LG
|
Lijia Yu, Shuaitong Liu, Gaojie Jin, Xinyu Li, Xiao-Shan Gao |
Prompt engineering has become an indispensable tool for using large language models (LLMs), turning LLMs into task-specific experts without changing their weights. Despite notable theoretical advances in prompt engineering, the theory for the more practical ha...Prompt engineering has become an indispensable tool for using large language models (LLMs), turning LLMs into task-specific experts without changing their weights. Despite notable theoretical advances in prompt engineering, the theory for the more practical hard or discrete prompts is largely open. In this paper, we try to fill this gap either by providing a complete solution or by making substantial progress on the three core theoretical questions regarding hard prompts. First, we show that determining the existence of a hard prompt for a transformer to solve a downstream task is NP-complete and that finding an optimal hard prompt is NP-hard, which is the first computational complexity result for hard prompting, as far as we know. Second, we show that, unlike soft or continuous prompts, hard prompts have essential limitations: hard prompts are not complete; short hard prompts do not significantly enhance the ability of transformers; and long hard prompts exhibit the "prompt dominating answer phenomenon," meaning that, with high probability, the same answer is given for all queries of the same length. On the other hand, linear hard prompts do not have the limitations of short or long prompts. Third, we provide a tight bound on the size of the task in terms of the prompt length for the performance of prompts on the finite task to generalize to the entire data distribution, leading to a necessary and sufficient condition for generalizability. This is the first result on generalization for prompting, as far as we know. Our findings not only offer the first theoretical insights into hard prompts but also provide provably reliable practical guidance for real-world LLM usage.
|
| 1119 |
Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction Obstruction
2609.32319
|
cs.LG
|
Kingsuk Maitra, Shagun Sood Morteza Hosseini, Suman Gunnala, Vikram Gupta |
We circumvent the Sanford-Hsu-Telgarsky (SHT) single-layer induction obstruction within a linear, one-step, causal, bilinear, symplectically consistent design class on the post-RoPE substrate, by lifting attention onto a symplectic phase space, mirroring Haire...We circumvent the Sanford-Hsu-Telgarsky (SHT) single-layer induction obstruction within a linear, one-step, causal, bilinear, symplectically consistent design class on the post-RoPE substrate, by lifting attention onto a symplectic phase space, mirroring Hairer's lift of Stormer-Verlet. The lift exits the premise of the SHT counting argument rather than the bound itself. The reframing exhibits the obstruction as a filter-order gap: a one-layer bilinear score realises a $z$-transform of joint order $(0,0)$, whereas the induction discriminator requires key-side order $\geq 1$. Applying the symplectic upper shear $M_\gamma:(q,p)\mapsto(q+\gamma p,p)$ to the post-RoPE query and key streams closes it. We prove this lift is unique within the factorised subclass, exactly symplectic at operator level, and requires post-RoPE placement; and in an explicit $T_4$-only Gaussian reduction we derive a closed-form two-branch induction phase transition, held out at $r=0.9876$ with zero fitted parameters. That law is an analytically solvable limit, not a robust prediction: restoring the $T_3$ channel moves $\gamma_c$ at $d_k=64$ from 1.030 to 0.569 and removes the crossover. Deployability follows by exact derivation: the KV cache is unchanged, prefix reuse and speculative decoding are preserved, overhead is $6d$ FLOPs per token per layer, INT8 headroom grows by at most $\log_2(1+2\gamma)$ bits, fused kernels are unmodified, and no parameters are added. At 91.3M parameters a supercritical sweep locates an emergence band: induction forms 3/3 seeds at $\gamma=0.80$ in a mean of 717 steps, against 2/3 seeds and 2700 steps at $\gamma=0$. Adverse results are reported as directly: a key-only half-lift reaches 0.949 against 0.811 for the symmetric operator, so if induction accuracy is the objective, the half-lift is the better construction. Forty-one notebooks and result files ship as ancillary material.
|
| 1120 |
Superposed Inference for Hyperdimensional Computing
2609.32320
|
cs.LG
|
Quanling Zhao, Nilesh Prasad Pandey, Ye Tian, Tajana Rosing |
Hyperdimensional computing (HDC) is attractive for efficient and robust learning, but conventional inference still encodes every query independently, repeatedly paying the cost of high-dimensional projection. We introduce SupHDC, a new inference paradigm that ...Hyperdimensional computing (HDC) is attractive for efficient and robust learning, but conventional inference still encodes every query independently, repeatedly paying the cost of high-dimensional projection. We introduce SupHDC, a new inference paradigm that processes multiple queries through a shared encoding computation. SupHDC assigns lightweight random slot keys, superposes the keyed queries before encoding, and uses slot-specific classifiers to recover their individual predictions. A random-feature kernel view explains why exact recovery of each hypervector is unnecessary: inference only needs to preserve the class evidence that determines the prediction. Across ten datasets, SupHDC achieves 1.39x analytical speedup with no average accuracy loss, and up to 2.08x speedup with only a 2.67 percentage-point mean accuracy loss. On a Raspberry Pi~5, it delivers 2.01x measured wall-clock speedup with a 2.26 percentage-point loss in mean prediction accuracy. SupHDC shows that high-dimensional redundancy can be used not only for robustness, but also as capacity for shared inference.
|
| 1121 |
Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality
2609.32322
|
cs.LG
|
Linhao Wang, Yiyan Fan, Dongjin Huang |
World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performan...World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performance when their errors occur on different state dimensions. We introduce Decision-Relevant Prediction Error (DRPE), which measures prediction error on the state dimensions that affect decisions. We also develop an iso-error evaluation protocol that varies error allocation while keeping total error fixed. In a factored gridworld with known state relevance and a standardized planner, we evaluate 55 controlled and learned models across different error levels and allocations. Total prediction error is weakly related to planning success (Spearman $\rho=-0.25$), whereas DRPE is strongly predictive ($\rho=-0.84$; $-0.98$ within the controlled family). Models with only a 1\% difference in total error can differ by 60 percentage points in planning success (97\% vs 37\%). The relevant error also depends on the task, with model rankings reversing across tasks at the same total error. Deeper imagination further amplifies decision-relevant errors, while learned models exhibit systematic bias on rare but decision-critical events. We formalize sufficient conditions under which DRPE correctly ranks models and total prediction error cannot.
|
| 1122 |
Active Feature Acquisition With Incomplete Training Data
2609.32325
|
cs.LG
|
Reza Rezvan, Valter Sch\"utz, Han Wu, Linus Aronsson, Morteza Haghir Chehreghani |
In many prediction tasks, acquiring all features can be a prohibitively expensive or outright impossible task. Further, in many cases a static subset of features may not be enough to solve the problem sufficiently across various instances. Active Feature Acqui...In many prediction tasks, acquiring all features can be a prohibitively expensive or outright impossible task. Further, in many cases a static subset of features may not be enough to solve the problem sufficiently across various instances. Active Feature Acquisition (AFA) addresses these problems by formalizing the trade-off between feature cost and predictive performance during sequential feature selection. However, prior AFA work largely assumes access to complete training data, an assumption that is often violated in practice. Here we study AFA with Incomplete Training Data (AFA-ITD), showing that under missing completely at random (MCAR) data, one-step acquisition values remain unchanged, whereas multi-step values can decrease. We analyze three approaches to learning from incomplete data: aliasing, filtering, and generative restoration. We show that filtering can require a number of training instances scaling exponentially with the dimension, whereas generative restoration scales exponentially with the acquisition budget. We empirically test our theory on a controlled experiment and across common AFA datasets and find that missingness mainly damages methods that exploit multi-step acquisitions and that generative restoration is able to recover lost performance in many experiments. Code is available at https://github.com/Linusaronsson/AFA-Benchmark/tree/missing-data.
|
| 1123 |
Continual Data Unlearning in Diffusion Models via Transition-based Regularization
2609.32328
|
cs.LG
|
Sunbeom Jeong, Sehwan Kim, Sangwoo Hong, Jungwoo Lee |
Data unlearning in diffusion models aims to remove the influence of specific training examples without suppressing the broader concepts they represent. However, when deletion requests arrive sequentially, updates for new requests can degrade generative utility...Data unlearning in diffusion models aims to remove the influence of specific training examples without suppressing the broader concepts they represent. However, when deletion requests arrive sequentially, updates for new requests can degrade generative utility and undermine earlier deletions. We propose a continual data unlearning framework that uses completed deletion transitions as directional references to regularize future updates. For each request, we record changes in denoiser responses on the same fixed noisy inputs before and after unlearning. Rather than matching full post-deletion responses, we apply a one-sided penalty that discourages reversal along the recorded directions relative to the post-deletion references, while leaving orthogonal response changes and progress beyond these references unpenalized. To keep storage independent of the number of requests, we maintain a fixed-capacity bank of representative transition records. Records are selected based on the local sensitivity of progress along their recorded directions to parameter updates, allowing them to be retained even when their penalties are inactive. Empirical evaluations show that the proposed framework achieves a better balance between deletion persistence and generative utility than existing unlearning baselines as requests accumulate, using only a small transition memory.
|
| 1124 |
When the Merge Coefficient Stops Mattering: Proximity Regularized Merging for Continual LoRA Adaptation
2609.32332
|
cs.LG
|
Yixuan Liu, Yuhao Sun, Sen Song, Jin Li |
Rehearsal-free continual learning with parameter-efficient adapters can be cast as a sequence of task-vector write-in operations: for each new task, a low-rank adapter is learned and merged into a running model. We propose Proximity Regularized Merging (PRM), ...Rehearsal-free continual learning with parameter-efficient adapters can be cast as a sequence of task-vector write-in operations: for each new task, a low-rank adapter is learned and merged into a running model. We propose Proximity Regularized Merging (PRM), a minimal modification to sequential LoRA merging that adds a proximal penalty during task-vector training without changing the subsequent write-in rule. PRM acts as a robust task-vector regularizer: in the reported Base->+Prox diagnostics, it improves AAA across multiple write-in rules, backbones, and class-incremental settings, while its fixed-coefficient variant remains competitive with strong coefficient-based baselines. Mechanistically, matched-prefix norm controls and proximal-strength sweeps show that proximal training shrinks the task-vector radius, lowers Fisher-weighted interference, broadens the coefficient plateau, and exposes a stability-plasticity trade-off. Together, these results suggest that the effectiveness of sequential LoRA merging depends not only on how much of a task vector is written in, but also on whether the task vector itself has been trained to be mergeable.
|
| 1125 |
A Comparative Analysis of Attention versus State-Space Models for In-Context Learning
2609.32341
|
cs.LG
|
Enes Arda, Semih Cayci, Atilla Eryilmaz |
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develo...Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.
|
| 1126 |
Analog-Friendly Predictive Coding without Activation Derivatives
2609.32350
|
cs.LG
|
Francesco Innocenti |
Predictive coding (PC) is a local, energy-based alternative to backpropagation (BP) whose iterative inference dynamics make it attractive for implementation on analog hardware. However, standard nonlinear PC requires evaluating the derivative of the activation...Predictive coding (PC) is a local, energy-based alternative to backpropagation (BP) whose iterative inference dynamics make it attractive for implementation on analog hardware. However, standard nonlinear PC requires evaluating the derivative of the activation function during both inference and learning, which can be difficult to realise physically. Here, we introduce \textit{activation-matched Bregman PC}, replacing standard squared-error energies with Bregman divergences matched to the activation function. This formulation eliminates activation derivatives and, when combined with inference via mirror descent, yields local inference and learning rules requiring only weighted sums, local prediction errors, state integration, and the activation function. In digital experiments, Bregman PC performs comparably to standard PC and BP on classification and generative tasks, while preserving characteristic learning dynamics of PC and its convergence to BP under stable large-model parameterisations. These results provide a more analog-friendly formulation of nonlinear PC while retaining its key computational properties.
|
| 1127 |
Editable Map-Conditioned Trajectory Generation for Human Mobility Simulation
2609.32360
|
cs.LG
|
Takayuki Mizuno, Shouji Fujimoto, Mikito Hiruki, Atushi Ishikawa |
Geospatial simulation of infrastructure interventions requires mobility generators that respond directly to edited maps, yet many data-driven generators do not expose the map as an editable condition. We formulate this task as map-conditioned autoregressive ge...Geospatial simulation of infrastructure interventions requires mobility generators that respond directly to edited maps, yet many data-driven generators do not expose the map as an editable condition. We formulate this task as map-conditioned autoregressive generation of human mobility: a road raster conditions a decoder that emits nominal 31.25 m mesh-cell tokens at one-minute intervals. The mesh-local vocabulary supports held-out and locally edited maps without retraining or vocabulary changes. We instantiate a ResNet-50 visual-prefix configuration and a Vision Transformer (ViT) cross-attention configuration, trained from scratch on 87,400 smartphone-derived trajectories from 874 meshes in Ishikawa Prefecture, Japan; 219 meshes are held out. We evaluate map sensitivity by comparing correct-map and within-split shuffled-map generations with held-out real trajectories. On the 110-mesh test split, for the ResNet-50 configuration, correct-map generations are closer than shuffled-map generations on 60% of meshes under Hausdorff-based energy distance (p = 0.021), while DTW is directional but inconclusive (57%, p = 0.074); correlation with real density is 0.38 with the correct map versus 0.01 with shuffled maps. The ViT configuration shows weaker trajectory-level sensitivity and smaller density gains. An illustrative bridge-removal edit changes generated continuations without retraining. Together, these results support the feasibility of editable-map human-mobility simulation.
|
| 1128 |
DiffPTS: Rethinking Diffusion ELBO for Probabilistic Time Series Forecasting
2609.32363
|
cs.LG
|
Weiwei Ye, Dongyuan Li, Hangchen Liu, Haotong Jiang, Yoshihide Sekimoto |
Probabilistic time series forecasting requires modeling and predicting complex and time-varying distributions. Recently, Denoising Diffusion Probabilistic Model (DDPM)-based approaches have shown promise by equipping the dif- fusion process with pretrained mea...Probabilistic time series forecasting requires modeling and predicting complex and time-varying distributions. Recently, Denoising Diffusion Probabilistic Model (DDPM)-based approaches have shown promise by equipping the dif- fusion process with pretrained mean and variance estimators to accommodate distributional shift. However, these methods typically follow the standard DDPM framework and consider only partial components of the evidence lower bound (ELBO), treating the training of estimators as designed regression tasks separate from the variational inference framework. To address this, we rethink the ELBO under the Location-Scale Noise Model (LSNM) and find that it naturally induces a Gaussian negative log likelihood objective for the estimators and inherently defines a joint training objective that unifies recent diffusion paradigms for probabilistic forecasting. Building on this principled ELBO reformulation, we propose Diff- PTS, a general framework that enables end-to-end optimization of all components within the ELBO. Across multiple benchmarks, DiffPTS consistently outperforms recent models, achieving state-of-the-art performance with an average CRPS/MSE reduction of over 14.53%/16.55% compared to existing diffusion-based methods. The code is available at https://github.com/wwy155/DiffPTS.
|
| 1129 |
Graph Memory: Spectral Associative Memory via Dirichlet Energy
2609.32365
|
cs.LG
|
Zhaoyang Shi |
Dense associative memories have traditionally focused on storing and retrieving vector-valued patterns. Many modern machine learning problems, however, are naturally graph-structured, requiring memory mechanisms for relational patterns, graph diffusion geometr...Dense associative memories have traditionally focused on storing and retrieving vector-valued patterns. Many modern machine learning problems, however, are naturally graph-structured, requiring memory mechanisms for relational patterns, graph diffusion geometries, community structures, and graph-based inductive biases. We propose a spectral dense associative memory for storage and retrieval of graph data, extending the classical vector-valued memories. Retrieval is performed through a log-sum-exp energy induced by Dirichlet energy with spectral norm distances, producing a softmax-weighted average of the stored Laplacians that remains a valid graph Laplacian. We prove exponential storage capacity and exponentially decaying retrieval error. Beyond graph retrieval, we establish theoretical guarantees for spectral quantities central to graph learning, including eigenvalues, eigenspaces, and diffusion operators. Experiments on synthetic graph data, real-world airline network, protein conformation data and wearable sensor data demonstrate robust graph retrieval while preserving the graph geometry of the data. Our framework provides a new associative memory paradigm for graph-structured data and bridges dense associative memory with modern graph learning and generative AI.
|
| 1130 |
STRIDE: State-Transition Representation via Increment Dynamics and Evolution
2609.32367
|
cs.LG
|
Yuchen Xiong, Siming Huang, Jianfeng Sun |
We introduce STRIDE (State-Transition Representation via Increment Dynamics and Evolution), which defines states through derivative fingerprints and learns local functions for state transitions (qpairs), recasting continuous forecasting as transition predictio...We introduce STRIDE (State-Transition Representation via Increment Dynamics and Evolution), which defines states through derivative fingerprints and learns local functions for state transitions (qpairs), recasting continuous forecasting as transition prediction. Trailing convolution windows estimate local joint-transition frequencies, whose lagged differences form a high-dimensional increment trajectory. Proper orthogonal decomposition (POD) gives coordinate paths, jointly forecast by sparse dynamics with memory. Recombining their forecasts and applying history-anchored inversion recovers future transition distributions; sampled state paths select local functions to generate continuous forecasts. Three-seed experiments compare STRIDE against fourteen baselines across nine benchmark families. Five independent Markov and hidden-state baselines cover all 230 evaluated tasks, with 220 complete whole-horizon pairs. Against these comparators, system-weighted late-Energy win fractions range from 73.6% to 78.5% on 190 multi-step pairs; against DLinear, the fraction is 77.3% on 72 paired multi-step tasks. Matched controls examine the intermediate representation. On 64 independently initialized Aizawa trajectories, late-Energy reductions against four matched controls range from approximately 24% to 56%, with all four prespecified contrasts passing Holm correction. These results connect transition-statistic prediction to continuous probabilistic forecasting, with substantial long-horizon gains in the matched Aizawa study.
|
| 1131 |
Measurement Boundaries in LLM Financial Agent Evaluation: Fixed-Tape Execution and Multi-Defect Auditing
2609.32379
|
cs.LG
|
Weicheng Xue |
What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on t...What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on three synthetic settings that share one 24-day upward phase mixes the execution rule with fresh model responses and portfolio feedback: the parsed decision paths agree in only $19.8\%$ of $450$ pairs. Replaying each stored response tape through both execution destinations gives a narrower result. Conditional on those responses, stressed execution changes total return by $-0.0170$ (95\% interval $[-0.0230,-0.0117]$), or $10.4\%$ of the idealized baseline, and ten seed clusters do not resolve the model ranking. Study~B corrects an incomplete answer key and replaces legacy tasks with matched zero-, one-, and two-defect tasks under an explicit multi-label prompt. The drop in target violation recall from one to two defects is positive in five of six combinations of auditor and source (median $0.267$), with three surviving Holm correction. Yet the auditor that includes both target labels most often has micro-precision $0.149$, emits findings on $98/100$ zero-defect tasks, and returns the exact dual-defect set in only $21/100$ cases. Target recall by itself therefore gives a poor account of audit quality on this construction. The studies address different limits: what an execution comparison estimates, and what target recall captures. Together, they show how fixed conditions and diagnostic controls bound the claims a score can support.
|
| 1132 |
Attribution Without a Second Pass: Inline Per-Sample Gradient Provenance at ~1% Overhead
2609.32380
|
cs.LG
|
Amit Nautiyal |
Data attribution methods used in practice (TRAK, LoGRA, EK-FAC) are post-hoc: after training they make a second pass over the training set to recompute per-sample gradients, repeated per checkpoint when ensembled. Traceprop avoids that pass by recording projec...Data attribution methods used in practice (TRAK, LoGRA, EK-FAC) are post-hoc: after training they make a second pass over the training set to recompute per-sample gradients, repeated per checkpoint when ensembled. Traceprop avoids that pass by recording projected per-sample gradients inline on the training backward pass. A Kronecker-factored sketch scales from a single tracked layer to every layer without materializing a dense projection matrix. On LoRA fine-tunes of GPT-2 and Pythia models up to 2.8B on one NVIDIA L4, inline logging costs 0.30-1.08% of wall-clock time at last-block scope and stays under 1% (0.79%) even when tracking every layer of Pythia-1B. Against LogIX, the closest inline-capable competitor, the factored sketch is 2.0-4.1x cheaper at equal storage, a gap that grows with tracked scope and is significant at every scope tested, while matching or exceeding LogIX's attribution quality at matched storage. Building the attribution-ready store inline is 60-242x cheaper than one post-hoc pass and 301-1211x cheaper than a five-checkpoint TRAK ensemble, with recorded gradients matching autograd exactly. Because each stored gradient carries source-file lineage, the same pass also produces EU AI Act Article 26 audit trails.
|
| 1133 |
TimeES: Probabilistic and Deterministic Time Series Forecasting via Evolutionary Spectra
2609.32384
|
cs.LG
|
Weiwei Ye, Renhe Jiang, Hangchen Liu, Dongyuan Li, Yoshihide Sekimoto |
Real-world time series are inherently non-stationary, with trends, periodic patterns, and uncertainty evolving over time. While the Fourier domain offers a natural lens to model time series, current deep learning approaches do not explicitly model evolution an...Real-world time series are inherently non-stationary, with trends, periodic patterns, and uncertainty evolving over time. While the Fourier domain offers a natural lens to model time series, current deep learning approaches do not explicitly model evolution and randomness in the Fourier spectra, which limits their ability to accurately predict both the expected trajectory and its uncertainty in non-stationary time series. Motivated by Evolutionary Spectra (ES) theory, we propose TimeES, a general framework that enables probabilistic and deterministic forecasting via the evolutionary spectra theory. Specifically, we derive a parameterizable evolutionary spectra formulation, recasting non-stationary random process modeling as learning an evolving representation modulated by random variables. Furthermore, we reduce the complexity of the estimated spectra from O(NM) to O(NK), where K << M/2, by exploiting Hermitian symmetry and spectral energy sparsity for frequency selection. Based on a simple linear backbone, our proposed TimeES achieves consistent state-of-the-art performance across both deterministic and probabilistic forecasting tasks, with high efficiency and interpretability. Code is available at: https://github.com/wwy155/TimeES.
|
| 1134 |
Using Machine Learning to Investigate Predictors of Fasting Blood Glucose: Insights into Circadian Timing and Age Interactions
2609.32386
|
cs.LG
|
Viktoriya Bu-Dager, Silvia Cirstea |
Impaired glucose regulation is a major contributor to metabolic dysfunction and type 2 diabetes. This study developed an interpretable machine-learning framework to predict log-transformed fasting blood glucose using metabolic, hormonal, lifestyle, demographic...Impaired glucose regulation is a major contributor to metabolic dysfunction and type 2 diabetes. This study developed an interpretable machine-learning framework to predict log-transformed fasting blood glucose using metabolic, hormonal, lifestyle, demographic, nutritional, and circadian variables from the National Health and Nutrition Examination Survey 2017--2020 pre-pandemic dataset. After merging multiple NHANES sub-datasets, data processing used a leakage-resistant pipeline in which imputation, scaling, and one-hot encoding were performed only after dataset splitting and within training folds. Elastic Net, LASSO, and XGBoost models were evaluated using 94 candidate predictors and engineered circadian interaction terms. Performance was assessed using mean absolute error, root mean squared error, coefficient of determination, calibration, and Shapley Additive Explanations. The final interaction-augmented XGBoost model achieved strong performance on the independent test set, with a mean absolute error of 0.0804, a root mean squared error of 0.1148, and a coefficient of determination of 0.7761, using 10 predictors. Glycohemoglobin was the dominant predictor, followed by insulin, diabetes diagnosis, gamma-glutamyl transferase, age, race, and gender. Among the engineered interaction terms, sleep midpoint multiplied by age was consistently retained in repeated random-split analyses, although its contribution remained modest relative to dominant glycaemic predictors. These findings support further investigation of circadian-age interactions in metabolic health.
|
| 1135 |
Recovery-Directed Symbolic Distillation of Neural Likelihoods
2609.32409
|
cs.LG
|
Kiant\'e Fernandez, Xinwei Li |
Amortized neural likelihoods enable computationally expensive inference for models with analytically intractable or unspecified likelihoods, but their black-box nature limits interpretability. We introduce a symbolic distillation pipeline that converts trained...Amortized neural likelihoods enable computationally expensive inference for models with analytically intractable or unspecified likelihoods, but their black-box nature limits interpretability. We introduce a symbolic distillation pipeline that converts trained neural likelihoods into explicit, interpretable expressions optimized for efficient parameter estimation. Our approach uses a recovery-directed objective to guide symbolic regression toward expressions that preserve parameter-recovery accuracy rather than merely approximating the likelihood function. Candidate expressions are evaluated on held-out datasets and selected using a criterion that jointly accounts for expression complexity, parameter-recovery performance, and distributional distance from the learned likelihood. We evaluate the pipeline on the diffusion decision model, a classical cognitive model, whose analytically tractable likelihood provides ground truth for controlled evaluation. The proposed recovery-directed objective improves parameter recovery over standard symbolic-regression objectives. The resulting symbolic likelihoods enable over 100 times faster parameter evaluation than both neural likelihoods and, when available, the exact likelihood, while maintaining a manageable loss in precision. We further demonstrate these computational benefits in Bayesian hierarchical inference on empirical data. Our pipeline provides a lightweight interface for integrating symbolic distillation with existing neural-likelihood estimation methods and can be adapted to a range of simulation-based inference settings.
|
| 1136 |
HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration
2609.32426
|
cs.LG
|
Inwoo Tae, Yoontae Hwang, Yongjae Lee |
For graph node classification, calibrated class probabilities are needed when confidence scores, usually the maximum predicted class probability, are used to rank predictions, defer uncertain nodes to human review, or control risk. Existing post-hoc calibrator...For graph node classification, calibrated class probabilities are needed when confidence scores, usually the maximum predicted class probability, are used to rank predictions, defer uncertain nodes to human review, or control risk. Existing post-hoc calibrators either apply one global temperature or use graph-aware modules without a principled structural form. We study how local graph structure should enter node-level calibration. Our first results show that a logit-only temperature rule is insufficient when nodes with identical logits but different local homophily require different optimal temperatures. We then analyze a population-concentration contextual stochastic block model with Gaussian features and a one-layer linear GCN. Under equidistant class means, the Bayes posterior over class-template scores is a temperature-scaled softmax whose inverse-temperature is governed by a homophily-dependent signal strength. In the positive-signal homophilic regime, the resulting temperature decreases approximately inversely with normalized local homophily. This law motivates Homophily-aware Temperature Scaling (HoTS), a simple post-hoc calibrator that assigns each node a positive scalar temperature from entropy-based logit concentration and estimated local homophily. HoTS has three temperature parameters, preserves the predicted class, and learns the strength of the structural correction from calibration data. Across 18 node-classification benchmarks, two GNN backbones, and eight calibration baselines, HoTS achieves the best mean Expected Calibration Error (ECE) of 4.79%, the best average rank, and the most reliable confidence ranking in selective classification. Code is available at https://github.com/inu0104/HoTS.
|
| 1137 |
Beyond the Manifold Hypothesis: Hybrid Spectral Parameterizations for Flow Matching
2609.32432
|
cs.LG
|
S\'egol\`ene Martin, Anne Gagneux, Quentin Bertrand, R\'emi Emonet, Mathurin Massias |
Flow matching and diffusion can be trained to predict different quantities, most commonly the data $x_1$, the source noise $x_0$, or the velocity~$v$. Although theoretically equivalent, these can lead to substantially different performances. We identify two ma...Flow matching and diffusion can be trained to predict different quantities, most commonly the data $x_1$, the source noise $x_0$, or the velocity~$v$. Although theoretically equivalent, these can lead to substantially different performances. We identify two main drivers for these differences: the source--data signal-to-noise ratio, and the information bottleneck induced by the neural architecture. We show that, beyond intrinsic data dimension, the factor affecting the optimal parametrization the most is a certain signal-to-noise ratio in each data covariance direction. From this analysis, we introduce new \emph{spectral hybrid} parameterizations that adapt across time and data covariance directions; we show that these are optimal for Gaussian data. We also show that architecture-induced compression changes which parameterization is easier to learn, with $v$-prediction being more sensitive to discarded directions than $x_1$-prediction. Experiments across architectures and source scales show that our spectral parameterizations are robust across regimes, can substantially accelerate optimization, while incurring essentially no additional training cost compared with standard parameterizations.
|
| 1138 |
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
2609.32444
|
cs.LG
|
Tianrun Yu, Kaixiang Zhao, Shangzhe Li, Yuxiao Yang, Porter Jenkins |
We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different p...We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrepancy in policy updates, we introduce calibrated importance sampling (CIS). CIS is motivated by an empirically supported logit-displacement characterization that expresses the mismatch as an additive displacement $\varepsilon_t$ in log-odds, determined by the per-logit perturbation before the softmax, whose distribution is approximately invariant to token confidence. This characterization motivates a confidence-aware truncation: large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases. Theoretically, we show that CIS replaces the unbounded second moment that governs the error of exact importance sampling with a term bounded by a constant, at the cost of a bias controlled by the truncated excess. In evaluation across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines. Diagnostic analyses show that CIS places less truncation bias on low-confidence tokens than truncated importance sampling, while upward clipping of small importance weights reduces held-out accuracy.
|
| 1139 |
Length-Independent State Tracking Under a Parallel Scan
2609.32447
|
cs.LG
|
Julien Brandoit, Arthur Fyon, Thomas Braipson, Tom Clara, Florent De Geeter |
Learning robust and scalable finite-state tracking is fundamental to sequence processing. While linear recurrent neural networks (RNNs), linear attention, and state space models enable scalable parallel training through affine recurrences, their theoretical ex...Learning robust and scalable finite-state tracking is fundamental to sequence processing. While linear recurrent neural networks (RNNs), linear attention, and state space models enable scalable parallel training through affine recurrences, their theoretical expressivity guarantees assume idealized arithmetic and do not extend to finite precision, where the parallel scan that makes them fast is itself a source of perturbation. We formalize finite-state tracking at finite precision and characterize length independence: tracking that stays correct at every sequence length, at a precision cost that does not grow with the length. We show that length-independent state tracking requires two competing dynamics within a single map: contraction to suppress numerical perturbations and separation to keep distinct states apart. We prove that affine recurrences, which offer a single rate at each step to serve both roles, realize at most definite automata at finite precision. Instead of treating scan compatibility as a restriction on the update map, we reinterpret it as a computational budget and introduce the Neural Finite-State Machine (NFSM): a nonaffine, scan-compatible RNN built for length-independent finite-state tracking. On synthetic benchmarks spanning abelian and nonabelian groups, noninvertible monoids, and textual state-tracking tasks, affine baselines fail on every nondefinite task, most of them within a few hundred steps. A single NFSM layer instead learns the exact transition tables of every algebraic task, which certifies correctness beyond the tested lengths, and a stack of NFSMs keeps perfect accuracy on the textual tasks at every tested length.
|
| 1140 |
Write Back the $\Delta$: Revisiting the Same Tokens with Fresh Representations
2609.32457
|
cs.LG
|
Wencheng Ye, Anning Hu, Xiangdong Zhang, Tianyi Wang, Yikang Li |
Transformers process information strictly forward through depth, preventing deeper computation from revisiting and refining earlier representations. To augment the standard forward pass, existing approaches either re-execute depth, incurring additional computa...Transformers process information strictly forward through depth, preventing deeper computation from revisiting and refining earlier representations. To augment the standard forward pass, existing approaches either re-execute depth, incurring additional computation, or modify the residual stream using predefined directions, limiting their instance-level adaptation. Recently, inference-time feedback offers a direct mechanism for recycling endogenously produced computation by writing deeper residual states back to earlier layers, yet what should be fed back remains unclear. We argue that the depth increment Delta, capturing newly accumulated computation between two layers, provides a more effective, composable, and scalable feedback signal than the full state. Building on this observation, we introduce ReFlux, a learnable feedback graph that dynamically selects and composes increment-carrying routes. ReFlux supports synchronous feedback to the same token and streaming feedback to subsequent tokens. Extensive experiments across various models, corpora, and benchmarks show that synchronous ReFlux consistently reduces perplexity across ten language-modeling corpora, and improves accuracy by 2.1-2.3 points, with gains reaching 4.7 points on multi-hop reasoning. Streaming ReFlux further retains most of these gains while preserving the base model's 1x theoretical backbone FLOPs. These results establish ReFlux as an efficient paradigm for unlocking the latent computational potential of LLMs, allowing them to revisit the same tokens with fresh representations. Code implementation can be found at https://github.com/gooogleshanghai/reflux.
|
| 1141 |
Adapting Nonstationary Multi-output Gaussian Processes to Bayesian Optimization
2609.32464
|
cs.LG
|
Zikai Xie |
Multi-objective Bayesian optimization (MOBO) commonly relies on independent Gaussian processes (GPs) with stationary kernels, limiting its ability to represent nonstationary structure and share information between objectives. However, expressive nonstationary ...Multi-objective Bayesian optimization (MOBO) commonly relies on independent Gaussian processes (GPs) with stationary kernels, limiting its ability to represent nonstationary structure and share information between objectives. However, expressive nonstationary GPs do not necessarily make reliable BO decisions. We study this mismatch for the multi-output low-rank nonstationary (MO-LRN) GP: strong training fit can coexist with large off-design errors and optimistic acquisition predictions. We introduce MOLRN-BO, which combines a regularized shared-spectral surrogate with objective-specific residuals, prequential mean correction and tempered covariance scaling, and Pareto-local qLogEHVI optimization with periodic global search. Experiments on 12 deterministic bi-objective benchmarks show that MOLRN-BO substantially improves upon the original MO-LRN and achieves the best average problem ranks for final normalized hypervolume and normalized inverted generational distance among nine evaluated algorithms. It also achieves the strongest adverse-tail performance while remaining competitive with the leading baselines in anytime optimization. Ablation studies further show that the shared spectral construction improves off-design prediction, the local--global decision policy improves optimization performance, and hierarchical calibration reduces systematic candidate bias. These results demonstrate that nonstationary multi-output surrogates can deliver strong and robust MOBO performance when their structure and use are explicitly adapted to the demands of sequential optimization.
|
| 1142 |
PolyStepOR: Learning to Decide Without Optimal Decisions
2609.32465
|
cs.LG
|
Viet The Nguyen, Gunther Gust, An Thai Le |
Decision-focused learning (DFL) trains predictors for downstream decision quality, but often relies on optimal reference decisions that are expensive to obtain. We present PolyStepOR, which trains directly from realized decision costs without pre-computed opti...Decision-focused learning (DFL) trains predictors for downstream decision quality, but often relies on optimal reference decisions that are expensive to obtain. We present PolyStepOR, which trains directly from realized decision costs without pre-computed optima and extends to in-constraint predictions through repair or infeasibility penalties. To handle piecewise-constant losses, PolyStepOR perturbs predictor parameters, evaluates the resulting decisions, and uses optimal transport to favor lower-cost directions, requiring no derivatives. Without task-specific tuning, PolyStepOR performs strongly on classical optimization benchmarks and competitively on predicted-constraint and real-world problems. Theoretically, we characterize decision-preserving perturbations and boundary detection, bound sensitivity to cost errors, and establish stationarity guarantees for a smoothed objective. PolyStepOR thus replaces optimal reference decisions and derivatives with forward evaluations.
|
| 1143 |
Bison: Cross-Dataset Learning for Unseen-Compound Perturbation Prediction
2609.32467
|
cs.LG
|
Yunfan Liu, Kasra Ghorbani, Yufei Huang, Zicheng Liu, Jiangbin Zheng |
Predicting transcriptional responses to unseen compounds is limited by fragmented chemical coverage and heterogeneous experimental platforms and gene panels. To assess molecular generalization across these settings, we build on Chem-PerturBridge to benchmark e...Predicting transcriptional responses to unseen compounds is limited by fragmented chemical coverage and heterogeneous experimental platforms and gene panels. To assess molecular generalization across these settings, we build on Chem-PerturBridge to benchmark eight datasets with 16,771 compounds, withholding test compounds from every training dataset. This comparison reveals that high overall response agreement can coexist with weak prediction of drug-specific differences, despite reproducible signals across repeated measurements. To exploit complementary chemical supervision while targeting these differences, we introduce Bison: a shared gene representation connects native panels, while two discrete diffusion models compose context-dependent responses with molecular deviations learned through matched drug-contrast supervision. A single Bison model jointly trained across all eight datasets achieves the highest mean overall-response and drug-contrast Pearson correlations on the full benchmark in comparison with 11 methods trained independently per dataset. Compared with dataset-specific training of the same architecture, joint training increases mean drug-contrast correlation by 27.4\%, with gains across all eight datasets and improvements in overall response prediction. These results demonstrate how matched drug contrasts turn complementary screens into shared molecular supervision for unseen-drug response prediction while preserving native gene measurements.
|
| 1144 |
On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models
2609.32470
|
cs.LG
|
Shuoyuan Wang, Beier Luo, Hao Zeng, Chengyao Yu, Songxin Zhang |
Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by t...Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by the model's pre-RL confidence distribution, which we term confidence prior. In this work, we reveal that off-the-shelf LRMs exhibit a confidence prior heavily concentrated on a few high values, which persists throughout RL. Theoretically, we prove that this concentration suppresses policy gradient updates for rarely sampled confidence values and inflates the lower bound on expected Brier risk. To overcome this exploration bottleneck, we propose CalibSFT, a plug-and-play supervised fine-tuning stage that shapes a calibrated confidence prior with broad support before RL. For each question, CalibSFT constructs confidence targets combining its success rate with response-level correctness, which provably preserves proper-scoring optimality, and then balances training responses across the confidence spectrum to enable diverse confidence exploration during RL. To learn from incorrect responses without imitating their reasoning, CalibSFT introduces correctness-conditional supervision, guiding confidence across all responses while supervising reasoning only on correct ones. Across 16 mathematical and general reasoning benchmarks, incorporating CalibSFT reduces calibration errors and improves discrimination across five representative RL algorithms while preserving comparable accuracy. Furthermore, CalibSFT delivers practical benefits for downstream selective prediction and model routing. Our code is available at https://github.com/ml-stat-Sustech/verbalized-confidence-training.
|
| 1145 |
Fast and Precise Learned Charged-Particle Trajectory Regression at the Large Hadron Collider
2609.32479
|
cs.LG
|
Jonathan Renusch, Benjamin Huth, Daniel Murnane, Eleni Xochelli, Do\u{g}a Elitez |
We propose a training recipe that treats charged-particle trajectory parameter regression on high-energy physics detector data as a sequence-modeling task. Kalman filters and linearized least-squares fits have been the classical standard approach for this task...We propose a training recipe that treats charged-particle trajectory parameter regression on high-energy physics detector data as a sequence-modeling task. Kalman filters and linearized least-squares fits have been the classical standard approach for this task: they are optimal estimators for sparsely sampled linear-Gaussian data and are commonly used for trajectory parameter regression (fitting). The classical fitting techniques implemented for this domain reach a final precision of one part in $10^5$ through detailed modeling of detector geometry, material, detection effects and precise numerical integration of the equations of motion through the detector's inhomogeneous magnetic field. With this study, we demonstrate that using a bidirectional gated linear recurrent encoder, one is able to reproduce the full precision of classical track fitting techniques. Using a custom kernel, we also achieve significantly higher throughput during GPU inference, compared to classical fitting software running on similarly priced multi-core CPU servers representing typically employed hardware. Such a speedup would lead to considerable cost savings for the pattern recognition at the Large Hadron Collider. To our knowledge, this is the first end-to-end learned track fit to reach the full precision and, at the same time, offer the opportunity to reduce the computing costs.
|
| 1146 |
What Should Federated LoRA Share? FedSAIL via Input-aware Subspace Alignment
2609.32485
|
cs.LG
|
Junye Du, Shuaida He, Long Feng |
Federated low-rank adaptation (LoRA) requires identifying an update structure that is shared across heterogeneous clients. Prior work reports strong similarity among trained LoRA projection matrices across clients; however, such agreement may be largely induce...Federated low-rank adaptation (LoRA) requires identifying an update structure that is shared across heterogeneous clients. Prior work reports strong similarity among trained LoRA projection matrices across clients; however, such agreement may be largely induced by common initialization and collapses toward random overlap under independent initialization. More crucially, relying solely on parameter similarity inherently ignores the influence of local input regime. To uncover a more robust shared structure, we introduce an input-aware action matrix that weights the adapter update by the second-moment statistics of local layer inputs. Empirically, while parameter similarity vanishes, the leading right singular directions of this action matrix remain strongly aligned across clients. This shared geometry preserves task-conditioned differences and naturally varies across network depths. Motivated by these findings, we propose Federated Subspace-Guided Action-Informed Learning (FedSAIL). Instead of averaging weights, FedSAIL estimates a shared action subspace to regularize local training while preserving client-specific coefficients. Across several benchmarks, our approach consistently improves predictive performance over competing federated LoRA methods while reducing communication cost significantly.
|
| 1147 |
Elastic Selective Spectral Hybrids for Train-Once, Export-Many Budgeted Inference
2609.32486
|
cs.LG
|
Dachuan Song, Chuchu Chen, Xuan Wang |
Deploying a language model under different computing and latency budgets calls for compact models with different quality-cost trade-offs. Towards this end, elastic spectral state space models provide ordered and temporally decomposed channels that can be trunc...Deploying a language model under different computing and latency budgets calls for compact models with different quality-cost trade-offs. Towards this end, elastic spectral state space models provide ordered and temporally decomposed channels that can be truncated, but they are associated with linear time-invariant filters that cannot selectively preserve relevant past information or forget irrelevant information as the context evolves. To address this, we introduce the Elastic Selective Spectral Hybrid (ESSH), which realizes each Hankel spectral channel as an independent recurrent unit using fitted damped rotation modes. It also features an input-dependent decay and write/read gates that make temporal retention and state update input-dependent while preserving channel-wise truncation and a structured recurrence for efficient execution. ESSH combines these selective spectral mixers with sliding-window attention and jointly trains multiple capacities by reducing spectral-channel count and feed-forward width at different rates through a two-rate capacity map with full-model distillation. The resulting models support chunked parallel training and fused recurrent decoding while avoiding computation for discarded channels. At full capacity, ESSH achieves language-modeling quality comparable to similarly sized independently trained models, while smaller exports exhibit a smooth quality-cost trade-off. We validate the effectiveness of the proposed framework using language understanding, retrieval, cross-domain text, and DNA experiments, by assessing quality retention and the trade-off against independently trained and elastic baselines. At the 1.53B model configuration, fused batch-one decoding takes 1.37 ms per token on a B300, providing a 2.14-2.80x speedup over the tested Mamba-2 and Mamba-3 implementations and 3.03x over Transformer++ at matched parameter counts.
|
| 1148 |
Neural Dynamics as the Composition of Quantized Units
2609.32487
|
cs.LG
|
Jacopo Minniti, Aravinth Kulanthaivelu, Richard Sproat |
Deep learning is commonly interpreted at two levels: the macroscopic, through aggregate trends in loss summarized by scaling laws, and the microscopic, through neurons, features, and circuits. A central challenge is understanding how these levels connect, so t...Deep learning is commonly interpreted at two levels: the macroscopic, through aggregate trends in loss summarized by scaling laws, and the microscopic, through neurons, features, and circuits. A central challenge is understanding how these levels connect, so that we can explain how elementary computations compose and collectively shape macroscopic behavior. To this end, we study an intermediate abstraction in which training is described as the ordered acquisition of quanta: reusable computations acquired suddenly and binary-activated across examples to reduce loss. By approximating population-gradient updates, we derive quanta's acquisition dynamics. This yields an acquisition priority governed by demand, how frequently a computation is required across examples, and conditional complexity, how difficult that computation is to acquire given those already available. In a Boolean compositional task, we derive predictions for acquisition order and show how staggered discrete acquisitions can produce smooth aggregate loss and, under certain geometries of quanta composition, give rise to scaling laws. We then train a Transformer to map numerals to English number names and recover candidate quanta from its checkpoint trajectory. From these units, we construct a model that preserves much of the Transformer's behavior while exposing interpretable latent computations and acquisition dynamics consistent with the theory. Separately, the quanta structure can serve as training targets to improve transformer generalization. Together, these results suggest the quanta abstraction can provide useful computational atoms for studying a variety of macroscopic phenomena.
|
| 1149 |
SoFT: Soft Targets for Generalizable LLM Fine-Tuning
2609.32493
|
cs.LG
|
Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi, Zhongwei Xie |
Distillation enables student language models to acquire new capabilities from expert teachers. However, integrating knowledge from multi-teacher, multi-domain demonstrations into a single student remains challenging. We study supervised fine-tuning (SFT) in th...Distillation enables student language models to acquire new capabilities from expert teachers. However, integrating knowledge from multi-teacher, multi-domain demonstrations into a single student remains challenging. We study supervised fine-tuning (SFT) in this setting, where students must acquire diverse capabilities while maintaining generalization beyond the training tasks. Our experiments reveal varying trade-offs between in-distribution learning and out-of-distribution generalization across SFT methods, motivating more explicit control over this balance. To this end, we propose soft-target fine-tuning (SoFT) to balance learning from teacher demonstrations with retaining the Base model's existing capabilities. SoFT sets a minimum target probability for each demonstrated token while making the smallest KL change to the Base distribution. The resulting objective couples learning from demonstrations with adaptively weighted regularization toward the Base model. We further use domain-specific gradient budgets to control this balance and determine a probability threshold for each trajectory. Experiments on mixed-domain reasoning and agentic tasks show that SoFT achieves the best overall performance among the compared methods, with improvements in both in-distribution capability acquisition and out-of-distribution generalization.
|
| 1150 |
Rondo: Unsupervised Discovery of Recurring Temporal Structure
2609.32500
|
cs.LG
|
Yingtian Shi, Ankith Chandra, Thomas Pl\"otz |
Many real-world time series data exhibit structural properties at multiple scales, from short, recurring units to complex sequences composed of these units. Unsupervised discovery of both these components and structure enables the design of intelligent systems...Many real-world time series data exhibit structural properties at multiple scales, from short, recurring units to complex sequences composed of these units. Unsupervised discovery of both these components and structure enables the design of intelligent systems that help interpret temporal data, thereby limiting the amount of costly human annotations required. Existing modeling approaches typically overlook the hierarchical structure inherent to many time series, treating recurring patterns at different temporal scales as independent structures. Moreover, most assume access to the complete data sequence and treat discovery as a static process, limiting their ability to evolve as new observations arrive. We introduce Rondo, an unsupervised approach for modeling recurring hierarchical structure in continuous temporal streams. By explicitly constructing vocabularies of reusable units and their recurring compositions, Rondo captures structure shared across complex temporal patterns while refining and expanding its discoveries as the stream evolves. Evaluations on temporal sequences spanning diverse domains and data modalities show that Rondo outperforms existing unsupervised recurrence-discovery baselines, with particularly pronounced advantages in limited-data and continual-stream settings. These capabilities provide a stronger foundation for recurring-pattern discovery, scalable behavior understanding, and adaptive intelligent systems operating on long, unlabeled temporal streams.
|
| 1151 |
Fisher Simplicity in Kolmogorov-Arnold Networks and Multilayer Perceptrons
2609.32503
|
cs.LG
|
Ami Tavory, Meir Feder |
Kolmogorov-Arnold Networks (KANs) are motivated in part by interpretability: their learned edge functions can be inspected, pruned, and reduced to symbolic structure. In a fixed-basis KAN, this makes a small or zero basis coefficient look like a certificate of...Kolmogorov-Arnold Networks (KANs) are motivated in part by interpretability: their learned edge functions can be inspected, pruned, and reduced to symbolic structure. In a fixed-basis KAN, this makes a small or zero basis coefficient look like a certificate of simplicity, much as a dead rectified linear unit (ReLU) marks unused computation in a multilayer perceptron (MLP). Fisher nullity gives a precise statistical notion: a parameter direction is Fisher-simple exactly when perturbing it is invisible under the task distribution. We study when these architectural and Fisher notions agree. For a dead ReLU unit, they agree: the closed activation region makes the associated score directions vanish. For a fixed-basis KAN, they do not. In the single-layer Gaussian case, the coefficient Fisher matrix is a basis Gram matrix under the input distribution and is independent of the fitted coefficients. In a multilayer KAN, Fisher simplicity is graph-path based: the data must reach a basis atom and its perturbation must propagate through the downstream network. We encode these two conditions in an effective edge measure and, under local dictionary independence and effective-measure nondegeneracy, show that zero effective exposure exactly identifies Fisher-null directions within an edge. Controlled diagnostics confirm that zero coefficients can preserve rank while effective path disconnections remove the predicted directions. Coefficient magnitude alone is therefore not a Fisher-based pruning criterion for KANs.
|
| 1152 |
What Do Latent Predictive Vehicle Representations Retain? Measuring State, Geometry, and Local Response
2609.32512
|
cs.LG
|
Enzo Nicol\'as Spotorno, Josafat Leal Filho, Ant\^onio Augusto Fr\"ohlich |
Models of vehicle dynamics learned from logged states and commands complement physics-based models, and latent world models, which predict in a learned representation, are used to plan and train controllers in other domains. Vehicle controllers are usually spe...Models of vehicle dynamics learned from logged states and commands complement physics-based models, and latent world models, which predict in a learned representation, are used to plan and train controllers in other domains. Vehicle controllers are usually specified in physical terms: costs, limits, and references depend on position, yaw angle, speed, and yaw rate, and the optimizer compares or differentiates predicted outcomes across nearby commands. A latent model placed in such a controller must therefore let these quantities be recovered and must change its predictions with commands as the vehicle does, and prediction error on its own latent targets measures neither. We contribute a measurement protocol for action-conditioned latent predictors with a physical readout that separately tests retention, physical-neighborhood organization, forecasting, and local response to command perturbations, using an untrained-encoder reference and three matched response paths that locate errors in the representation or the predictor. In a case study of a temporal joint-embedding predictive model trained on signals logged in IPG CarMaker, the representations retain the measured planar outputs, though an untrained encoder of the same architecture retains them slightly better; future-command input improves one-second forecasts with retention nearly unchanged; and responses to small command pulses diverge from the simulator already in latent coordinates, raising regret when choosing among nearby commands in all comparisons. Updating the predictor on responses corrects them locally at a cost in forecast accuracy. Measuring retention, forecasting, and local response separately is thus what qualifies a predictive latent as a candidate model for control, and the protocol provides the basis for its closed-loop evaluation.
|
| 1153 |
Does Transolver really need a Transformer?
2609.32525
|
cs.LG
|
Shizheng Wen, Siddhartha Mishra |
The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the p...The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empirical and theoretical analysis to elucidate the mechanisms which are responsible for model performance. To this end, we perform careful ablations on a challenging suite of nine 3D fluid dynamics benchmarks to find that replacing token attention with a constant linear map does not affect the accuracy. Thus, Transolver does not need a Transformer at all. However, removing the global mixing (slicing/deslicing) or doing it only once leads to performance collapse. We leverage the theory of averaging neural operators to explain and corroborate our findings by showing that just slicing/deslicing, in conjunction with pointwise MLPs, already suffices for universal approximation of continuous operators and attention is redundant in this context. Finally, we provide a novel FlashAttention-style efficient implementation of the key slicing/deslicing module of Transolver. This flashslice kernel streams slice and deslice over the points without materializing the heavy slice-weight tensor, while reproducing the best available implementation to floating point error. At the same time, it leads to very significant memory and compute savings, particularly at large slice counts.
|
| 1154 |
What Does a ProcGen Generalization Gap Measure? Action Rules, Convergence, and the Missing Random Floor
2609.32532
|
cs.LG
|
Abhisek Keshari |
A generalization gap in reinforcement learning (return on training levels minus return on held-out levels) is usually reported without a reference point. We argue that the missing reference is a measured random floor: the return of a uniform-random policy on t...A generalization gap in reinforcement learning (return on training levels minus return on held-out levels) is usually reported without a reference point. We argue that the missing reference is a measured random floor: the return of a uniform-random policy on the same levels under the same harness. On ProcGen, the floor changes what several standard numbers mean. On identical checkpoints and levels across eight environments, switching between sampled and greedy (argmax) test-time actions moves held-out return in both directions, and greedy evaluation takes three environments to or below the floor: in miner, the sampled policy scores 4.9x the floor on held-out levels while its argmax scores below it. Raw policy entropy places seven of eight environments short of convergence, but 35-65% of that entropy is spread across actions with identical effects; after merging them, one to three remain short, and against the floor only heist has learned nothing that transfers. An audit of twelve prior ProcGen codebases finds that all eleven with held-out evaluation sample test-time actions for their policy-gradient agents, nine by default rather than explicit choice, and six report running in-loop averages rather than evaluating a fixed checkpoint. Applied to our own case study, the same checks grade down a statistically significant encoder effect and rule out a within-encoder train-vs-test CKA statistic. We recommend that every reported gap state its action rule, use a matched and seeded protocol, and report the random floor on both level sets.
|
| 1155 |
Prioritizing Repeated LLM Evaluation for Hidden Failure Discovery
2609.32547
|
cs.LG
|
Keita Broadwater, Akin Broadwater |
Large language models are commonly evaluated by generating a small number of stochastic responses for each prompt in a benchmark. Because inference budgets are limited, this shallow evaluation may fail to observe low-probability but operationally important fai...Large language models are commonly evaluated by generating a small number of stochastic responses for each prompt in a benchmark. Because inference budgets are limited, this shallow evaluation may fail to observe low-probability but operationally important failures. A prompt that produces no failures in a small sample may therefore appear reliable despite having a nonzero latent probability of failure under repeated inference. We formulate LLM reliability evaluation as a budget-constrained discovery problem in which each prompt is associated with an unknown per-generation failure probability. We propose a budgeted discovery framework that first performs shallow evaluation across the prompt set and then uses trial-level failure outcomes together with prompt-derived representations to learn a feature-based ranking of failure propensity. The resulting scores prioritize prompts with zero observed shallow failures for deeper evaluation, concentrating the deep-evaluation budget where hidden failures are more likely to be discovered. We evaluate this approach on AIRBench and StrongREJECT across multiple model and system-prompt conditions. The central empirical test asks whether models fit without access to deep-evaluation outcomes can rank prompts with zero observed shallow failures according to their likelihood of producing failures under deeper evaluation. On AIRBench, the highest-ranked 10\% of unresolved prompts achieves 2.54x hidden-failure lift for Qwen 2.5 7B and 1.87x for Gemma 3n E4B, recovering 25.4\% and 18.7\% of subsequently observed hidden failures, respectively, compared with 10\% expected under random allocation. Semantic-neighborhood and feature-ablation analyses further show that this predictive signal can be recovered from multiple representations of prompt content and relationships.
|
| 1156 |
Compositional Objectives: Learning Structure in Structure
2609.32566
|
cs.LG
|
Pranavchandra Vivekananda, Sumukh Bettadapura, Ajan Subramanian |
Intelligence is defined in many ways. One of these definitions defines intelligence as the pursuit of learnable novelty. However, learnable novelty can be meaningless without the ability to compose the learned structures to take action and achieve goals. Learn...Intelligence is defined in many ways. One of these definitions defines intelligence as the pursuit of learnable novelty. However, learnable novelty can be meaningless without the ability to compose the learned structures to take action and achieve goals. Learnable novelty builds on epiplexity, which is a way to measure learnable structure in data through a bounded observer. In this paper, we investigate a closed-form spectral approximation to compute epiplexity. We use a fixed-trace constraint and find that the epiplexity objective prefers a more uniform distribution of spectral mass rather than concentrating it in a small number of directions. However, a representation may spread information across many directions without organizing that information into features useful for a particular task. To address this gap, we propose a compositional objective whose observer measures the relationships between the parts and interactions of an image. We compare it with the original spectral objective given only the masked parts. In our ImageNet training runs, the spectral objective with masked parts produces an almost maximally spread representation while achieving the strongest frozen-feature classification performance on most evaluations, more than doubling the linear-probe accuracy of the whole-image baseline. Across multiple image benchmarks, changing what the observer sees matters more than adding relation and interaction tokens. At the same time, our prediction-oriented compositional objective produces substantially better held-out observer prediction but relatively weaker classification, revealing that spectral diversity, predictability, and downstream utility are distinct properties. These results suggest that the usefulness of spectral spreading depends not only on how much structure is preserved, but on which relationships the observer makes available to the objective.
|
| 1157 |
CAESAR: Clustering via Autonomous Embedding-Space Agglomerative Reorganization
2609.32570
|
cs.LG
|
Ilan Bacry, R\'emi Devaux, Antoine Jardin |
Clustering algorithms that operate on nearest-neighbor graphs, such as FINCH (First Integer Neighbor Clustering Hierarchy), depend heavily on the quality of the embedding space they are given. However, pretrained vision and language model embeddings are not op...Clustering algorithms that operate on nearest-neighbor graphs, such as FINCH (First Integer Neighbor Clustering Hierarchy), depend heavily on the quality of the embedding space they are given. However, pretrained vision and language model embeddings are not optimized for this purpose. We propose CAESAR, a method that reorganizes a pretrained embedding space: a reorganization network is trained to pull mutual nearest neighbors together and push non-neighbors apart, yielding a reorganized embedding space substantially better suited to clustering. CAESAR offers a second major advantage: it never requires the number of clusters $K$. This matters because in realistic unsupervised settings, $K$ is typically unknown and discovering it is often part of the problem, yet most strong clustering methods take it as an input. We therefore design the entire CAESAR pipeline to infer $K$ rather than assume it is known. Empirically, reorganizing the embeddings consistently improves clustering over the raw space on both text and image datasets. Since the few deep clustering methods that also infer $K$ do not release their code, we complement controlled comparisons with methods that infer $K$ on the same embedding space by comparisons with strong deep clustering methods that are given the true $K$, giving them a substantial oracle advantage. Even so, CAESAR outperforms all of them on text, achieves the best results on the most challenging image benchmark and remains competitive on the others. Reorganizing pretrained embeddings thus emerges as a simple and powerful route to clustering realistic data, where classes overlap and the number of clusters is unknown.
|
| 1158 |
Trapped by Their Own Rollouts: Understanding Aggregation--Rollout Feedback in Federated On-Policy Distillation
2609.32573
|
cs.LG
|
Jinqian Chen, Jihua Zhu, Chang Liu |
On-policy distillation (OPD) is a promising approach to language-model adaptation, aligning teacher supervision with the student's own generated trajectories. When adaptation prompts are distributed across clients, can this process benefit from federated colla...On-policy distillation (OPD) is a promising approach to language-model adaptation, aligning teacher supervision with the student's own generated trajectories. When adaptation prompts are distributed across clients, can this process benefit from federated collaboration? We study federated OPD and find that substantial collaboration gains can be obscured by learning-rate sensitivity: FedAvg can perform no better than independent local training at a small learning rate, yet recover a clear advantage at a larger rate. We explain this phenomenon through the student's dual role as learner and generator of future training data. An aggregation-induced optimization lag can delay access to useful teacher supervision, which in turn slows subsequent learning. Our theory establishes this aggregation--rollout feedback in a solvable model with a common optimum and stable updates, and identifies two coupled roles of learning rate: learning from current supervision and reaching future supervision. Guided by this analysis, we propose FedTOPS (Federated Teacher-guided On-Policy Scaling), which reuses teacher feedback on current trajectories to adapt the FedAvg update magnitude under clientwise predictive-change constraints. Across six mathematical reasoning benchmarks, FedTOPS improves macro Avg@8 over FedAvg by 4.56--14.57 percentage points across the evaluated student models and local learning rates.
|
| 1159 |
DimPO: Dimensionality Reduction for Attention using Preference Optimization
2609.32579
|
cs.LG
|
Vojt\v{e}ch Lanz, Yufei Cui, Prasanna Parthasarathi |
A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-w...A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-weighted keys provide a better signal than matching the full attention distribution with KL divergence, especially in long-context settings. We introduce DimPO, which combines listwise preference optimization with a lightweight top-k cross-entropy term for head-fidelity. DimPO is trained offline from the attention patterns of a frozen language model, with one map per layer, shared by the query and the keys and trained separately from the other layers. Across LLaMA3.2-3B, LLaMA3.1-8B, Qwen2.5-7B, and Qwen3-4B Instruct models, pairwise preference objectives outperform the triplet baseline and retain 98% of the original score on short-context tasks when projecting to half the dimension on the last 40% of the layers. With more projected layers or on long-context RULER, they degrade rapidly. In contrast, KL and DimPO, which use every key during training, retain about 95% of the original RULER 4k score on the 8B model when projecting up to 50% of the layers. KL-based projections remain closer to the original attention distribution and attention output, yet DimPO achieves better downstream performance. Beyond 50% of projected layers, DimPO increasingly outperforms KL on tasks including SQuAD, common-word extraction, frequent-word extraction, and variable tracking. These results suggest that under dimensionality reduction, preserving the ordering and concentration of task-relevant attention can matter more than reproducing the full attention distribution.
|
| 1160 |
Age of Learning: Temporal Persistence of Prediction Errors as a Learning Signal
2609.32593
|
cs.LG
|
Chenyang Wang, Stefan Forsstr\"om, Roger Olsson, Di Yuan, Qing He |
Current machine learning algorithms primarily rely on instantaneous signals such as loss, margin, and prediction confidence to characterize model behavior. These signals indicate how difficult a prediction is at the current optimization step, but they do not c...Current machine learning algorithms primarily rely on instantaneous signals such as loss, margin, and prediction confidence to characterize model behavior. These signals indicate how difficult a prediction is at the current optimization step, but they do not capture how long the model has remained incorrect. We study this temporal dimension of learning and introduce Age of Learning (AoL), a learning-state variable that measures the persistence of prediction errors over time. AoL increases while an error remains unresolved and resets when a correct prediction is achieved, thereby distinguishing persistent under-learning from transient mistakes. We develop AoL-based training strategies for both offline and streaming settings. In offline learning, sample-level AoL is accumulated over training and aggregated into class-level states that guide adaptive reweighting and resampling. In streaming learning, where full historical access is unavailable, we maintain lightweight class-level AoL states using current and buffered observations. Across long-tailed classification settings, AoL improves or matches standard training baselines, with larger benefits when learning difficulty persists over time. Multi-seed streaming experiments further show reproducible gains under temporally stable imbalance. Analysis of class frequency, loss, and margin shows that AoL is related to conventional difficulty measures but captures additional information about error duration. These results suggest that temporal persistence provides a useful complementary signal for characterizing and controlling learning dynamics in imbalanced and non-stationary environments.
|
| 1161 |
Not Every Term Adds New Structure: Sobolev Novelty for Symbolic Regression
2609.32597
|
cs.LG
|
Boxiao Wang, Kai Li, Yuheng Jing, Tianyi Liu, Chen Li |
Symbolic regression (SR) aims to discover compact and meaningful mathematical equations from data, but searching the vast combinatorial space of symbolic structures remains challenging. Existing methods typically guide this process using expression-level objec...Symbolic regression (SR) aims to discover compact and meaningful mathematical equations from data, but searching the vast combinatorial space of symbolic structures remains challenging. Existing methods typically guide this process using expression-level objectives, such as fitting error, which assess a candidate equation as a whole but provide little information about whether an individual term contributes genuinely new structure or is largely redundant with the rest of the expression. We introduce \textbf{Sobolev Novelty}, a term-level measure of structural independence for symbolic equations. For each term, we construct an empirical Sobolev signature from its function values and exact derivatives over the observed inputs, and quantify how much of this behavior cannot be reconstructed by the remaining terms. We further derive a theory-calibrated threshold, yielding a principled and tuning-free criterion for identifying structurally novel terms. Using this threshold, 92.6\% of terms in benchmark ground-truth equations exhibit sufficient structural novelty, compared with only 38.3\% on average for expressions produced by 15 SR methods, revealing a substantial gap between scientific equations and current SR solutions. As a lightweight plug-in, Sobolev Novelty can be incorporated into diverse SR paradigms to support term pruning, search guidance, LLM feedback, and data selection, yielding consistent performance gains and demonstrating broad applicability.
|
| 1162 |
AnchorRep: Defending LLMs Against Cross-Model Adversarial Transfer via Representation Repulsion
2609.32602
|
cs.LG
|
Gal Wertheizer, Rom Himelstein, Tomer Peretz, Avi Mendelson |
Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability ...Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability across models, yet existing defenses are not designed for this cross-model threat. We find that cross-model transfer aligns with shared internal representation geometry, making it a natural defense target. AnchorRep targets this geometry directly with a lightweight LoRA adapter that pushes the defended model's internal representations of harmful prompts away from those of a frozen anchor model on the same prompts. Training uses a small set of harmful prompts and no adversarial examples. Across five models and four architectural families, AnchorRep reduces cross-model attack success rate to <=1.1% on 2,000 transferred attacks (0% on two), including the largest drop on Mistral (36% -> 1.1%). Existing defenses can reduce transfer, but only at high cost either inducing up to 77% degenerate benign output or increasing over-refusal by up to 18%. Because such degenerate benign outputs are not captured by standard refusal-based metrics, we introduce the Benign Garble Rate to quantify them. Our results suggest that cross-model robustness can be achieved by shaping representation geometry, without requiring attack-specific training
|
| 1163 |
Learning the Graph and the Embedding Together: Classifier-Independent Rewiring for Heterophilic Node Classification
2609.32613
|
cs.LG
|
Harshit Kumar, Sujan Chakraborty, Priyanka Saha, Pritam Kar, Saptarshi Bej |
Graph neural networks lose much of their advantage on heterophilic graphs, where connected nodes often carry different labels. Graph rewiring is a popular remedy, but rewiring methods are usually evaluated with a single classifier, which makes it hard to tell ...Graph neural networks lose much of their advantage on heterophilic graphs, where connected nodes often carry different labels. Graph rewiring is a popular remedy, but rewiring methods are usually evaluated with a single classifier, which makes it hard to tell whether the gains come from the new topology or from that particular pairing. We propose an affinity-guided rewiring method that estimates the graph and the node representation together. It alternates, in the spirit of expectation maximisation, between training a lightweight graph neural network on the current graph and re-weighting candidate edges under a modularity objective with a pseudo-label homophily term. Candidate edges come from a compact pool scored by a contrastively learned node similarity and a neighbourhood-distribution affinity. The method returns two classifier-independent outputs: a rewired graph and a node embedding learned on it. Across six heterophilic benchmarks and five downstream classifiers, it improves accuracy over the original graph with normalised features in 23 of 30 classifier-dataset combinations, with a mean gain of 5.8 points, and reduces the accuracy spread between classifiers about fourfold. A controlled ablation shows that the two outputs are each useful and play complementary roles: the embedding contributes most of the accuracy gain, while the rewired graph makes different classifiers agree. A fully unsupervised variant, which uses no labels during rewiring, retains most of the improvement. The rewired graphs are also more homophilic and improve label propagation and community detection.
|
| 1164 |
Stabilizing the Dynamic Low-Rank Training
2609.32615
|
cs.LG
|
Zhonghan Xu, Ling Wang, Junhao Chen, Jianwei Zhao, Jinwei Yang |
Training neural networks directly in a low-rank parameterization is an appealing route to reducing memory, compute, and storage simultaneously during both training and inference. Dynamic low-rank training (DLRT), which confines weights to a rank-$r$ manifold v...Training neural networks directly in a low-rank parameterization is an appealing route to reducing memory, compute, and storage simultaneously during both training and inference. Dynamic low-rank training (DLRT), which confines weights to a rank-$r$ manifold via the Galerkin projection of the gradient flow, is particularly attractive because it identifies efficient subnetworks on the fly without specialized initialization or post-factorization. However, DLRT fails to find trainable networks under high compression. In this paper, we derive the gradient flow of the best rank-$r$ approximation and point out that the offset of DLRT comes from a curvature-coupling term which is large and thus non-negligible under aggressive compression. Guided by this analysis, we propose a stable dynamic low-rank training method, named SDLRT, which maintains a lightweight compensation buffer that reinjects the top neglected singular directions. Additionally, we introduce a negative feedback on the truncation tolerance to stabilize each layer's rank. Experimentally, SDLRT reliably finds trainable subnetworks where DLRT collapses and as a PEFT adapter on DeBERTa-v3, it achieves the best average score on SuperGLUE at only $2.8\%$ parameter overhead over LoRA.
|
| 1165 |
Intuition vectors
2609.32619
|
cs.LG
|
Shahar Haim, Daniel C. McNamee |
Large self-supervised vision models learn representations that support scene segmentation and the semantic decomposition of physical objects. We ask whether their representational geometry supports transfer to visual reasoning problems without any task-specifi...Large self-supervised vision models learn representations that support scene segmentation and the semantic decomposition of physical objects. We ask whether their representational geometry supports transfer to visual reasoning problems without any task-specific fine-tuning. We hypothesized that relational representations may bridge perception and abstract reasoning by encoding similarities and transformations among visual inputs such that an intuitive, implicit form of reasoning may be performed via latent vector arithmetic. Specifically, we examine DINOv3, MAE, and random pixel projections on abstract and naturalistic Bongard problems, ARC-AGI-1 and ARC-AGI-2, and novel ARC-GEN instances. On both Bongard benchmarks, the accuracy of a simple nearest-centroid readout of frozen visual embeddings is within four percentage points of the task-specific baselines reported with the original benchmarks. In ARC, latent difference vectors summarizing demonstration input-output transformations, which we refer to as intuition vectors, show greater alignment with test vectors from the same task, whereas those from unrelated tasks are near orthogonal. This latent geometry is operational: transporting a query along its intuition vector consistently improves exact-output retrieval, reaching 70.7 on ARC-AGI-2 evaluation. Across 397,000 ARC-GEN instances from 794 tasks, single-pair intuition vectors identify the generating task with approximately 87\% leave-one-out accuracy. These findings suggest that latent vector arithmetic over frozen visual representations supports implicit rule inference across varied problem domains without a generative model component, indicating that inferring an abstract transformation and generating its instance-specific consequence may be separable capacities.
|
| 1166 |
Extremely Fast and Compact Binary Graph Representations via Randomized Operator Sketching
2609.32641
|
cs.LG
|
Srajan Agarwal, Megha P, Bikas C Das, Zakaria Laskar, Saptarshi Bej |
Graph neural networks typically rely on dense, floating-point node representations, which can impose substantial memory and computational costs. Binary graph hashing offers an alternative by encoding node information as compact bit strings. However, existing a...Graph neural networks typically rely on dense, floating-point node representations, which can impose substantial memory and computational costs. Binary graph hashing offers an alternative by encoding node information as compact bit strings. However, existing approaches either sacrifice global topological information for computational efficiency or incur substantial generation costs. We introduce an ultra-fast, entirely algebraic hashing method that constructs binary node representations directly from graph structure, without requiring node features or gradient-based training. Our method approximates a high-order structural transition matrix using randomized column sampling inspired by the Nystr\"om method and combines it with an efficient label-safe semantic propagation mechanism. The resulting continuous representations are discretized through column-wise thresholding to obtain compact binary codes. Experiments on ten node classification datasets show that the proposed method consistently improves classification accuracy over existing feature-free binary baselines while requiring sub-second code generation on many datasets. The resulting binary representations are also naturally suited to event-driven computation, making them compatible with neuromorphic spiking neural networks and gradient-free learning rules. These results demonstrate that simple algebraic approximations can provide an efficient alternative to learned pipelines for discrete graph representation learning.
|
| 1167 |
Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution
2609.32659
|
cs.LG
|
Ningfeng Yang, Tor M. Aamodt |
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed...Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise, they either introduce additional hyperparameters or memory overhead, or cannot consistently improve model accuracy. In this work, we propose optimization with $\textbf{C}$onstrained $\textbf{E}$mpirical $\textbf{W}$eight dis$\textbf{T}$ribution (CEWT), the first hyperparameter-free memory-overhead-free oscillation suppression method that consistently improves QAPT performance: an optimizer post-update step that projects weights to the nearest point in weight space whose empirical distribution (histogram) matches a zero-mean Gaussian. Our key insight is many quantizers are designed with the implicit assumption that the to-be-quantized data are permutations of samples from a zero-mean Gaussian, and this assumption is not true during QAPT. By enforcing the zero-mean Gaussian prior as a hard constraint, CEWT can suppress this detrimental noise. Empirical results on various combinations of SOTA quantizers and hypersphere optimizers suggest, that with a geomean increase of 4% in training time, CEWT can consistently reduce the pre-training perplexity (by an average of 2.5 and up to 21 points) of low-precision (down to 1-bit activations and weights and up to 610M parameters) LLaMA/GPT models without introducing any hyperparameters or storage overhead. Code is available at https://github.com/1733116199/cewt
|
| 1168 |
Equivariant Neural Primal-Dual Assignment for Maximum Common Edge Subgraphs
2609.32661
|
cs.LG
|
Jiaqing Xie, Yanchao Li, Zhuo Yang, Yuxin Wang, Tianfan Fu |
Maximum common edge subgraph (MCES) matching finds a partial vertex correspondence between two labeled graphs that preserves as many labeled edges as possible. Molecular similarity search requires matching many graph pairs, making the cost of repeated queries ...Maximum common edge subgraph (MCES) matching finds a partial vertex correspondence between two labeled graphs that preserves as many labeled edges as possible. Molecular similarity search requires matching many graph pairs, making the cost of repeated queries important. The strongest baseline attains accurate MCES solutions but trains a separate network for each pair. We introduce Equivariant Neural Primal-Dual Assignment (ENPDA), which learns a shared matching policy and applies it to new pairs without further training, answering queries roughly three orders of magnitude faster and recovering its training cost after a few dozen queries. The policy recomputes exact objective marginals for candidate matches and learns corrections and step sizes that update their scores. Target prices respond to competition when several source vertices favor the same target. Four update rounds and a Hungarian projection produce a partial one-to-one matching. We prove per-pair guarantees that hold for any network parameters. In exact arithmetic, reordering either graph permutes the assignment and price states, the projected matching is one-to-one, and repaired prices give a valid MCES upper bound. Subtracting the preserved-edge count bounds the optimality gap; combined with structural caps, these certificates prove global optimality for 60 of 291 native test pairs. On three molecular benchmarks with disjoint train/validation/test splits, ENPDA improves over an analytic counterpart with the same update and projection budget by 7.4-8.6 accuracy points; after one second of refinement search, 2.5-3 points of the gain remain. Transferred without fine-tuning to edge-deletion tasks from social and protein graphs, the policy gains 9.1-17.6 points over the analytic counterpart. When output matchings must keep aromatic rings intact, ENPDA recovers more reference bonds than the baselines on all three datasets.
|
| 1169 |
Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL
2609.32665
|
cs.LG
|
Qinwei Ma, Jingzhe Shi, Simin Fan, Ling Li, Mengdi Wang |
ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to...ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.
|
| 1170 |
STAMP: Predicting Out-of-Distribution Generalization without Target Data
2609.32672
|
cs.LG
|
Md Kawsher Mahbub, Milon Biswas |
Predicting whether a trained model will generalize under distribution shift remains difficult, especially when target-domain data are unavailable. We introduce STAMP (Semantic Temporal Augmented Model Prediction), a source-only, target-label-free criterion tha...Predicting whether a trained model will generalize under distribution shift remains difficult, especially when target-domain data are unavailable. We introduce STAMP (Semantic Temporal Augmented Model Prediction), a source-only, target-label-free criterion that estimates out-of-distribution (OOD) performance from paired source-domain images. STAMP computes the output-space correlation ratio $\eta^2=S_B/S_T$ by contrasting semantically stable pairs with random pairs: higher $\eta^2$ indicates that model outputs vary with semantic identity rather than nuisance variation. On 44 chest X-ray models spanning CNNs, ViTs, MetaFormers, foundation models, and SSL/VLM probes, temporal STAMP attains Spearman correlations of $0.844$--$0.855$ with macro AUROC on VinDr-CXR, CheXpert, and MIMIC-CXR; a class-matched variant improves single-class RSNA from $0.311$ to $0.663$. STAMP attains the best average source-only medical ranking and outperforms the target-domain ATC and AoTL estimators without any target data. On 27 ImageNet models, temperature-scaled STAMPTS attains $\rho{=}0.984$ on ObjectNet and $\rho\geq0.905$ on four additional distribution shifts, with partial correlations of $0.662$--$0.949$ after controlling for ImageNet accuracy. Requiring approximately 12s per model on one GPU, STAMP is a practical pre-deployment model-selection and auditing tool.
|
| 1171 |
SIFT: Enhancing Time Series Foundation Models via Semantic Invariance and Structural Fidelity Fine-Tuning
2609.32676
|
cs.LG
|
Yi Tang, Tengxue Zhang, Yang Shu, Chenjuan Guo, Chenchen Sun |
Time Series Foundation Models (TSFMs) have achieved remarkable zero-shot performance through extensive pre-training on massive time series datasets. Nevertheless, due to the low-dimensional properties and diverse structural patterns of time series data, perfor...Time Series Foundation Models (TSFMs) have achieved remarkable zero-shot performance through extensive pre-training on massive time series datasets. Nevertheless, due to the low-dimensional properties and diverse structural patterns of time series data, performing naive fine-tuning on TSFMs often leads to overfitting and falling into the mean-prediction trap. To address these challenges, we propose SIFT, a robust adaptation method that enhances time series foundation models by preserving Semantic Invariance and structural Fidelity throughout the fine-Tuning process. We employ semantic-invariant adversarial augmentation, which utilizes semantic spectrum decomposition to partition the semantic space and then generates perturbations within the non-core semantic subspace to bolster the model's robustness against these perturbations, mitigating overfitting. We implement a component-based structural fidelity enhancement, which facilitates component-wise mixup and imposes a reconstruction objective to improve the model's ability to preserve structural fidelity, alleviating the mean-prediction trap. Extensive experiments on representative TSFMs covering 10 real-world datasets demonstrate that SIFT can significantly enhance the performance of TSFMs.
|
| 1172 |
Self-Evolving Time-Series Forecasting Agents with Episodic Memory and Online Policy Learning
2609.32689
|
cs.LG
|
Junyi Wang, Yilin Wang, Wen Wu, Chao Zhang |
LLM-based agents are increasingly used for time-series forecasting because they can organise contextual information, perform multi-step analysis, and guide the sequence of actions required to complete forecasting tasks. Most existing agents focus only on the c...LLM-based agents are increasingly used for time-series forecasting because they can organise contextual information, perform multi-step analysis, and guide the sequence of actions required to complete forecasting tasks. Most existing agents focus only on the current forecasting instance. However, in real-world deployments, forecasting commonly operates online, with new forecasts issued from the currently available history as the forecast origin advances and the ground-truth targets of earlier instances progressively become available. These targets provide feedback on the actions taken in earlier instances, yet existing agents generally do not preserve or utilise this information to adapt their subsequent actions. To address this limitation, we introduce FASE, a Feedback-Aware Self-Evolving forecasting agent that converts such feedback into task-specific experience for subsequent forecasting instances. FASE combines episodic memory, which retrieves relevant completed instances, with online policy learning, which summarises the feedback accumulated across instances into ranking guidance. The proposed framework is evaluated on 29 dataset configurations selected from the GIFT-Eval benchmark. Across these 29 configurations, FASE attains the strongest aggregate point forecasting performance among the evaluated methods and reduces the normalised MAE by 9.1% relative to the best individual foundation model baseline. The results further indicate that the cumulative advantage of FASE increases as delayed feedback accumulates. Together, these findings demonstrate that FASE can continually self-evolve through feedback from completed forecasting instances without updating the parameters of the LLM.
|
| 1173 |
BiasReducer: Adaptive Bias Mitigation for Reward Models
2609.32720
|
cs.LG
|
Shuang Liu, Yongliang Miao, Yanguang Liu, Haoyi Xiong, Mengnan Du |
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct r...Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.
|
| 1174 |
Scaling Properties of Same-Family On-Policy Distillation
2609.32722
|
cs.LG
|
Yuntai Bao, Qinfeng Li, Guoqing Jiang, Liwei Chen, Zhiheng Qin |
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distilla...*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
|
| 1175 |
Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication
2609.32727
|
cs.LG
|
Chenyu Lu, Zijun Chen, Nian Si |
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, cover...We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback--Leibler and $f_k$-divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP to be weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for estimating the robust optimal average reward and $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for learning an $\epsilon$-optimal policy. Here, $p_{\wedge}$ is the smallest positive nominal transition probability and $u_{\delta}^{\ast}$ is a robust optimal bias function. We further provide an almost-tight explicit upper bound on $\operatorname{Span}(u_{\delta}^{\ast})$. Finally, we validate the predicted $n^{-1/2}$ convergence rate through numerical experiments.
|
| 1176 |
Predicting the Financial Impact of Supply Chain Risk for Major AI-Related Semiconductor Firms: A Heterogeneous Graph Patch Transformer Approach
2609.32741
|
cs.LG
|
Jianna Hur, Sagar Samtani |
Modern semiconductor production relies on a globally distributed, multi-tier supply chain in which financial stress at one firm spreads with a delay and eventually affects the revenue, inventory, and profitability of the companies that design AI chips. Most fi...Modern semiconductor production relies on a globally distributed, multi-tier supply chain in which financial stress at one firm spreads with a delay and eventually affects the revenue, inventory, and profitability of the companies that design AI chips. Most firms see only their direct partners, and prior predictive research has mainly targeted market-based risk measures, so few tools forecast how supply chain stress will appear in reported financials. In this study, we propose a heterogeneous graph patch transformer that forecasts these quarterly changes one and two quarters ahead. Learning from a 15,186-company network over 60 quarters, the proposed model fuses quarterly fundamentals with macro-trade, event, and disaster signals through learned gates, carries risk across supplier, customer, ownership, and headquarters relations through typed, direction-specific propagation, and encodes the propagated histories with patch-based tokenization. In preliminary experiments on 116 focal semiconductor firms, the proposed model achieves the lowest error on every target at both horizons, and its profitability advantage widens at the two-quarter horizon. These forecasts can help supply chain managers and investors act before disruptions appear in reported financials.
|
| 1177 |
Benchmarking EEG Foundation Models at Scale: Lessons from 20,000 Evaluations
2609.32743
|
cs.LG
|
Zhige Chen, Shu Peng, Chengxuan Qin, Rui Liu, Rui Yang |
Electroencephalography (EEG) foundation models (FMs) promise transferable neural representations, yet their advantages over strong supervised baselines and their prospects for further scaling remain unclear. To address these questions, we introduce EEG-Arena, ...Electroencephalography (EEG) foundation models (FMs) promise transferable neural representations, yet their advantages over strong supervised baselines and their prospects for further scaling remain unclear. To address these questions, we introduce EEG-Arena, an open-source benchmark covering 30 EEG FMs and 25 supervised baselines evaluated on 57 downstream tasks from 23 public datasets. Through more than 20,000 evaluations across five experimental protocols, we assess downstream performance, pretraining benefits, model size scaling, pretraining data scaling, and robustness to channel configuration. We find that (1) EEG FMs outperform strong task-specific supervised baselines on most evaluated tasks, particularly under non-bipolar settings; (2) compared with architecture-matched supervised training from scratch, pretraining improves both early optimization and final downstream performance, with larger and more consistent gains as more labeled downstream data become available; (3) existing EEG FMs do not exhibit a consistent positive relationship between parameter count and downstream performance; (4) under a fixed architecture, increasing the pretraining data scale yields sustained downstream gains; and (5) channel-flexible FMs achieve higher absolute performance than channel-constrained models across most evaluated channel configurations. Together, these findings demonstrate the downstream value of EEG FMs and identify pretraining data expansion as a promising direction for further progress. To support continued research, we release EEG-Arena as an open-source evaluation framework that provides shared infrastructure for reproducible benchmarking, model comparison, and community-driven development.
|
| 1178 |
Reuse or Relearn? A Spectral View of Earth Observation Foundation Models
2609.32756
|
cs.LG
|
Mehmet Ozgur Turkoglu, Valerio Marsocci, Dominik J. M\"uhlematter, Dominik Senti, Konrad Schindler |
Foundation models are rarely used as generic, frozen feature extractors; instead, they are fine-tuned for the target downstream application. This practice is particularly prevalent in Earth observation (EO), and it raises a question that downstream accuracy al...Foundation models are rarely used as generic, frozen feature extractors; instead, they are fine-tuned for the target downstream application. This practice is particularly prevalent in Earth observation (EO), and it raises a question that downstream accuracy alone cannot answer: does fine-tuning reuse the pretrained representation, or does it relearn a new one? We study this with spectral diagnostics that compare a model before and after adaptation, quantifying how well its dominant singular subspaces are preserved, how broadly the weight update is distributed, and how large it is. Using natural image models such as CLIP and DINO as a reference, we find that, under the evaluated fine-tuning settings, EO models undergo far larger, higher-rank updates and retain much less of their pretrained structure, so their downstream performance is often obtained with substantial changes to the pretrained weight structure. The diagnostics further provide insight into how cheaply a model can be adapted: where the pretrained subspaces are preserved, adapting a small fraction of the parameters can match full fine-tuning, and where they are not, it can fall behind. More broadly, foundation models, and EO foundation models in particular, should be assessed not only by benchmark accuracy, but also by how reusable their pretrained representation is.
|
| 1179 |
The Extender: A Log-Structured Transformer
2609.32759
|
cs.LG
|
Jakob Eriksson (UIC) |
We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a conc...We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each layer $\ell$ emits both a residual update $\delta_\ell$ which is added to $\mathbf{h}$, and a much smaller extension $\epsilon_\ell$ which is appended to $\mathbf{x}$. While both the FFN and $\mathbf{q}$ see $\mathbf{h}$, the attention $\mathbf{kv}$ projections take only $\mathbf{x}$ as input. As a result, the fully extended $\mathbf{x}$ contains the complete input for the $\mathbf{kv}$ projections of all layers, reducing the persistent attention memory footprint from $2Ld_{model}$ to $\sum|\epsilon_\ell|$. We find that with $|\epsilon_\ell|=32$, the Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters. For our 1664-wide, 924M model, the Extender's persistent attention memory footprint is $104\times$ smaller than MHA. The memory savings grow with model width.
|
| 1180 |
AnchorMixGAN: Anchor-Aligned Generative Semi-Supervision for DDoS Detection in Cloud-Integrated IoT Networks
2609.32764
|
cs.LG
|
Jin Yang, Xufeng Liu, Yong Hu, Xueyang Wang, Honglu Yang |
Detecting distributed denial-of-service (DDoS) attacks in cloud-integrated IoT networks is difficult when labeled traffic is scarce. Generative semi-supervised learning can supplement the available training data, but prediction shifts induced by synthetic view...Detecting distributed denial-of-service (DDoS) attacks in cloud-integrated IoT networks is difficult when labeled traffic is scarce. Generative semi-supervised learning can supplement the available training data, but prediction shifts induced by synthetic views may affect the targets assigned to real unlabeled flows. We propose AnchorMixGAN, a generative semi-supervised framework that addresses this problem through anchor-aligned target construction. Its Anchor-MAS module treats each real unlabeled flow as an anchor and creates alternative views by replacing one field group at a time with values from generated traffic. A frozen reference classifier predicts the anchor and its views; averaging and sharpening these predictions produces a soft target for the original flow. The flow and its target are then mixed with a labeled example using MixUp, allowing the detector to learn from both the original labeled records and the mixed examples. We analyze how reference-classifier error, view construction, and sharpening affect the target, and derive a bound on the resulting change in cross-entropy at a fixed detector prediction. At the reported 90% training setting with 20% of the training records labeled, AnchorMixGAN attains accuracies of 97.3%, 97.4%, and 96.5% on NSLKDD, BoT-IoT, and CICIoT2023, respectively, exceeding the corresponding MixGAN results by 1.6, 1.0, and 4.4 percentage points.
|
| 1181 |
Structuring Relations Among Learning Paradigms via Protocol--Objective--Resource Reductions
2609.32766
|
cs.LG
|
Junwei Su, Changjie Wang, Dongyang Chang |
Modern machine learning spans supervised, transfer, continual, meta-learning, and related regimes that often reuse the same hypothesis classes, architectures, and optimizers but differ in information access, objectives, memory, adaptation, and sample accountin...Modern machine learning spans supervised, transfer, continual, meta-learning, and related regimes that often reuse the same hypothesis classes, architectures, and optimizers but differ in information access, objectives, memory, adaptation, and sample accounting. This makes it difficult to determine whether one paradigm is genuinely distinct, a special case of another, or part of a broader structural hierarchy. We introduce a protocol-objective-resource (POR) framework that separates representational capacity from these design choices. A paradigm is specified by an environment class, observation protocol, admissible learners, performance functional, and resource accounting rule. POR reductions combine environment embeddings, learner compilers, threshold maps, and calibrated resource overheads. Our main theorem shows that such reductions imply worst-case complexity domination on embedded comparison classes, transferring upper bounds forward and lower bounds backward; under labeled-example accounting, this yields sample complexity domination. We also show that calibrated nontrivial accuracy regimes are necessary to avoid vacuous comparisons, and that strengthening the objective can strictly increase minimax sample complexity even with unchanged protocols and learner classes. Instantiating the framework for supervised, transfer, continual, and meta-learning yields canonical special-case relations: continual contains transfer, transfer contains supervised, and meta-learning contains supervised under aligned raw-example accounting. We further derive a non-exact episode-to-example reduction for episodic meta-learning and capture within-paradigm refinements such as replay memory and task identifiers. The framework thus provides a unified language for structuring learning paradigms and transferring complexity guarantees across them.
|
| 1182 |
Continual Learning via Self-Probe Gradients
2609.32771
|
cs.LG
|
Dongkyu Cho, Rumi Chunara, Sungmin Cha |
Adapting pretrained models to new data can cause catastrophic forgetting of previously learned behavior. When only a few past samples remain, they give continual learning methods sparse and narrow evidence about what to preserve. We show that language models c...Adapting pretrained models to new data can cause catastrophic forgetting of previously learned behavior. When only a few past samples remain, they give continual learning methods sparse and narrow evidence about what to preserve. We show that language models can expand this evidence through self-probing, in which the frozen model generates new inputs from the retained samples and records its own predictions on them. Unlike prior work that replays such data as training examples, our method, CPLUS uses self-probe and past-sample gradients to scale down parameter updates that conflict with prior behavior. Experiments with five language models on four benchmarks show three results. First, the same probes preserve more prior behavior as gradient signals than as replay data. Second, CPLUS learns the new data while consistently reducing forgetting more than existing baselines, especially when past data are scarce, and this protection extends to benchmarks not used for training. Third, we observe that CPLUS also becomes more effective as models grow: within the Qwen3 model family, it recovers an increasing share of the forgetting caused by standard fine-tuning.
|
| 1183 |
Beyond Gaussian Assumptions: Distribution-Aware Channel Capacity for Effective Connectivity
2609.32774
|
cs.LG
|
Jianan Jian, Jacob Kang, Nurahmed Multezem, Benjamin Li, Nan Xu |
Effective-connectivity estimation from brain signals often relies on Gaussian residual modeling, which enables tractable estimation but can discard informative distributional structure and distort inferred directed interactions when empirical residuals are non...Effective-connectivity estimation from brain signals often relies on Gaussian residual modeling, which enables tractable estimation but can discard informative distributional structure and distort inferred directed interactions when empirical residuals are non-Gaussian. We show across multiple modalities, species, and experimental conditions that both brain signals and fitted channel residuals frequently deviate from Gaussianity. We therefore introduce a distribution-aware, information-theoretic measure of effective connectivity based on channel capacity under general residual distributions. To estimate the resulting capacity from empirical, potentially non-Gaussian residuals, we develop a dual-flow min-max estimator based on normalizing flows, in which a generator searches over admissible input distributions under a power constraint while an observer estimates output entropy. We provide a theoretical characterization of the estimator, showing that the observer objective recovers differential entropy up to a KL approximation term, that the formulation reduces to classical Gaussian capacity as a special case, and that residual entropy can alter achievable information rates beyond variance; game-gap and error analyses further characterize optimization and approximation sources. In brain-like simulations with known directed connectivity, Dual-flow achieves the highest AUROC and AUPRC across ten conditions spanning diverse network topologies, hidden drivers, feedback, and heterogeneous hemodynamics, compared with Gaussian capacity, Granger causality, VAR-LiNGAM, and GIMME. Applied to multimodal brain signals, the method reveals time- and condition-resolved directed interactions consistent with known neurobiological circuitry. Together, these results establish a principled distribution-aware framework for effective-connectivity estimation beyond Gaussian residual modeling.
|
| 1184 |
Robust Bayesian Optimization with Q-Exponential Surrogates
2609.32775
|
cs.LG
|
Richard Cornelius Suwandi, Zhidi Lin, Feng Yin, Abdelhak M. Zoubir |
Bayesian optimization (BO) is a widely used framework for optimizing expensive black-box objectives, but standard BO methods often use Gaussian process (GP) surrogates whose Gaussian assumption is sensitive to outliers and heavy-tailed noise. We introduce q-ED...Bayesian optimization (BO) is a widely used framework for optimizing expensive black-box objectives, but standard BO methods often use Gaussian process (GP) surrogates whose Gaussian assumption is sensitive to outliers and heavy-tailed noise. We introduce q-ED-BO, a robust BO method whose surrogate follows a univariate q-exponential (q-ED) distribution, preserving GP-BO's closed-form posterior mean and variance while a shape parameter q controls the tail behavior, recovering the GP at q = 2 and growing heavier-tailed with wider confidence bounds as q decreases. This tractability yields a closed-form q-upper confidence bound (q-UCB) with sublinear regret, and an exact closed-form q-expected improvement (q-EI) that generalizes EI to the heavy-tailed predictive, recovering classical EI at q = 2. Experiments on beamformer and adaptive filter tuning with impulsive outliers show that q-ED-BO matches or exceeds existing baselines on clean data, and under corruption, improves the strongest baseline by approximately 0.7 dB in output SINR and 1.1 to 1.2 dB in misalignment reduction.
|
| 1185 |
Learning When to Recur: Token-Adaptive Recursion for Imbalanced Ophthalmic Domain Incremental Learning
2609.32785
|
cs.LG
|
Nanxi Yu, Kang Li, Ye Du, Xiaowei Hu, Weihua Yang |
Domain incremental learning is essential for adapting ophthalmic deep learning models to sequential clinical domains while preserving diagnostic expertise. Existing domain incremental learning methods predominantly address the domain shift induced by style var...Domain incremental learning is essential for adapting ophthalmic deep learning models to sequential clinical domains while preserving diagnostic expertise. Existing domain incremental learning methods predominantly address the domain shift induced by style variations. However, they often overlook the severe class imbalance inherent in real-world clinical scenarios, such as clinical referral systems. Institutions in these systems encounter drastic fluctuations in class priors, resulting in label distribution shift, a critical form of domain shift that triggers severe catastrophic forgetting. To address these challenges, we propose ToRe, a rehearsal-free and parameter-efficient framework that leverages frozen ophthalmic foundation models for robust incremental adaptation. ToRe employs a parameter isolation strategy to decouple domain-specific optimization paths, thereby helping mitigate catastrophic forgetting driven by both label distribution shift and style variations. Simultaneously, it introduces token-adaptive recursion that adaptively allocates additional computational depth across tokens, allowing simple tokens to exit the recursion loop early while subjecting complex tokens, such as those associated with lesions, to deeper recursive processing. This mechanism enhances the feature representations for minority classes, thereby supporting generalization throughout the domain incremental learning process. Extensive evaluations on nine heterogeneous datasets demonstrate that ToRe consistently outperforms state-of-the-art methods in overall performance across the three benchmarks, while maintaining near-zero forgetting. Together, these results support the applicability of ToRe to dynamic and imbalanced clinical environments. The code is available at https://github.com/Nancyolo/ToRe
|
| 1186 |
Getting Motif-ated: Controllable AI Compositions from Injected Motif Prompts
2609.32788
|
cs.LG
|
Chao Peter Yang, Cynthia Rudin, Yue Jiang, Simon Mak, Stephen Ni-Hahn |
Deep learning has transformed symbolic music generation by borrowing the training paradigms of large language models, with systems such as NotaGen now producing complete, stylistically convincing classical scores from a short prompt. These systems could become...Deep learning has transformed symbolic music generation by borrowing the training paradigms of large language models, with systems such as NotaGen now producing complete, stylistically convincing classical scores from a short prompt. These systems could become powerful creative partners, helping musicians generate endless possibilities. However, current systems expose almost no control handles on the music itself. In principle, control handles could be built into a foundation model trained from scratch, but this is rarely practical without massive amounts of quality annotated data and compute resources. We therefore present MotiGen, a recipe for retrofitting pretrained symbolic music models to use new instruction prompts. MotiGen injects a musical motif as a structured prompt line, reinforces it with a scalar attention bias toward the motif tokens, and learns the association with a two-phase curriculum. First, it learns from focused excerpts cropped around motif occurrences in the training data, then full scores including the motifs. Our experiments show that our model composes with the prompted motif in over 92.3\% of generated pieces. Generated pieces using a variety of motifs are included in our sample site: https://motigen-site.github.io/.
|
| 1187 |
FedHV: Low-Overhead Hypervolume Weighting for Federated Multi-Objective Optimization
2609.32790
|
cs.LG
|
Amirardalan Dehghanpour, Seyed Mohammad Azimi-Abarghouyi, Christopher G. Brinton |
Task-wise federated multi-objective optimization (FedMOO) trains a shared model for competing prediction objectives under heterogeneous data, partial participation, and communication constraints. Existing methods commonly derive task weights from gradient or u...Task-wise federated multi-objective optimization (FedMOO) trains a shared model for competing prediction objectives under heterogeneous data, partial participation, and communication constraints. Existing methods commonly derive task weights from gradient or update geometry. This requires task-specific information or iterative server-side optimization. We introduce FedHV, which maps reference-relative objective slacks to closed-form inverse-slack weights. Each client optimizes one weighted loss and returns objective estimates with its model update. The protocol adds exactly 2m auxiliary scalars per participating client, yielding Theta(d + m) total per-client communication, compared with the Theta(md) task-specific communication of FSMGDA, and requires no additional synchronization stage. We analyze the resulting one-round-delayed weights under client heterogeneity, multi-step local updates, partial participation, and finite-sample objective reports. Under a fixed-horizon positive-slack reference condition, with the prescribed horizon-dependent step size and vanishing report error, FedHV achieves an O(T^(-1/2)) rate for the average squared log-hypervolume gradient norm; persistent report error determines the resulting stationarity neighborhood. The same bound controls the squared Pareto-stationarity residual. Across six Dirichlet-partitioned non-IID settings from four vision benchmark families and three training seeds, FedHV exceeds FSMGDA and FedCMOO in mean accuracy in five settings. Among these methods and uniform scalarization, it achieves the highest worst-task accuracy in four settings and improves the difficult CIFAR-10 objective in both CIFAR10-MNIST settings.
|
| 1188 |
Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency
2609.32797
|
cs.LG
|
Xinhan Yang, Lu Lu, Shancong Mou |
Derivative-informed training improves neural operators by directly supervising their input-output sensitivities, which is crucial when neural operators are used as differentiable surrogates for inverse problems, PDE-constrained optimization, design, and contro...Derivative-informed training improves neural operators by directly supervising their input-output sensitivities, which is crucial when neural operators are used as differentiable surrogates for inverse problems, PDE-constrained optimization, design, and control. However, existing methods rely on offline-generated derivative labels, making data generation slow, storage-intensive, and difficult to adapt across datasets, resolutions, or perturbation bases. We propose sketched tangent consistency loss (sTCL), an on-the-fly derivative-informed training objective that, for PDEs with a known and differentiable residual, enforces sensitivity consistency directly from the governing equation without offline tangent labels or neural-operator architecture changes. sTCL uses randomly sketched input perturbations to provide a lightweight derivative-level physics constraint during training. However, raw forward-sensitivity residual penalties can fail for stiff, ill-conditioned, indefinite, or coupled saddle-point tangent operators. To address this, we introduce lightweight operator-aware loss-conditioning mechanisms selected by a simple tangent-operator decision rule. Across Helmholtz, nonlinear diffusion-reaction, Burgers, Allen-Cahn, and Navier-Stokes, with the same neural-operator backbone for all methods, the PDE-specific sTCL losses achieve solution and Jacobian accuracy comparable to offline derivative-informed training (DIFNO) while eliminating the offline derivative-data generation and storage pipeline. These results show that on-the-fly derivative-informed training need not merely amortize offline tangent-solve cost into training; with appropriate sketching and loss design, sTCL provides an effective drop-in path to derivative-informed neural operators. Code is available at https://github.com/yang9579/Derivative-informed-traning-on-the-fly.
|
| 1189 |
Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
2609.32814
|
cs.LG
|
Ilya Koziev, Ivan Oseledets |
Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra constructio...Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix multiplication with a sparser interaction table over the same weight blocks, we construct a family with quadratic arithmetic in the matrix dimension when the physical block size remains fixed, and derive finite-shape constraints for GPU execution. The construction is provably optimal for its bilinear rank by the Alder--Strassen bound and can be realized as row-typed rectangular projections compatible with causal masking and KV-cached decoding. We provide an empirical test of this approach by training two approximately 110M-parameter decoder-only Transformer LMs from the same recipe and 12.3B-token budget, differing only in their feed-forward layer: one uses ordinary dense matrix multiplication and the other uses the associative-algebra product. Across four prompt domains, the algebraic model achieves a 6.2--7.8\% increase in end-to-end generation throughput, while obtaining lower scores on all three reported downstream metrics. We treat these results as a feasibility and trainability check for the proposed approach at small scale, leaving further investigation to future work.
|
| 1190 |
Retimed Bellman Flows: Escaping the Impossible Triangle of Velocity Bootstrapping
2609.32828
|
cs.LG
|
Boyang Xu, Shengzhe Chen, Hao Yan |
Flow critics learn return distributions by transporting Gaussian noise to Bellman endpoints via continuous velocity fields. While velocity bootstrapping stabilizes training by querying a successor teacher, existing methods face a structural dilemma: on straigh...Flow critics learn return distributions by transporting Gaussian noise to Bellman endpoints via continuous velocity fields. While velocity bootstrapping stabilizes training by querying a successor teacher, existing methods face a structural dilemma: on straight paths, no residual-free same-time affine mapping can preserve Gaussian initial noise while maintaining an unbiased target. To overcome this limitation, we introduce Retimed Bellman Flows (ReBF). ReBF queries the teacher critic at a dynamically shifted earlier flow time, aligning intermediate student and teacher trajectories. By combining this retimed clock with fresh, decoupled noise generation, ReBF constructs a provably conditionally unbiased velocity target that preserves the Bellman fixed point and contracts under Wasserstein distances. Empirically, ReBF reduces $W_1$ distance to ground-truth return distributions by up to $7.7\times$ on synthetic MRPs and outperforms existing flow critics across 38 challenging OGBench and D4RL offline reinforcement learning tasks.
|
| 1191 |
UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models
2609.32831
|
cs.LG
|
Wanqi Yang, Yuexiao Ma, Mei Xie, Xiawu Zheng, Shiwei Liu |
Unified multimodal models combine understanding, generation, and editing within a single network, offering a promising foundation for versatile multimodal applications. However, growing multimodal contexts make KV cache storage and access increasingly costly. ...Unified multimodal models combine understanding, generation, and editing within a single network, offering a promising foundation for versatile multimodal applications. However, growing multimodal contexts make KV cache storage and access increasingly costly. Existing KV cache compression methods are typically tailored to specific tasks and single-modality caches, while overlooking changes in cache importance across tasks and timesteps. However, in unified multimodal models, each task involves multiple KV cache types, and both their composition and dynamics differ across tasks. As a result, a single compression policy overlooks task- and type-specific requirements, leading to the loss of critical information and degraded quality across tasks. Based on these findings, we propose UniCache, a training-free framework for task- and type-aware KV cache compression. UniCache identifies the cache segments activated by each task and assigns suitable compression policies through offline calibration. It coordinates their parallel execution under a shared storage budget through attention-guided allocation and task-aware temporal scheduling. Experiments show that UniCache achieves $5\times$ KV cache compression for understanding and editing and $2.5\times$ for generation with negligible quality loss, while increasing throughput by up to $1.78\times$ in long-context settings, significantly improving the practicality of scaling unified multimodal models to longer context.
|
| 1192 |
Over-the-Air Federated Learning in Heterogeneous Mobile Wireless Networks
2609.32832
|
cs.LG
|
Ming Xiang, Nicol\`o Michelusi, Yonina C. Eldar, Lili Su |
Over-the-air computation has emerged as a scalable and efficient solution for deploying federated learning algorithms in wireless networks by exploiting waveform superposition for simultaneous model aggregation. Most existing work struggles with heterogeneous ...Over-the-air computation has emerged as a scalable and efficient solution for deploying federated learning algorithms in wireless networks by exploiting waveform superposition for simultaneous model aggregation. Most existing work struggles with heterogeneous fading channels. These approaches either enforce unbiased updates from all devices or allow partial device contributions, requiring careful tuning of the convergence bound to mitigate bias under specific fading models. However, the former significantly amplifies receiver noise due to the weakest channel, whereas the latter is sensitive to fading model mismatch and converges only to a biased objective. To tackle these challenges, we propose FedOAG, which employs algorithmic components to automatically satisfy energy constraints via gradient normalization and evenly mix devices' updates through implicit gossiping. Importantly, FedOAG does not require transmission from all devices, nor does it rely on a specific fading model or knowledge of time-varying statistical channel distributions. We show that FedOAG converges to a stationary point of an unbiased non-convex objective at the best possible rate $O(1/\sqrt{T})$ for any stochastic first-order method. We corroborate our analysis with numerical experiments over dynamic wireless conditions on real-world datasets.
|
| 1193 |
Permutation-Equivariant Flow Matching for Alignment-Free Neural Weight Generation
2609.32833
|
cs.LG
|
Arkadi Piven, Yam Eitan, Guy Bar-Shalom, Fabrizio Frasca, Daniel Cremers |
A trained neural network can be represented by a parameter vector in high dimensions. Learning distributions over these vectors enables the generation of new models across various tasks and architectures. A central challenge is permutation symmetry: permuting ...A trained neural network can be represented by a parameter vector in high dimensions. Learning distributions over these vectors enables the generation of new models across various tasks and architectures. A central challenge is permutation symmetry: permuting hidden neurons can produce distant parameter vectors representing the same function. This introduces variations that a generative model must account for when learning from trained networks. Existing methods typically address this using networks derived from a common base model or costly approximate neuron alignment. We instead parameterize a flow-matching velocity field with a permutation-equivariant Graph Meta Network, enabling direct learning from independently trained networks without alignment. Extensive experiments show that our method closely reproduces the joint statistics of accuracy, functional similarity, and weight similarity of independently trained collections, providing evidence of generation beyond checkpoint memorization. A single conditional model also generates task-specific networks on heterogeneous architectures and generalizes to unseen hidden-width configurations. On a tabular domain-shift task, intermediate conditioning produces individual networks with performance comparable to logit ensembles across both domains. Taken together, our results show how permutation equivariance enables learning from diverse collections of independently trained networks without permutation alignment.
|
| 1194 |
HamiFormer: Dual-Expert Diffusion Fields with Affine Symplectic Maps
2609.32838
|
cs.LG
|
Haoxiang Huang, Xiang Liu, Shuwei Wang, Jingheng Ma, Sen Cui |
Predicting smooth dynamics and collisions requires modeling continuous evolution and abrupt state changes. We introduce HamiFormer, a dual-expert diffusion field combining whole-window denoising with residual-corrected Hamiltonian propagation. Their mixed-stat...Predicting smooth dynamics and collisions requires modeling continuous evolution and abrupt state changes. We introduce HamiFormer, a dual-expert diffusion field combining whole-window denoising with residual-corrected Hamiltonian propagation. Their mixed-state feedback attenuates the direct contribution of inherited autoregressive error: each mixed state guides subsequent propagation within the jointly refined window. Parallel Local Affine Scan (PLAS) amortizes iterative refinement across rectified-flow steps and evaluates derivatives in parallel across physical time. PLAS's affine symplectic maps achieve lower solver error and runtime than sequential explicit Euler in our evaluation. A Regime Model Tree specializes residuals and routing to balance typical-state accuracy against large tail errors. Our analysis gives conditions for physically consistent refinement and warm-start tracking, and finite-window error bounds under diffusion feedback. In 192-step evaluations, HamiFormer reduces normalized phase-space MSE by 26.3% against PhysiFormer on HamiBalls-1 and 21.4% against DiT on HamiBalls-2, with comparable model capacities. Disjoint-interval comparisons show the lowest late-horizon position and momentum errors among baselines on both datasets. Project page: https://hamiformer.github.io/.
|
| 1195 |
Transfer Learning for Edge Classification on Dynamic Text-Attributed Graphs
2609.32849
|
cs.LG
|
Tyler Bonnet, Marek Rei |
Learning transferable representations for dynamic text-attributed graphs (DyTAGs) requires models to capture underlying interaction dynamics that persist across domains. However, existing methods tend to overfit to domain-specific structural, temporal, and sem...Learning transferable representations for dynamic text-attributed graphs (DyTAGs) requires models to capture underlying interaction dynamics that persist across domains. However, existing methods tend to overfit to domain-specific structural, temporal, and semantic patterns, limiting edge classification performance under distribution shifts. To expose and address this, we formally establish a leave-one-domain-out (LODO) transfer learning protocol for edge classification on DyTAGs. Under this protocol, we demonstrate that state-of-the-art self-supervised methods for dynamic graph learning perform poorly when transferred to unseen domains. Strikingly, existing methods underperform a structurally and temporally unaware Bag of Events (BoE) model we introduce, which inputs only unordered sequences of node and edge text features. Proceeding from the BoE, we propose Spatio-Temporal Semantic Alignment (STSA), which integrates a spatio-temporal encoder that fuses representations of time deltas and node occurrence frequencies into a unified manifold. STSA is trained with a Contrastive Semantic Forecasting objective, which anchors edge representations to a multi-domain textual latent space initialized by a pretrained language model, providing a robust prior that outperforms BoE and all existing methods we evaluate.
|
| 1196 |
Logic Gate Networks and Lookup Table Networks as Lightweight Hardware Classifiers for Inter-patient ECG Arrhythmia Classification
2609.32854
|
cs.LG
|
Wout Mommen, Lars Keuninckx, Siddharth Patil, Paul Detterer, Achiel Colpaert |
Deep Differentiable Logic Gate Networks (LGNs) and Lookup Table Networks (LUTNs) offer a promising approach for very low power inference due to their use of simple binary logic operations instead of arithmetic. In this work, we generalize the logic gates of LG...Deep Differentiable Logic Gate Networks (LGNs) and Lookup Table Networks (LUTNs) offer a promising approach for very low power inference due to their use of simple binary logic operations instead of arithmetic. In this work, we generalize the logic gates of LGNs to more than two input pins, naturally arriving at networks consisting of $N$-input LUTs. To obtain a differentiable expression for training the $N$-LUT entries, we adopt the Boolean equation of a $2^N$:1 multiplexer (MUX) and optimize its input parameters during training. We investigate the applicability of LGNs and LUTNs to inter-patient ECG arrhythmia classification using the MIT-BIH data set. The proposed models achieve up to 94.41\% accuracy and a $j\kappa$ index of 0.683 on a four-class task, showing a competitive performance compared to existing CNN-, SVM- and SNN-based methods. Our LGNs and LUTNs only require an estimated 2.89k to 6.17k FLOPs, including preprocessing and readout, which is three to six orders of magnitude less than state-of-the-art methods. We verified our design, which consists of the preprocessing pipeline and a 6-LUTN classifier, by implementing it on a Xilinx Zynq-7000 ZedBoard. The complete system consumes a dynamic energy of 8.25 $\mu$J/inference, of which only 0.46 nJ is utilized by the LUTN classifier. These results show that both LGNs and LUTNs can be employed as lightweight hardware-based classifiers for inter-patient ECG arrhythmia classification.
|
| 1197 |
Muon Under Gradient Noise and the Limits of Orthogonalization Near Optima
2609.32861
|
cs.LG
|
Xiaohui Xie |
Muon replaces the momentum buffer of each weight matrix by its orthogonal polar factor. We ask what this orthogonalization does near an optimum, where minibatch noise dominates the gradient. Under a Gaussian noise model, the expected Muon update becomes a scal...Muon replaces the momentum buffer of each weight matrix by its orthogonal polar factor. We ask what this orthogonalization does near an optimum, where minibatch noise dominates the gradient. Under a Gaussian noise model, the expected Muon update becomes a scaled gradient step, so a linear method with a suitably matched learning rate reproduces Muon's first-order mean response. The stochastic update is a different matter: after the response is matched, Muon retains a nonlinear residual that is uncorrelated with the input noise and contributes additional covariance. A Hermite expansion shows how momentum acts on this residual. Its higher-order components decorrelate faster than the linear component, so momentum suppresses the residual's accumulated covariance relative to the linear part, but never eliminates it. In a local quadratic surrogate that evaluates the residual on the stationary noise buffer, the residual adds stationary covariance and raises the stationary loss floor at every stable step size, while leaving the contraction dynamics unchanged. Simulations of the full nonlinear recursion on quadratics and measurements on frozen transformer gradients support each step of this picture. Together, the results make a theoretical case for replacing orthogonalization by response-matched momentum SGD once optimization becomes noise-dominated.
|
| 1198 |
Staying on the Attractor: Supervising Neural Surrogates of 3D Turbulence Where They Leave It
2609.32864
|
cs.LG
|
Yilong Dai, Shaswata Mitra, Raj Patel, Yiming Sun, Shengyu Chen |
Neural surrogates are trained to predict 3D turbulent flows in place of direct numerical simulation (DNS). For chaotic flows, the goal is short-term pointwise accuracy followed by long-term physical and statistical fidelity. However, small prediction errors ca...Neural surrogates are trained to predict 3D turbulent flows in place of direct numerical simulation (DNS). For chaotic flows, the goal is short-term pointwise accuracy followed by long-term physical and statistical fidelity. However, small prediction errors can carry a surrogate away from the flow's attractor. Off-attractor states are poorly represented in training data, leaving their evolution weakly constrained. The learned dynamics can then amplify deviations and lead to blow-up, freezing, or statistical drift. We propose off-attractor supervision (OAS) to supervise neural surrogates where they leave the attractor. OAS teaches the model how the true Navier-Stokes dynamics would evolve from these states. Each selected state is paired with its own future computed by DNS. Three generators select a few hundred states for relabeling. The first collects states from the surrogate's own rollouts. The second uses surrogate attacks to target freezing, excessive amplification, and violations of incompressibility and energy balance. The third perturbs training states along an amplified direction and a strongly damped random direction of the dynamics. All attacks run on the surrogate alone, and DNS relabeling is performed offline once per selected state. Experiments on $128^3$ turbulence show that OAS increases the median time to failure from 21 to 721 steps. The compared baselines achieve medians of at most 110 steps, and the advantage holds across training seeds. OAS also achieves the lowest pointwise error at step 15 and the best long-horizon statistics among the compared methods. OAS integrates physical models into neural simulation by extending supervision from fixed reference trajectories to states where the surrogate is likely to fail. This principle can guide the development of more reliable scientific surrogates when deployment takes models beyond the coverage of their training data.
|
| 1199 |
Refreshing Less, Selecting Better: Reusing Stale Gradient Features for Efficient Influence-Based Data Selection
2609.32888
|
cs.LG
|
Jianchang Su, Yifan Zhang, Wei Zhang |
Gradient-based data selection methods such as LESS score each candidate by the alignment between its gradient and a target validation gradient, and recomputing per-example gradient features at every new checkpoint dominates their cost. Across three selection s...Gradient-based data selection methods such as LESS score each candidate by the alignment between its gradient and a target validation gradient, and recomputing per-example gradient features at every new checkpoint dominates their cost. Across three selection seeds, two model families, two candidate pools, and two target tasks, features cached at a post-warmup checkpoint and paired with fresh validation gradients preserve the ranking 40 optimizer steps later with Spearman correlation from 0.952 to 0.991, while the top-10% subset they induce misses 10 to 22% of the examples that full recomputation selects. We therefore propose Cached Diverse Influence Selection (CDIS), which recomputes gradient features for the top-ranked fraction $p$ of candidates under the stale scores, fits an affine calibration on the recomputed examples, and selects the final subset under source and length quotas. A refresh fraction at or above the selection fraction recovers the exact top-$k$ subset whenever the calibrated stale scores have bounded error, and the budget curves follow this rule: at $p=0.3$ the recovered top-$k$ subset coincides with full recomputation in every setting, allocating the same budget per stratum recovers the stratified subset at 0.92 to 1.00, and gradient-stage wall-clock drops 3.5 to 3.6 times. Iterating the cache over four checkpoints keeps top-$k$ overlap at 0.98 or higher at 1.9 gradient features per example against 4 for full recomputation, and stale-to-recomputed agreement on the refreshed examples provides a free check for unsafe reuse. Downstream, unconstrained top-$k$ selection collapses to a single data source and scores 13 points below random selection on GSM8K. CDIS scores 12 points above random selection with paired confidence intervals that exclude zero and trails full recomputation by 4.3 points, one training-run standard deviation, at 3.4 times lower selection cost.
|
| 1200 |
Measurement-Gated Provenance Attenuation for Frozen EEG Representations
2609.32889
|
cs.LG
|
Anuar Aimoldin, Yankai Chen, Ayana Mussabayeva, Nurdaulet Akhanov, Xue Liu |
Frozen EEG representations retain acquisition signatures as well as neural activity. Source predictability alone does not identify what should be removed: it can reflect measurement effects or genuine biological and population differences, which should not be ...Frozen EEG representations retain acquisition signatures as well as neural activity. Source predictability alone does not identify what should be removed: it can reflect measurement effects or genuine biological and population differences, which should not be erased. We propose Measurement-Gated Provenance Attenuation (MGPA), built on one principle: measurement evidence determines where correction may act, and preserved information determines what it should aim for. Paired measurement contrasts define a gate outside which nothing changes; inside it, the source score is moved to the value the preserved coordinates already predict: for a fixed affine score, this keeps the same information as any target set by those coordinates and needs the least expected squared movement. Closed-form and critic-guided iterative constructions apply it without source identity or encoder retraining. Three studies test the principle at increasing distance from its assumptions. Under controlled reference changes, where the source-task association is known, MGPA brings source to near chance with task performance unchanged, whereas erasing what predicts source (LEACE) lowers frozen-task AUROC from .753 to .656 while barely touching source; ablations attribute the attenuation to the gate's directions and 2.7x less movement to the conditional target. Across recordings from different devices and electrodes, iterative correction lowers source accessibility while preserving or improving task performance. Finally, one iterative map selected on one task and reused unchanged on existing heads for two others raises their worst-association AUROC (lowest over device-label shifts) by .057 and .019 over LEACE, at a cost to those heads while the training association holds. A reusable correction shows its value in how an existing predictor behaves once acquisition cues stop being reliable, not only in what a probe can read.
|
| 1201 |
A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks
2609.32890
|
cs.LG
|
Maciej Cicho\'n, Bart{\l}omiej Dmitruk |
Language models are increasingly evaluated as vulnerability detectors, and scores reported for similar models differ widely between papers. We measured how much of that difference evaluation protocol accounts for, with model outputs held fixed. In a paired tes...Language models are increasingly evaluated as vulnerability detectors, and scores reported for similar models differ widely between papers. We measured how much of that difference evaluation protocol accounts for, with model outputs held fixed. In a paired test, a model must flag a vulnerable function and clear its version after a fixing commit. Three choices that published evaluations make differently were varied one at a time: metric, verdict extraction and output budget. Seven frontier and large open models were evaluated on five released pair benchmarks and a set pooled for this work under one protocol, and 61 open models of 1.5B to 36B parameters on the pooled set. Function-level F1 follows how often a model flags both functions of a pair (Spearman $+0.86$ over 42 combinations) and is nearly unrelated to pair-level correctness ($+0.16$). On the pair score, extraction changes a model's number by $+0.001$ at the median and budget by $+0.02$ with an interval through zero, whereas the model changes a benchmark's number by up to 0.18 and the benchmark a model's by up to 0.16; a function-level score therefore measures flag rate more than model. For 37 of 68 models the difference between correct and reversed pairs is within its 95% interval of zero, the value for a null model that flags each function at a fixed rate, while both-flagged and both-cleared rates exceed that null by 0.055 on median, and for 64 of 68 both functions of a pair receive one answer more often than independence predicts: verdicts are determined by the text common to both functions. On length-matched pairs a linear probe on activations separates 0.78 by within-pair ranking, against 0.64 for a tf-idf baseline and 0.5 for length; the generated verdict is near chance for three of six models and at 0.55 to 0.57 for the other three, and a prompted logit is at chance for all six.
|
| 1202 |
Optimizing H-Graph Hybridization for Diffusion-Guided RRT
2609.32897
|
cs.LG
|
Omer Talmi |
Sampling-based motion planners guided by diffusion models produce high-quality trajectories in a single run, yet the stochastic diversity available at inference time is left largely unexploited. We present two inference-time diversification strategies for a fi...Sampling-based motion planners guided by diffusion models produce high-quality trajectories in a single run, yet the stochastic diversity available at inference time is left largely unexploited. We present two inference-time diversification strategies for a fixed, pretrained DiTree model, combined via H-Graph hybridization, and evaluate them on a holonomic AntMaze robot across 15 maze scenarios. The first, factorial diversity, sweeps the random seed and Diffusion Goal Bias (DGB) parameter, the second, refinement-only diversity, sweeps the diffusion refinement strength (RS) that controls how much an RRT-generated trajectory is edited. Because a single-run baseline only partially succeeds, we additionally compare H-Graph results with pool-based statistics. H-Graph improves the mean pool length of the factorial and refinement-only diversities by 18.8% and 14.5%, respectively. In addition, it also improves the best individual candidate's lengths by 9.7% and 6.8%, respectively. And last, compared with the successful baseline's trajectory length, it improves the results by 18.2% and 19.9%, respectively. These results show that inference-time parameter variation is a reliable, training-free source of path diversity, and that H-Graph hybridization reliably converts this diversity into shorter, higher quality trajectories.
|
| 1203 |
When Less Compute Is More: Adaptive Early Exit Improves Pretrained Outlier Detection
2609.32898
|
cs.LG
|
Tianyang Zhou, Leman Akoglu |
Pretrained tabular foundation models process every dataset at a fixed depth, with inference costs growing with dataset size. To address this, we present the first study of depth-adaptive early-exit for pretrained outlier detection models. While early-exit is t...Pretrained tabular foundation models process every dataset at a fixed depth, with inference costs growing with dataset size. To address this, we present the first study of depth-adaptive early-exit for pretrained outlier detection models. While early-exit is typically motivated by efficiency, we uncover a surprising benefit: exiting at the optimal intermediate layer can also improve detection performance on diverse real-world benchmarks by 4.7-7.3% on average, consistent across three distinct foundation models. First, we investigate the factors driving these gains, and identify a key mechanism: context pollution, i.e., the presence of outliers among in-context samples. Our analysis reveals that nearby in-context samples exert increasing influence on query predictions at greater depths, consistent with a retrieval-based view of these models. In effect, early-exit alleviates the adverse effects of retrieving accurate-yet-polluted neighbors, with gains of 13-21% when context pollution matches the natural outlier rate. Motivated by these findings, we pretrain a plug-in router to select a dataset-specific exit layer, using query outlier labels as privileged information available only during router training. The router operates post hoc, leaving the base model parameters and prediction head unchanged. Experiments on three large real-world benchmarks show that, on clean context, the router recovers up to 45% of the oracle gain with up to 1.8x speedup across three pretrained backbones, with larger gains as context pollution increases.
|
| 1204 |
Predicting the Next State Is Not Enough: JEPA Representations for Lean Theorem Proving
2609.32908
|
cs.LG
|
Aarnav Choudhary |
Neural theorem provers must both propose tactics and decide which valid successor states to explore. We study whether one-step Lean transitions provide a self-supervised signal for branch ordering. A JEPA-style model predicts latent successor representations a...Neural theorem provers must both propose tactics and decide which valid successor states to explore. We study whether one-step Lean transitions provide a self-supervised signal for branch ordering. A JEPA-style model predicts latent successor representations and scores only kernel-validated, nonterminal successors generated by a fixed pretrained ByT5 proposer. JEPA achieves higher Top-1 than matched InfoNCE on a same-theorem ranking diagnostic(50.18% versus 31.55%), but averages 282.3 of 987 solved theorems across three seeds versus 308 for proposer ordering, while requiring more tactic checks. In this setting, accurate one-step transition ranking is therefore insufficient as a long-horizon search value. The controlled evaluation separates representation from proposal quality and treats kernel-checked proof completion as the primary endpoint.
|
| 1205 |
Allspark: Weak to Strong Transfer via Alternating Chain of Thought
2609.32913
|
cs.LG
|
Kaizhao Liang, Junxiong Wang, Chen Liang, Zhendong Wang, Qiang Liu |
Recent progress in frontier models has renewed interest in large-scale reinforcement learning (RL), but the cost of generating large-model rollouts makes even testing RL recipes expensive. We ask whether reasoning improvements learned by a small, weak model ca...Recent progress in frontier models has renewed interest in large-scale reinforcement learning (RL), but the cost of generating large-model rollouts makes even testing RL recipes expensive. We ask whether reasoning improvements learned by a small, weak model can benefit a larger, stronger model without using the strong model's rollouts during training. We introduce Allspark, a training and inference framework for weak-to-strong transfer through alternating chains of thought. A weak teacher is trained alongside a frozen copy of the same model; the two alternate reasoning segments, and the frozen model produces the final answer. At inference time, a stronger student replaces the frozen training partner, while both models remain fixed. Because they communicate through text, the teacher can steer students from different model families and with different tokenizers. We study Allspark at two scales: controlled Qwen experiments across math and reasoning, and larger-scale Inkling experiments on ARC-AGI-2. The Inkling experiments show accuracy gains in within-family and cross-family settings, including transfer to Kimi and Nemotron, with benefits that vary across inference settings. These findings motivate reusing a trained weak teacher across strong students and examining the resulting accuracy--token tradeoff.
|
| 1206 |
Adaptive Latent Capacity for World Models
2609.32921
|
cs.LG
|
Idan Achituve, Lior Dikstein, Idit Diamant, Arnon Netzer, Hai Victor Habi |
We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns ...We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.
|
| 1207 |
The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning
2609.32927
|
cs.LG
|
Cristina V. Lopes, Yuangang Li, Justin Tian Jin Chen, Alberto Krone-Martins, Iris Ma |
Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating t...Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating the hypothesis that robust reasoning benefits from separating values from the types that control their manipulation. Can making this separation an architectural primitive improve the learnability and generalization of logical mechanisms? We introduce \textbf{STRAT} (\textbf{ST}ratified \textbf{R}egisters \textbf{A}nd \textbf{T}ypes), which partitions the residual stream into orthogonal Data and Type subspaces and uses Type-based attention and gating to govern Data transformations. Controlled arithmetic ablations identify three failure modes associated with data-control interference: the Linear Trap, Gradient Wall, and Open Gate Trap. Mechanistic analysis reveals interpretable logical structure, and in arithmetic, STRAT reduces median OOD error 35-fold relative to a Transformer baseline. On each of 11 datasets spanning 10 tasks, STRAT outperforms the Transformer baseline in mean accuracy, by 26 percentage points on average, with both models trained from 10 base examples per dataset using identical task-specific augmentation where applicable. Under distribution shift, STRAT's mean accuracy drops by only 2.39 percentage points, compared with 11.75 for the Transformer.
|
| 1208 |
Phenomenon-Graph JEPA: Label-Efficient Representation Learning for Contactless Cardiorespiratory Sensing
2609.32928
|
cs.LG
|
Constantino \'Alvarez Casado, Nhi Nguyen, Mohammad Rakibur Rahman, Le Nguyen, Manuel Lage Ca\~nellas |
Millimeter-wave (mmWave) radar and RGB-D cameras can record cardiac and respiratory waveforms continuously and without contact, but labeled recordings remain scarce because every label requires a supervised acquisition session. Self-supervised pretraining can ...Millimeter-wave (mmWave) radar and RGB-D cameras can record cardiac and respiratory waveforms continuously and without contact, but labeled recordings remain scarce because every label requires a supervised acquisition session. Self-supervised pretraining can exploit the unlabeled signals, yet contrastive methods depend on signal transformations and negative pairs whose validity is uncertain for cardiorespiratory data, where time warping changes breathing rate and distant windows can share the same physiological state. We present Phenomenon-Graph JEPA, a joint-embedding predictive architecture that learns from four processed one-dimensional streams without negative pairs or synthetic augmentation in its base configuration. Each stream is encoded by a temporal convolutional branch and a band-limited spectral branch. During pretraining, the model predicts stopped target embeddings along typed edges, which connect streams assigned to the same physiological phenomenon, and forward in time within a state episode. We treat this physiological typing as a testable hypothesis and compare it with wrong-edge and all-pairs prediction graphs. In the OMuSense-23 dataset, pretraining improves label-efficiency area over matched supervised training by 3.91 percentage points (95% interval 2.08 to 5.80, Holm-adjusted p = 0.006), and by 3.74 points under a second configuration evaluated on the same test participants. However, the wrong-edge and all-pairs controls do not establish a benefit from physiological typing. Optional Takens-inspired delay coordinates improve a validation comparison with learned history, whereas two wrist-only WESAD protocols do not establish a pretraining advantage. The study therefore separates the measured benefit of predictive representations from the physiological prior used to organize their training.
|
| 1209 |
Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linear Propagation
2609.32929
|
cs.LG
|
Kiarash Banihashem, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Silvio Lattanzi, Danny Mittal |
Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution t...Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution through temporal GNN architectures without focusing on efficient dynamic maintenance, or are restricted to linear propagation models based on Personalized PageRank. In this work, we study how to efficiently maintain node representations for non-linear GNN propagation under edge insertions and deletions. The propagation has no learned parameters, and only a classifier applied afterward is trained. For a broad class of standard activation functions, we develop a residual-based dynamic algorithm that selectively propagates local errors via push operations, maintaining an approximation to the evolving fixed point without full recomputation. We prove that our method achieves amortized $O(1/\epsilon)$ update time per graph change under a degree-normalized error guarantee. Our approach uses a potential-based analysis in a degree-scaled norm and, in contrast to prior work on the linear case, requires no randomness assumptions on either the update sequence or the input vector. For the linear special case, we additionally provide an exact dynamic algorithm via low-rank matrix inverse updates. Experiments on benchmark datasets show that incorporating non-linearity improves accuracy while preserving efficient update performance, yielding a scalable and theoretically grounded method for maintaining this propagation on dynamic graphs.
|
| 1210 |
Counting on Thinking: Tracing Evidence Integration in Language Models
2609.32932
|
cs.LG
|
Jingming Xue, Robert C. Wilson, Huadong Xiong |
Finite computational resources force a tradeoff between automatic System 1 processes and costly System 2 thinking. Large language models (LLMs) can spend extra computation on hard problems, yet direct answers struggle even with counting, an elementary operatio...Finite computational resources force a tradeoff between automatic System 1 processes and costly System 2 thinking. Large language models (LLMs) can spend extra computation on hard problems, yet direct answers struggle even with counting, an elementary operation humans and animals perform automatically. We ask why this requires thinking in LLMs. Evidence integration has long been used in psychology and neuroscience to probe decision-making. Our evidence-integration task presents one letter per conversational turn and asks which of two target letters appeared more often. A running count difference solves the task optimally by weighting every letter equally; tokens at each turn could represent and update this difference. Direct responses instead weighted evidence unevenly, with strong recency effects, and assigned less probability to the correct answer as difficulty increased. Thinking improved performance and made integration weights nearly uniform, yet final-query attention remained concentrated on the sequence ends in both modes. Reasoning trajectories showed models revisiting input, recounting letters, and checking intermediate counts that informed the answer, suggesting that thinking constructs the accumulated count that direct responses lack rather than reading out one already formed. Reasoning-token costs grew with the number of letters far more than with coherence. Outcome feedback did not bring this computation into direct responses: under in-context reinforcement learning (ICRL), performance deteriorated over repeated games and recency effects strengthened, yet models grew more confident. Humans and animals amortize such computations into automatic processes, whereas current LLMs still pay for them with thinking on every trial. Which operations learning can make directly available remains central to how future models allocate computation.
|
| 1211 |
Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs
2609.32933
|
cs.LG
|
Nam Phuong Tran, Trinh Ha Mai Huynh, Tuyen Pham Le, Van-Truong Nguyen, Quan Nguyen |
In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforc...In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has established such guarantees in exact-gradient or tabular online settings, yet scalable results for structured large-state problems remain open. We develop a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs). Our analysis separates contraction of the regularised primal-dual dynamics from actor approximation and statistical errors in policy evaluation under online exploration. This enables model-free on- and off-policy learning with structured function approximation: optimistic policy evaluation avoids explicit transition-model construction, while a compact parametric actor avoids maintaining mixtures or histories of past policies. We instantiate the framework for linear CMDPs and general function approximation, obtaining representation-dependent complexity and improved target-accuracy dependence over prior optimistic regularised primal-dual analyses. We further validate the stabilising effect predicted by our theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpart shows larger oscillations.
|
| 1212 |
The Impact of Stochasticity on the Rashomon Effect in Machine Learning
2609.32934
|
cs.LG
|
Andrea Apicella, Francesco Isgr\`o, Andrea Pollastro, Roberto Prevete |
Neural network training is inherently stochastic, with factors such as weight initialization leading to distinct models despite comparable predictive performance. This phenomenon is commonly associated with the Rashomon effect, which describes the existence of...Neural network training is inherently stochastic, with factors such as weight initialization leading to distinct models despite comparable predictive performance. This phenomenon is commonly associated with the Rashomon effect, which describes the existence of multiple near-optimal models for the same task. Although the Rashomon effect has received increasing attention, it remains unclear whether different sources of training stochasticity contribute similarly or differently to its manifestations. In this work, we present an empirical study of the Rashomon phenomenon along three complementary dimensions: solution-space multiplicity, predictive multiplicity, and decision-basis multiplicity. These dimensions are quantified through the size of the empirical Rashomon set, predictive ambiguity, and agreement between XAI attribution maps, respectively. By independently controlling three standard sources of stochasticity, namely weight initialization, mini-batch data ordering, and dropout, we isolate their respective contributions to each dimension of the Rashomon phenomenon. Experiments on tabular and image classification benchmarks reveal that these sources affect the three dimensions in different ways. In particular, larger empirical Rashomon sets do not necessarily correspond to greater predictive disagreement or lower explanation agreement, indicating that solution-space, predictive, and decision-basis multiplicity capture complementary rather than interchangeable aspects of the Rashomon effect. Overall, our results show that training stochasticity influences not only predictive performance but also the stability of predictions and explanations, highlighting the importance of identifying the specific sources of stochasticity responsible for different manifestations of the Rashomon phenomenon when assessing the reliability, reproducibility, and interpretability of neural network models.
|
| 1213 |
TwinS-GCN: Spectral conjugate for Spectral Graph Convolutional Networks
2609.32940
|
cs.LG
|
Chun Hei Michael Chan, Flavia Petruso, Dimitri Van De Ville |
Graph convolutional networks propagate information by repeated local aggregation through a graph shift operator; i.e., a $K$-layer network reaches $K$ hops neighborhood. On the one hand, such spreading can lead to oversmoothing. On the other hand, long-range d...Graph convolutional networks propagate information by repeated local aggregation through a graph shift operator; i.e., a $K$-layer network reaches $K$ hops neighborhood. On the one hand, such spreading can lead to oversmoothing. On the other hand, long-range dependencies demand the depth. Transporting information on long distances and without attenuation requires the shift to distinguish a direction of flow, which a symmetric operator cannot perform but a directed one can fulfill. A natural way to extract pure-directionality is to take the skew-symmetric part of the shift operator through the Cartesian split, which, however, generally does not commute with the shift itself, meaning that the filters built on it are not shift-invariant. We instead use the spectral conjugate; i.e., the image of the operator under $\tau:z\mapsto \bar{z}$, which commutes with the shift and splits it into a dissipative and a non-dissipative part. Two filter families follow: a sum filter, whose non-dissipative component transports signal without energy loss, and a ratio filter, ratio in the pair of components rather than polynomial in the shift. Both arise from non-holomorphic kernels, placing them outside the holomorphic class underlying classical spectral convolution. On the directed cycle, the ratio filter becomes an IIR filter with global impulse response, for which we prove a long-range reach gap against every degree-$K$ polynomial filter. Chebyshev reparameterization gives stable vertex-domain layers with real coefficients, yielding TwinS-GCN, which solves graph transfer tasks at reduced depth and is competitive with state-of-the-art graph convolutional networks on node classification benchmarks.
|
| 1214 |
Generative Priors Conditioned on Natural Language for Bayesian Inversion in PDEs
2609.32941
|
cs.LG
|
Pengyu Zhang, Mark Girolami, Arnaud Vadeboncoeur |
Inferring quantities of interest (QoI) from data is a central task in Science and Engineering. In such contexts, we often have access to both quantitative data and qualitative data. Quantitative data may be represented by noisy sensor measurements, simulation ...Inferring quantities of interest (QoI) from data is a central task in Science and Engineering. In such contexts, we often have access to both quantitative data and qualitative data. Quantitative data may be represented by noisy sensor measurements, simulation data, re-analysis data; qualitative data may be in the form of text descriptions of experimental setups, expected experiment outcomes, and human-perceived system behaviours. The task we address in this paper is the following. Given a training set of paired qualitative text and quantitative QoI data, we learn to exploit the inherent correlation between the two modalities to learn a highly informative data-driven natural-language-conditional Bayesian prior, such that when presented with a new physical system, we can coherently combine (i) the training dataset, (ii) qualitative text describing the new system, and (iii) a small number of noisy sensor readings from that new system, to perform inference and uncertainty quantification (UQ) over the QoI. To achieve this task, we develop two parallel approaches, one uses conditional diffusion and the other conditional autoencoders, and compare both against classical Bayesian methodology, unconditional generative models and deterministic supervised methods. Each approach has specific strengths and tradeoffs; conditional autoencoder offers theoretical tractability, allows for fast posterior sampling, and provides better-calibrated UQ, whereas conditional diffusion is explored for greater expressiveness and capturing complex posteriors with irregular QoI fields. The approach is tested on the steady-state heat equation, damped Helmholtz equation, and UK weather reanalysis data.
|
| 1215 |
Optimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial Inventory
2609.32949
|
cs.LG
|
Mengxiao Zhang, Yingfei Wang, Haipeng Luo |
We study online dynamic pricing with censored demand, where an arbitrary inventory level is revealed before pricing and may adapt to past observations, while demand follows an unknown, price-dependent distribution that is stationary over time. For a horizon of...We study online dynamic pricing with censored demand, where an arbitrary inventory level is revealed before pricing and may adapt to past observations, while demand follows an unknown, price-dependent distribution that is stationary over time. For a horizon of $T$ rounds, Xu et al. [2026] achieved $\widetilde{\mathcal{O}}(\sqrt{T})$ regret under restrictive structural assumptions including linear demand, price-independent additive noise, and conditions relating inventory levels to the noise support. Our first contribution is to extend this framework to a substantially more general and statistically harder nonparametric setting, requiring only the natural assumption that expected sales are nonincreasing in price and allowing nonlinear demand curves and price-dependent noise. For this model, we first propose a simple baseline, Double-Grid-UCB, which discretizes both price and inventory and achieves $\widetilde{\mathcal{O}}(T^{3/4})$ expected regret using separate revenue estimates for each price-inventory grid pair. Then, we develop Threshold-UCB, which improves the expected regret to $\widetilde{\mathcal{O}}(T^{2/3})$. Unlike Double-Grid-UCB, Threshold-UCB reuses sales observations across inventory levels through shared estimates of demand-tail probabilities, allowing the same data to support revenue upper bounds for multiple inventories rather than a single inventory bin. We also complement this upper bound with an $\Omega(T^{2/3})$ lower bound via a reduction from stochastic posted pricing, establishing its minimax optimality. Finally, extensive experiments across inventory processes, demand functions, and noise models demonstrate consistently superior performance of Threshold-UCB over benchmark algorithms.
|
| 1216 |
Constrained Flow Policy Updates: A Generalized Schr\"odinger Bridge View
2609.32952
|
cs.LG
|
Boyang Li, Matthew Kim, Sylvia Herbert |
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. Reward and safety can induce multimodal action distributions, challenging the prevailing primal-dual methods: Gaussian actors may collapse onto a s...Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. Reward and safety can induce multimodal action distributions, challenging the prevailing primal-dual methods: Gaussian actors may collapse onto a single suboptimal mode, and optimization over the nonconvex Lagrangian landscape can be unstable. Diffusion and flow policies can represent such distributions, but recent work with a diffusion actor relies on estimating and matching the score of an augmented-Lagrangian target policy. Instead, we differentiate the augmented objective directly through the generation path of a flow policy, so no score needs to be estimated. Because a flow policy lacks a readily available action log-density for entropy regularization, we build on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and propose Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL. We formulate its update as a constrained one-ended generalized Schr\"odinger bridge and show that, for each source draw, this path-space problem is exactly an entropy-regularized problem in action space. At positive noise, its solution reweights the reward-only action distribution only where the estimated cost exceeds a threshold set by the Lagrange multiplier. As the noise vanishes, the optimal value converges to that of a least-energy map objective that the flow policy optimizes directly. Across seven Safety-Gymnasium tasks, RAFALE achieves competitive reward with mean final cost within budget on every task, whereas strong baselines trade one for the other; ablations support the necessity of both its augmented objective and its flow actor.
|
| 1217 |
Efficient Message Passing for Partial Differential Equation Priors
2609.32956
|
cs.LG
|
Anna Kazachkova, Leonhard Hennicke, Rainer Schlosser, Ralf Herbrich |
Prior information for real-world physical quantities is most elegantly expressed via partial differential equations (PDEs). In this paper, we propose a novel way to solve PDEs using probabilistic inference on a factor graph. In general, factor graphs provide a...Prior information for real-world physical quantities is most elegantly expressed via partial differential equations (PDEs). In this paper, we propose a novel way to solve PDEs using probabilistic inference on a factor graph. In general, factor graphs provide a natural way to encode prior knowledge into a model as explicit factors; here, this knowledge is provided by a governing PDE, which narrows the solution space, while observed data further shape the posterior over the parameters. The approximate parameter posterior is inferred using message passing based on moment matching, without posterior sampling or global gradient-based optimization. We demonstrate our approach on the first-order advection and the second-order semi-linear Fisher-KPP equations, where it achieves predictive accuracy comparable to a standard baseline while providing structured predictive uncertainty. Moreover, the inferred posterior marginal means and uncertainty structure match more closely those obtained using Hamiltonian Monte Carlo than the evaluated mean-field variational inference baseline, while requiring up to 10x less training time in our experiments, with inference speed comparable to variational inference.
|
| 1218 |
Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
2609.32961
|
cs.LG
|
Ritul Satish, Prasoon Sinha, Akiho Kawada, Neeraja J. Yadwadkar |
As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle deci...As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle decisions about what to compress, when to compress, and how much to remove into fixed policies. A systematic characterization is needed to disentangle these decisions and reveal how each affects task success and execution cost. We systematically vary these decisions across three open-weight models on SWE-bench Verified and Terminal-Bench 1.0. Across nearly 35,000 agent runs, we measure task success, token use, end-to-end latency, and estimated cost. We find that fewer tokens need not mean faster or cheaper execution: on Terminal-Bench with Qwen, policies using roughly one-third as many tokens can take 20-80% longer than the uncompressed agent. Policies with similar overall success can solve different tasks, while the same policy can perform quite differently across models. Our results motivate evaluating compression by its effects on agent execution and tailoring policies to the task, model, and workload.
|
| 1219 |
Clipped or Unclipped? Finite-Sample Trade-offs for Averaged SGD under Heavy-Tailed Noise
2609.32962
|
cs.LG
|
Alexandra Suvorikova, Egor Gladin, Darina Dvinskikh, Artem Agafonov, Mohammad Alkousa |
Gradient clipping is widely used to stabilize training, but it need not improve the statistical accuracy of averaged SGD, even under heavy-tailed noise. We derive a finite-sample comparison of clipped and unclipped Polyak-Ruppert averaged SGD under finite cond...Gradient clipping is widely used to stabilize training, but it need not improve the statistical accuracy of averaged SGD, even under heavy-tailed noise. We derive a finite-sample comparison of clipped and unclipped Polyak-Ruppert averaged SGD under finite conditional $p$-th moments, $p\ge2$. Our main result gives explicit accuracy and confidence conditions under which, for $p>2$, the Gaussian term dominates the unclipped deviation bound, so clipping need not improve its leading order. By balancing clipping bias and concentration, we obtain a bound in which the heavy-tail correction depends logarithmically rather than polynomially on the inverse failure probability. At $p=2$, this improves the confidence dependence of the leading bound. We establish sharpness of the unclipped heavy-tail term through an exact one-dimensional quadratic recursion and extend the comparison to projected convex SGD. We also prove concrete costs of clipping: every fixed finite threshold increases asymptotic variance on a scalar Gaussian quadratic, while whole-gradient clipping can shift the limiting point under asymmetric noise.
|
| 1220 |
Self-Confirming Superposition Traps in Reinforcement Learning
2609.32966
|
cs.LG
|
Dai Shi, Andi Han, Feng Chen, Yiqun Duan, Junbin Gao |
Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally op...Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirming superposition trap, every optimal code assigns overlapping directions to features that rarely occur together under the current policy. An alternative action brings them together, causing interference that lowers its return and reinforces avoidance, although refitting to that action would yield more return at the same capacity. We characterize the dimensions admitting a trap in a tied two-step model and show separately that equal feature frequencies, continued visitation, and independent controller learning need not prevent it. Because fitting weights errors by visitation, an avoided action can lose its return advantage at little cost to the objective. In a finite-action model, we bound this distortion and derive a replay condition: sufficient training weight on the best separately adapted action preserves its ranking despite residual error. Neural PPO experiments show how the feedback develops during learning: agents initialized toward different actions develop different interference patterns, opposite mean return rankings, and different final policies at the same capacity. We therefore test whether retaining access to neglected states can improve control. Keeping these states in training reduces measured interference and improves sequential return, with gains even when the encoder is frozen. Related interventions on state access, replay weights, and feature overlap improve control on MiniGrid and DMControl. For agents that learn through a world model, protected fitting improves DreamerV3--Crafter's cumulative training scores at unchanged capacity.
|
| 1221 |
Adaptive Ensemble Selection for Noisy Labels on Tabular Data
2609.32976
|
cs.LG
|
Faizaan Ali, Inwon Kang, Oshani Seneviratne |
Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science wo...Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label noise is critical for building reliable models. We propose a data-centric reasoning module for AI data science systems that automatically diagnoses dataset quality and selects appropriate cleaning strategies. Given a dataset, a meta-model predicts weights over a diverse set of detectors, including confidence-based, neighborhood-based, and distributional methods. Across benchmark datasets with controlled noise, our approach achieves performance comparable to a Confident Learning baseline on average, with dataset-dependent gains and losses, particularly in heterogeneous regimes. We further show that detector effectiveness is systematically linked to dataset properties. These results demonstrate the value of descriptor-driven, data-centric ensembling as a component of AI-assisted data-science pipelines for robust dataset assessment and model reliability.
|
| 1222 |
Feasible Flow Matching for Graph Reconstruction via Within-Sampling Primal-Dual Guidance
2609.32980
|
cs.LG
|
Haoming Chen, Nicolas Zilberstein, Santiago Paternain, Santiago Segarra |
Graph reconstruction from partial observations often comes with structural side information, such as degree bounds, triangle counts, or an edge-density band. Prior-Informed Flow Matching (PIFM) reconstructs graphs by transporting a local prior toward the graph...Graph reconstruction from partial observations often comes with structural side information, such as degree bounds, triangle counts, or an edge-density band. Prior-Informed Flow Matching (PIFM) reconstructs graphs by transporting a local prior toward the graph distribution, but it provides no mechanism to incorporate this side information. We put forth Constrained Primal-Dual PIFM (CPD-PIFM), which augments the sampler with Lagrange multipliers that evolve along each trajectory. The multipliers respond to constraint violations at a predicted endpoint and guide subsequent sampling steps without retraining. We prove that the sampler inherits PIFM's permutation equivariance and bound its expected terminal slack by a term that decays as the inverse square root of the number of steps, plus two approximation terms. On three link-prediction benchmarks and nine combinations of datasets and constraints, CPD-PIFM raises feasibility by 11-26 percentage points and remains competitive with fixed guidance without selecting a separate multiplier for each constraint.
|
| 1223 |
What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use
2609.32991
|
cs.LG
|
Yixiao Chen, Ke Cheng, Jiangtao Guan, Shuo Huang, Yue Liu |
What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data princip...What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data principle connects them: localize the missing operation, preserve its causal relation, vary shortcut-bearing context, and re-audit the residual. Formation-sensitive selection and prerequisite ordering accelerate a binding-matching-transport path; a brief early prefix from the same training multiset retains a validation advantage through 100B tokens. Availability counterfactuals then distinguish writing content from invoking available memory, while paired supervision and context-opportunity ranking improve matched route decisions and long-context answer likelihood. A continuous 350M-model experiment connects the three interventions on the same facts: early circuit training improves subsequent learning, and the complete sequence outperforms stage-replacement controls on facts withheld from Use teaching. Independent query surfaces and opposed-source decisions expose conditional arbitration as the remaining frontier. Together, these results show why a change in the limiting operation calls for a change in supervision, not merely a new ranking of difficult examples.
|
| 1224 |
Saturation-Insensitive Dueling Bandits with General Function Approximation
2609.33011
|
cs.LG
|
Chenggong Zhang, Xuheng Li, Qiwei Di, Weitong Zhang, Quanquan Gu |
We study contextual dueling bandits with general function approximation under the Bradley-Terry-Luce (BTL) preference model. A key challenge in this setting is the saturation of the preference model: when the current reward model can already distinguish two ac...We study contextual dueling bandits with general function approximation under the Bradley-Terry-Luce (BTL) preference model. A key challenge in this setting is the saturation of the preference model: when the current reward model can already distinguish two actions with high confidence, the resulting preference feedback becomes weakly informative, making it difficult to further improve reward estimation. Consequently, existing sample-complexity analyses often depend on the inverse-derivative factor $1 / \sigma'[\Delta_{r^\ast}]$ which can be prohibitively large when the link function $\sigma$ saturates for large reward gaps $\Delta_{r^\ast}$. To address this issue, we introduce `SI-CDB`, an algorithm that selects opponent arms using a carefully designed heuristic for arm selection. This design enables saturation-insensitive reward learning and recovers the near-optimal dependence for linear reward classes, eliminating the unfavorable $1/\sigma'(\cdot)$ factor. The core of our analysis is a localized Eluder dimension framework tailored to dueling bandits with general function approximation. Our theoretical results also explain why two-arm regret analysis is crucial for improving single-arm performance in dueling bandits.
|
| 1225 |
Relative Generalization Invariance of LLM Pretraining
2609.33016
|
cs.LG
|
Fengzhuo Zhang, Shuche Wang, Shenggui Li, Tianyu Ruan, Jianliang He |
Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear....Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by studying relative generalization. We introduce Relative Generalization Invariance (RGI), the invariance of the validation-loss difference between any two tokens across models. We show that RGI approximately holds across a wide range of optimizers and moderate architectural variations, suggesting that these choices induce an approximately uniform shift in token-wise losses. In contrast, changing the training data stream can substantially alter relative generalization. We further show that RGI cannot be explained by the neural tangent kernel or mean-field regimes alone and prove that it can emerge in an overparameterized quadratic model. Overall, our work identifies RGI as a new phenomenon in LLM pretraining that helps distinguish the effects of optimizers and architectures from those of training data.
|
| 1226 |
Low-Rank Single-Index Bandits with Unknown Links: From Matrices to Tensors
2609.33025
|
cs.LG
|
Zhongxuan Liu, Yue Kang, Thomas C. M. Lee |
Low-rank matrix and tensor bandits exploit structured interactions but typically assume a known reward link. Recent single-index bandit methods accommodate unknown links without directly exploiting matrix or tensor rank. We address this gap by studying stochas...Low-rank matrix and tensor bandits exploit structured interactions but typically assume a known reward link. Recent single-index bandit methods accommodate unknown links without directly exploiting matrix or tensor rank. We address this gap by studying stochastic matrix and tensor bandits with an unknown shared Lipschitz link and a low-rank index parameter under known regular candidate distributions and finite-variance noise. For monotone links, T-ESTOR combines robust, rank-adaptive Stein estimation with epoch-based greedy selection. Under exact selected-score access and a uniformly positive selected-design Stein signal, it achieves square-root regret with dimension dependence determined by the low-rank structure. For every admissible design, the monotone lower bound matches the rank, dimension, and horizon dependence up to logarithmic factors at large horizons, for fixed menu size and model/design constants. For nonmonotone links under a nonzero base-law Stein signal, T-BSTOR combines structured estimation with robust bin-based learning and attains the optimal $\widetilde{O}(T^{2/3})$ horizon rate for fixed dimensions, menu size, and model/design constants. Synthetic and CCLE-based experiments illustrate the benefits of structured estimation relative to vectorized and competing single-index baseline methods.
|
| 1227 |
\L{}ukasiewicz Neural Networks Extended: Residual Architectures and Crystallization Strategies for Interpretable Rule Extraction
2609.33028
|
cs.LG
|
Carlos Leandro |
A feed-forward neural network whose weights are integers and whose activation is the truncated identity implements, neuron by neuron, the connectives of \L{}ukasiewicz many-valued logic. This exact correspondence --- established theoretically by Castro and Tri...A feed-forward neural network whose weights are integers and whose activation is the truncated identity implements, neuron by neuron, the connectives of \L{}ukasiewicz many-valued logic. This exact correspondence --- established theoretically by Castro and Trillas and developed into a training algorithm by Leandro --- enables \emph{symbolic knowledge extraction}: training produces not a black-box model but a logical formula. Two obstacles have limited the approach to shallow architectures and small datasets: crystallization (forcing weights to integers) succeeds only probabilistically under the original Levenberg--Marquardt training scheme, and the theoretical guarantees break down as networks grow deeper. This paper addresses both obstacles. First, we prove that \emph{residual connections} (skip connections of the kind used in ResNets) extend \L{}ukasiewicz neural networks to arbitrary depth while preserving the symbolic correspondence \emph{at merge neurons} by construction: merge neurons in a \L{}ukasiewicz residual block automatically satisfy the neuron-classification proposition, regardless of the inner layer weights; inner-layer neurons are trained toward representability by the crystallization strategy. Second, we analyse three crystallization strategies --- Levenberg--Marquardt (corrected), straight-through estimation (STE), and proximal regularization --- characterizing their theoretical guarantees, failure modes, and interpretability trade-offs.
|
| 1228 |
What Must a World Model Distinguish for Planning?
2609.33030
|
cs.LG
|
Rongzhe Wei, Hans Hao-Hsun Hsu, Peizhi Niu, Yifan Li, Pan Li |
World models simulate the consequences of action candidates, but good planning need not preserve every physical distinction required for accurate prediction. We formalize this gap through a hierarchy of mechanism, response, and decision sufficiency. Given a ca...World models simulate the consequences of action candidates, but good planning need not preserve every physical distinction required for accurate prediction. We formalize this gap through a hierarchy of mechanism, response, and decision sufficiency. Given a candidate set, the planning query determines which physical variations matter and how precisely they must be preserved: coarse decisions can discard much of the information needed for prediction, whereas fine decisions may require nearly the same resolution. In practice, planners often adaptively search to construct candidates, and information unnecessary for final selection may still be needed to discover good candidates. What a world model must preserve therefore depends on the query, the candidate set, and the planner. We study these effects in a collision system, nonlinear dynamics, and robotic planning. These varying requirements raise a design question: where should query information enter the planning system? A model that jointly generates actions and outcomes conditioned on the query achieves lower regret than an action-conditioned world model on seen objectives, but this advantage largely disappears when generalizing to unseen objectives. Motivated by this, we propose a modular design in which the query determines where to look and an action-conditioned model predicts what will happen, allowing the same predictions to be reused across objectives.
|
| 1229 |
Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
2609.33041
|
cs.LG
|
Kristi Topollai, Anna Choromanska |
Learning-rate warmup is a standard technique in language-model training, yet its duration remains largely heuristic. Common approaches use either a fixed number of updates or a fixed fraction of the training horizon, two choices that imply very different scali...Learning-rate warmup is a standard technique in language-model training, yet its duration remains largely heuristic. Common approaches use either a fixed number of updates or a fixed fraction of the training horizon, two choices that imply very different scaling as training gets longer. When should warmup stay fixed, and when should it grow with the horizon? We address this question with a quadratic model whose modes respond differently to the peak learning rate. Warmup slows progress in directions that already contract well at the peak rate, but can remove persistent error in directions near the stability edge, with higher peak rates shifting the balance toward longer warmup durations. This yields a compact horizon scaling law that captures regimes ranging from essentially no warmup, through fixed-duration warmup, to durations that grow with the training horizon, and explains how the preferred regime changes with peak learning rate. Because the law captures the tradeoff between giving up early progress and improving the trajectory that follows, it can be fit using shorter runs and used to predict warmup at substantially longer horizons. Together, our results explain several familiar properties of warmup through a single tradeoff and suggest treating warmup duration as a horizon-dependent hyperparameter rather than a fixed training heuristic.
|
| 1230 |
Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers
2609.33047
|
cs.LG
|
Kristi Topollai, Anna Choromanska |
Matrix optimizers such as Muon transform each momentum matrix through an approximate orthogonalization, typically implemented by a small number of Newton-Schulz matrix multiplications. The quality and cost of this approximation depend strongly on the singular-...Matrix optimizers such as Muon transform each momentum matrix through an approximate orthogonalization, typically implemented by a small number of Newton-Schulz matrix multiplications. The quality and cost of this approximation depend strongly on the singular-value spectrum of its input, yet existing implementations use the same fixed polynomial routine for every layer and throughout training. We show that this uniform treatment is unnecessary: the computations in the Newton-Schulz method already reveal enough information to make the method adaptive. The Gram matrices formed inside Newton-Schulz iterations yield spectral moments through inexpensive scalar reductions, requiring no additional matrix multiplications. From these moments, we recover an estimate of the empirical singular-value distribution and use it to select a polynomial routine specialized to the current matrix. This turns Newton--Schulz orthogonalization into a spectrum-adaptive procedure that responds to differences across both layers and training time. On saved momentum matrices, spectral estimation substantially reduces orthogonalization error at a fixed iteration budget or reaches the same accuracy with fewer iterations, and in GPT pretraining up to 1B parameters it lowers the validation loss of two matrix optimizers. Our results suggest that matrix-function operations inside optimizers need not be designed for a conservative worst-case spectrum: they can cheaply measure the spectrum they are already processing and specialize computations accordingly.
|
| 1231 |
DevelopmentODE: Structured Neural ODEs for Early Brain Development Dynamics Across a Decade
2609.33048
|
cs.LG
|
Kaiqiao Han, Haitao Chen, Bryan Quah, Xiaoda Wang, Janelle Liu |
Understanding how individual brain development unfolds over childhood requires modeling developmental trajectories from sparse longitudinal observations. Long-term neurodevelopmental forecasting is challenging because each child is typically observed at only a...Understanding how individual brain development unfolds over childhood requires modeling developmental trajectories from sparse longitudinal observations. Long-term neurodevelopmental forecasting is challenging because each child is typically observed at only a few irregularly spaced visits, while developmental dynamics vary across individuals and age. Generic continuous-time models accommodate irregular timing but often absorb these factors into a single flexible transition function, providing little structure for how population progression, individual variability, and developmental age shape the dynamics. We propose DevelopmentODE, a structured continuous-time framework that organizes population- and subject-specific variation within a shared developmental geometry while allowing the governing dynamics to evolve with age. The model builds this geometry around a developmental canal representing the population trajectory, whose local direction provides a reference for organizing subject-specific variation. Subject deviation velocities are constrained relative to this direction, while a shared nonlinear deviation field captures individual developmental motion without disrupting population-level progression. DevelopmentODE further models developmental non-stationarity through ordered age-dependent deformations of the shared vector field, progressively adapting a common dynamical structure as age changes, while elapsed time determines the integration horizon. This formulation uses population-level developmental structure to guide learning from sparse individual trajectories while allowing dynamics to evolve smoothly with age. We evaluate DevelopmentODE on longitudinal fMRI by predicting future functional connectivity of the same child from earlier observations. DevelopmentODE consistently outperforms competing baselines across short- and long-horizon predictions.
|
| 1232 |
SketchSSM: Write to the Full State, Read from a Compact Sketch
2609.33051
|
cs.LG
|
Omin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer, Sehoon Kim |
Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and va...Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and values, but each new query still requires a full-state read even though the state remains unchanged between state updates. We observe that low-rank state-weighted query approximation accurately preserves state-read outputs. Although future queries are unknown, the basis vectors used to approximate them can be fixed offline. Based on this observation, we introduce SketchSSM, which preserves full-state updates while approximating reads. At each state update, SketchSSM reads the full state once to precompute outputs for these basis vectors, storing them in a compact sketch. Each subsequent decode step combines the sketch vectors with query-dependent coefficients to reconstruct the output without a full-state read. Across four Mamba-2-, GDN-, and KDA-based models, SketchSSM reduces state-access traffic by approximately 10x while largely preserving average accuracy across four decode benchmarks and recall on four RULER retrieval tasks. On one NVIDIA B300, linear-attention kernel speedups over the standard vLLM baseline reach 7.78x, 5.22x, and 5.20x for Mamba-2, GDN, and KDA, respectively, with up to 2.64x higher decode throughput on Nemotron 3 Super.
|
| 1233 |
Deep Learning Techniques for Phoneme Recognition in Italian Children' s Speech
2609.33060
|
cs.LG
|
Nicola Barbaro, Cristina Gena, Francesco Petriglia, Andrea Meirone, Alessandro Mazzei |
Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pr...Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-minute dataset of Italian child speech collected through a range of standardized diagnostic tests for children aged 3.5-6.5. Broca was optimized to handle phonetic variability in children's speech, including tone, accent, and speech errors, and achieved a state-of-the-art weighted Phoneme Error Rate of 13.36% on Italian speech. Remarkably, this performance was obtained using less than three hours of child-specific data, underscoring the model's efficiency and robustness in low-resource clinical settings. This work demonstrates that accurate, vocabulary-independent speech-to-IPA transcription can be achieved with minimal data, paving the way for more accessible, data-efficient tools to support speech assessment and diagnosis.
|
| 1234 |
Zero-Storage Procedural Neural Synthesis via Boundary Dynamics: Formal Verification in Lean 4 and Bare-Metal Gauntlet Validation
2609.33066
|
cs.LG
|
Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i} |
Contemporary neural inference architectures rely on dense floating-point weight matrices stored in high-bandwidth memory (VRAM), incurring severe memory-wall bottlenecks and preventing native execution inside deterministic virtual machines like the Ethereum Vi...Contemporary neural inference architectures rely on dense floating-point weight matrices stored in high-bandwidth memory (VRAM), incurring severe memory-wall bottlenecks and preventing native execution inside deterministic virtual machines like the Ethereum Virtual Machine (EVM). Verifying termination and arithmetic invariants for recursive dynamical systems over continuous domains is generally undecidable in the Blum-Shub-Smale model. Here, we present the formal verification and bare-metal empirical validation of WERR (Waves & Errors) and Phase III Orbital Error Dynamics (OED), a non-tensor decision paradigm that procedurally synthesizes non-linear decision boundaries on demand from a 24-byte coordinate seed $\Theta = (c_x, c_y, \text{zoom})$ along the boundary of the Mandelbrot set ($\partial\mathcal{M}$). By projecting the recurrence $z_{n+1} = z_n^2 + c$ onto the modular residue ring $\mathbb{Z}/9\mathbb{Z}$ and the fixed-point domain $\mathbb{Q}_{16.16}$, we establish ten machine-verified theorems in Lean 4 (v4.34.1) with Mathlib4 and zero unproven conjectures (sorry): proving $\mathcal{I}_3 = \{0,3,6\} \subset \mathbb{Z}/9\mathbb{Z}$ ideal closure, universal fuel-bounded halting ($\le 9$ and $\le 12$ steps), absence of $\mathbb{Q}_{16.16}$ square overflow below $2^{63}-1$, non-constant boundary escape sensitivity, and a parametric EVM gas bound ($\le 22,557 \le 24,000$ gas). Evaluated on a 40-core Dual Intel Xeon server, the vectorized 36-iteration CPU kernel processes 100,000 decisions in 6.49 s (15,397 decisions/s, 0 Bytes VRAM, 15.15x speedup), while a sigmoidal outlier gate suppresses 100.00% of adversarial spikes while preserving 89.60% of clean baseline signals.
|
| 1235 |
KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
2609.33074
|
cs.LG
|
Changxin Ke, Rui Zhang, Zixiang Fang, Zhenghong Li, Yuanbo Wen |
High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality ...High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities, and the inherent trade-off between kernel correctness and performance. To address these challenges, we propose KernelZero, a co-evolution framework that continuously improves GPU kernel generation through two specialized models: a Proposer that generates Torch modules from API sets and a Coder that translates them into CUDA or Triton kernels. KernelZero uses a frontier-driven module generation mechanism to continuously produce capability-aligned training modules based on the Coder's current weaknesses. It further introduces Correctness-Aware Group Relative Policy Optimization (CA-GRPO), which optimizes performance only after correctness becomes sufficiently reliable. By alternating the optimization of the Proposer and Coder, KernelZero forms an automatic curriculum that enables targeted and training-efficient capability improvement. Empirically, KernelZero-7B surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton. On KernelBench Level 1 and 2, it achieves CUDA pass@1 scores of 75.8% and 69.6%, respectively, with pass@10 reaching 100% and 97%. On Triton, it achieves pass@1 scores of 77.2% and 72.5%, respectively.
|
| 1236 |
The limits of exactness: On the failure of automatic differentiation in physics-informed machine learning
2609.33078
|
cs.LG
|
Ameya D. Jagtap |
Automatic differentiation (AD) lets neural networks compute derivatives of governing equations to machine precision, and this precision has made it the computational backbone of physics-informed machine learning. Yet exactness in the mathematical sense is not ...Automatic differentiation (AD) lets neural networks compute derivatives of governing equations to machine precision, and this precision has made it the computational backbone of physics-informed machine learning. Yet exactness in the mathematical sense is not the same as fidelity to the physics. Here I argue that a derivative can be numerically perfect and still be the wrong derivative for the problem at hand, because AD, by construction, has no notion of the physical structure a solution must obey. Convection and its associated directionality, diffusion, and dispersion are only the most visible instances of a much longer list that spans all branches of computational science and engineering, including conservation, thermodynamic consistency, symmetry, symplectic structure, positivity, monotonicity, and boundedness. Recognizing this broader gap reframes how the field should build the next generation of PDE-driven neural surrogates.
|
| 1237 |
From Constitutions to Control: Interpretable Rewards for Aligning Language Models
2609.33086
|
cs.LG
|
Johann D. Gaebler, Calvin Isley, Max Lamparth, Stephen Casper, Sharad Goel |
Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgmen...Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgments, obscuring what drives the resulting reward, while principle-based methods specify high-level values without fully operationalizing them. To address this gap, we develop a rubric-based framework to transform a general-purpose constitution into an interpretable and tunable reward model, using constitution-guided AI feedback to estimate initial weights for the constituent rubric items. We then reweight those dimensions to construct modified rewards for training. Across experiments on political alignment and safety-helpfulness tradeoffs, reweighting individual dimensions predictably changes targeted behaviors largely independently while navigating tradeoffs between conflicting alignment objectives. We show that the same framework can mitigate label bias encoded in preference judgments -- including sycophancy and demographic bias -- by reducing their influence on the training reward. Our results demonstrate that constitution-derived, interpretable rewards can translate high-level alignment principles into more transparent and controllable model behavior.
|
| 1238 |
Geometry-Aware Operator Families for Structured Representation Learning
2609.33089
|
cs.LG
|
Zuyuan Zhang, Fei Xu Yu, Tian Lan |
The geometry of latent representations governs which components should interact and how information should propagate, making geometry-aware operator design a fundamental ingredient of structured deep representation learning. However, existing neural architectu...The geometry of latent representations governs which components should interact and how information should propagate, making geometry-aware operator design a fundamental ingredient of structured deep representation learning. However, existing neural architectures typically rely on generic operator templates or geometry-specific constructions, creating a need for a unified framework that can derive admissible operators directly from fixed structural information while remaining adaptive to changing contexts. We introduce \emph{Geometry-Induced Operator Families} (GIOF), a general framework that converts fixed geometry into a structured family of propagation operators and dynamically selects an appropriate member of this family according to the current context. GIOF first transforms geometry-derived interaction channels into reusable generator bases, then combines them through a context-dependent selector and adaptive propagation scale, and finally realizes the selected operator through stable continuous-time propagation and a bottleneck residual layer. We establish theoretical guarantees covering parameter compression, identifiability, stability, locality, compositional structure, and oversmoothing behavior, while controlled experiments validate these mechanisms and experiments on PEMS-BAY and METR-LA achieve the lowest mean MAE across all reported regional-outage settings, improving over the strongest retained baseline by 2.4\%--8.8\% at 30\% missing sensors.
|
| 1239 |
How Linear Attention Remembers
2609.33093
|
cs.LG
|
Kichang Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko |
Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must sh...Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must share and repeatedly update the same memory. We study how this recurrent state functions as a memory system. Using an analytical decomposition together with controlled causal interventions in pretrained GLA and GDN models, we trace how recalled information is written, retained, and later accessed. We find that fact-specific information enters recurrent memory through concentrated, content-dependent writes and is later accessed through concentrated query-time read pathways. Multiple facts can remain selectively accessible within the same state, yet their internal representations exhibit cross-fact causal coupling rather than independent KV-like storage. As memory load increases, both recall and targeted editability degrade, whereas elapsed context alone has a substantially smaller effect within the tested regime. Causal interventions on subsequent writes further show that interference is shaped by their overlap with existing memory. Finally, in hybrid architectures that combine recurrent layers with full attention, the runtime memory directly supporting recall shifts predominantly to the full-attention KV state. Together, these results reveal how fixed-size recurrent memory supports selective recall despite shared storage, while exposing the interference and capacity limits that distinguish it from token-addressable KV memory.
|
| 1240 |
CARVE: Breaking Data Barriers in Chip Placement by Harnessing Reusable Expertise
2609.33106
|
cs.LG
|
Jiefu Zhang, Haixiang Sun, Yang Xu, Vaneet Aggarwal, Zishen Wan |
Pretrained macro-placement policies can reduce repeated optimization across circuits, but deployment often exposes them to unfamiliar designs when the original training data are unavailable. Repeatedly fine-tuning a single serving model can overwrite earlier i...Pretrained macro-placement policies can reduce repeated optimization across circuits, but deployment often exposes them to unfamiliar designs when the original training data are unavailable. Repeatedly fine-tuning a single serving model can overwrite earlier improvements, while simply saving checkpoints does not determine where they can be reliably reused. We introduce Continual Adaptation through the Reuse of Validated Expertise (CARVE), a framework that represents accumulated expertise as a frozen base policy, immutable specialists, and task-specific credentials obtained through local validation. For a new task, CARVE first checks existing specialists and trains a new specialist from the frozen base only when none qualifies. Under fixed task distributions and validation rules that control cumulative error, we establish expected-performance guarantees for repeated reuse. For bounded losses, we also derive matching worst-case bounds on the local samples needed for reliable reuse. In macro placement, a reuse-first follow-up reduces recorded training time by 58.5% (9.66 to 4.01 hours), while mean HPWL gain changes only from 8.41% to 7.86%. In a simulated receiving deployment, imported specialists are reused on six of seven new IBM circuits with no receiver-side training, achieving a 5.76% mean HPWL gain. Navigation studies provide complementary evidence on repair retention and repeated adaptation.
|
| 1241 |
D-JEPA: Design-Recoverable JEPA Representation with Swappable Physics Decoders
2609.33110
|
cs.LG
|
Nitin Nagesh Kulkarni, Aashwin Anand Mishra, Yin Yu, Peter Lyu |
Joint-Embedding Predictive Architectures (JEPAs) provide a framework for learning compact representations without directly reconstructing high-dimensional observations. However, in parameterized physical systems, learned representations can entangle geometry w...Joint-Embedding Predictive Architectures (JEPAs) provide a framework for learning compact representations without directly reconstructing high-dimensional observations. However, in parameterized physical systems, learned representations can entangle geometry with operating conditions and task-specific physical responses, limiting their reuse across prediction tasks. We introduce D-JEPA (Design-recoverable JEPA), a geometry-centric JEPA that computes a compact representation from geometry alone and reuses it across operating conditions and physical response spaces through lightweight physics-specific decoders. An explicit design-recoverability objective encourages the geometry latent to preserve information about the underlying design variables, enabling the representation to support design analysis and optimization. We further identify a case-level collapse failure mode in which target representations become nearly invariant across distinct geometries despite low reconstruction error, and mitigate it using case-level variation constraints and auxiliary target reconstruction. Across four 3D aerodynamic, hydrodynamic, and structural benchmarks, D-JEPA maintains or improves full-field prediction accuracy while achieving near-perfect linear recoverability of design parameters. The frozen geometry representation can be reused at held-out operating conditions and transferred to a structural response task with fewer trainable parameters. Finally, the representation supports differentiable design optimization, with designs validated using high-fidelity CFD, preserving the predicted ranking of candidate designs. These results demonstrate that separating a reusable geometry representation from physics-specific prediction provides a practical representation for scientific surrogate modeling and design.
|
| 1242 |
Simulation-Free Learning of GP-SDEs from Irregular Observations
2609.33112
|
cs.LG
|
Zhidi Lin, Yuhao Liu, Ying Li, Edwin Fong, Petar Djuri\'c |
Gaussian process stochastic differential equations (GP-SDEs) provide a flexible Bayesian model for unknown continuous-time state dynamics with uncertainty quantification, but learning and inference from noisy and irregular observations remain computationally c...Gaussian process stochastic differential equations (GP-SDEs) provide a flexible Bayesian model for unknown continuous-time state dynamics with uncertainty quantification, but learning and inference from noisy and irregular observations remain computationally challenging. To address this issue, we propose GP-SDE Matching, a simulation-free variational framework for Bayesian GP drift learning and continuous-time state smoothing. We analytically marginalize the sparse GP posterior to derive a tractable drift-matching objective that accounts for both the posterior mean and uncertainty of the unknown drift. To handle irregular observations, we further introduce an irregular-time-aware variational state posterior that incorporates the actual observation times during both encoding and continuous-time marginal querying. Experiments on the stochastic Lorenz--63 system demonstrate substantially improved drift recovery and state reconstruction under irregular observations, while five system identification benchmarks show robust forecasting under increasing observation sparsity and competitive performance against existing latent-SDE and state-space methods.
|
| 1243 |
Mycelium: A Generalizable Cross-Grid Multi-Task Model for Electrical Distribution Systems
2609.33120
|
cs.LG
|
Zhengyang Wei, Shourya Bose, Helgi Hilmarsson, Elena Carnio, Dhruv Suri |
Electrical distribution grid operations require inference across heterogeneous networks from sparse, noisy, and incomplete time series measurements. In this work, we identify challenges and explore solutions towards a unified model that can perform diverse tas...Electrical distribution grid operations require inference across heterogeneous networks from sparse, noisy, and incomplete time series measurements. In this work, we identify challenges and explore solutions towards a unified model that can perform diverse tasks grounded in the physics of the electric grid and generalize to unseen distribution networks. We define a unified grid ontology that represents variable sized distribution networks as heterogeneous graphs while preserving native topology, asset types, and electrical relationships across networks. We develop a physics based data simulation pipeline that combines reference and procedurally generated distribution networks with network reconfigurations, fault scenarios, and configurable sensing conditions. We present Mycelium, a heterogeneous graph transformer with structure aware communication edges and electrical reference features that encode network position and nominal phase orientation, together with task specific temporal readouts which generate per task outputs. We train Mycelium on reference as well as synthetic grids, and study its generalization on benchmark networks completely excluded from training and validation. Mycelium is observed to outperform task specific neural baselines on most reported benchmark metrics. Architectural ablations and the aforementioned studies reveal Mycelium's capability to learn representations of the underlying physics which serves to enhance cross-task performance, thereby addressing a significant challenge in unified grid models.
|
| 1244 |
Policy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online Adaptation
2609.33127
|
cs.LG
|
Yuheng Huang, Yunpeng Qing, Yixiao Chi, Yilun Kong, Changqing Zou |
Offline-to-Online Reinforcement Learning (O2O RL) has emerged as a practical paradigm that pre-trains the policy using static offline datasets and subsequently adapts the policy through online interactions. Existing O2O methods primarily address the transition...Offline-to-Online Reinforcement Learning (O2O RL) has emerged as a practical paradigm that pre-trains the policy using static offline datasets and subsequently adapts the policy through online interactions. Existing O2O methods primarily address the transition through value calibration, while generally treating the offline-trained policy as a given initialization. We instead study O2O adaptation from the perspective of network plasticity, asking whether the offline-trained policy remains sufficiently adaptable for online learning. Controlled experiments show that prolonged optimization on static offline data progressively reduces network plasticity even after offline performance has largely saturated, and that lower plasticity is associated with weaker subsequent online improvement. Motivated by these observations, we propose REstoring plasticity via Fresh Initialization and policy Transfer (REFIT), a lightweight model-level method for the O2O transition. Before online fine-tuning, REFIT distills the offline policy into a freshly initialized student while temporarily freezing a random subset of student units, transferring the learned offline behavior to a more plastic policy initialization. Extensive experiments on D4RL and OGBench demonstrate that REFIT consistently achieves higher aggregate performance than existing O2O plug-in methods across both Cal-QL and IQL backbones, while plasticity diagnostics and ablations provide further evidence of restored network plasticity.
|
| 1245 |
Flow-Matching-Based Protein Structure Tokenizer Made Efficient and Easy
2609.33129
|
cs.LG
|
Zhe Zhang, Yikai Zhang, Jiangtao Feng, Ya-Qin Zhang, Wei-Ying Ma |
As the bridge between protein modality and discrete modeling, protein structure tokenization still largely relies on heavily engineered training objectives tailored to specific downstream tasks and large training datasets, which hinders its transfer to broader...As the bridge between protein modality and discrete modeling, protein structure tokenization still largely relies on heavily engineered training objectives tailored to specific downstream tasks and large training datasets, which hinders its transfer to broader application scenarios. To address this issue, we propose ProFiT, a lightweight flow matching tokenizer. With simple training strategies that encourage healthy codebook utilization, ProFiT can be trained efficiently and naturally learns semantically meaningful representations without any manual semantic alignment, while achieving reconstruction quality and generalization that match or surpass those of substantially larger tokenizers. We conduct extensive evaluations across a wide range of settings and demonstrate that ProFiT is a plug-and-play tokenizer adaptable to diverse downstream tasks. This study further reveals the significant potential of the flow matching tokenizer paradigm. Our code is publicly available at https://github.com/QDKStorm/ProFiT.
|
| 1246 |
ILP-BO: Integer Linear Programming-Based Black-Box Optimization
2609.33131
|
cs.LG
|
Hyakka Nakada, Shu Tanaka |
Black-box Optimization (BO) is a powerful framework for optimizing expensive objective functions or unknown functions with a limited number of evaluations. A central step of standard BO such as Bayesian optimization is the optimization of a surrogate-based acq...Black-box Optimization (BO) is a powerful framework for optimizing expensive objective functions or unknown functions with a limited number of evaluations. A central step of standard BO such as Bayesian optimization is the optimization of a surrogate-based acquisition criterion, which is commonly performed using nonlinear optimization or heuristic search. Therefore, conventional black-box optimization generally does not guarantee global optimality in candidate selection. In this study, we propose Integer Linear Programming-based Black-box Optimization (ILP-BO), a quasi-Bayesian optimization framework that transforms kernel-based surrogate optimization over discrete domains into an Integer Linear Programming (ILP) problem. The key idea is to represent nonlinear kernel functions exactly on finite discrete distance levels by introducing binary one-hot auxiliary variables. This transformation converts the nonlinear surrogate into a linear objective with linear constraints and binary variables. To incorporate exploration while preserving the linear structure, we further introduce a Hamming-distance margin that excludes neighborhoods around previously observed points. We derive the proposed formulation for several standard kernels and obtain an analytical upper bound on the Hamming-distance threshold based on the measure in the binary search space. The resulting candidate-selection problem can be solved by integer programming solvers with certificates of optimality. Thus, our methodology has the potential to serve as a highly transparent black-box optimization framework. Experiments on synthetic and discrete optimization benchmarks show that ILP-BO achieves competitive optimization performance compared with practical Bayesian optimization methods.
|
| 1247 |
Leaky Students: Membership Inference against On-Policy Distillation
2609.33136
|
cs.LG
|
Zhexi Lu, Mingzhi Zhu, Stacy Patterson, Lei Yu |
On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks privat...On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplied to the teacher during distillation remains poorly understood. To the best of our knowledge, we present the first systematic study of membership inference in this setting. We find that fresh student trajectories expose sparse membership signals that fixed reference-answer losses often miss. These signals are mixed with probability changes caused by training on other records. We introduce Leaky, which samples fresh trajectories from the target model and compares its token log-probabilities with the maximum across matched reference models trained without the candidate records. It applies Leaky ReLU to the resulting gaps, preserving positive gaps and downweighting negative gaps as an approximate correction for incidental positive gaps in non-members. Across fifteen targets spanning mathematics, medical question answering, and code generation, Leaky outperforms all evaluated baselines and achieves mean AUROC 0.875, compared with 0.614 for the strongest baseline on each target in the main evaluation. On the same sampled trajectories, the strongest baseline achieves mean AUROC 0.826. These results show that students trained through OPD can expose the membership of records used for teacher supervision, even when fixed reference-answer losses provide little evidence of membership.
|
| 1248 |
Beyond the Training Horizon: Mechanisms and Limits of Length Generalization in Looped Transformers
2609.33144
|
cs.LG
|
Jia Liang, Xi Jin, Liangming Pan |
Looped Transformers can generalize to reasoning chains longer than those encountered during training, but the computations enabling this behavior and limiting its extent remain unclear. We mechanistically compare two looped-Transformer configurations, which we...Looped Transformers can generalize to reasoning chains longer than those encountered during training, but the computations enabling this behavior and limiting its extent remain unclear. We mechanistically compare two looped-Transformer configurations, which we call the Matched-Recurrence Looped Transformer (MR-Loop) and Decoupled-Recurrence Looped Transformer (DR-Loop), reflecting their respective recurrence-training schemes. We evaluate polynomial iteration, finite-state composition, and knowledge-graph traversal using detailed mechanistic analysis. Attention analysis, intermediate-state decoding, and causal interventions reveal distinct mechanisms learned under final-answer supervision. MR-Loop updates an intermediate state at a fixed readout while advancing relation selection through adjacent-token interactions and a transferable progress cue. DR-Loop instead propagates intermediate states across relation positions, forming an advancing computational frontier. However, both mechanisms become unreliable at greater depths: MR-Loop exhibits degradation of its readout state and progress cues, while DR-Loop exhibits declining reliability of state propagation. Limited self-correction allows local errors to persist and compound. Across both models, we uncover a common representational principle: recurrent states encode not only task-relevant content but also its computational status, whether that content remains in a form that can support subsequent computation. Transferable live-consumed and fresh-aged residual directions causally control whether represented information can participate in subsequent computation, including beyond the training horizon. We further show that length generalization need not rely on faithful step-by-step reasoning, as Looped Transformers can exploit task structure without explicitly representing every intermediate state.
|
| 1249 |
CFLoRA: Federated Fine-tuning of LLMs with Complementary Factors for Error-free Aggregation
2609.33147
|
cs.LG
|
Yanan Ma, Qiyuan Chen, Zihan Fang, Xianhao Chen, Yuguang Fang |
Federated low-rank adaptation (LoRA) enables collaborative fine-tuning of large language models without centralizing private client data. Its factorized update, however, creates a structural mismatch in federated averaging: averaging the two LoRA factors separ...Federated low-rank adaptation (LoRA) enables collaborative fine-tuning of large language models without centralizing private client data. Its factorized update, however, creates a structural mismatch in federated averaging: averaging the two LoRA factors separately does not equal averaging their products. Existing exact methods resolve this issue mainly by freezing an entire factor or alternating factors across rounds, but none can update factors simultaneously without aggregation errors or expanding communication ranks. To address this fundamental problem, we present \texttt{CFLoRA}, a federated LoRA scheme that partitions latent LoRA channels into two complementary sets in every communication round. By ensuring that columns and rows are complementary across factors, we eliminate bilinear terms in matrix multiplications, making federated aggregation exact. Crucially, our framework also supports clients with heterogeneous rank budgets. Convergence analysis validates \texttt{CFLoRA} achieves $\mathcal{O}(1/\sqrt{T})$ convergence rate of the \textit{original} LoRA objective in homogeneous-rank cases. Extensive experiments with RoBERTa on the GLUE benchmark and with LLaMA-3.2-3B-Instruct on commonsense reasoning tasks demonstrate that \texttt{CFLoRA} achieves superior performance and training efficiency compared to state-of-the-art federated LoRA baselines.
|
| 1250 |
Convergence of Practical Muon
2609.33152
|
cs.LG
|
Haonan Wang, Yu Wu, Minghui Liwang, Xinlei Yi, Yiguang Hong |
Muon is emerging as a promising alternative to AdamW for large-scale neural network training, yet theoretical understanding of its practical implementation remains incomplete, as existing analyses often simplify or omit two key components: (i) practical Newton...Muon is emerging as a promising alternative to AdamW for large-scale neural network training, yet theoretical understanding of its practical implementation remains incomplete, as existing analyses often simplify or omit two key components: (i) practical Newton--Schulz iterations with empirically tuned polynomial coefficients $(3.4445,-4.7750,2.0315)$; and (ii) decoupled weight decay for regularization. In this paper, we provide an optimization interpretation and establish convergence for practical Muon, jointly accounting for both components. Specifically, we interpret practical Muon as right-preconditioned optimization of the original loss with a dynamic weighted $\ell_2$ regularizer that vanishes as stationarity is approached, so that the optimization target remains the original objective. We then establish, to our best knowledge, the first convergence guarantee for practical Muon in the stochastic nonconvex setting, with an $\mathcal{O}(T^{-1/4})$ convergence rate in terms of the expected Frobenius norm of the gradient, improving the dimension dependence of the best known AdamW's convergence rate by a factor of $\sqrt{d}$, where $T$ is the iteration horizon and $d$ is the parameter dimension. Experiments further support the theoretical convergence results.
|
| 1251 |
Downstream-Aware Context Selection for Online In-Context Reinforcement Learning
2609.33166
|
cs.LG
|
Ruihan A. Li, Shangtong Zhang, Rohan Chandra |
In-context reinforcement learning (ICRL) enables large language model agents to adapt to new environments using their interaction history without updating model parameters. However, repeatedly conditioning on growing histories can lead to substantial token cos...In-context reinforcement learning (ICRL) enables large language model agents to adapt to new environments using their interaction history without updating model parameters. However, repeatedly conditioning on growing histories can lead to substantial token cost. We propose a bounded-history context-management framework that predicts the task-dependent downstream effect of removing historical interactions to guide history selection and determine a decision-dependent context budget. Formally, our framework uses the full rolling history as a reference. The predictor evaluates removal effects, defines a deletion ordering, and applies a shared selection criterion to determine how much history to retain at each decision. We evaluate the method in closed-loop SUMO driving under held-out in-distribution, unseen-domain, and unseen-route settings, and in ScienceWorld under a continual ICRL protocol. Relative to a baseline using the full context, our method reduces total token usage by 25.7%, 25.8%, and 23.2% across the three driving settings while maintaining comparable closed-loop driving performance. In ScienceWorld, it reduces total token usage by 52.1% compared to full context and uses 30.2% and 37.8% fewer tokens than the Recent and Similarity baselines, respectively, while maintaining performance.
|
| 1252 |
When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability
2609.33169
|
cs.LG
|
Xingjian Li, Yi Han, Jianhua Z. Huang |
Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping...Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping its values. We ask when this cut matters. Holding the forward computation fixed and varying only derivative edges, we measure parameter gradients, the updates the optimizer applies, and continued training in a Transformer vessel-trajectory model and a quadrotor tracking policy. In the vessel model, detaching the key-value cache shrank the gradient to about a tenth of its norm, with little rotation, when gradients flowed through all earlier physical states, but barely changed it under one-step physical credit. In this strongly clipped regime the optimizer, not the gradient, set how far updates differed: global-norm clipping removed most of the gradient difference between memory-cut graphs, whereas AdamW turned a 2% gradient difference between two placements of the cut into update differences of up to 31% at the step where the placement was switched. In a quadrotor trained from initialization with 0.20 m/s velocity noise, removing memory raised tracking error by 43% and cutting memory gradients raised it by 32%; at low noise the cut's mean cost exceeded the value of memory. Two-step truncation segments gave no measurable gain, although with hidden velocity a two-step window captured most of the value of memory; eight-step segments removed half to three quarters of the cost. Switching the cut on only for the last fifth of training understated its cost about threefold at 0.20-0.30 m/s, but not at low noise or with hidden velocity. These results suggest measuring the cost of a memory cut by training with it from initialization, and comparing backward graphs by the updates the optimizer applies rather than by raw gradients.
|
| 1253 |
Perturb-and-Solve: Efficient Learned-Operator Conditioning for Latent Diffusion Inverse Problems
2609.33171
|
cs.LG
|
Abduragim Shtanchaev, Arip Asadulaev, Luiza Labazanova, Aidar Alimbayev, Karim Salta |
Latent diffusion models serve as powerful priors for solving inverse problems in image restoration, such as deblurring, inpainting, and super-resolution. Current methods have a trade-off between generality and efficiency. Solvers that are restricted to a fixed...Latent diffusion models serve as powerful priors for solving inverse problems in image restoration, such as deblurring, inpainting, and super-resolution. Current methods have a trade-off between generality and efficiency. Solvers that are restricted to a fixed set of degradation operators are fast and efficient. Methods that support arbitrary degradation operators are slow and require gradients through the diffusion network. To break this bottleneck, we introduce PASEO (Perturb-And-Solve for Efficient Operator conditioning), a method that uses a small (1M parameters) learned network to degrade diffusion model predictions in latent space. PASEO supports learned degradation operators without back-propagating through the diffusion network. We efficiently sample reconstructions from an approximate posterior by combining the diffusion model's prediction with the observed image. We do this by adding noise and solving linear equations based on a local linear approximation of the learned network, without building or inverting large covariance matrices. Across super-resolution, deblurring, and inpainting on FFHQ and COCO, PASEO achieves strong perceptual quality while running up to 9x faster and using up to 34% less peak memory than the tested baselines, with the same or fewer model evaluations.
|
| 1254 |
When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM
2609.33189
|
cs.LG
|
Jiaxuan Guo, Jingxin Yang, Jiaqi Ye, Youran Sun, Shuo Xin |
When the human reference scores zero on a metric, the released GTRS-Dense label generator for NAVSIM marks every candidate trajectory in the scene as passing it. NAVSIM's authors introduced this human-reference forgiveness to avoid penalizing contextually just...When the human reference scores zero on a metric, the released GTRS-Dense label generator for NAVSIM marks every candidate trajectory in the scene as passing it. NAVSIM's authors introduced this human-reference forgiveness to avoid penalizing contextually justified maneuvers when scoring one trajectory, and warned that it could overlook important failures. In label generation it sets a whole column of 16,384 candidate targets to passing. To measure the consequences for supervision, we re-run the generator with the overwrite disabled and compare the pre-overwrite targets with the released labels on all 103,288 navtrain scenes. The rule erases a candidate distinction that the training loss reads on 11,237 of them (10.8793%). Firing usually changes most of a column: lane keeping carries 9,982 of the 13,042 forgiven loss columns, and its median forgiven column had 14,391 of 16,384 candidates failing before the overwrite. On held-out navtest scenes forgiven on lane keeping, the released lane-keeping head's median AUC against the pre-overwrite outcome is 0.7095; on unforgiven scenes matched on failing-candidate count it is 0.9807. For the Hydra-MDP checkpoint released with GTRS, whose configuration takes the same label file, the two values are 0.6627 and 0.9761. Continuing the released GTRS-Dense checkpoint for 300 optimizer steps with three paired seeds, we observe the forgiven-scene AUC 0.1086-0.1251 higher with pre-overwrite than with published targets, and a narrower gap between matched groups, still above zero. Scoring with forgiveness disabled, we observe lane keeping higher by 2.478-3.524 points on navtest scenes forgiven on any of five loss metrics, with lower adjacent-frame plan consistency. Both changes are larger there than on the rest. EPDMS, scored the same way, does not separate the two target sets.
|
| 1255 |
Apparent Compression, Real Stability: The Intrinsic Dimension of Learning a Quantum Wavefunction
2609.33193
|
cs.LG
|
Lu Wei, Yufeng Wang, Chenfeng Cao, Haibin Ling |
How many directions in weight space does training need? The intrinsic dimension answers this with the smallest number of random directions in which training still reaches a target accuracy, and small values have motivated parameter-efficient methods such as Lo...How many directions in weight space does training need? The intrinsic dimension answers this with the smallest number of random directions in which training still reaches a target accuracy, and small values have motivated parameter-efficient methods such as LoRA. We measure it for variational Monte Carlo (VMC), which trains a neural network to represent the ground state of a quantum many-body system. VMC is a demanding test, because the network generates its own training samples and every gradient is noisy, and a revealing one, because the exact answer is known and every run can be scored. We train only a small latent vector that a frozen random map turns into the network's weights, with no change to the standard natural-gradient optimizer. We find that a small dimension can be misleading, while the stability it brings is real. On a magnet with a hard sign pattern, a network that cannot represent signs reaches its best energy in 8 of 28,642 directions, but only because no such network can go lower; once signs are learnable, neither the signs nor the magnitudes are cheap. The dimension rises across a quantum phase transition, so it tracks how difficult a state is at far less compute than fitting a scaling law, yet it never falls below a floor set by the random subspace itself, even where the ground state is nearly trivial. Training in the subspace, in contrast, never diverged in our experiments, whereas full-parameter training with the same settings did, and a control with matched solvers attributes the difference to the reduced dimension.
|
| 1256 |
Orthogonal Witness Control for Muon Optimization via Sigmoid Spectral Reshaping
2609.33194
|
cs.LG
|
Dat Phi Van, Ngo Vu Minh, Tuc Nguyen, Thin Nguyen, Ngoc-Thanh Dinh |
Matrix-valued optimizers such as Muon exploit the spectral structure of neural network updates through Newton--Schulz orthogonalization, but their near-flattening of the singular spectrum discards relative magnitude information across gradient modes. We introd...Matrix-valued optimizers such as Muon exploit the spectral structure of neural network updates through Newton--Schulz orthogonalization, but their near-flattening of the singular spectrum discards relative magnitude information across gradient modes. We introduce \emph{Soren} (\textbf{S}pectral \textbf{O}rthogonal \textbf{Re}shapi\textbf{n}g), a matrix-valued optimizer that preserves the singular subspaces of the gradient while applying a bounded, monotone sigmoid transformation to its singular values. This smoothly compresses dominant modes without fully flattening the spectrum. We interpret Soren as a positive-definite preconditioned gradient method and establish convergence guarantees under relative smoothness and metric Polyak--{\L}ojasiewicz geometry. To avoid explicit singular value decomposition, we further develop a finite-depth Soft Newton--Schulz (SNS) polynomial realization of the sigmoid spectral map and characterize how its spectral approximation affects the induced convergence geometry. Experiments across LLM pre-training, supervised fine-tuning, and direct preference optimization demonstrate the effectiveness and robustness of Soren against established optimizers.
|
| 1257 |
SMORE: Stability-Promoting Mesh-Agnostic Model Reduction for Time-Dependent PDEs
2609.33205
|
cs.LG
|
Yangyuan Li, Weichao Li, Shaowu Pan |
High-fidelity simulations of time-dependent partial differential equations (PDEs) are computationally expensive, motivating data-driven reduced-order surrogates for many-query tasks such as uncertainty quantification, design optimization, data assimilation, an...High-fidelity simulations of time-dependent partial differential equations (PDEs) are computationally expensive, motivating data-driven reduced-order surrogates for many-query tasks such as uncertainty quantification, design optimization, data assimilation, and optimal control. However, existing surrogate models often exhibit poor temporal stability, which can lead to unstable rollouts and exploding gradients during backpropagation, especially in multistep long-horizon forecasting. To address this, we propose SMORE, a mesh-agnostic framework for model order reduction of time-dependent PDEs. Its latent dynamics are trained with Lyapunov-guided stability regularization, which promotes stable long-horizon rollouts. We provide theoretical guarantees under the stated structural assumptions. Beyond forecasting PDE evolution, the learned latent dynamics, which are interpretable and linear or linear-quadratic, could bring benefits for downstream tasks such as data assimilation and optimal control. Moreover, our framework is capable of predicting continuous PDE solution fields from sparse measurements of the initial condition. We evaluate SMORE on a range of problems, including wave propagation, the Navier-Stokes equations, and the shallow water equations. Our results show that it improves long-horizon rollout generalization and empirical robustness, and achieves competitive accuracy at comparable parameter budgets relative to competitive baselines including DINo, FNO, CNO, and Transolver.
|
| 1258 |
RMB: Reward Model Boosting Mitigates Reward Hacking
2609.33221
|
cs.LG
|
Jiabin Fan, Dezhi Ye, Yongchang Hao, Lili Mou |
Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while ...Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with respect to the true human preference, due to the imperfection of the proxy. To address this, we propose Reward Model Boosting (RMB), a novel approach that enhances the robustness and reliability of the reward signal for RLHF. RMB first trains a set of reward models with a diversity-promoting regularizer. This encourages each model to learn complementary aspects of the reward landscape. Then, RMB learns a lightweight aggregator in the principle of boosting to aggregate the outputs of the diverse reward models into a more accurate and robust reward signal. Our extensive experiments demonstrate that RMB significantly improves reward accuracy on both in-distribution and out-of-distribution datasets, substantially mitigating the reward hacking issue and ultimately improving RLHF performance.
|
| 1259 |
MTLiquid: Enabling Efficient Multi-Task Learning using Liquid Neural Networks for Lightweight Healthcare Monitoring Systems
2609.33232
|
cs.LG
|
Rachmad Vidya Wicaksana Putra, Fahad Abdul Rauf, Muhammad Shafique |
Continuous-time sensing and monitoring with timely and accurate decision-making are critical for many real-world applications. In healthcare monitoring systems, physiological signals are often available or sampled at irregular time intervals, hence requiring c...Continuous-time sensing and monitoring with timely and accurate decision-making are critical for many real-world applications. In healthcare monitoring systems, physiological signals are often available or sampled at irregular time intervals, hence requiring continuous-time processing to provide accurate prediction. Moreover, such systems often need to solve multiple detection/prediction tasks to provide a comprehensive patient review from different physiological aspects for more accurate decision-making. To solve this, continuous-time neural networks (CTNNs) can be employed. However, state-of-the-art works typically solve only one task at each network, thereby limiting their efficiency gains. To address this limitation, we propose MTLiquid, a novel methodology to enable efficient multi-task learning in continuous-time processing for healthcare monitoring systems through effective network design and training strategy. MTLiquid employs: (1) multiple input and output heads to accommodate different tasks, while sharing the same backbone network across tasks; as well as (2) an effective training strategy that leverages a loss-weighting technique to balance learning updates across different tasks and a proportional data presentation technique to address imbalanced dataset sizes. Experimental results for mortality prediction (P12) and sepsis early detection (P19) tasks for ICU patients show that, MTLiquid achieves strong performance (AUROC: 0.84 for P12 and 0.94 for P19) comparable to the state-of-the-art single-task learning in both continuous-time networks (AUROC: 0.84 for P12 and 0.95 for P19) and discrete-time networks (AUROC: 0.79-0.82 for P12 and 0.92-0.94 for P19), while incurring significantly smaller memory cost by 44%-94%. These results highlight the potential of our MTLiquid methodology to enable lightweight continuous-time healthcare monitoring systems for better decision-making.
|
| 1260 |
The Price of Locality: Why Forward-Forward Underperforms Backpropagation?
2609.33240
|
cs.LG
|
Zhaoxian Wu, Haichuan Liu, Tianyi Chen |
The Forward-Forward Algorithm (FFA) replaces backpropagation (BP) with layer-wise local contrastive objectives, eliminating the backward pass and the need to retain intermediate activations, yet suffers a persistent performance gap with BP that worsens with de...The Forward-Forward Algorithm (FFA) replaces backpropagation (BP) with layer-wise local contrastive objectives, eliminating the backward pass and the need to retain intermediate activations, yet suffers a persistent performance gap with BP that worsens with depth. This paper diagnoses two structural failure modes: an optimization floor arising from concurrent local updates; and a geometric collapse of layer representations driven by the local update mechanism. On the optimization side, we prove that the FFA loss satisfies the Polyak--Lojasiewicz inequality at each layer; however, simultaneous layer updates induce inter-layer representation-distribution drift, so each layer optimizes against a moving input distribution and incurs an error floor. On the representational side, the pairwise similarity kernel of layer representations contracts exponentially toward rank one as depth increases, collapsing the diversity of per-layer error signals. This collapse bounds FFA's effective learning capacity, which measures the diversity of gradient information across layers, independently of depth, whereas BP's chain-rule signal preserves per-layer diversity, yielding a capacity that scales with depth.
|
| 1261 |
Minimax-Optimality of Posterior Sampling for Reinforcement Learning
2609.33246
|
cs.LG
|
Taewon Goo, Kihyuk Hong |
Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. ...Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. Exact vanilla PSRL is minimax optimal in leading-order Bayesian regret under arbitrary correlated priors. The difficulty is that a posterior-sampled transition model is coupled with its own continuation value. We overcome this with a common empirical transition reference that isolates the resulting value mismatch and a Bellman-based variance argument that controls it without an extra leading-order state-space factor. For finite-horizon, time-inhomogeneous tabular MDPs with unknown stochastic rewards, this yields the minimax $\widetilde{O}(\sqrt{SAH^3K})$ regret rate under arbitrary joint priors over rewards and transitions. The same proof principle gives the minimax $\widetilde{O}(d\sqrt{H^3K})$ rate for linear-mixture MDPs under arbitrary joint parameter priors.
|
| 1262 |
Feedback-Robust AI for Patient Knowledge Graphs
2609.33248
|
cs.LG
|
Mohammed Sameer Syed |
Patient knowledge graphs from bedside monitoring should type their relations and state whether the data support their signs. In anesthesia and intensive care, clinicians titrate drugs and ventilation in response to the physiology, so temporal relations mix the...Patient knowledge graphs from bedside monitoring should type their relations and state whether the data support their signs. In anesthesia and intensive care, clinicians titrate drugs and ventilation in response to the physiology, so temporal relations mix the patient's response with the clinician's policy. We introduce ClosedLoopBench: 29 relations with signs fixed by physics, pharmacology or clinical practice, on 3,442 VitalDB surgical cases (12,653 h) with negative-control action streams. When each patient's actions are replaced by another patient's, six of 12 estimators declare on average 11-18 of their 19-29 distinct relation estimates significant without calibration, and after calibration cross-correlation and Granger tests still assign ventilator rate -> end-tidal CO2 the sign of the clinician's policy. We propose feedback-robust patient graphs that combine concept nodes with evidence pointers, typed relations admitted against negative controls, and beat-level couplings. On VitalDB under null streams, our graphs contain 0.06-0.10 false concept-level relation instances per graph, versus 10-12 for correlational construction. Patient-specific estimates of 11 slow drug and ventilator responses predict later data no better than population estimates, whereas the pulse-arrival-time-systolic-pressure slope is negative in 94.3% of 2,884 cases and patient-specific (early-late correlation 0.67 [0.63, 0.70]).
|
| 1263 |
Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts
2609.33257
|
cs.LG
|
Cheng Peng, Ruixi Luo, Zhi Chen, Wei Tang |
Large language models (LLMs) are increasingly used to forecast real-world events, but access to the strongest individual forecaster may be costly or otherwise constrained. We study weak-to-strong forecast aggregation: can individually weaker LLM forecasters be...Large language models (LLMs) are increasingly used to forecast real-world events, but access to the strongest individual forecaster may be costly or otherwise constrained. We study weak-to-strong forecast aggregation: can individually weaker LLM forecasters be aggregated to outperform a stronger forecaster? Using ForecastBench (Karger et al., 2025), we evaluate 70 LLM forecasters across 16 comparison groups, each with more than 1,000 shared subquestions, yielding 1,121 weaker-model pairs. Within each group, we identify the strongest individual by test Brier score and evaluate aggregates composed exclusively of weaker forecasters, with aggregation weights learned on separate training data. We find substantial evidence of weak-to-strong improvement. Learned linear pooling identifies a weaker pair that matches or outperforms the strongest individual in 11 of 16 groups and comes within 5% of its Brier score in all 16 groups. We also find that these improvements do not rely on having a near-best constituent and are generally accompanied by good calibration. Additional analyses show that adding more models does not consistently improve performance, and competitive weaker-model aggregates also remain available under practical constraints.
|
| 1264 |
GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences
2609.33259
|
cs.LG
|
Abdul Monaf Chowdhury, MD Sameer Iqbal Chowdhury, Shifat E Arman, Md Mehedi Hasan |
In offline goal-conditioned reinforcement learning (GCRL), divide-and-conquer scales to long horizons by joining two shorter segments at a subgoal. However, under stochastic dynamics, the base case of this rule values the luckiest trajectories through the data...In offline goal-conditioned reinforcement learning (GCRL), divide-and-conquer scales to long horizons by joining two shorter segments at a subgoal. However, under stochastic dynamics, the base case of this rule values the luckiest trajectories through the data. The subgoal must also lie on a shared trajectory, so a state-goal pair that no trajectory connects gets no value update at all. To address both, we present Grounded Transitive RL (GTRL), an offline GCRL value learning algorithm that grounds the divide-and-conquer update with a one-step TD target. Over a single step, TD is correct, as its target averages over the successors and needs no subgoal. GTRL adds this target to the composition rather than replacing it, so every pair receives an update, and the composition still carries the long horizon. GTRL also corrects the bias from hindsight relabeling by reweighting each goal against how reachable it was from other successors. We evaluate our algorithm on nineteen OGBench tasks spanning stochastic, deterministic, and stitching environments, where it achieves the highest average success rate. Code will be released soon.
|
| 1265 |
Towards Identifiable Representations under Misspecified Structure
2609.33273
|
cs.LG
|
Yuke Li, Yujia Zheng, Ziyi Chen, Kun Zhang, Heng Huang |
The presence of noise that depends on the latent variables poses a fundamental challenge to identifiability. Existing results rely on conditional independence among the observations given the latent variables. We study a more general \emph{misspecified structu...The presence of noise that depends on the latent variables poses a fundamental challenge to identifiability. Existing results rely on conditional independence among the observations given the latent variables. We study a more general \emph{misspecified structure}, where this conditional factorization does not hold, and establish both precise and approximate identifiability guarantees. We characterize structural misspecification as a perturbed factor analysis problem. For precise identifiability, we establish subspace identifiability under spectral separation and controlled perturbation, followed by component-wise identifiability under structural sparsity. When the precise condition is not guaranteed, we derive an approximate subspace-identifiability theorem. Based on these results, we develop an unsupervised variational estimator for recovering latent variables. Experiments demonstrate the effectiveness of the proposed framework.
|
| 1266 |
Domain Generalization under Sampling Pattern Shifts in Irregular Time Series
2609.33279
|
cs.LG
|
Changhun Kim, Joohyung Lee, Kwanhyung Lee, Donghwee Yoon, Grigorios Chrysos |
Irregularly sampled multivariate time series (ISMTS) are prevalent in real-world applications, where both observation times and available measurements can vary substantially across domains. While recent models increasingly exploit such sampling information for...Irregularly sampled multivariate time series (ISMTS) are prevalent in real-world applications, where both observation times and available measurements can vary substantially across domains. While recent models increasingly exploit such sampling information for prediction, its robustness under sampling pattern shifts remains underexplored. We introduce HAR-C, to the best of our knowledge the first controlled benchmark for sampling pattern shifts in ISMTS, and show that sampling shifts alone can substantially degrade performance, induce sampling-specific shortcuts, and remain challenging for existing domain generalization (DG) methods. Motivated by these findings, we propose PRISM, a DG framework that first learns complementary feature-centric and sampling-centric representations without task labels, and subsequently performs robust supervised training across diverse sampling variations to discourage brittle shortcut reliance. Extensive experiments on controlled and real-world ISMTS benchmarks demonstrate that PRISM consistently improves robustness to unseen sampling shifts over existing methods. Our code is available at https://anonymous.4open.science/r/PRISM.
|
| 1267 |
Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs
2609.33298
|
cs.LG
|
Fansheng Zhang, Shengran Guo, Zexiao Wang, Liang Yuan, Jiyuan Chen |
In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified pre...In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.
|
| 1268 |
BITS: Rethinking Fair and Comprehensive Evaluation for Irregular Time Series Forecasting
2609.33303
|
cs.LG
|
Kangjia Yan, Linfeng Wang, Tianen Shen, Xiangfei Qiu, Ruitong Zhang |
Despite recent progress in irregular time series forecasting, the field still lacks a unified benchmark for fair and comprehensive evaluation. Existing evaluations are often conducted on a limited set of datasets with inconsistent experimental protocols and pr...Despite recent progress in irregular time series forecasting, the field still lacks a unified benchmark for fair and comprehensive evaluation. Existing evaluations are often conducted on a limited set of datasets with inconsistent experimental protocols and predominantly error-based metrics, rendering it difficult to compare and assess methods fairly and comprehensively across diverse settings. To eliminate these limitations and accelerate progress, we propose BITS, a standardized, reproducible, and extensible benchmark for advancing research on irregular time series forecasting. BITS covers eleven datasets from nine domains with diverse irregularity characteristics, and it characterizes the datasets according to their missing rate, missing pattern complexity, sampling irregularity, and skewness. Further, it offers a unified pipeline for data preprocessing, model integration and evaluation, and reporting. It accommodates regular and irregular time series forecasting methods, including time series foundation models, under consistent settings, incorporating both error-based and non-error-based evaluation metrics. Findings include that method performance varies substantially across irregularity characteristics, with no single modeling strategy consistently dominating. We also find that using error-based or non-error-based metrics can yield different model rankings, highlighting the need for multi-dimensional evaluation. The code can be found at https://anonymous.4open.science/r/BITS-8F2E/.
|
| 1269 |
ZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph Models
2609.33314
|
cs.LG
|
Zhongjian Zhang, Xiao Wang, Busheng Zhang, Bo Yan, Xingtong Yu |
Zero-shot graph models (ZGMs), which learn transferable knowledge from source graphs and directly apply to unseen target graphs without any adaptation, have achieved promising performance and attracted considerable attention. Despite their proliferation, exist...Zero-shot graph models (ZGMs), which learn transferable knowledge from source graphs and directly apply to unseen target graphs without any adaptation, have achieved promising performance and attracted considerable attention. Despite their proliferation, existing ZGMs are predominantly evaluated on clean graphs, while existing graph robustness benchmarks mainly focus on supervised settings, leaving a fundamental question largely unexplored: How robust are ZGMs when their unseen target graphs are exposed to adversarial manipulation? In this paper, we answer this question by proposing ZeroGAR, the first systematic benchmark for evaluating the adversarial robustness of ZGMs. ZeroGAR evaluates 13 representative ZGMs from 3 different paradigms on 8 graph datasets across 4 domains, covering both in-domain and cross-domain transfer under structural, textual, and node injection attacks with multiple perturbation budgets. It further investigates whether existing graph defenses remain effective in the zero-shot setting. Extensive experiments reveal that strong clean zero-shot performance does not guarantee adversarial robustness, with three key findings: (1) Vulnerability patterns are related to model prediction mechanisms: GNN-based methods are particularly vulnerable to structural and node injection attacks, whereas LLM-based methods are more vulnerable to textual attacks; (2) Stronger LLM backbones introduce a structure-text robustness trade-off; (3) Existing graph defense methods do not consistently improve zero-shot robustness and may compromise clean performance. We hope that ZeroGAR will facilitate rapid, equitable evaluation and inspire further innovative research in ZGM security.
|
| 1270 |
Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning
2609.33336
|
cs.LG
|
Xuanlin Chen, Ziyue Wang, Xunlan Zhou, Yuan-yih Shang, Qiang Wu |
Model-based imitation learning (MBIL) improves real-environment interaction efficiency by optimizing policies on imagined rollouts from a learned world model. However, the gap between model-induced and real-environment occupancies makes policy learning sensiti...Model-based imitation learning (MBIL) improves real-environment interaction efficiency by optimizing policies on imagined rollouts from a learned world model. However, the gap between model-induced and real-environment occupancies makes policy learning sensitive to model error. Conservative MBIL mitigates model exploitation during policy optimization, but when real-environment interactions are collected by the same conservative policy, uncertain regions around the expert distribution remain insufficiently sampled. Generic uncertainty-driven exploration, on the other hand, may allocate interaction to novel but task-irrelevant dynamics. We propose REcoverability-CONditioned Exploration for Model-Based Imitation Learning (RECON). RECON separates conservative policy learning from active data collection by maintaining a main policy for task execution and an explorer for real-environment interaction. The explorer is optimized based on epistemic uncertainty conditioned on recoverability estimated from multi-step main-policy imagination, focusing data collection on unknown states from which the main policy can still return toward expert behavior. Experiments on locomotion, navigation and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness, indicating that RECON directs real-environment interaction toward recovery regions around the expert distribution that are underexplored by prior methods, and thereby learns a world model better suited for imitation.
|
| 1271 |
Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning
2609.33337
|
cs.LG
|
Boyang Li, Matthew Kim, Sylvia Lee Herbert |
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Proces...Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates that only enforce safety on average. To address this limitation, hard, state-wise constraints are introduced and often imposed through Hamilton-Jacobi (HJ) reachability. Yet such constraints require solving different objectives in the feasible and infeasible regions: reward maximization in the former, recovery toward the feasible regions in the latter. The resulting target action distributions are inherently multimodal, and this structure poses a fundamental challenge for the Gaussian or deterministic actors used in existing HJ-based safe RL, which often collapse onto suboptimal modes. Diffusion policies provide the expressiveness needed to represent such distributions, and recent work on Q-score matching offers a route to training them for online RL by score regression -- but has been applied only to reward maximization. We propose Safe Score Matching (SSM), an off-policy actor-critic method that adapts Q-score matching to hard-constrained safe RL by gating a two-branch score target with HJ reachability: inside the feasible set, the denoising process degenerates to Q-score matching on actions classified as viable by the HJ critic; outside, a recovery branch biases denoising toward regions with lower worst-case violation. On quadrotor and fixed-wing trajectory-tracking and stabilize-and-avoid benchmarks, SSM attains the best or near-best task performance with low false-safe rates, whereas the primal-dual baseline admits more unsafe behavior and reachability-based baselines tend to be more conservative; on Safety-Gymnasium velocity tasks, SSM attains the lowest cost with competitive reward.
|
| 1272 |
RAEGL: Risk-Aware Evidence-Gated Learning for Selective Contextual Routing under Temporal Shift
2609.33340
|
cs.LG
|
Yifan Guo |
Contextual specialization can improve forecasting accuracy, but a correction selected on one historical interval may become unreliable under temporal distribution shift. To address this issue, we propose RAEGL, a Risk-Aware Evidence-Gated Learning framework fo...Contextual specialization can improve forecasting accuracy, but a correction selected on one historical interval may become unreliable under temporal distribution shift. To address this issue, we propose RAEGL, a Risk-Aware Evidence-Gated Learning framework for selective contextual forecasting. RAEGL retains a validated global predictor by default and activates a contextual residual only when pre-deployment evidence supports its use. The framework separates candidate selection from gate calibration and jointly evaluates randomization significance, practically meaningful gain, and temporal stability. Experiments on real-world audits and controlled panels show how RAEGL can prevent harmful contextual deployment while making conservative opportunity costs explicit. In a reconstructed Our World in Data audit, exact fallback avoids RMSE degradations of 0.0960 and 0.0239 caused by two validation-selected corrections. In a sealed World Development Indicators evaluation, a region-based correction passes the randomization test but is withheld because its gain is only 0.000092, its country-clustered 95% confidence interval crosses zero, and only 0.02% of bootstrap replicates reach the practical threshold. In controlled panels, the stability- and support-aware extension activates in 97.2% of strong, stable-context runs while rejecting all high-drift settings. These results support RAEGL as an auditable, evidence-based mechanism for managing contextual deployment risk and as a conservative alternative to validation-driven contextual selection.
|
| 1273 |
MultiEcho: An Experimental Science of Learned Worlds
2609.33347
|
cs.LG
|
Meng Zhu, Airui Zhang |
World models can be studied as experimental systems with response laws of their own. We introduce MultiEcho, a framework for estimating these laws through controlled counterfactual interventions, delimiting their applicability, and separately testing their phy...World models can be studied as experimental systems with response laws of their own. We introduce MultiEcho, a framework for estimating these laws through controlled counterfactual interventions, delimiting their applicability, and separately testing their physical correspondence. Across nine simulated physical systems and seven frozen model configurations, three-reference estimators predict complete intervention responses and recover intervention parameters. Estimator selection uses discovery data only; frozen fits are evaluated on validation and confirmation contexts. The experiments distinguish response predictability, intervention readability and physical accuracy. Responses can be locally describable yet poorly match physical effects in the same target coordinates. Event-window, visibility and camera interventions reveal conditional applicability, and paired generator configurations show reduced readability under a scene prompt with stronger guidance. Magnitude sweeps expose small image errors alongside large relative effect errors. An exact-reset material experiment separates registered visible-response success from fixed-readout failure on material-dependent futures at matched positions and velocities. Exact finite-scale identities resolve odd and even response errors; first-order remainder bounds specify when refined calibration converges. MultiEcho provides an experimental basis for studying learned-world laws independently of, and in relation to, physical laws.
|
| 1274 |
KoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution Snapshots
2609.33350
|
cs.LG
|
Wanfeng Lu, Yutong Zhang, Keyi Zhou, Chenxin Ge, Wei Lin |
Learning population dynamics from temporally sparse, unpaired distribution snapshots is a fundamental challenge in developmental biology. Recent approaches based on neural differential equations and flow matching can interpolate between observed population sna...Learning population dynamics from temporally sparse, unpaired distribution snapshots is a fundamental challenge in developmental biology. Recent approaches based on neural differential equations and flow matching can interpolate between observed population snapshots, but may struggle to extrapolate beyond the training horizon and often lack an explicit mechanism for modeling developmental branching. We propose KoopCell, a unified generative framework based on Koopman-Mori-Zwanzig theory that jointly learns representations and predictive linear latent dynamics. Theoretically, using the weak continuity equation, we derive a closed-form least-squares estimator for the Koopman generator from distribution snapshots and establish convergence guarantees under suitable assumptions. To model branching dynamics, we further develop KoopCell-M, which incorporates non-Markovian memory into the latent Koopman dynamics through a Markovian embedding. Experiments on synthetic systems and three scRNA-seq datasets demonstrate the ability of our framework to recover Koopman spectra, model branching through memory, and scale to predicting high-dimensional gene expression distributions, achieving state-of-the-art performance among the evaluated methods.
|
| 1275 |
How Much Imprecision is Enough Imprecision in my Classifier? A Practical Elicitation Procedure
2609.33352
|
cs.LG
|
Victor F. Lopes de Souza, S\'ebastien Destercke, Abdelhak Imoussaten |
Set-valued classifiers, whether derived from precise probabilities and an adapted cost function, from convex sets with a robust inference mechanism, or from conformal methods, are routine options to obtain more robust, trustworthy predictions. However, there i...Set-valued classifiers, whether derived from precise probabilities and an adapted cost function, from convex sets with a robust inference mechanism, or from conformal methods, are routine options to obtain more robust, trustworthy predictions. However, there is a lack of operational tools to measure how robust or imprecise a given user is ready to be when receiving predictions, that is how much precision he/she is ready to let go in exchange of more accuracy. This is why we propose, in this paper, practical and operational elicitation procedures to measure the user proneness to set-valued predictions. The effectiveness of the iterative elicitation procedure in converging to the target parameter value is demonstrated on both tabular and image datasets drawn from standard machine learning benchmarks. The results show that the procedure also presents the user with a small number of instances, highlighting the practicality of the approach for real-world applications aimed at identifying the decision maker's optimal behavior when faced with imprecision.
|
| 1276 |
The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection
2609.33370
|
cs.LG
|
Farhan Shahriyar Hossain, Taufikur Rahman Fuad, Md Abrar Jahin, Md Rizwan Parvez |
Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers ...Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers are copied from earlier papers, and most anomalies are minority classes relabeled as anomalous. We ask how much of the reported ranking these conventions decide. We re-run two recent methods, DEMO and NSReg, together with OUTPOST, a small first-order detector built for this study. All three use one protocol with identical seeds and splits on eight graphs (seven for the baselines, which cannot run on ogbn-mag), ten seeds each, and every run is scored under both the best-epoch rule and a deployable validation rule. Before the runs that test them, we registered 40 predictions. Three findings hold. First, the rule changes the leader: under the best-epoch rule, OUTPOST and NSReg each lead three of seven graphs, while under the validation rule, NSReg leads five. Second, the best-epoch bonus depends on how the benchmark was built: 0.045--0.080 AUC-ROC on the three small relabeled-class graphs and 0.002--0.014 on the three real fraud graphs. Third, pseudo-labeling in OUTPOST is worth 0.038--0.065 AUC-ROC on the same three graphs but gives no benefit on any real fraud graph. We also show that a 0.002 tie band for hyperparameter selection lies below the paired standard error on all six graphs tested, even at ten seeds. Twelve of our 40 predictions were falsified, and we report them. We close with a short reporting checklist.
|
| 1277 |
TNF based Spectral Embedding for Effective Application of Supervised Machine Learning Techniques in Automobile Insurance Fraud Detection
2609.33376
|
cs.LG
|
Rohan Yashraj Gupta, Lalith Srikanth Chintalapati, Satya Sai Mudigonda, Pallav Kumar Baruah, Raghunatha Sarma Rachakonda |
Fraud detection is an important area of research in the insurance business due to its financial implications. The primary aim of a fraud detection model is to identify fraud and non-fraud cases with high accuracy along with other important metrics such as Sens...Fraud detection is an important area of research in the insurance business due to its financial implications. The primary aim of a fraud detection model is to identify fraud and non-fraud cases with high accuracy along with other important metrics such as Sensitivity, Specificity, Precision, F1-score, False Positive Rate, False Discovery Rate, AUC etc. To achieve this, we need to explore a suitable classification model to identify fraud and non-fraud cases. In this work, we have used auto insurance data set and explored classification models such as Decision Tree (DT), Random Forest (RF), XGBoost, LightGBM and Gradient Boosting Machine (GBM). To overcome the problem of data imbalance, we have employed MWMOTE and TGAN techniques. We have used Topological Node Feature(TNF) based spectral embedding for low dimensional data representation along with some popular embedding methods like MDS, Isomaps and t-SNE. After studying all the 65 possible combinations of these models, we have proposed an innovative method for effective automobile insurance fraud detection. For the given dataset, our results show that using a combination of MWMOTE as a data imbalance handling technique (Phase I), TNFSE2 as data embedding (Phase II) and Random Forest as classification (Phase III) provides the best result in comparison to all other combinations. This work also highlights the efficacy of TNF based spectral embedding in automobile insurance dataset
|
| 1278 |
Optimal Transport Dropout for Structured Predictive Uncertainty
2609.33377
|
cs.LG
|
Giacomo Lorenzon, Francesco Regazzoni |
Deterministic neural networks and neural operators provide point predictions with no intrinsic measure of reliability. Yet, predictive uncertainty may stem from irreducible outcome variability, finite data, or limitations of the chosen model class. Monte Carlo...Deterministic neural networks and neural operators provide point predictions with no intrinsic measure of reliability. Yet, predictive uncertainty may stem from irreducible outcome variability, finite data, or limitations of the chosen model class. Monte Carlo dropout offers a computationally convenient way to construct a predictive distribution through stochastic feature masking, without training multiple independent networks or explicitly inferring a posterior over model parameters. However, its perturbation law is largely prescribed a priori and typically factorised across latent coordinates. We introduce Optimal Transport Dropout (OTD), which instead learns the predictive mapping and the law of its latent perturbations jointly. Starting from a simple independent reference distribution, OTD transports latent perturbations through a learnable flow and propagates them through the predictive neural network, thereby inducing a structured predictive law. Training uses the strictly proper Energy Score, while a kinetic-action term geometrically regularises the transport. Synthetic benchmarks show that OTD captures multimodal predictive distributions, generates meaningful dispersion when the model is misspecified, and exhibits contracting dispersion as more training data or greater model capacity are provided. For a field-valued partial differential equation surrogate, predictive dispersion strongly aligns with the spatial pattern of prediction errors. On this task, compared with Monte Carlo dropout, OTD yields more accurate predictions and better-calibrated, substantially narrower intervals. On real-world regression benchmarks, it further shows competitive accuracy and better probabilistic predictions compared to established baselines. OTD therefore offers a way to learn structured predictive uncertainty without explicit posterior inference or ensembles of independently trained predictors.
|
| 1279 |
From Grey-Box to Green-Box: When can Physics-Informed Machine Learning Reduce Carbon Footprints in Structural Health Monitoring?
2609.33387
|
cs.LG
|
Daisy R. Bradley, Nathan A. Hinchliffe, Daniel J. Pitchforth, Matthew R. Jones, Elizabeth J. Cross |
Machine learning plays an increasingly vital role in engineering, but the corresponding increase in compute time is not without environmental cost. Physics-informed machine learning or "grey-box" models have been developed to overcome some of the limitations o...Machine learning plays an increasingly vital role in engineering, but the corresponding increase in compute time is not without environmental cost. Physics-informed machine learning or "grey-box" models have been developed to overcome some of the limitations of traditional black-box learners, utilising the physical insight that an engineer would have about the structure they are modelling and have shown promising results in the structural engineering field among many others. This work explores whether an additional advantage could be a reduced environmental impact, considering the relationship between training data quantity and training time, linking this duration to carbon emissions from computing. In a structural health monitoring context, four physics-informed machine learning approaches - spanning Gaussian processes and neural networks - are evaluated: residual modelling, input augmentation, hybrid modelling, and constrained learning. The emissions for training each of the models to reach a given error threshold is compared, and in most examples, shown to be lower for the physics-informed models (with input augmented models being an exception). This reduction in training emissions further compounds the environmental savings achieved by collecting and storing less data. Although promising results, we cannot expect a silver bullet and the case studies demonstrate that a trade-off is needed between the increased complexity that comes from introducing physics into a machine learner, against the gain from reduced training data requirements.
|
| 1280 |
Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents
2609.33391
|
cs.LG
|
Mingju Chen, Can Lv, Jinrong Liu, Huan Zhang, Heng Chang |
Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify ...Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD
|
| 1281 |
AutoHGNN: Robust and Efficient Neural Architecture Search for Hypergraph Neural Networks
2609.33392
|
cs.LG
|
Sirui Li, Pietro Li\`o b, Xinsheng Li, Baisong Liu, Chengbin Peng |
Hypergraph neural networks have achieved significant success in recent years. However, manual architecture crafting is labor-intensive and often fails to capture complex, higher-order relations, making the automation of hypergraph neural network structure desi...Hypergraph neural networks have achieved significant success in recent years. However, manual architecture crafting is labor-intensive and often fails to capture complex, higher-order relations, making the automation of hypergraph neural network structure design crucial. To improve the automation and adaptability of hypergraph learning, this paper proposes AutoHGNN, a neural architecture search framework tailored for hypergraph neural networks. First, we introduce a Hyper-Interaction Module (HIM) into the search space to address the mismatch between conventional graph neural network designs and hypergraph data. Second, we propose Hypergraph Stable Topological Distance (HyperSTD) as a structural selection criterion to identify architectures that best preserve the intrinsic structural affinities of the original hypergraph during differentiable search. Extensive experiments on various benchmark datasets demonstrate that AutoHGNN consistently outperforms manually designed and automatically searched baselines in classification accuracy and time efficiency, proving that the discovered architectures are significantly more effective.
|
| 1282 |
StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC Networks
2609.33408
|
cs.LG
|
Mustafa Bora \c{C}elik, Ceren \c{C}elik, Orhan Gazi |
In Integrated Sensing and Communications (ISAC), radar sensing must operate under chirp subsampling with up to 90\% missing data. An attention-based baseline, limited to a 52~ms buffer, collapses toward maximum uniform entropy ($H=2.584$ bits) as sparsity incr...In Integrated Sensing and Communications (ISAC), radar sensing must operate under chirp subsampling with up to 90\% missing data. An attention-based baseline, limited to a 52~ms buffer, collapses toward maximum uniform entropy ($H=2.584$ bits) as sparsity increases, failing to capture long-range gait-cycle context. We propose StarBOA, which replaces attention with a causal Mamba state-space model that updates incrementally on a per-window basis without re-scanning past reconstructions. By maintaining a persistent state, StarBOA integrates over $100\times$ more temporal history at no additional per-step computational cost. StarBOA outperforms the baseline's published results across all sparsity levels, with SSIM gains increasing from $+0.0379$ at 50\% missing data to $+0.2472$ at 90\%. Each window is processed in 1.53~ms with zero lookahead, demonstrating efficient causal reconstruction under extreme chirp subsampling.
|
| 1283 |
FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward
2609.33410
|
cs.LG
|
Sriman Achanta |
Autoregressive decode repeatedly streams a growing KV cache, making attention a major cost at long context. Existing high-performance kernels use online softmax, which discovers a row's normalization reference as it scans keys. Earlier contributions therefore ...Autoregressive decode repeatedly streams a growing KV cache, making attention a major cost at long context. Existing high-performance kernels use online softmax, which discovers a row's normalization reference as it scans keys. Earlier contributions therefore remain provisional and may require rescaling. We argue that the reference need not be discovered: softmax is invariant to a common shift, so the reference only has to keep the weights in range. We present FoldAttention, an additive formulation of softmax attention that fixes a finite reference $Z_i$ before scanning the KV cache. Each weight $2^{s_{ij}-Z_i}$ is then final when computed, so contributions add across disjoint key ranges and their quotient equals softmax attention in real arithmetic. We use this property to develop two techniques for Hopper decode: (1) final weights gate key and value reads before the bytes are fetched, and a per-call depth $T$ cuts keys below $2^{-T}$ while keeping their mass, and (2) additive partials compose split KV and shared-prefix cascades without rescaling. On H100 at $T=16$, FoldAttention decodes seven real-model generations 1.36-2.30$\times$ faster than the fastest BF16 baseline, and up to 3.09$\times$ faster across MHA and GQA shapes, at an error within 1.5% of the lowest BF16 error on six of the seven; reading every key, it is 1.14-1.30$\times$ faster at matched error. We validate on Qwen3-8B that a whole decode step is up to 1.46$\times$ faster while likelihood and long-context accuracy match those under BF16 kernels. The same principle makes the backward deterministic: CTAs round bounded partial gradients onto an integer grid declared before the reduction and add them in any order. FoldAttention thereby removes the determinism tax: its deterministic backward is up to 1.84$\times$ faster than deterministic FlashAttention-3/4 and 1.05$\times$ faster than the fastest nondeterministic kernel.
|
| 1284 |
Investigating the Effect of k-NN Preprocessing on Developing Graph Neural Networks: A Fairness-Based Perspective
2609.33416
|
cs.LG
|
Nikolaos Zafeiropoulos, Emmanouil Mavrikos, George E. Tsekouras |
In this paper, a methodology to design fair graph convolutional neural networks (GCNs) is developed and tested over several application data sets. The graphs that are used as inputs to the network are constructed by a k-nearest neighbor-based preprocessing pro...In this paper, a methodology to design fair graph convolutional neural networks (GCNs) is developed and tested over several application data sets. The graphs that are used as inputs to the network are constructed by a k-nearest neighbor-based preprocessing procedure, while fairness issues are considered in terms of the equalized odds criterion. To effectively incorporate the above heterogenous information, the equalized odds criterion is directly embedded into the model's optimization objective through an additional fairness-driven loss functional term. The proposed methodology investigates how varying the neighborhood size in the k-NN algorithm during graph construction influences both the classification performance and the fairness of the resulting models. Extensive experimentation is conducted on three real-world tabular datasets with known biases, evaluating the interplay between graph structure and fairness enforcement. The results demonstrate that the choice of the value of the parameter k critically impacts the performance trends, either steadily improving or peaking at intermediate values depending on dataset characteristics, while the application of fairness constraints significantly mitigates disparities in false positive and false negative rates across groups defined by the protected variable at hand, without incurring major sacrifices in overall accuracy. This study highlights the importance of jointly optimizing the graph construction process and fairness objectives in GCN-based learning, providing a systematic approach toward building more equitable and effective graph-based models.
|
| 1285 |
A Light Bilevel Refinement Aligns Self-Supervised Representations for Stronger Task-Specific Learning
2609.33424
|
cs.LG
|
Gustav Wagner Zakarias, Zheng-Hua Tan |
Self-supervised pretraining learns representations that are broadly transferable across downstream tasks, yet direct fine-tuning can be suboptimal due to misalignment between self-supervised and downstream task objectives, potentially degrading pretrained feat...Self-supervised pretraining learns representations that are broadly transferable across downstream tasks, yet direct fine-tuning can be suboptimal due to misalignment between self-supervised and downstream task objectives, potentially degrading pretrained features beneficial to the downstream task. The BiSSL framework addressed this by introducing a transitional training stage formulated as a bilevel optimization problem, in which the downstream task objective guides the self-supervised learning process in refining pretrained representations to better facilitate subsequent fine-tuning. However, BiSSL relies on conventional bilevel optimization solving techniques whose costly implicit hypergradient approximations render the method increasingly impractical for contemporary model architectures. To make it efficient and scalable, we introduce BiSSLight, which combines M-FAC-based implicit gradient approximation with parameter-efficient fine-tuning via LoRA, enabling efficient application at larger scales that were previously impractical. Evaluation across multiple downstream tasks and contemporary model architectures shows that BiSSLight consistently improves downstream performance, with gains becoming more pronounced as model size increases despite stronger baselines. The method is highly computationally efficient, reducing computation time by more than a factor of ten compared to its predecessor on a ViT-H backbone.
|
| 1286 |
MoGround: Measuring and Mitigating Modality Distraction in Vision-Language Models
2609.33431
|
cs.LG
|
Luca Zhou, Bo Zhao, Rose Yu, Emanuele Rodol\`a, Roberto Dess\`i |
We release MoGround, a vision-language dataset spanning four visual domains in which the answer to every question is guaranteed to be available from exactly one modality. This guarantee enables us to measure modality distraction, the failure in which a model a...We release MoGround, a vision-language dataset spanning four visual domains in which the answer to every question is guaranteed to be available from exactly one modality. This guarantee enables us to measure modality distraction, the failure in which a model answers a question correctly from one modality alone and then flips to a wrong answer once irrelevant content from the other modality is added. Existing probes rarely establish single-modality answerability this way, making it hard to isolate distraction in the first place. Across seven open-source VLMs, we find that modality distraction is not universal but model-dependent. The weaker-grounded modality is the more distracted one (r = +0.86), and distraction scales inversely with grounding strength (r = -0.90). The single-modality guarantee also enables a mitigation method that needs to distinguish between relevant and irrelevant context. Trained on one split of MoGround alone, a weight-space robustness vector reduces distraction on all seven models by 9% to 51%, at a cost of only 0.1 average points of accuracy on standard multimodal tasks.
|
| 1287 |
SchemaMem: Schema-Indexed Recurrent Memory for Delayed State Retrieval
2609.33436
|
cs.LG
|
Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung |
Attention provides direct access to past representations, but retaining an ever-growing history is costly. Recurrent models bound persistent state, yet must preserve selected information while processing subsequent inputs. We introduce SchemaMem, an attention-...Attention provides direct access to past representations, but retaining an ever-growing history is costly. Recurrent models bound persistent state, yet must preserve selected information while processing subsequent inputs. We introduce SchemaMem, an attention-based recurrent memory architecture combining chunk-local attention with a persistent, schema-indexed phase state. Learned schema embeddings provide a shared representational reference for reading and writing. Reads use the current state, whereas writes use the layer input and static schema embeddings, excluding direct feedback from that layer's own state. Chunk-boundary commits aggregate bounded phase increments through forward computation. The same parameters also support full-history attention training before and during recurrent training. We studied selective updates, preservation, and delayed retrieval in a controlled address--value task, comparing three-layer models with approximately matched parameter counts and persistent-state dimensions. Across nine address/value settings and three training seeds, SchemaMem has higher mean written-value retention at four times the maximum training delay than both baselines, which are trained toward a higher in-range accuracy target. Updated-value recovery favors SchemaMem in all nine settings against Mamba-3 and seven against Gated DeltaNet. Defaults consistently favor Gated DeltaNet over SchemaMem at that delay, and SchemaMem requires substantially more optimization steps. These results identify a promising retention--optimization trade-off in schema-indexed recurrence.
|
| 1288 |
Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning
2609.33444
|
cs.LG
|
Toyota Li, David Zhao, Alan Zhao |
A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting...A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting motivations. It is yet opaque what, if anything, they share. We substantiate that each is the solution of one divergence-constrained reward-maximization problem, and they are differentiated only by the convex generator that defines the constraint. Under the unified modeling framework, we unravel the relaxations that prior art made during building the advantage-embedded regression target: approximating the KKT condition and posterior normalizer for the linear and exponential tilt shapes DiffusionNFT and FlowAWR respectively, while preserving the exact sparsemax projection onto the probability simplex for linear tilt leads to another superior model type in this work. Beyond the theoretical underpinnings, we further empirically investigate the design space and shed light on the training recipe for regression-style diffusion RL. Retaining the merits discovered during our exploration gives rise to DiffusionRFT, our paradigm that converges faster, trains more stably, and attains the top performance.
|
| 1289 |
A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation
2609.33457
|
cs.LG
|
Md Tanveer Hossain Munim, Bijoy Ahmed Saiem, Al-Amin Sany, Tanzima Hashem |
Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial dis...Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration: sharp observations promote many blocks above threshold, while attenuated predictions from squared-error regression leave the same blocks below it. We repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system. The transform cannot reverse pixel ordering, so any contrast it reproduces cannot establish improved spatial ranking. On SEVIR, two released checkpoints of one architecture differ by -29.5% in extreme-threshold CSI before the control and by +5.3% after it. Across 450 pairwise contrasts among 6 systems, the difference in pooled frequency-bias deviation is associated with how far the CSI contrast moves under the control (r = +0.796), and 51 contrasts reverse sign. At CasCast's published extreme-event operating point, the cascade-over-backbone CSI gap falls from 0.1601 to 0.0339, a 78.8% reduction; the remaining gap stays positive. The effect persists when the transform is fitted on a window before the test period, and calibration also reveals advantages hidden by a better-calibrated baseline. On geostationary infrared imagery the relative gain grows as events become rarer, crowd counting reproduces the bias-gain relationship under patch-sum pooling, and semantic segmentation, where frequency bias is already near one, shows little average change. The confound therefore requires both a fixed operating point and a training regime that leaves the output miscalibrated there. We recommend reporting pooled frequency bias and a symmetric held-out FreeKnob Audit alongside rare-event pool-and-threshold scores.
|
| 1290 |
Geometric Identification in Predict-Then-Optimize Learning
2609.33472
|
cs.LG
|
Jiaxiao Xu, Changhong Mou, Keji Liu, Dinghua Xu, Yeyu Zhang |
Decision-focused surrogates can recover downstream decisions without identifying the quotient report. We characterize the equality set of the convex Smart Predict-then-Optimize surrogate (SPO+) population risk. Under central symmetry, the centered mean class i...Decision-focused surrogates can recover downstream decisions without identifying the quotient report. We characterize the equality set of the convex Smart Predict-then-Optimize surrogate (SPO+) population risk. Under central symmetry, the centered mean class is the unique Bayes minimizer exactly when every nonzero effective displacement makes the old optimizer leave the shifted optimal face with positive probability. This condition separates face crossing from selected-oracle disagreement and gives quantitative local coercivity. Without symmetry, strict crossing alone need not identify the mean; selection balance with reflected crossing restores quotient-report identification, and conditional versions extend the result to measurable predictors. These are population statements, without finite-sample report-recovery or generic transfer-regret guarantees. Closed-form mechanisms reproduce the analytic identities and rates. Portfolio, complete-matrix KuaiRec, and Energy/Storage studies measure predictive fidelity, shifted regret, and fitted-report geometry. A known data-generating process (DGP) companion retains their application geometries while isolating conditional-mean recovery and crossing, without testing the original observational assumptions.
|
| 1291 |
How Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional Coverage
2609.33482
|
cs.LG
|
Qianyi Chen, Bo Li |
Conformal prediction provides distribution-free finite-sample marginal coverage, but post-hoc calibration data may be too scarce to learn how uncertainty varies across inputs. Meanwhile, abundant covariates can often be labeled cheaply by domain models or gene...Conformal prediction provides distribution-free finite-sample marginal coverage, but post-hoc calibration data may be too scarce to learn how uncertainty varies across inputs. Meanwhile, abundant covariates can often be labeled cheaply by domain models or general-purpose language models. We study whether these synthetic labels can improve conditional coverage when only a small trusted sample is available. Building on score-quantile regression, we introduce prediction-powered quantile learning: a synthetic-labeled pool estimates pinball risk, paired trusted and synthetic outcomes correct its bias, and an independent trusted split performs final conformalization. Profiling pinball risk over scalar corrections reveals that population conditional-coverage error is its functional gradient; the corresponding Hessian removes global shifts and weights remaining shape error by boundary density. Composing this geometry with prediction-powered learning yields a three-resource expansion and a benefit--cost rule for synthetic power. Across eight regression benchmarks, synthetic-powered quantile learning substantially improves downstream conditional coverage while preserving marginal validity and producing more compact prediction sets. A human-rating study finds similar gains from external LLM labels and exposes a quality--quantity--cost tradeoff.
|
| 1292 |
What masking geometry works best for EEG foundation models?
2609.33487
|
cs.LG
|
Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi, Arnaud Delorme, Thomas Moreau |
EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it de...EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it determines what the network must predict and from which context. Yet it has never been ablated in isolation, as each new model bundles a new masking strategy with a new backbone and objective. In this paper, we formalize the design choices for spatio-temporal masking strategies and train various models with a single pipeline under varying masking configurations across two SSL frameworks (MAE and JEPA). We then systematically evaluate the resulting 58 pre-trained models on the 12 datasets of OpenEEGBench under a linear probe. Both frameworks agree on an optimal masking configuration and on shared failure modes. Outside these, performance is robust: 11 MAE and 9 JEPA configurations are statistically indistinguishable from the best. We further identify a novel JEPA-specific failure mode, tagged bias-inflation collapse, invisible to standard detectors. With a well-chosen mask, our pipeline reaches REVE-level downstream performance at a fraction of REVE's pre-training compute.
|
| 1293 |
Predicting Block-Coordinate Performance via Cross-Curvature
2609.33489
|
cs.LG
|
Shengkun Zhu, Jinshan Zeng, Zhiqiang Kou, Yongxin Tong, Yang Liu |
Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depe...Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depends on both the objective geometry and the number of iterations. We develop a unified theory for comparing Jacobi (JC), Gauss--Seidel (GS), and partially sequential deterministic block-gradient updates. Our analysis expresses the one-step loss difference through cross-block curvature, with an $O(\eta^3)$ remainder, where $\eta$ is the learning rate. We derive a signed loss comparison after $K$ iterations with $O(K\eta^3)$ error under regularity conditions and $\eta K\le T$ for fixed $T$, identifying the better method when the predicted difference exceeds this error. We evaluate these formulas along observed training trajectories across different machine learning settings. Over 500 iterations, our theory correctly identifies the lower-loss method in 98.0\% of iterations for the neural network, 83.4\% for federated learning, and 97.6\% for LoRA. Applying the loss recursion at each step using the measured parameter difference raises these rates to 100.0\%, 93.2\%, and 99.6\%, respectively.
|
| 1294 |
Chameleon: Dynamic Format Adapter for Efficient Diffusion
2609.33496
|
cs.LG
|
Arnab Sanyal, Sandeep Chinchali |
Post-training quantization (PTQ) is the standard way to run modern diffusion models on memory-constrained accelerators, yet every existing diffusion PTQ scheme fixes the $\mathit{number\ format}$ in advance and only tunes the scale, zero point, or per-layer bi...Post-training quantization (PTQ) is the standard way to run modern diffusion models on memory-constrained accelerators, yet every existing diffusion PTQ scheme fixes the $\mathit{number\ format}$ in advance and only tunes the scale, zero point, or per-layer bit-width. At a fixed bit-width the best format depends on the distribution being encoded, and that distribution differs across weight channels, across layers, and along the diffusion timestep, where activation distributions slide from heavy-tailed and noise-dominated to tightly clustered and structured. We propose Chameleon, a PTQ framework that holds the bit-width fixed and treats the format itself as a discrete variable, chosen per weight channel and per (layer, timestep bucket) activation tensor. Activation formats come from {INT8, FP8 E4M3, FP8 E5M2, MXFP8, MXINT8}, selected ahead of time from two cheap statistics (empirical kurtosis and the closed-form diffusion SNR) and stored in a lookup table; weight formats come from {INT8, MXINT8} at 8 bits or {INT4, NF4, FP4 E2M1, MXINT4, MXFP4} at 4 bits, selected offline by reconstruction error. An architectural fork adapts the same selection layer to multi-step UNets, single-step distilled models, and Diffusion Transformers. Across SDXL, SDXL-Turbo, and PixArt-$\alpha$ on COCO-2014, Chameleon achieves the best FID in all six backbone $\times$ bit-width settings, with CLIP within 0.24 of the FP16 reference and the best of all quantized methods at $W_{4}A_{8}$.
|
| 1295 |
Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
2609.33497
|
cs.LG
|
Tamim Zoabi, Ameen Ali, Lior Wolf |
Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting fu...Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing action-conditioned JEPAs ask one embedding to serve both perception and control. We introduce H-JEPA, which separates the two. A wide perceptual code is regularized toward a well-scaled isotropic geometry with a Bures-Wasserstein prior, and a fixed orthonormal slice of that code is the control state, which inherits the code's covariance without any objective of its own. The state evolves under phase-conditioned dissipative port-Hamiltonian dynamics whose input port has orthonormal columns. Port-inverse consistency (PIC) reads the executed action back through the transpose of that port. We show that this readout is exactly the rollout error projected onto the port directions, so PIC is a parameter-free reweighting of prediction error and not an auxiliary action decoder. Untying the readout from the port breaks this identity and loses half of the gain. H-JEPA matches or exceeds reconstruction-free baselines, including the action-decoding Delta-JEPA, on four pixel-based control benchmarks after at most $10$ training epochs, and its largest gain is on OGB-Cube ($91.9$ against $79.3$ percent). Ablations on PushT and OGB-Cube separate the contributions of the structured predictor, PIC, the prediction horizon, the state rank, and the anti-collapse prior.
|
| 1296 |
The cost of useful natural gradient updates
2609.33499
|
cs.LG
|
Subhransu S. Bhattacharjee, Dylan Campbell, Rahul Shome |
What information is needed to turn a natural-gradient direction into a useful finite update? Under a population Kullback-Leibler (KL) budget, we call a step useful if it is feasible and loses at most a fraction $\varepsilon$ of the best feasible gain along the...What information is needed to turn a natural-gradient direction into a useful finite update? Under a population Kullback-Leibler (KL) budget, we call a step useful if it is feasible and loses at most a fraction $\varepsilon$ of the best feasible gain along the direction. We construct a four-state exponential family whose laws share their initial gradient, scalar Fisher information and natural gradient, yet two laws have disjoint useful-step sets. With these quantities supplied exactly and the law otherwise known only through draws, the family's worst-case sample complexity is $\Theta(\log(1/\delta)/(p\varepsilon^2))$ for small $\varepsilon$, where $p$ scales rare-state probabilities and $\delta$ is the failure probability. The budget is fixed and the optimal gain stays bounded away from zero, so the step length, not the direction, carries this cost. For succinctly described event-tilt models, returning a useful step is NP-hard even with the exact natural gradient and efficient exact sampling. Recovering the unit natural gradient to constant error is also NP-hard even in a two-parameter logistic family with Fisher condition number at most 3. We also give matching sample bounds for event tilts, sample bounds for damped Fisher solves and a population-KL certificate for affine classifiers. In frozen-feature classifier heads, stopping at a sampled KL boundary succeeds in about half of the trials, and a 10% KL margin raises joint success above 93% at a KL budget of 0.01. Thus, knowing where to move is not enough: how far to move can carry an update's entire cost.
|
| 1297 |
Pulseflow: PPG Counterfactual Generation Via Latent Transport
2609.33501
|
cs.LG
|
Hung Manh Pham, Dong Ma, Bin Zhu, Pan Zhou |
Photoplethysmography (PPG) has become an important modality for continuous cardiovascular monitoring, including atrial fibrillation (AF) detection. However, labeled AF recordings remain limited in many clinical settings, making model adaptation difficult when ...Photoplethysmography (PPG) has become an important modality for continuous cardiovascular monitoring, including atrial fibrillation (AF) detection. However, labeled AF recordings remain limited in many clinical settings, making model adaptation difficult when only limited target data are available. Generative modeling offers a natural way to alleviate this scarcity by synthesizing additional AF signals. Existing approaches, however, mainly generate samples that match the target condition without explicitly modeling how an observed source recording should be transformed, making it difficult to leverage abundant source recordings from a specific population or cohort for targeted augmentation. We introduce PulseFlow, a source-conditioned counterfactual generation framework that combines conditional representation learning with invertible latent transport to edit cardiac rhythm while retaining information from the source. Experiments across two clinical cohorts demonstrate effective rhythm transformation, measurable source correspondence, and improved AF classification under limited labels.
|
| 1298 |
Source Anchoring for Physical Consistency in Flow Matching Models
2609.33510
|
cs.LG
|
Giulia Romoli, Filippo Ruffini, Paolo Soda |
Deep generative models are used to solve partial differential equations and model distributions of physical system states, but ensuring that the generated samples satisfy the governing laws remains challenging. Projection-based flow-matching methods enforce ph...Deep generative models are used to solve partial differential equations and model distributions of physical system states, but ensuring that the generated samples satisfy the governing laws remains challenging. Projection-based flow-matching methods enforce physics by correcting the flow from an unconstrained noise distribution. These corrections shift the generated samples away from the distribution of target solutions, especially in high noise regions. To address this limitation, we propose Source Anchoring for Physical Consistency (SAPC), a Functional Flow Matching method that encodes the physical constraints into the source noise before generation begins. We evaluate SAPC on five systems governed by partial differential equations, covering six tasks with linear and non-linear dynamics, and compare results against five baselines and the unconstrained backbone. Anchoring the source reduces the need for large corrections that drive samples onto admissible but off-distribution states, and SAPC reproduces the target distributions most accurately on every evaluated task, while matching the constraint precision of the best projection-based baselines. Ablation experiments show that this gain arises from pairing source projection with a matched training objective that regresses toward the projected source. These results identify the source distribution as a key design choice for physically consistent generative modelling.
|
| 1299 |
SLP-ProbHard: Probabilistic Hard-Constrained Learning via Structural Latent Parameterization
2609.33515
|
cs.LG
|
Wondesen Teshome Bekele, Marco D'Oria |
Many probabilistic predictors must satisfy exact structure in every stochastic realization, yet common hard-constraint approaches form predictions in ambient coordinates and then correct or project them. We introduce SLP-ProbHard, a cross-family, representatio...Many probabilistic predictors must satisfy exact structure in every stochastic realization, yet common hard-constraint approaches form predictions in ambient coordinates and then correct or project them. We introduce SLP-ProbHard, a cross-family, representation-centered framework for probabilistic hard-constrained learning when explicit structural parameterizations are available. Its core object, a Structural Feasible Latent Parameterization (SFLP), combines a structural latent law $Z \sim P^Z_\theta(\cdot\mid x)$ with a feasible map $Y=h_\phi(x,Z)$ that satisfies the constraint for every latent realization. Together these components define the predictive law itself, including its support and boundary probabilities, rather than serving as a final feasibility wrapper. We study how feasible coordinates and maps affect stochastic dimension, dependence, calibration, expressiveness, and computation. Experiments use Gaussian latent laws and fixed geometry-derived maps across affine equalities, ordering and simplex constraints, nonlinear manifolds, and three structural representations of seven-basin hydrological flow-duration-curve (FDC) data. In an official-source affine comparison with ProbHardE2E/DPPL, both methods achieve zero practical constraint violations. SLP-ProbHard uses 8 instead of 11 stochastic coordinates and improves MSE/MAE, while DPPL yields better marginal CRPS and closer-to-nominal coverage; a paired test detects no Energy Score difference across ten seeds. Real-world affine and nonlinear FDC representations reduce 13 to 7 and 14 to 8 ambient versus computational coordinates, respectively. Exact feasibility alone thus does not determine a predictive law, motivating direct structural generation when meaningful feasible coordinates are available.
|
| 1300 |
LLM4Trust: Exploring the Capabilities of Large Language Models for Trust Evaluation
2609.33521
|
cs.LG
|
Jie Wang, Yanbo Sun, Zheng Yan, Jiahe Lan, Elisa Bertino |
Trust evaluation plays a critical role in cybersecurity by supporting risk mitigation and decision-making. A variety of trust evaluation methods have been proposed, with learning-based approaches offering high accuracy and automation. However, they often requi...Trust evaluation plays a critical role in cybersecurity by supporting risk mitigation and decision-making. A variety of trust evaluation methods have been proposed, with learning-based approaches offering high accuracy and automation. However, they often require substantial ground truth, suffer from low training efficiency, lack support for basic trust properties, and provide limited explainability. Large Language Models (LLMs) offer a compelling alternative due to their strong zero-/few-shot reasoning abilities and broad knowledge. To this end, we propose LLM4Trust, the first benchmark framework that systematically explores the capabilities of LLMs for trust evaluation. We first construct diverse trust graphs to model five basic trust properties and design corresponding property understanding tasks. We then assess the ability of eight representative LLMs to understand these properties under nine prompt methods. Based on this exploration, we identify the most effective LLM-prompt combinations and apply them to five real-world datasets for validating LLMs' trust evaluation capability. During this process, we propose two strategies to extract key information from large-scale trust graphs, addressing the context window limitations of LLMs. Extensive experiments show that LLMs can effectively understand basic trust properties and have great potential for real-world trust evaluation, particularly under limited supervision. However, they remain vulnerable to attacks targeting trust graphs and demonstration examples used in few-shot prompting, and incur high inference costs. Accordingly, we propose a defense mechanism and batch inference to improve the robustness and efficiency of LLM-based trust evaluation. The source code of LLM4Trust is available at https://github.com/Jieerbobo/LLM4Trust
|
| 1301 |
Discovering Symmetries in Neural Network Parameter Spaces
2609.33527
|
cs.LG
|
Bo Zhao, Nima Dehmamy, Robin Walters, Rose Yu |
Parameter space symmetries are important for understanding neural networks' loss landscape, training dynamics, and generalization. However, systematically identifying these symmetries remains a challenge. In this paper, we formalize data-dependent parameter sy...Parameter space symmetries are important for understanding neural networks' loss landscape, training dynamics, and generalization. However, systematically identifying these symmetries remains a challenge. In this paper, we formalize data-dependent parameter symmetries and characterize loss invariance and the group-action axioms through infinitesimal conditions, which provide objectives for jointly learning group generators and nonlinear action maps. Our framework systematically uncovers parameter symmetries, including previously unknown ones. To study larger networks, we establish conditions under which subnetwork symmetries extend to the full model. The same construction gives an explicit family of finite-batch symmetries, providing both analytical examples and a foundation for discovery through small subnetworks. Using the infinitesimal characterization and subnetwork construction, we implement a framework for automated discovery of parameter symmetries, and successfully uncovered symmetries in various architectures, including pretrained transformer models.
|
| 1302 |
Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving
2609.33536
|
cs.LG
|
Jiantong Jiang, Yue Yang, Peiyu Yang, Feng Liu |
Large language model (LLM) serving is increasingly constrained by the GPU memory consumed by key-value (KV) caches. Existing compression, eviction, and offloading techniques alleviate this pressure, but serving runtimes typically treat only the configured targ...Large language model (LLM) serving is increasingly constrained by the GPU memory consumed by key-value (KV) caches. Existing compression, eviction, and offloading techniques alleviate this pressure, but serving runtimes typically treat only the configured target KV representation as execution-ready. Under memory pressure, this target-only contract can turn KV shortage into request stalls and preemptions. We present ElasticKV, a mixed-fidelity KV runtime built on the observation that target fidelity need not gate execution. ElasticKV introduces a compact intermediate KV state, making fidelity a runtime-managed execution property. To realize this state in a paged serving runtime, ElasticKV combines (i) a pair-structured layout that turns fidelity reduction into reusable GPU capacity, (ii) a dual-mode attention backend that directly consumes the compact state while preserving the native target-only path, and (iii) pressure-aware fidelity management that adapts KV fidelity to memory pressure. Our extensive evaluation across diverse workloads, model families and scales, and GPU platforms demonstrates the effectiveness and generality of ElasticKV. Under high concurrency, ElasticKV achieves 3.8-4.0$\times$ lower time-to-first-token (TTFT) and 9.1$\times$ lower P90 TTFT than vLLM while preserving generation quality.
|
| 1303 |
TerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the Edge
2609.33548
|
cs.LG
|
Houssem Sifaou, Prabodh Katti, Bipin Rajendran, Osvaldo Simeone |
Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. Howev...Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. However, for BitNet architectures, a family of LLMs with ternary {-1,0,1\} weights and 8-bit activations, fine-tuning requires updating full-precision latent weights, and thus the memory footprint of MeZO no longer matches that of inference. A promising solution is to finetune only a subset of the latent weights, but existing sparse zeroth-order (ZO) methods either ignore the ternary structure or require first-order gradient information to build a sparse mask, which is at odds with the purpose of ZO fine-tuning. We propose TerMeZO, a sparse MeZO scheme that exploits the geometry of the ternary quantizer itself to identify the latent weights that are more likely to change values during fine-tuning, at no additional data or memory cost. Our convergence analysis shows that TerMeZO can converge faster than full-parameter MeZO, owing to its optimized reduction of the fine-tuning effective dimension. We run extensive experiments on BitNet models ranging from 1B to 3B parameters, spanning classification, instruction-following, and mathematical reasoning tasks. TerMeZO matches or exceeds the performance of full-parameter MeZO while substantially reducing the fine-tuning memory footprint.
|
| 1304 |
Fine Until Fine-Tuned: Repeated Solutions Make Reasoning Fragile
2609.33559
|
cs.LG
|
Ely Sheikh |
Recipes such as s1 and LIMO teach a model to reason with little data by showing it the same thousand or fewer worked solutions many times over. Judged when that training ends, the repetition looks harmless. But reasoning models are often trained again, and we ...Recipes such as s1 and LIMO teach a model to reason with little data by showing it the same thousand or fewer worked solutions many times over. Judged when that training ends, the repetition looks harmless. But reasoning models are often trained again, and we find that repetition leaves their reasoning fragile to that next stage, even when the stage has nothing to do with reasoning. We fine-tuned Qwen3.5-9B-Base on its own correct solutions to competition math problems, either drilling a few hundred of them about eight times each or showing many more once; with the same amount of training, both solve about 95% of held-out problems. A single pass of ordinary instruction tuning leaves the once-trained model where it was, while the drilled one falls to 86.0%, and harsher later stages take it to 59.3% or below. A third model that visited the drilled problems just as often, with a new solution at every visit, was unharmed, so the damage comes from seeing the same texts again rather than from having few problems. The break recurs with a stronger model's traces, in further training runs and on other models and tasks. It is also cheap to undo: the reasoning is suppressed rather than erased, and five updates of reasoning training bring almost all of it back, as does brief training on the reasoning format with almost no mathematics. Fresh solutions prevented the damage, and so did replaying 6.25% of the original solutions in a gentler later stage, so our claim concerns later training without such replay. Sharpening alone does not explain the break, since a model sharpened three-quarters as much without repetition was unharmed. On a skill the base model could not perform within a token budget, repetition mainly cost learning.
|
| 1305 |
GraphSelect for Budgeted Representation Selection in Multimodal Graph Inference
2609.33561
|
cs.LG
|
Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li |
Multimodal graph predictors combine text, images, and relations to classify connected entities. How much of this input is needed to preserve their predictions? We study budgeted representation selection, which chooses a subset of candidate text and image vecto...Multimodal graph predictors combine text, images, and relations to classify connected entities. How much of this input is needed to preserve their predictions? We study budgeted representation selection, which chooses a subset of candidate text and image vectors under a separate capacity for each modality. Predictions from the complete candidate input define the classes to preserve. The challenge is that a representation's contribution depends on the other selected inputs, while graph propagation extends its effects across nodes. Our empirical study shows that candidate rankings change with the selected input, while predicted probabilities remain informative after the class stops changing. Updating scores improves selection, and exchanging inputs can improve a subset whose capacity is already filled. These findings lead to GraphSelect, which starts from individual candidate gains and refines the subset through jointly evaluated exchanges. It screens promising removals and additions, accepts an exchange when it reduces the prediction loss, and updates the scores. Experiments on six graphs show higher mean objective recovery than six attribution and explanation methods adapted to the selection task. Across nine trained architectures on two graphs, retaining 20% of the candidate representations per modality gives a mean accuracy drop of 0.10 percentage points relative to full candidate input, preserving classification performance with substantially fewer text and image representations.
|
| 1306 |
OOD Generalization as a Bifurcation Problem
2609.33562
|
cs.LG
|
Nguyen-Thanh-Luong Doan, Quang-Vu Nguyen, Tang-Phu-Quy Le, Cong-Phap Huynh |
Systematic out-of-distribution (OOD) generation remains a critical bottleneck for continuous-time generative models. While standard joint classifier-free guidance (CFG) routinely fails to synthesize unobserved concept combinations, exact decomposed scoring gen...Systematic out-of-distribution (OOD) generation remains a critical bottleneck for continuous-time generative models. While standard joint classifier-free guidance (CFG) routinely fails to synthesize unobserved concept combinations, exact decomposed scoring generalizes robustly at the cost of severe computational overhead. In this work, we reveal that compositional binding is not a uniform process but a highly localized phase transition. We identify the semantic bifurcation window - the precise temporal interval where joint and decomposed vector fields meaningfully diverge. Exploiting this dynamic, we propose surgical guidance, a hybrid sampling strategy that restricts exact multi-pass scoring strictly to this critical window. On an OOD bi-digit MNIST testbed, surgical guidance achieves state-of-the-art compositional fidelity at a fraction of the inference cost, yielding a +5.3% absolute improvement in pairwise accuracy over the joint baseline by intervening during just the first 15% of the diffusion trajectory. Furthermore, our empirical analysis uncovers a fundamental topological divide: diffusion models (SDEs) force conceptual resolution immediately at peak noise, whereas Conditional Flow Matching (ODEs) delays structural binding until intermediate features emerge, establishing a new temporal framework for accelerating large-scale generative decoding.
|
| 1307 |
MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning
2609.33563
|
cs.LG
|
Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt |
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (...World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents' local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.
|
| 1308 |
Correct then Forecast: Observer State-Space Models for Time Series Forecasting
2609.33566
|
cs.LG
|
Alexis-Raja Brachet, Guillaume Clavier--Fr\'emond, Abdelhakim Ziani, Pierre-Yves Richard, C\'eline Hudelot |
Time series forecasting requires extrapolating the dynamics of an observed process beyond the last available measurement. Yet recurrent forecasting models typically treat observations as inputs that directly control their latent dynamics. It leads to a regime ...Time series forecasting requires extrapolating the dynamics of an observed process beyond the last available measurement. Yet recurrent forecasting models typically treat observations as inputs that directly control their latent dynamics. It leads to a regime change when these observations become unavailable at prediction time. Following a state-estimation perspective, we introduce Observer State-Space Models (OSSMs), a class of recurrent models that interprets the observed input time series as measurements of an underlying autonomous dynamical system. OSSMs explicitly separate latent-state propagation from measurement assimilation: a single transition governs the dynamics across both context and forecasting intervals, while available observations correct the estimated state through an observer. This formulation naturally exposes classical control-theoretic properties, including observability and convergence of the state estimation error. We further show that conventional and recent SSMs can be recovered as particular instances of our OSSM framework, thereby providing a unified interpretation of their recurrent dynamics and revealing modeling inconsistencies. We perform experiments across several benchmarks showing that OSSM achieves substantial improvements while maintaining the same parameter count and training setup as the corresponding SSM baseline. These results support a simple principle for recurrent forecasting: observations should correct the estimated latent state, rather than control the dynamics used to propagate it.
|
| 1309 |
HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training
2609.33570
|
cs.LG
|
Xinrui Chen, Mengyang Li, Ou Wu, Ji Zhang |
Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite su...Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-management methods set state fidelity from execution cost, tensor properties, or generic compression sensitivity, without explicitly incorporating GRPO's analytic update structure into state-fidelity allocation. We formalize this dependence as policy-update exposure, linking the current GRPO loss coefficients to state-level approximation sensitivity. These coefficients are available before backward without an additional backward pass. We introduce HiLoRe, which allocates graph-attributed recovery units among high-precision storage, low-precision compression, and deterministic recomputation using measured recovery utility and update-conditioned approximation risk. It combines high-precision storage and deterministic recomputation with low-precision recovery under a calibrated risk budget. Across five model-task settings with 2K responses and memory < 1.10 times GC's per-GPU actor-update peak, HiLoRe's actor-update throughput gains reach 13.5% over GC and 7.9% over the fastest evaluated baseline, with paired mean downstream-score differences below 0.6 percentage points.
|
| 1310 |
Compressing Value Predictions for Learning-Augmented Metrical Task Systems
2609.33580
|
cs.LG
|
Sizhe Li, Yecheng Li, Kun He |
Learning-augmented algorithms for metrical task systems (MTS) can exploit predictions of canonical dual values, but existing formulations typically require a prediction for every state. We study whether these predictions can be compressed to a small set of rep...Learning-augmented algorithms for metrical task systems (MTS) can exploit predictions of canonical dual values, but existing formulations typically require a prediction for every state. We study whether these predictions can be compressed to a small set of representative states while retaining their algorithmic value. We introduce landmark-compressed value predictions, in which the predictor reports predicted dual values only at $m$ landmarks and the remaining values are reconstructed by a Lipschitz extension. Our algorithm achieves additive excess cost $O(T\,r(L) + \sum_t \delta_t)$, where $r(L)$ is the covering radius of the landmarks and $\delta_t$ measures prediction error up to additive shifts; local and value-dependent bounds refine this guarantee. For sparse landmark sets on unit-spaced finite lines, we prove a matching $\Omega(T r_m)$ lower bound for every randomized algorithm using fixed landmarks, even with advance access to their entire exact absolute-value table. The prediction interface also matters: on two states with one landmark, exact absolute values permit horizon-independent excess, whereas exact relative values force worst-case expected excess linear in $T$. We give PAC guarantees for learning compressed prediction tables, with efficient empirical-risk minimization for fixed landmarks. Our results connect metric coverage, prediction interfaces, and learning guarantees for compressed predictions in online MTS.
|
| 1311 |
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration
2609.33586
|
cs.LG
|
Shangzhen Zhu, Muyan Hu, Tomasz Kozlowski |
On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated acc...On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characterize what a pretrained model does need by approximating softmax at inference in ten frozen decoder-only models (0.5B-72B). The number of positions the softmax map assigns probability to and within-row resolution can be cut substantially, yet uniform weighting of the same positions is damaging. Where a fixed resolution budget is placed matters as much as its size, with resolution near the row maximum consistently favored. Perturbations matched on scalar distortion produce model-dependent responses of opposite sign. These findings motivate Rowmax-PoT, a coarse logarithmic weight representation anchored at each row maximum, and Rowmax-H15, its hardware specialization in FlashAttention-4. On B200, the patched FP8 attention forward is 12.4% faster at causal 8K and 25.8% faster at non-causal 8K in host-side call-latency measurements; board energy per forward falls by 8.4% at causal 16K. Measured separately on the BF16 kernel path at 2K, Rowmax-H15 increases perplexity by 0.091-0.492% across five models from three families.
|
| 1312 |
Pretraining Transformers with Quantized Softmax in Attention
2609.33591
|
cs.LG
|
Shangzhen Zhu, Muyan Hu, Tomasz Kozlowski |
Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation....Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interval attention, which approximates the exponential using K+1 grid values. We vary per-row grid calibration, interpolation versus hard rounding, and the placement of a straight-through surrogate relative to normalization. We derive the corresponding backward rules, including calibration derivatives, and compare these choices in pretraining experiments matched on model, data, and optimizer. Detaching the row extrema leaves the forward computation unchanged but produces a delayed increase in validation loss. With hard rounding at K=4, min-max calibration and a pre-normalization surrogate incur a large loss gap; changing either choice substantially reduces it. At 124M parameters and 2.5B training tokens, fixed-window calibration with a post-normalization surrogate yields a validation loss gap of +0.019 nats relative to softmax at K=4, and with a pre-normalization surrogate yields +0.004 nats at K=16.
|
| 1313 |
Short-Length Code Designs for Integrated Sensing and Communications: A Deep Learning Approach
2609.33605
|
cs.LG
|
Muah Kim, Shuangyang Li, Tayyebeh Jahani-Nezhad, Rafael F. Schaefer, Giuseppe Caire |
Integrated sensing and communication (ISAC) enables joint communication and sensing using a shared waveform, but its signal design is challenging due to the inherent trade-off between the two objectives, particularly in the short blocklength regime. This paper...Integrated sensing and communication (ISAC) enables joint communication and sensing using a shared waveform, but its signal design is challenging due to the inherent trade-off between the two objectives, particularly in the short blocklength regime. This paper proposes an autoencoder (AE)-based framework for ISAC waveform design in noncoherent settings. We derive a modified Cram\'er-Rao bound for multi-target delay estimation and analyze the maximum-likelihood decoding rule for noncoherent communication under correlated fading. These results reveal structural connections and trade-offs between communication and sensing objectives in waveform design. Based on this analysis, the AE learns waveform representations that jointly optimize both functionalities, with a tunable parameter controlling the trade-off. Simulation results show that the proposed design outperforms conventional schemes in both communication reliability and sensing accuracy, especially under short blocklength and fading conditions.
|
| 1314 |
Hierarchical Response Preservation for Continual Adaptation of Zero-Shot Graph-Text Models
2609.33607
|
cs.LG
|
Haopeng Zhang, Yuhan Wang, Yubing Su, Yingxin Chen, Xiao Wang |
Pretrained graph-text models align graph representations with textual semantics, enabling recognition of unseen classes and transfer across graph domains. However, as graph data and classes continually arrive, models should learn from new supervision while ret...Pretrained graph-text models align graph representations with textual semantics, enabling recognition of unseen classes and transfer across graph domains. However, as graph data and classes continually arrive, models should learn from new supervision while retaining their zero-shot transfer capabilities and historical task knowledge. Two challenges arise: (i) new classes can overturn historical predictions despite preserved distinctions among historical classes, and (ii) overly strict response preservation can stall learning of new classes. To address these challenges, we propose Hierarchical Response Preservation (HiRP). HiRP represents this competition through a hierarchical response that keeps each historical-class probability and sums new-class probabilities, preserving historical distinctions and aggregate competition while allowing distinctions within the new class group to adapt. It further uses the geometry induced by this response to guide constrained updates, retaining useful adaptation directions while controlling response drift. Across three class-incremental settings, HiRP achieves absolute gains of 1.84-7.95 percentage points in average accuracy over the strongest compared baseline in each setting, while mitigating zero-shot transfer degradation.
|
| 1315 |
You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement
2609.33609
|
cs.LG
|
Jiarong Wen, Qi Wang, Yun Qu, Yixiu Mao, Heming Zou |
In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinator...In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likelihood proxies to implicitly assess ICL quality. Making repeated queries to the target LLM with these strategies can incur substantial costs. This work simplifies selection by framing it as a constrained local search problem and presents local demonstration editing (LDE). Starting with an initially retrieved set of demonstrations, LDE employs a single structured edit to explore its surrounding neighborhood while balancing performance gains with search costs. Technically, LDE is reduced to a policy search problem, for which we train a small LLM, referred to as Jev-LDE. This model as the System-1 modifies the retrieved demonstration set by performing actions such as \texttt{Keep}, \texttt{Delete}, or \texttt{Replace} elements, all within a framework of reinforcement learning with verifiable rewards. At test time, Jev-LDE executes a single edit of the retrieved demonstration set, followed by one inference from the target LLM, avoiding the need for iterative context scoring or subset searches. Across standard classification benchmarks, various target LLMs with Jev-LDE as the plug-and-play module consistently improve ICL performance, and Jev-LDE shows transferability to held-out benchmarks and models without retraining. These findings indicate that the LDE approach offers an efficient and adaptable method for harnessing the ICL capabilities of target LLMs.
|
| 1316 |
Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss
2609.33620
|
cs.LG
|
Yi Ren, Wenlong Deng, Guanzhe Hong, Clare Lyle, Yarin Gal |
Modern language models are likely to be updated throughout their lifetime rather than trained once and frozen. Each update therefore participates in a recurring cycle: decide which experience to learn from, understand what that update changes, and remain capab...Modern language models are likely to be updated throughout their lifetime rather than trained once and frozen. Each update therefore participates in a recurring cycle: decide which experience to learn from, understand what that update changes, and remain capable of learning from what comes next. We show that these challenges are governed by the same evolving update--behavior interaction. We derive a token- and layer-wise decomposition of how learning from one token changes another prediction. By separating the softmax force, shared readout geometry, and residual connections, it exposes two interaction channels and yields a forward-computable approximation. Following this interaction through time reveals a unified picture of continual adaptation. Positive interaction identifies useful experience; negative interaction produces either concentrated collision or accumulated erosion; over longer horizons, updates reshape the shared geometry mediating future learning signals, reducing their transmission. These predictions lead to effective data selection, mechanism-specific controls for interference, and a readout-based diagnostic of future learnability whose degradation predicts the benefit of restoring the readout. Across models and training regimes, the same local interaction thus explains both what an update changes now and how learning today changes what can be learned tomorrow. This view connects data attribution, forgetting, and plasticity loss as distinct regimes of the same evolving learning dynamics.
|
| 1317 |
Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning
2609.33628
|
cs.LG
|
Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang, Jinyuan Jia |
Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. Howe...Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such as GPT-6-Luna, a major challenge is the cold-start problem: every attack attempt by the attacker LLM fails and thus receives zero reward, providing no signal for learning. In this work, we propose a curriculum learning-based method to address the cold-start problem. In particular, we propose to train the attacker LLM against a sequence of increasingly robust target LLMs, with each stage warm-starting from the attacker LLM obtained in the previous one. However, simply training against a weak target (e.g., GPT-4o-mini) may not sufficiently prepare the attacker LLM to obtain useful learning signals against a frontier LLM (e.g., GPT-5.6-Terra). Instead, we find that the design of the curriculum is critical: after each stage, the attacker LLM needs to partially succeed against the next target LLM such that it can learn from successful attempts to attack the new target. Our extensive evaluation shows that our method can effectively red-team frontier LLMs, achieving an attack success rate (ASR@10) of 93.8\% and 45.0\% against GPT-5.6-Luna and GPT-5.6-Terra on AgentDyn, whereas state-of-the-art RL methods such as RL-Hammer and PISmith achieve 0\% ASR under the same setting. Moreover, we find that the attacker LLM transfers across targets, e.g., an attacker LLM trained to defeat one strong LLM (GPT-5.6-Terra) also succeeds against six other frontier LLMs (e.g., GPT-6-Luna) it was never trained on. Our code is available at https://github.com/albert-y1n/PIForge.
|
| 1318 |
SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models
2609.33640
|
cs.LG
|
Xinmiao Wang, Ruijie Wang, Menghui Wang, Jiawei Chen, Haoyue Deng |
Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned ...Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows that safety robustness must hold across input modalities while balancing safety, over-refusal, and utility. To address these challenges, we construct SafeMolBench, a molecular multimodal safety-alignment benchmark with 3702 samples covering 618 unique hazardous molecules and safe molecular tasks, organized into hazardous-harmful, hazardous-allowed, and utility-replay subsets to support unified training and evaluation of safety, over-refusal, and utility. Based on SafeMolBench, we propose SafeMol, a parameter-efficient safety alignment framework that jointly optimizes lightweight modules across text-only and graph-conditioned inputs, uses MMD for distribution-level representation alignment to reduce modality-induced discrepancies, and explicitly models molecular hazardousness and harmful operational intent. Experiments on SafeMolBench show that SafeMol reduces attack success by several tens of percentage points while largely maintaining low over-refusal and preserving molecular-task utility.
|
| 1319 |
From Distributions to Stochastic Processes: Neural Approximation of Measure-Valued Maps
2609.33649
|
cs.LG
|
Yichen Wang, Ziyi Wang, Wenlian Lu, Chenghuang Shen, Jianfeng Liu |
Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learnin...Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learning and extend it to mappings between stochastic processes. For continuous operators on $W_2$-compact families of finite-dimensional probability laws, we establish uniform neural approximation in the 2-Wasserstein metric using finite law statistics, a simplex-valued neural map, and a shared atomic output support that guarantees valid probability measures. We further extend this principle to probability laws on separable Hilbert spaces through finite-rank orthogonal projections. These results establish the representational feasibility of learning transformations between probability laws rather than deterministic vectors or functions. To demonstrate practical relevance, we study two problems naturally defined at the distribution level: prediction of first-passage-time distributions for an Ornstein--Uhlenbeck process and nonlinear response-path laws of a Duffing oscillator. Because the theory is model-agnostic and broader than any single practical architecture, the experiments use task-adapted neural models rather than reproducing the theoretical construction exactly. In both problems, the proposed models outperform a fixed-feature MLP baseline and distribution-space kernel regression. These experiments complement the theory by demonstrating the practical learnability of distribution-to-distribution transformations in random systems.
|
| 1320 |
FuseAlign: Forced Alignment in the Wild
2609.33650
|
cs.LG
|
Mithilesh Vaidya, Stephen Bailey, Sumukh Badam, Matthew Bendel, Xingzhe He |
Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by ...Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.
|
| 1321 |
Towards Eliminating Catastrophic Forgetting in the Curriculum Learning of Math Reasoning Tasks
2609.33655
|
cs.LG
|
Zengyan Yang, Yangyang Wu, Kai Huang, Pengfei Lyu, Tianyi Zhang |
Curriculum learning has found broad application across numerous domains. Nevertheless, its effectiveness is intrinsically curtailed by catastrophic forgetting, driven by the shifts in model parameter distributions between curriculum tasks. In this paper, we in...Curriculum learning has found broad application across numerous domains. Nevertheless, its effectiveness is intrinsically curtailed by catastrophic forgetting, driven by the shifts in model parameter distributions between curriculum tasks. In this paper, we investigate the phenomenon of catastrophic forgetting in this training paradigm, building on the established efficacy of curriculum learning. Our theoretical analyses of parameter update dynamics demonstrate that catastrophic forgetting in curriculum learning stems from the divergence of task optima, which is generally essential to the faster convergence of curriculum learning; therefore, forgetting cannot be completely eliminated. Based on this finding, we augment the training process and propose IV-EWC, which incorporates Elastic Weight Consolidation (EWC) into the curriculum learning objective to curb catastrophic forgetting in mathematical reasoning, a prototypical curriculum learning scenario. IV-EWC employs the influence function to construct a representative validation set from the curriculum's training data, which is used to drive dynamic regularization during training. We further present an extended theoretical analysis to show that EWC-based regularization methods mitigate catastrophic forgetting in curriculum learning, thereby providing theoretical support for IV-EWC. Empirical evaluations on three backbone models and three benchmarks indicate that curriculum learning exhibits catastrophic forgetting. IV-EWC alleviates this issue, reducing forgetting by 162% on average relative to vanilla curriculum learning and yielding positive backward transfer, as evidenced by improved performance on easier tasks after subsequent training on challenging tasks.
|
| 1322 |
Scalable Attribution and Control of Model Behavior During Training
2609.33667
|
cs.LG
|
Sleem Abdelghafar |
Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individu...Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individual contributions difficult to distinguish. We address this ambiguity through mutual information, accounting for interference within the batch by quantifying how much the combined behavioral change reveals about each example's contribution. We show that this mutual information is a logarithmic function of Behavioral Gradient Uniqueness (BGU). BGU gives the information measure its geometric interpretation. Our Batch-Space Ghost (BS-Ghost) algorithm makes these scores practical inside the training loop through shared computation in batch space, without storing model-sized example gradients. On a complete 1,000-example Qwen2.5-7B-Instruct workload, our BS-Ghost implementation adds 27 seconds (8.0%) to 5.5 minutes of ordinary training. Removal and retraining demonstrate that BGU identifies data that causally shapes final behavior. At each training step, signed information identifies which examples strengthen or weaken the target behavior, explaining how behavior develops during training. Signed information also enables cheap intervention during training: it predicts how changing example weights will affect behavior in the next update. We then use these predictions to choose weights that steer behavior toward a desired target. This makes our framework a practical foundation for scalable oversight and verification of training pipelines and processes, helping evaluators assess model alignment, understand how it develops during training, and guide interventions that shape ongoing learning.
|
| 1323 |
Reachability is not enough: Diagnosing long-range behavior in GNNs
2609.33674
|
cs.LG
|
Filippo Maria Bianchi |
Graph neural networks (GNNs) are often called long-range because their architecture can connect distant nodes, but this does not show whether they use distant information correctly. We introduce a framework that measures how strongly inputs at each graph dista...Graph neural networks (GNNs) are often called long-range because their architecture can connect distant nodes, but this does not show whether they use distant information correctly. We introduce a framework that measures how strongly inputs at each graph distance affect predictions and separates limitations due to architecture, finite approximation, training, and numerical execution. Our analysis shows that local message-passing can spread influence slowly, so a finite implementation may rely mainly on nearby inputs even when the ideal computation uses the whole graph. We also explain why mathematically equivalent filters can differ in how easily they are learned and how reliably they run. Across controlled tasks, models with similar architectural reach use distant information very differently, while low average error can hide failures on distant interactions. Together, these results show that long-range capability depends on learning to use information at the distances required by the task and preserving that use during computation.
|
| 1324 |
Benign Overfitting for General Norms and Distributions
2609.33675
|
cs.LG
|
Daniel Barzilai, Ohad Shamir |
Understanding why predictors can generalize despite interpolating noisy training data is a central puzzle in machine learning. Most work on such "benign overfitting" studies minimum-2-norm linear regression, reflecting the inductive bias of gradient descent. H...Understanding why predictors can generalize despite interpolating noisy training data is a central puzzle in machine learning. Most work on such "benign overfitting" studies minimum-2-norm linear regression, reflecting the inductive bias of gradient descent. However, modern optimizers such as Adam and Muon use non-Euclidean update geometries, favoring solutions associated with other norms. Analyzing regression for non-Euclidean norms is substantially more difficult, with known results essentially limited to Gaussians. In this paper, we develop a method to analyze benign overfitting in linear regression for general norms and general (sub-Gaussian) distributions. As a special case, we prove that minimum-p-norm interpolation with p>1 can benignly overfit even for non-Gaussian distributions, under suitable conditions. Perhaps surprisingly, for the 1-norm, benign overfitting does not hold in general for well-behaved (but non-Gaussian) distributions, showing that existing positive 1-norm results rely crucially on Gaussianity. Our proof analyzes the geometry of the dual optimization problem, using concentration and central limit tools to show it is approximately Euclidean in many high-dimensional cases.
|
| 1325 |
T-MoXAI: A Hierarchical Explainability Framework for Temporal Multimodal Data
2609.33685
|
cs.LG
|
Ali Inha, Mo Vali, Saaliha Vali, Pietro Li\`o, Meen-Yau Thum |
Artificial Intelligence (AI) models for temporal multimodal data have potential in healthcare and agriculture, but their opacity can limit trust and adoption. We introduce T-MoXAI (Temporal Multimodal eXplainable AI), a hierarchical framework explaining (1) wh...Artificial Intelligence (AI) models for temporal multimodal data have potential in healthcare and agriculture, but their opacity can limit trust and adoption. We introduce T-MoXAI (Temporal Multimodal eXplainable AI), a hierarchical framework explaining (1) when timepoints influence predictions, using temporal Shapley values; (2) which modalities contribute at those moments, using attention analysis; and (3) what features or image regions drive decisions, using gradient based attribution. A transformer based architecture handles irregular temporal sequences and heterogeneous data, generating all three explanation levels in under one second for interactive decision support. We evaluate the framework on two real world tasks: predicting IVF treatment outcomes from ultrasound sequences and clinical measurements (AUC 0.660 despite significant class imbalance), and forecasting wheat yield from temporal RGB imagery and phenotypic traits ($R^2$ 0.265 amid substantial environmental variability). Ablation studies indicate that temporal modelling is critical in both domains: removing it reduces performance to the equivalent of random guessing. Temporal ROAR experiments provide evidence that the explanations reflect the model's reasoning process. With a unified, domain agnostic architecture and open source implementation, T-MoXAI provides a baseline for temporal multimodal XAI, addressing fragmentation in the field and supporting applications where understanding decisions is as important as predictive accuracy.
|
| 1326 |
Dynamic Kuramoto-Hodge Operators for PDEs on Complex Geometries and Topologies
2609.33693
|
cs.LG
|
Xiang Li, Yue Song |
Learning PDE operators on complex domains requires capturing interactions among fields on vertices, edges, and faces, alongside global responses shaped by topology. Existing neural operators accommodate irregular geometries but often overlook these distinct fi...Learning PDE operators on complex domains requires capturing interactions among fields on vertices, edges, and faces, alongside global responses shaped by topology. Existing neural operators accommodate irregular geometries but often overlook these distinct field supports or their condition-dependent coupling. We introduce the Dynamic Kuramoto--Hodge Operator (DKHO), which combines topology-constrained interactions with learned coordination. DKHO encodes conditions on their native cochain supports, evolves Kuramoto-inspired relation states through the boundary and coboundary operators that compose the Dirac operator, and decodes non-harmonic and harmonic responses in orthogonal Hodge subspaces. Topology thus determines where information can flow, while learned dynamics adapts how it is exchanged to each PDE instance. Across porous-medium Darcy flow, torus transport--diffusion, and cavity magnetostatics, DKHO-large reduces prediction error by approximately 61% on average over leading baselines, while DKHO-small remains competitive using only 11.5--24.3% as many parameters. These results suggest that coupling topological structure with adaptive dynamics provides an effective inductive bias for accurate and parameter-efficient PDE operator learning on complex geometries and topologies.
|
| 1327 |
Geometric Inductive Biases for Semi-Supervised Equalization: The Constellation-Aware Transformer
2609.33695
|
cs.LG
|
Avi Caciularu |
Decoding signals over unknown channels with minimal pilot overhead is a critical challenge in next-generation communications. Existing deep learning approaches typically rely on generic encoders that struggle to model long-range temporal dependencies or effici...Decoding signals over unknown channels with minimal pilot overhead is a critical challenge in next-generation communications. Existing deep learning approaches typically rely on generic encoders that struggle to model long-range temporal dependencies or efficiently capture the channel's physical properties from scarce data. We argue that standard architectures suffer from agnostic estimation gaps, as they must implicitly learn the constellation geometry that is already known. We introduce the Constellation-Aware Transformer (CAT), a novel architecture that explicitly injects geometric inductive biases into the equalization process. CAT is composed of a stack of custom TransFIRmer blocks, which use an "early interaction" paradigm to co-process received signals and ideal constellation symbols. Each block features a split Feed-Forward Network that applies a Finite Impulse Response (FIR)-inspired filter for deconvolution and a parallel MLP for geometric refinement. We show that this design is structurally aligned with the optimal linear (MIMO Wiener) receiver: its attention can implement a matched-filter bank, and its bidirectional FIR branch provides the non-causal filtering that block MMSE equalization requires. In the semi-supervised setting, CAT needs fewer pilots than VAE and standard Transformer baselines: on two of our three ISI channels, it reaches a lower SER with 64 pilots than they do with 128.
|
| 1328 |
A Spectral Theory of Compositional Learning
2609.33708
|
cs.LG
|
Hugo Rydel |
How does compositional reasoning emerge during learning? We address this question by mathematically analyzing the learning dynamics of deep linear networks. We train these networks in structured synthetic environments and derive a theory linking the structure ...How does compositional reasoning emerge during learning? We address this question by mathematically analyzing the learning dynamics of deep linear networks. We train these networks in structured synthetic environments and derive a theory linking the structure of experience to compositional learning. Our theory predicts when compositional inferences emerge, whether they are identifiable from the available evidence, and how new linking evidence can rapidly unlock previously unavailable inferences. These results provide a qualitative explanation for several phenomena observed in human cognition. They account for why a composition can fail despite knowing its premises, why similar compositions can emerge at different times, and how a single linking fact can suddenly enable many new inferences. Taken together, these findings establish a mathematical link between the statistical structure of experience and the development of compositional reasoning.
|
| 1329 |
DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation
2609.33711
|
cs.LG
|
Ao Yu, Weibo Gao, Heng Zhou, Linan Yue, Rui Li |
On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignore...On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.
|
| 1330 |
Theory Guided and Interpretable Neural Operator Design for Partial Differential Equation Learning
2609.33715
|
cs.LG
|
Zeyuan Song, Zheyu Jiang |
Accurate numerical solutions of partial differential equations (PDEs) are crucial in numerous science and engineering applications. In this work, we introduce a novel neural PDE solver named AFDONet, which incorporates neural operator learning and adaptive Fou...Accurate numerical solutions of partial differential equations (PDEs) are crucial in numerous science and engineering applications. In this work, we introduce a novel neural PDE solver named AFDONet, which incorporates neural operator learning and adaptive Fourier decomposition (AFD) theory for the first time into a specifically designed variational autoencoder (VAE) structure, to solve a general class of nonlinear PDEs on smooth manifolds. AFDONet is the first neural PDE solver whose architectural and component design is fully guided by an established mathematical framework (in this case, AFD theory), turning neural operator design from an art to a science. Thus, AFDONet also exhibits exceptional mathematical explainability and groundness, and enjoys several desired properties. Furthermore, AFDONet achieves outstanding solution accuracy and competitive computational efficiency in several benchmark problems. In particular, thanks to its deep connections with AFD theory, AFDONet shows superior performance in solving PDEs on i) arbitrary (Riemannian) manifolds, and ii) datasets with sharp gradients. Overall, this work presents a new paradigm for designing explainable neural operator frameworks.
|
| 1331 |
ALDER: Discovering the Laws of a World by Acting in It
2609.33728
|
cs.LG
|
Teng Cao, Yu Deng, Quentin Delfosse, Kristian Kersting |
Reliable world models should not only predict future states but express how actions change the world in an explicit, transparent and testable form, such as equations. Yet methods that rely on a fixed set of trajectories cannot distinguish equally good competin...Reliable world models should not only predict future states but express how actions change the world in an explicit, transparent and testable form, such as equations. Yet methods that rely on a fixed set of trajectories cannot distinguish equally good competing hypotheses, while searches over a fixed set of predefined candidates cannot discover equations outside the initial hypothesis space. We introduce ALDER (Action-guided Law Discovery, Evaluation, and Revision), a method that actively proposes novel experiments to test and revise models. Specifically, ALDER proposes parametric equations; a numerical optimizer fits their coefficients; an independent verifier tests these candidates on held-out data. To distinguish between competing valid hypotheses, a cost- and safety-aware selector queries interventions, in the form of novel experiments. The resulting counterexamples update the evidence ledger and guide the next structural revision, while incompatible laws are discarded. Across an in-house benchmark, ODE equation discovery tasks, and robotic experiments, ALDER discovers laws beyond its initial formula set, repairs failed model proposals, distinguishes fixed candidate models with fewer interactions, and improves out-of-distribution prediction. Furthermore, given a current state and a target, ALDER selects control actions by solving the inverse problem defined by its validated world model. Together, these results show that explicit equation-based world models can be tested and revised through interaction, then naturally used to guide goal-directed control.
|
| 1332 |
PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention
2609.33746
|
cs.LG
|
Kunming Shao, Jierun Chen, Yanli Wang, Ruoyu Wang, Haoli Bai |
At each decoding step a language model attends over the key-value (KV) cache of every earlier token, so at long context the attention call is bounded by memory bandwidth. Sparse attention reads only a subset of keys chosen by a cheap score estimate, and most m...At each decoding step a language model attends over the key-value (KV) cache of every earlier token, so at long context the attention call is bounded by memory bandwidth. Sparse attention reads only a subset of keys chosen by a cheap score estimate, and most methods give the unread tokens zero weight. The output then draws on only a small fraction of the KV cache, and accuracy drops at small budgets, most on tasks that aggregate information across the context. An inverted-file product-quantization (IVF-PQ) index over the cached keys computes an approximate score for every indexed token in order to rank them; after ranking, those scores approximate the attention logits of the tokens left out. PQ-HSA (hybrid sparse-approximate attention) attends the selected tokens with their original keys and values, and the unselected tokens, the background, enter the same softmax through those scores, summed per inverted list and multiplied by the list's mean value. At 128K and a 1-2% retrieval budget, PQ-HSA is more accurate than Quest and SnapKV on Llama-3.1-8B and Qwen3-30B-A3B and stays close to full attention in macro accuracy; with the same selector, the background term raises macro accuracy on the 8B model from 0.71 to 0.83. In the same 128K setting, inside vLLM on one NVIDIA H20, the decode attention call runs 1.6x faster than the FlashAttention-3 kernel; the speedup grows with context length, and a cost model fitted on 8B to 30B models gives the context length at which it begins. A vLLM plugin runs PQ-HSA on two engine versions without changes to the engine source; code is available at https://github.com/KunmingSHAO/pqhsa_release.
|
| 1333 |
Task-Aware Discretization of Differentiable Logic Gate Networks
2609.33747
|
cs.LG
|
Thore Gerlach |
Differentiable logic gate networks (DLGNs) enable gradient-based training of highly efficient Boolean networks by relaxing discrete logic gates during training and discretizing them for inference. Standard approaches make this discretization decision locally, ...Differentiable logic gate networks (DLGNs) enable gradient-based training of highly efficient Boolean networks by relaxing discrete logic gates during training and discretizing them for inference. Standard approaches make this discretization decision locally, typically through argmax selection and confidence- or entropy-based convergence criteria. We show that local discretization can be task-suboptimal even for globally optimal relaxed solutions, with high gate confidence providing no general guarantee, and derive bounds relating task-aware gate selection to tractable interventions in the relaxed network. Motivated by these results, we study first-order downstream task information for progressive discretization and characterize when this local approximation is reliable. Experiments on convolutional DLGNs reveal a strong locality dependence: first-order scores become unreliable when directly optimized over nonlocal interventions, but accurately assess local argmax decisions for progressive freezing.
|
| 1334 |
ResDiffFRG: Residual Diffusion for Multiple Appropriate Facial Reaction Generation
2609.33749
|
cs.LG
|
Shizhe Liu, Jiayan Gu, Xiangyu Kong, Siyang Song |
In dyadic human speaker-listener conversations, the listener's facial reactions allows the speaker to accurately perceive the listener's emotional states. Since human facial reactions are non-deterministic, the ability to generate multiple appropriate human-li...In dyadic human speaker-listener conversations, the listener's facial reactions allows the speaker to accurately perceive the listener's emotional states. Since human facial reactions are non-deterministic, the ability to generate multiple appropriate human-like facial reactions is crucial for realistic human-agent interactions. Although diffusion models are naturally suited to such one-to-many generation, existing diffusion-based Multiple Appropriate Facial Reaction Generation (MAFRG) methods attempt to denoise random Gaussian initialisations directly into multiple appropriate facial reactions (AFRs). These random initialisations are usually not well-aligned with the target listener facial reaction, which requires complex denoising trajectories from these initialisations, and subsequently creates substantial opportunities for deviations away from the range of trajectories leading to appropriate AFRs. Given the inherent mimicry between the human listener's and speaker's facial behaviours, we address the above denoising trajectory issue by leveraging this strong prior. Specifically, we propose ResDiffFRG, a novel diffusion-based MAFRG framework that explicitly anchors the diffusion process to the speaker behaviour by defining its diffusion target as the residual between the speaker anchor and an AFR. The denoiser only needs to model the comparatively small, reaction-specific residual needed to transform this anchor into an AFR, rather than reconstructing the complete reaction from an unstructured state. Extensive experiments show that ResDiffFRG achieves large improvements in correlation-based appropriateness over existing methods. Our denoising trajectory analysis showed that even at the start of the denoising trajectory, ResDiffFRG already achieves a higher facial-reaction correlation score than the Gaussian Diffusion baseline does after completing 60% of its denoising trajectory.
|
| 1335 |
Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos
2609.33754
|
cs.LG
|
Simeon Allmendinger, Domenique Zipperling, Burhanettin Bahadir Kibar, Niklas K{\"u}hl |
Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analyt...Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analytics. Federated learning enables collaboration without direct data sharing but does not resolve minority-class scarcity. Synthetic data generation can help, yet lightweight methods are interpolation-bound, while generative models require substantial data and computation. Existing collaborative generative approaches often rely on federated learning, imposing considerable organization-side training burdens. In this paper, we examine CollaFuse as a collaborative diffusion-based alternative for fraud detection and evaluate it across five fraud datasets. Compared with classical oversampling, local generative baselines, and centralized diffusion benchmarks, CollaFuse does not achieve the highest local fidelity but improves downstream fraud detection more consistently across most datasets. These findings suggest that synthetic data create analytical value less through local realism than through transferable cross-organizational structure.
|
| 1336 |
Oracle-Efficient Online Classification with Stochastic Inputs and Adversarial Outputs
2609.33760
|
cs.LG
|
Gon Buzaglo, Elad Hazan |
We consider contextual binary prediction with i.i.d. contexts from an unknown distribution and adaptively chosen losses. We show that a simple Follow-the-Perturbed-Leader algorithm with Gaussian perturbation for each observed context achieves the optimal $\wid...We consider contextual binary prediction with i.i.d. contexts from an unknown distribution and adaptively chosen losses. We show that a simple Follow-the-Perturbed-Leader algorithm with Gaussian perturbation for each observed context achieves the optimal $\widetilde O(\sqrt{T\log N})$ expected regret for a class of $N$ experts, while requiring one optimization-oracle call per round and no explicit enumeration of the class. For an infinite hypothesis class $\mathcal H$, the algorithm attains $\widetilde O(\sqrt{T\operatorname{VC}(\mathcal H)})$ regret. This resolves an open problem posed by Lazaric and Munos (2012), showing that hybrid classification is computationally as easy as statistical learning.
|
| 1337 |
StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences
2609.33761
|
cs.LG
|
Mustafa Ozaytac, Ozge Karadag Atas |
Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-d...Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-discriminator GAN with evolutionary weight adaptation, around a held-out protocol: the final two calendar years of each dataset are held out behind a 168 hour embargo, calibration is fitted on the training block only, and all metrics are computed on the held-out block. Evidence comes from 25 matched (location, seed) pairs across five Koppen-Geiger climates, tested with Wilcoxon signed-rank tests under Holm correction. Four results follow. First, isotonic calibration drives the Kolmogorov-Smirnov distance to within 2% of a per-location noise-and-shift floor for every architecture tested, including a deliberately weak RCGAN baseline, so calibrated marginal metrics cannot discriminate between architectures. Second, the sorted-representation discriminator is the only component whose removal significantly degrades cross-variable dependence (Kendall tau MAE +0.080, Holm p = 0.009), with a regime-dependent effect: near zero in Ankara, above 115% in Dubai and Yakutsk. A rank-transformed variant isolates the mechanism as quantile supervision of the marginals rather than copula matching. Third, physical constraint violations are injected by calibration, not the generator; projection removes them at negligible cost (deltaKS <= 0.003). Fourth, pooled metrics conceal a collapse of between-sequence weekly-mean variability, a proxy for seasonal and regime diversity, in TimeGAN that only sequence-level statistics expose. We recommend floor-referenced marginal evaluation, matched-pair testing, and sequence-level variance decomposition as minimum requirements for calibrated generative pipelines.
|
| 1338 |
Beyond Fixed Features: Architecture-Dependent Sensitivity to Node Representations under Heterophily
2609.33764
|
cs.LG
|
Priyanath Maji, Sidharth Gaur, Rajavinoth Paul Durai |
Graph Neural Networks (GNNs) perform well on homophilic graphs but struggle in heterophilic settings, where connected nodes often carry dissimilar labels. Existing evaluations typically compare architectures under a fixed node-feature representation, leaving u...Graph Neural Networks (GNNs) perform well on homophilic graphs but struggle in heterophilic settings, where connected nodes often carry dissimilar labels. Existing evaluations typically compare architectures under a fixed node-feature representation, leaving unclear whether conclusions about heterophily robustness remain stable as the input representation changes. We address this question by constructing parallel feature variants of two large-scale heterophilic benchmarks, Roman-Empire and Amazon-Ratings, pairing each graph with representations ranging from static fastText vectors to contextual Transformer embeddings and evaluating seven GNN architectures across these representations. We find that the effect of representation varies across architectures: on Roman-Empire, the contextual gain ranges from 2.38 percentage points for GCN-sep to 13.67 points for GAT, with H2GCN gaining 8.77 points. On Amazon-Ratings, where node text is limited to short product titles, GAT improves by 6.78 points from fastText to MPNet, while GCN-sep changes by only 0.20 points. These results show that architectural performance is conditional on node representation: the same representation change can produce different magnitudes of performance gain across architectures, so architecture and representation cannot be treated as independent evaluation factors. A rank-correlation analysis on these two benchmarks further shows that the relative ordering of architectures remains highly stable across representations, isolating differential sensitivity, rather than ranking instability, as the primary effect.
|
| 1339 |
DEALS: Decentralized Expertise-Aware Load Serving for Multi-Agent LLM Systems
2609.33768
|
cs.LG
|
Jingjuan Huang, Wenbin Wang, Yanchuan Yin, Alvaro Velasquez, Jia Liu |
Multi-agent systems (MAS) have recently emerged as an effective approach for coordinating large language model (LLM)-based agents to solve complex tasks through structured interactions. In practice, MASs often handle a stream of heterogeneous and complex tasks...Multi-agent systems (MAS) have recently emerged as an effective approach for coordinating large language model (LLM)-based agents to solve complex tasks through structured interactions. In practice, MASs often handle a stream of heterogeneous and complex tasks, requiring agents to decompose each task and then self-organize and self-evolve to adapt to incoming tasks while sharing execution resources. However, most early approaches to MASs rely on centralized controllers or fixed coordination patterns, which can limit scalability or adaptability. In contrast, existing decentralized and dynamic MASs often require training dedicated routers or invoking LLMs for agent selection, resulting in substantial computational costs and coordination overhead. To address these challenges and enable efficient task-level self-organization and self-evolution for task- and workload-level collaboration, we propose Decentralized Expertise-Aware Load Serving (DEALS), a decentralized and low-complexity framework that enables agents to self-organize and dynamically route concurrent tasks for processing. Specifically, each agent maintains local queues of incoming tasks, and its router decides whether to process a task locally or forward it to a neighbor based on differences in backlog and success rate. Meanwhile, executors process independent tasks concurrently within and across agents, and partially solved tasks can be resumed by other agents. Experiments show that DEALS not only improves performance along multiple dimensions (e.g., answer accuracy and task throughput) in both homogeneous and heterogeneous agent pools, but also balances agent expertise and workload in a self-organized manner, enabling effective decentralized coordination.
|
| 1340 |
Selecting Diverse SFT Traces Improves Post-RL Generalization
2609.33780
|
cs.LG
|
Dylan Zhang, Mingyuan Wu, Jinning Li |
Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a...Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. From one pool at one budget, with matched training recipes and checkpoints, selecting diverse rather than similar routes improves post-RL problem coverage across puzzles and mathematics, including on problems harder than those seen in either training stage. In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT. In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks. Pre-RL diagnostics suggest why: diverse SFT can produce both successful and failed attempts on more prompts despite slightly lower mean accuracy, giving group-relative RL more prompts with a learning signal. On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance. These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL.
|
| 1341 |
Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
2609.33791
|
cs.LG
|
Wenze Lin, Jiyuan Long, Jiale Zhao, Shenzhi Wang, Xitai Jiang |
Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL diverg...Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in this work, we find that KL divergence may not be necessary for OPD. We show that simply preserving the update direction is sufficient for effective OPD. As long as the update direction is toward the teacher, OPD works. More precisely, it is not the direction of every token, but the direction of a small subset of tokens where the teacher and student disagree strongly. We first show that simply assigning a reward of (+1) to tokens where the teacher probability is higher than the student probability and (-1) where it is lower, which merely encourages updates toward the teacher, reproduces almost the same training mode as OPD with reverse KL. We further show that only the direction of a small subset of tokens with large teacher-student disagreement is critical, and training works as long as their update direction is toward the teacher, even if other tokens are pulled away from the teacher. And as an application of these findings, we introduce Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to improve Multi-Teacher On-Policy Distillation (MOPD). Unlike MOPD, which routes each sample to a single teacher and may cause capability conflicts across domains, C-MOPD lets every sample be supervised by all teachers. Experiments show that C-MOPD consistently outperforms MOPD on both math and code benchmarks. Our code is available at https://github.com/LeapLabTHU/KL-Free-OPD.
|
| 1342 |
Binding Multiple Modalities via Multimodal Wasserstein Barycenter
2609.33800
|
cs.LG
|
Xiaole Tang, Jiayi Xu, Xiang Gu, Yan Yang, Jian Sun |
Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of ...Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of $n$-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to establish a unified semantic space around the WB embedding. Specifically, we project specific modalities to the WB, which minimizes the average Wasserstein distances to multimodal distributions and serves as the anchor for subsequent alignment. We then construct a barycenter simplex, whose volume is taken as a similarity metric for global alignment centered at the WB. Experiments show that BaryBind achieves competitive performance in text-video-audio retrieval, classification, videoQA, and cross-modal generation tasks, along with robustness under modality absence and scalability to more than three modalities. Code is released at https://github.com/xl-tang3/BaryBind.
|
| 1343 |
Diffusion Reward Models
2609.33803
|
cs.LGcs.AI
|
Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze Wang, Ziqing Qiao |
Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multim...Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many ways, and no single family covers all of them. To better fit this structure, we introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over $p(\mathbf{r}\mid x,y)$. Conditioned on a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector, placing no parametric assumption on the output distribution and naturally representing its multimodal structure. A single architecture handles both multi-attribute regression and pairwise preference data, and at inference $N$ samples form an empirical reward distribution that can be aggregated into a scalar, a variance, or quantiles. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, stays competitive with much larger discriminative, distributional, and generative RMs despite its modest training scale, and recovers multimodal reward structure where conventional heads collapse to a point. Uncertainty-aware rejection and lower-confidence-bound (LCB) aggregation further demonstrate that DRM can exploit distributional information beyond a scalar reward to improve reward-model decisions. Downstream RLHF experiments additionally show that using DRM as the training-time reward leads to improved policy performance, directly validating the practical benefit of diffusion-based reward modeling for RLHF training.
|
| 1344 |
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
2609.33804
|
cs.LG
|
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao |
Modeling spatiotemporal coupling is a key challenge in building physical intelligence across scales, from microscopic to macroscopic. Existing models capture such structure broadly through physics-motivated dynamical formulations or learning-motivated architec...Modeling spatiotemporal coupling is a key challenge in building physical intelligence across scales, from microscopic to macroscopic. Existing models capture such structure broadly through physics-motivated dynamical formulations or learning-motivated architectures. The former provide stronger priors but may constrain flexibility, whereas the latter are more flexible but leave the spatiotemporal coupling largely implicit. We therefore seek an approach that combines flexible learning with an explicit geometric bias for jointly modeling time and space. To this end, we propose Minkowski Positional Encoding (MinkowskiPE), which uses joint temporal and spatial coordinates to parameterize Lorentz transformations applied to query and key features. With MinkowskiPE, the query-key attention score depends on position only through the relative spacetime displacement between the two tokens and is therefore invariant to global translation of the coordinates. This paradigm retains the standard dot-product attention interface and remains compatible with efficient attention implementations. We evaluate MinkowskiPE on microscopic molecular dynamics and macroscopic video prediction tasks, achieving the best results on all nine multi-trajectory molecular evaluations and reducing KTH video-prediction MSE by 9.9% relative to the best baseline while using roughly one-tenth as many parameters.
|
| 1345 |
dOPT: Differentiating Conic Optimization via Geometric Reduction
2609.33828
|
cs.LG
|
Fengyu Yang, Connor W. Magoon, Tyler Watts, Shahar Z. Kovalsky |
Optimization layers enable the incorporation of structured constraints and decision problems into learning systems. Training such systems requires differentiating through the embedded optimization problem, which can be challenging for general conic programs. W...Optimization layers enable the incorporation of structured constraints and decision problems into learning systems. Training such systems requires differentiating through the embedded optimization problem, which can be challenging for general conic programs. We introduce dOPT, a solver-agnostic framework that, rather than differentiating the full conic formulation, reduces it at a computed primal-dual solution to an equality-constrained quadratic program that preserves the reference solution and its first-order sensitivity. The reduction captures the local first- and second-order conic geometry relevant to differentiation and remains well defined at singular configurations. Computing solution derivatives then requires a single symmetric linear solve, independently of the forward solver. We derive explicit reductions for convex NLPs, QPs, SOCPs, and SDPs. Numerical experiments validate the computed gradients and show favorable backward-pass scalability, with substantial speedups over existing differentiable conic optimization methods as problem size increases.
|
| 1346 |
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
2609.33848
|
cs.LG
|
Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang |
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to ...Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% $\to$ 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to $1.85\times$ and $1.78\times$ speedups over Colocate and Async, respectively.
|
| 1347 |
PI-NOMT: Physics-Informed Neural Optimal Mass Transport for Brain Fluid Dynamics
2609.33857
|
cs.LG
|
Mehmet Emin Acar, Vahit Bugra Yesilkaynak, Helene Benveniste, Gozde Unal |
Recovering hidden transport mechanisms from sparse spatiotemporal observations is a fundamental inverse problem in scientific machine learning. In brain tracer imaging, dynamic contrast-enhanced MRI (DCE-MRI) provides time-resolved measurements of tracer conce...Recovering hidden transport mechanisms from sparse spatiotemporal observations is a fundamental inverse problem in scientific machine learning. In brain tracer imaging, dynamic contrast-enhanced MRI (DCE-MRI) provides time-resolved measurements of tracer concentration, while the underlying velocity and source mechanisms governing tracer propagation remain unobserved. We formulate this problem as physics-informed latent-state inference, in which the transport field itself is the primary object of inference rather than an auxiliary variable used only to reconstruct observed densities. We propose Physics-Informed Neural Optimal Mass Transport (PI-NOMT), a framework that represents density, velocity, and source as continuous neural fields and combines a continuous neural density teacher, recursive differentiable advection--diffusion--source rollout, unbalanced optimal-transport regularization, and governing-equation supervision. Physical laws act as structural priors that constrain the space of admissible transport mechanisms, while observed tracer dynamics provide evidence for estimating the latent transport state. We evaluate PI-NOMT on a synthetic benchmark with known ground-truth transport and on DCE-MRI sequences from nine control rats. On the synthetic benchmark, PI-NOMT accurately recovers the prescribed velocity field, including its magnitude, direction, and integrated trajectories, rather than merely reconstructing endpoint densities. Across the nine rat datasets, the framework yields sub-percent local endpoint error, consistent physical speed scales, and low post-training PDE and incompressibility residuals. These results support physics-informed latent-state inference as a general framework for recovering hidden transport mechanisms from observed dynamic scalar fields.
|
| 1348 |
JET: Justification Evaluation in Transformer
2609.33874
|
cs.LG
|
Shenghao Ding |
JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across candidates. Experiments on desktop CPUs and consumer GPUs ass...JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across candidates. Experiments on desktop CPUs and consumer GPUs assess decision accuracy and execution cost. Qwen3.6-35B-A3B achieves 87.48% accuracy on the full MMLU test set and 3.69 requests per second on a separately timed MMLU subset. The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference. Controlled execution experiments show 2.18-2.23-fold speedups from prefix reuse and cache management, and a 30.8% reduction in process time from input preparation optimizations, with unchanged outputs. Optional reasoning has a task-dependent accuracy-throughput trade-off. These results support local decision inference from existing models.
|
| 1349 |
Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits
2609.33888
|
cs.LG
|
Idan Attias, Orin Levy, Alexander Ryabchenko, Yishay Mansour, Uri Stemmer |
Can vanilla policy optimization explore enough to achieve near-optimal regret in stochastic contextual bandits? We show that standard exponential policy updates driven by offline regression do so under realizability, without exploration bonuses or importance w...Can vanilla policy optimization explore enough to achieve near-optimal regret in stochastic contextual bandits? We show that standard exponential policy updates driven by offline regression do so under realizability, without exploration bonuses or importance weighting. For $A$ actions, $T$ rounds, and a finite prediction class $F$, vanilla PO achieves $\widetilde O(\sqrt{AT\log(|F|)})$ regret with high probability. Our analysis reveals an implicit exploration mechanism of independent interest: gradual policy updates prevent actions from losing probability too quickly, allowing the regression oracle to learn their expected losses. We further develop a batched version using only $O(\log T)$ regression calls and policy switches, and show how private regression oracles yield differentially private contextual bandit algorithms without composition across batches. For a finite class, this gives pure $\varepsilon_{\rm priv}$-DP and regret $\widetilde O\left( \sqrt{AT \log(|F|/\delta)}(1+\varepsilon_{\rm priv}^{-1/2}) \right)$. Finally, experiments across oracle-based contextual bandit algorithms, with and without privacy, demonstrate the practical effectiveness of policy optimization and the value of explicit exploration under stronger privacy constraints.
|
| 1350 |
Augmented Feature Boosting for Multicalibration
2609.33891
|
cs.LG
|
Ira Globus-Harris, Inbal Livni Navon |
Multicalibration requires a predictor's residuals to be unbiased not only globally, but also after conditioning on the predictor's own level sets and reweighting by a rich class of test functions. Standard boosting approaches in the distributional setting achi...Multicalibration requires a predictor's residuals to be unbiased not only globally, but also after conditioning on the predictor's own level sets and reweighting by a rich class of test functions. Standard boosting approaches in the distributional setting achieve this by repeatedly discretizing the predictor's range then auditing and repairing the resulting level sets. One consequence is that in practice, the algorithm's guarantees are sensitive to this parametrization of the rounding parameter. A natural theoretical question, then, is how to do discretization-free boosting which avoids this rounding within the boosting process itself. Here, we analyze an alternative feature-augmentation boosting paradigm inspired by Tax et al. (2026): at each round, a squared-loss oracle is called on hypotheses that receive the previous predictor's output as an additional feature, and only the final predictor is rounded to have a finite set of level sets to provide the multicalibration guarantee with respect to. We give a theoretical analysis of this procedure through the expressivity of the augmented hypothesis class, and show how the expressivity of this class yields a hierarchy of guarantees, including multiaccuracy, multicalibration, and the stronger notion of level-set multicalibration.
|
| 1351 |
MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models
2609.33893
|
cs.LG
|
Zhi Wen Soi, Giulio Segalini, Jian-Jia Chen, Lydia Chen |
Large audio-language models (LALMs) produce fluent responses about audio but often hallucinate by making plausible yet ungrounded claims. Existing audio hallucination benchmarks mainly measure response correctness, leaving it unclear whether an LALM hallucinat...Large audio-language models (LALMs) produce fluent responses about audio but often hallucinate by making plausible yet ungrounded claims. Existing audio hallucination benchmarks mainly measure response correctness, leaving it unclear whether an LALM hallucinates or simply fails to understand the audio. We challenge correctness-based evaluation by defining two hallucination categories: (i) context, where claims are not grounded in the audio; and (ii) knowledge, where claims about audio-related topics lack support from externally verifiable facts. We introduce MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering both categories. To evaluate open-ended responses, we propose a groundedness judge that uses reference rubrics and judge prompts guided by human annotations. We extensively evaluate ten state-of-the-art LALMs and show that hallucination remains substantial. Even a frontier model such as Gemini 3.7 Flash reaches a hallucination rate of 36.5%. We further adapt and benchmark four mitigation methods from multiple domains for LALMs. Despite some improvements, effective hallucination mitigation remains an open challenge. Finally, we call on the community to evaluate hallucination and benchmark mitigation methods with MISHAP-Bench.
|
| 1352 |
No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability
2609.33898
|
cs.LG
|
Yiyong Liu, Jun Sakuma, Michael Backes, Rui Wen |
Training efficiency has become the central driver of recent progress in foundation models. To overcome the massive computational and data requirements of large-scale training, researchers increasingly adopt strategies such as selective data sampling, efficient...Training efficiency has become the central driver of recent progress in foundation models. To overcome the massive computational and data requirements of large-scale training, researchers increasingly adopt strategies such as selective data sampling, efficient pre-training, and simplified reinforcement learning pipelines. While these strategies drastically reduce overhead, they prompt a critical, yet neglected question: Is efficiency achieved at the expense of model robustness and security? To our knowledge, we present the first systematic cross-domain investigation of the efficiency-vulnerability trade-off. Across vision and language models, we show that efficiency-oriented training increases susceptibility to adversarial and privacy attacks. We characterize this vulnerability by analyzing the models' internal geometry and functional representations, demonstrating that the evaluated efficient variants consistently exhibit sharper loss geometry together with systematic changes in representational structure. We further extend our analysis to "zero RL training", finding that models trained using simplified RL recipes exhibit substantially greater susceptibility to catastrophic forgetting and more pronounced overconfidence than those trained through conventional alignment pipelines. Our findings suggest that training efficiency is rarely a "free lunch"; rather, the mechanisms that minimize computation can inadvertently compromise safety. We conclude by calling for a paradigm shift toward multi-objective training that jointly optimizes for performance, cost, and security.
|
| 1353 |
Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
2609.33901
|
cs.LG
|
Soutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron, Nadav Dym |
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empiric...Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical foundations remain limited. In this work, we study when finite probe-based representations are sufficient for learning neural functionals. We establish general identification and universality results for probing, and show that using intermediate hidden representations can provide significantly more informative representations than relying only on final outputs. Motivated by these results, we introduce HIDDENPROBE, a simple architecture for learning from hidden probe responses. Across a range of neural functional benchmarks, including both MLPs and Transformers, HIDDENPROBE consistently improves over existing probing methods and achieves state-of-the-art performance. Our code is publicly available on GitHub.
|
| 1354 |
On the Two Faces of Adam in Separable Linear Classification
2609.33904
|
cs.LG
|
Chen Fan, Csaba Szepesv\'{a}ri |
We consider the behavior of deterministic, full-batch, bias-corrected Adam in separable linear classification with softmax parametrization under log-loss. In this setting, under a wide range of conditions Adam is known to approach max-norm-margin optimality wh...We consider the behavior of deterministic, full-batch, bias-corrected Adam in separable linear classification with softmax parametrization under log-loss. In this setting, under a wide range of conditions Adam is known to approach max-norm-margin optimality when its stability constant $\epsilon$ is zero, while with a positive $\epsilon$, it is known to approach Euclidean-margin optimality. Our main contribution is the quantitative description of Adam's behavior for small fixed positive $\epsilon$. We give sufficient conditions under which an Adam-trained classifier nearly maximizes the max-norm margin before the updates become gradient-like. We also show that the classifier reaches a fixed target Euclidean margin only much later. Specifically, we show that for polynomially decreasing stepsizes with exponent \(a\), where \(1/3<a<1\), the updates become approximately proportional to the negative gradient after $\Theta(\log(1/\epsilon)^{1/(1-a)})$ iterations. At that time, the classifier still nearly maximizes the max-norm margin. Reaching a fixed target Euclidean margin above that of every max-norm-optimal classifier, but below the optimum, is shown to require $\epsilon^{-\Theta(1)/(1-a)}$ iterations. Under inverse-linear stepsize decay (\(a=1\)), the update transition takes polynomially many iterations, whereas reaching the target margin takes exponentially many. Experiments support these predictions. The later change in the classifier can improve or worsen generalization after training error reaches zero, connecting the analysis to grokking and its reverse.
|
| 1355 |
JIVE: Jacobian-Informed Volume Expansion for Diverse Generative Sampling
2609.33906
|
cs.LG
|
Guangxun Zhang, Brian Cai, Boxuan Zhang, Chao Chen, Ruixiang Tang |
Generative models often suffer from mode collapse and limited sample diversity. While prior works attempt to mitigate this by jointly generating a batch of samples and repelling their trajectories, these heuristics do not explicitly maximize the diversity of t...Generative models often suffer from mode collapse and limited sample diversity. While prior works attempt to mitigate this by jointly generating a batch of samples and repelling their trajectories, these heuristics do not explicitly maximize the diversity of the resulting endpoints. We introduce JIVE, a training-free framework that enhances generative diversity by injecting velocity perturbations aligned with the leading right singular subspace of the generator's endpoint Jacobian. By leveraging this local geometric structure, JIVE provably maximizes endpoint diversity while preserving sample quality. To maintain practical efficiency, we compute these perturbation directions via matrix-free iterations rooted in classical numerical linear algebra, requiring only a small computational overhead. Across different benchmarks, JIVE boosts both pixel and feature-level diversity in few-step and one-step generation.
|
| 1356 |
Training Witnesses: Trusting the Training without Trusting the Trainer
2609.33915
|
cs.LG
|
Houjun Liu, Pratyusha Sharma |
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditional...Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of "auto-certified" training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.
|
| 1357 |
Optimizing the Phi-2 Small Language Model for Real-time Chatbot Applications Using Parameter-Efficient Fine-Tuning (PEFT) with QLoRA Quantization
2609.33927
|
cs.LG
|
PhanTan Khanh Nguyen, Ashfaq Ali Shafin, Khandaker Mamun Ahmed |
This study explores the optimization of the Phi-2 Small Language Models (SLMs) for real-time chatbot applications through Parameter-Efficient Fine-Tuning (PEFT) and Quantized Low-Rank Adaptation (QLoRA). QLoRA specifically refers to the integration of PEFT wit...This study explores the optimization of the Phi-2 Small Language Models (SLMs) for real-time chatbot applications through Parameter-Efficient Fine-Tuning (PEFT) and Quantized Low-Rank Adaptation (QLoRA). QLoRA specifically refers to the integration of PEFT with LoRA alongside a 4-bit quantization process, aimed at enhancing computational efficiency. These models, initially designed for high performance with minimal computational overhead, are further refined to address the constraints of mobile and edge computing environments. By integrating PEFT with QLoRA, the research aims to reduce memory usage significantly while maintaining, or potentially improving, the accuracy of model responses in real-time interactions. The effectiveness of these techniques was evaluated using the ROUGE metric system, which showed notable improvements in the summarization tasks performed by the models. This approach not only confirms the feasibility of using SLMs in resource-restricted environments but also opens up new avenues for deploying advanced AI-driven applications in real-time settings. The study's findings have significant implications for the development of efficient, scalable, and accessible AI technologies, paving the way for broader adoption in various industries.
|
| 1358 |
Diffusion-Based Rollouts as a Stabilization Mechanism for Long-Horizon Environmental Forecasting
2609.33930
|
cs.LG
|
Marina Vicens-Miquel, Amy McGovern, Aaron J. Hill, Efi Foufoula-Georgiou, Samuel S. P. Shen |
Extending forecast lead times while maintaining predictive skill remains a major challenge in environmental forecasting. We investigate diffusion-based rollouts as a stabilization mechanism for recursive forecasting using low-dimensional water-level time serie...Extending forecast lead times while maintaining predictive skill remains a major challenge in environmental forecasting. We investigate diffusion-based rollouts as a stabilization mechanism for recursive forecasting using low-dimensional water-level time series and high-dimensional precipitation fields. Across both modalities, diffusion suppresses recursive error growth, with the largest stabilization occurring where deterministic rollouts are most unstable. However, stabilization does not guarantee forecast fidelity. In the water-level experiments, forecasts progressively lose event-level fidelity as the rollout loses access to external predictive information, and trajectory-level comparisons show that diffusion can remain numerically stable while contracting toward central values and exhibiting reduced variability. In the precipitation experiments, which retain conditioning from numerical weather prediction throughout the rollout, diffusion better preserves spatial organization and event-detection skill. Together, these contrasting experiments indicate that diffusion can control recursive error amplification, while its practical benefit also depends on the predictive information available to constrain future evolution.
|
| 1359 |
How Strong Is the Evidence for the Artificial Hivemind? Reevaluating Evidence for the Open-Ended Homogeneity of Language Models
2609.33936
|
cs.LG
|
Rylan Schaeffer, Brando Miranda, Joshua Kazdan, Jessica Chudnovsky, Sanmi Koyejo |
Recent research argues that language models exhibit pronounced homogeneity in open-ended generation, framing such behavior as an Artificial Hivemind that poses a long-term threat to human creativity. We examine three of its central results. First, the flagship...Recent research argues that language models exhibit pronounced homogeneity in open-ended generation, framing such behavior as an Artificial Hivemind that poses a long-term threat to human creativity. We examine three of its central results. First, the flagship example is that model responses to "Write a metaphor involving time" collapse into two clusters. Visualization, spectral analysis, clustering, and language model labels all contradict this description. The labels record each response's vehicle, what it compares time to. Our responses and the original authors' own show one dominant vehicle plus a heavy tail of distinct minority vehicles. "Time" is one of our least diverse topics, so the example is a favorable case, not a representative one. Second, the paper measures homogeneity against an undemanding null: responses to unrelated prompts. Under a more demanding null (same-prompt responses expressing genuinely different ideas), 20%-32% of such pairs already exceed the paper's 0.8 convergence threshold. A residual effect survives this null. The paper's same-prompt pairs exceed 0.8 roughly two to three times as often as our different-idea pairs. Much of what the paper calls homogeneity is the shared geometry of answering the same prompt. The remaining measurements lack any null: no human baseline is collected, and the model-indistinguishability statistic has no null. Third, the paper concludes that inference-time interventions are inadequate for combating the Artificial Hivemind, writing that "more generalizable solutions are needed at the model training level." We show that this conclusion is unsupported in three ways, and that an inference-time intervention (prompting) reliably raises measured response diversity. We do not resolve whether the Artificial Hivemind is real. We show that the published evidence does not establish it.
|
| 1360 |
Behavioral Monitoring of JEPA World Models with Jacobian Centroids
2609.33940
|
cs.LG
|
Thomas Walker, Randall Balestriero, Richard Baraniuk |
Detecting failures in World Model (WM)-based planning requires monitoring whether the model is behaviorally aligned with the current task, which in turn requires studying its internal representations. Here, we show that centroids---sub-component Jacobian row-s...Detecting failures in World Model (WM)-based planning requires monitoring whether the model is behaviorally aligned with the current task, which in turn requires studying its internal representations. Here, we show that centroids---sub-component Jacobian row-sums---effectively identify the behavioral properties of WMs, complementing traditional activation-based knowledge signals. The centroids of a model are easily computed through Jacobian vector products and characterize how the model organizes the geometry of its input space, yielding an efficient perspective on internal representations, including the generation of task-relevant saliency maps. Evaluated on continuous control tasks using JEPA WMs, this behavioral view reveals a structural dissociation, where the encoder correctly represents the goal while the predictor remains behaviorally unresponsive. This failure mode directly predicts planning failure before any action is taken, allowing for goal resampling to recapture out-of-distribution success. Moreover, centroid-based methods outperform baseline methods as distribution-shift detectors. Together, these tools yield a behavioral monitoring stack that is operational and consequential under distribution shifts.
|
| 1361 |
GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning
2609.33977
|
cs.LG
|
Zhengao Li, Shuoqiu Li, Xiaofang Zhang, Yukai Jin, Gokcen Kestor |
Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet i...Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, leaving open whether adaptive allocation is of limited value for semi-structured pruning in general or only under the fine-grained N:M pattern. We examine this question with group-level sparsity, which partitions each weight matrix into regular groups, retains or prunes each group as a whole, and allows each layer's sparsity ratio to vary under a global budget. We propose GroupMask, which generates the group selectors of all layers with a lightweight hypernetwork, relaxes them with a Gumbel-Sigmoid parameterization and a straight-through estimator, and learns them through sparsity-budget regularization and self-distillation while keeping the pretrained weights frozen. On LLaMA-2-7B at 50% sparsity with the same $1\times256$ group size, learned layer-adaptive allocation reduces WikiText-2 perplexity from 10.02 to 8.30 and raises the average zero-shot accuracy from 0.455 to 0.496 relative to a uniform per-layer ratio. GroupMask obtains the lowest WikiText-2 perplexity on LLaMA-2-7B and the highest average zero-shot accuracy with Alpaca calibration among the evaluated baselines on five LLaMA and Qwen models. Our code is available at https://github.com/ZhengaoLi/GroupMask.
|
| 1362 |
DynGraphAgentBench: A Benchmark for Agentic Lifecycle Control in Dynamic Graph Anomaly Detection
2609.33980
|
cs.LG
|
Yuwei Han, Lingwei Wei, Wooseong Yang, Liangjie Huang, Liancheng Fang |
Dynamic graph anomaly detection requires repeated decisions as graph structure and class prevalence drift, yet detector benchmarks usually score a fixed pipeline after current labels are known. We introduce DynGraphAgentBench, an executable benchmark for agent...Dynamic graph anomaly detection requires repeated decisions as graph structure and class prevalence drift, yet detector benchmarks usually score a fixed pipeline after current labels are known. We introduce DynGraphAgentBench, an executable benchmark for agentic lifecycle control under delayed feedback. It comprises seven temporal graph datasets with node- and edge-level anomaly tasks, eleven selectable detectors, and eight chronological deployment windows per dataset. In each window, a controller sees only time-causal aggregate context, registered model cards, and its own matured history. It must choose a detector before current-window training or candidate scores exist. A sandboxed executor trains the chosen architecture on mature data, scores a hidden deployment window, and releases the outcome after a one-window delay. A deterministic verifier checks decision timing, leakage guards, legal actions, training scope, and persisted artifacts. We measure detection utility with average precision and capture at fixed review depth, and characterize adaptation through model switches and compute. Complete eight-window trajectories from two primary controllers and a no-memory reference on four datasets, together with three additional controllers on three datasets, expose useful, costly, and ineffective reactions to delayed evidence without granting an exhaustive current-window oracle.
|
| 1363 |
Future Information-Directed Sampling for Bayesian Nonstationary Bandits
2609.33981
|
cs.LG
|
Yichen Song, Alessio Russo, Aldo Pacchiano |
Exploration--exploitation is a central trade-off in bandit learning. While classical algorithms such as upper confidence bound methods and Thompson Sampling effectively balance this trade-off in stationary environments, their exploration strategies mainly redu...Exploration--exploitation is a central trade-off in bandit learning. While classical algorithms such as upper confidence bound methods and Thompson Sampling effectively balance this trade-off in stationary environments, their exploration strategies mainly reduce uncertainty about the current optimal arm, which can be insufficient in nonstationary settings where future optimal arms may differ substantially from current ones. In this paper, we propose Future Information-Directed Sampling (FIDS), a new algorithm for Bayesian nonstationary bandits that explicitly explores to gather information about future optimal arms. We show that FIDS achieves regret comparable to Thompson Sampling up to a small constant factor, while being able to exploit predictive information structures that conventional exploration objectives fail to capture. To address the practical difficulty of posterior inference, we further propose a supervised-learning-based approximation framework that learns the FIDS policy from offline data, and demonstrate its effectiveness on synthetic benchmarks.
|
| 1364 |
From HL to H+L-1 Parameters: A Hankel-Toeplitz Forecaster for Long-Term Time Series Forecasting
2609.33984
|
cs.LG
|
Chaoqi Zhang, Yu Wang, Haixu Tang |
Linear forecasters have shown competitive accuracy against Transformer-based models in long-term time series forecasting. We study how classical stationary prediction theory can guide parameter sharing for more compact linear forecasters. For centered second-o...Linear forecasters have shown competitive accuracy against Transformer-based models in long-term time series forecasting. We study how classical stationary prediction theory can guide parameter sharing for more compact linear forecasters. For centered second-order stationary processes with nonsingular history covariance, the minimum-MSE finite-window linear predictor factors into a Hankel cross-covariance matrix and an inverse Toeplitz covariance matrix. Shared lags and scale cancellation specify this predictor using $H+L-1$ autocorrelations for lookback $L$ and horizon $H$. Building on the innovations representation, our Hankel-Toeplitz Forecaster (HTF) learns one impulse response that defines both an inverse filter and a forecast map. We characterize the finite-history correction and, under summability assumptions, bound the excess risk of truncating the true filters. HTF uses $H+L-1$ trainable coefficients while allowing a full-rank forecasting matrix. Across seven benchmarks at $L=336$, its horizon-averaged MSE is within 1.2% of Dense Linear on each dataset with 75-229 times fewer trainable parameters.
|
| 1365 |
ICMAPE: In-Context Multiagent Pure Exploration
2609.33986
|
cs.LG
|
Xinyi Hu, Alessio Russo, Aldo Pacchiano |
In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literatu...In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent problems with well-specified models, while there is currently a gap for practical multi-agent methods that can perform active sequential testing. We fill this gap with ICMAPE, a Bayesian learning-based framework for decentralized multi-agent pure-exploration driven by inference objectives. ICMAPE converts the fixed-confidence identification objective into a reward derived from inference confidence, so that standard reinforcement learning machinery can be applied to decentralized pure exploration. It jointly learns a centralized neural inference network that estimates a posterior distribution over hypotheses from global trajectory data, and decentralized policies that select actions from local observation histories and learn when to stop collecting data once the target confidence is reached. On two synthetic benchmarks and a Maryland nitrate concentration monitoring task based on real-world data, ICMAPE-TD3 achieves target accuracy with fewer exploration steps.
|
| 1366 |
ASTRA: ADMM-Accelerated Topology Reconfiguration for Dynamic Satellite Constellations
2609.33993
|
cs.LG
|
Jo\~ao Norberto, Ricardo Ferreira, Cl\'audia Soares |
Dynamic topology reconfiguration is central to the reliability and efficiency of large satellite constellations, yet many existing approaches rely on idealized assumptions such as full constellation deployment or uniform orbital spacing. We present Adaptive Sa...Dynamic topology reconfiguration is central to the reliability and efficiency of large satellite constellations, yet many existing approaches rely on idealized assumptions such as full constellation deployment or uniform orbital spacing. We present Adaptive Satellite Topology via Regret-Aware learning (ASTRA), a theoretically-grounded framework for dynamic satellite topology reconfiguration that builds on an online learning formulation and makes it computationally practical. ASTRA combines an ADMM-based offline solver with efficient online updates for both online gradient descent and online conditional gradient, yielding markedly cheaper constrained updates than generic optimization pipelines. On the theory side, we show that for a relevant class of entry-wise nonzero utility matrices, the objective is strongly convex, which yields logarithmic static regret for online gradient descent, and we further instantiate known dynamic-regret guarantees under inexact ADMM inner loops. Empirically, ASTRA matches or improves topology quality, presenting a good trade-off with computational time on synthetic constellations, and it remains effective on real Starlink data under partial deployment and non-uniform spacing, where idealized structural assumptions break down. These results position ASTRA as an efficient and theoretically grounded approach to topology reconfiguration in realistic Low Earth Orbit networks.
|
| 1367 |
T-SNN: Temporal Simplicial Neural Network for EEG Decoding
2609.34002
|
cs.LG
|
Nikita Malik, Shubhajit Roy, Mohit Kataria, Isuru Herath, Suraj Yadav |
Decoding brain states requires models that capture both the evolution of neural activity and interactions among groups of brain regions. Existing EEG methods often treat recordings as multivariate time series or represent functional connectivity with pairwise ...Decoding brain states requires models that capture both the evolution of neural activity and interactions among groups of brain regions. Existing EEG methods often treat recordings as multivariate time series or represent functional connectivity with pairwise graphs, leaving dynamic higher-order interactions largely unmodeled. We introduce the Temporal Simplicial Neural Network (T-SNN), which represents EEG recordings as sequences of evolving simplicial complexes. By combining simplicial convolutions with recurrent updates, T-SNN jointly learns higher-order interactions and their temporal evolution. On the seven-class SEED-VII emotion recognition task, T-SNN outperforms convolutional, recurrent, graph-based, and Transformer methods in both trial-wise and cross-subject evaluations. Incorporating eye-movement features further improves performance, demonstrating the framework's potential for multimodal brain-state decoding.
|
| 1368 |
RICE-Alpha: Reliability-Informed Correction with Event Graphs for LLM-Agent Stock Forecasting
2609.34004
|
cs.LG
|
Tong Liu, Lanmiao Liu, Xiang Hu |
Equity-relevant news evolves through temporally dependent corporate events, making historical information useful only when event continuity, information availability, and transition reliability are modeled. Existing LLM-based financial agents incorporate histo...Equity-relevant news evolves through temporally dependent corporate events, making historical information useful only when event continuity, information availability, and transition reliability are modeled. Existing LLM-based financial agents incorporate historical evidence, yet they provide limited support for preserving issuer-specific chronology under point-in-time constraints and for identifying when historical transitions contribute information beyond the current forecast. We present RICE-Alpha (Reliability-Informed Correction with Event Graphs), a point-in-time stock-scoring framework that separates a history-aware multi-view Base Alpha from a reliability-calibrated residual correction derived from historical event continuation. A Multi-Tier Memory Layer grounds news interpretation in temporally eligible issuer-specific history, while a Typed Event Agent constructs event states whose successor relations are formed within issuers and pooled across firms only after valid local pairing. Matured transitions are calibrated by their empirical reliability, and the resulting graph signal is residualized against the Base Alpha and technical view to obtain the RICE Delta. On daily Nasdaq-100 and Hang Seng Index panels from 2024 to 2026, RICE-Alpha achieves the strongest results among the evaluated LLM-based agents and momentum across four predictive and four portfolio-level metrics. Its ICIR more than doubles that of the strongest baseline, while net Sharpe ratios reach 1.656 and 1.725 in the U.S. and Hong Kong, respectively. U.S. ablations further show significant reductions in IC and RankIC after Holm adjustment when major components are removed. These results indicate that historical event continuation adds incremental information when it is temporally grounded, reliability-calibrated, and introduced as a residual correction to a multi-view forecast.
|
| 1369 |
Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs
2609.34009
|
cs.LG
|
Seohyun Lee, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. Brinton |
Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher ...Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.
|
| 1370 |
When Known Physics Helps Neural PDE Models: Residual Constraints Out-Regularize Generic Priors for Nonlinear Dynamics
2609.34012
|
cs.LG
|
Zahra Farazpay, Aniruddha Bora |
Neural PDE surrogates increasingly incorporate structural priors, yet it is often unclear whether their gains arise from physics-specific information or simply from regularization and training choices. We evaluate several such priors under a common protocol ag...Neural PDE surrogates increasingly incorporate structural priors, yet it is often unclear whether their gains arise from physics-specific information or simply from regularization and training choices. We evaluate several such priors under a common protocol against a matched from-scratch neural operator baseline. Our central result is that a known-equation residual consistently outperforms the best generic regularizer at equal tuning budget. At fixed capacity this benefit appears across linear and nonlinear PDEs, but a capacity sweep reveals a sharp distinction: the advantage persists and grows for Burgers, KdV, and Allen-Cahn, while collapsing toward or below parity for linear heat and advection-diffusion. Thus, the durable value of the residual is specific to nonlinear operators. We further falsify a pre-registered hypothesis that the benefit is activated only by data sparsity: the residual remains advantageous even under full supervision. Its usefulness does, however, have a clear boundary. Under grid under-resolution, nonlinear coarse fields no longer satisfy the naive governing-equation residual, and enforcing it becomes actively harmful. In contrast, cross-family pretraining and in-context conditioning fail to outperform the strong from-scratch baseline in the regime studied. Together, these results identify when known physics provides non-redundant information to neural PDE models, when it does not, and when enforcing it introduces bias.
|
| 1371 |
LTV-CTDNet: Compositional Turning Decomposition for Short-Term Turning-Movement Forecasting
2609.34014
|
cs.LG
|
Md Atiqur Rahman Mallick, Kamrul Hasan, Robert T. White |
Short-term turning-movement forecasts can support signal control and corridor operations, but unconstrained neural networks may produce physically impossible negative counts or outputs that are not explicitly tied to an approach-demand total. This study introd...Short-term turning-movement forecasts can support signal control and corridor operations, but unconstrained neural networks may produce physically impossible negative counts or outputs that are not explicitly tied to an approach-demand total. This study introduces the Linear Temporal-Variable Compositional Turning Decomposition Network (LTV-CTDNet), a forecasting framework designed to combine competitive accuracy with structurally admissible outputs. LTV-CTDNet was evaluated using seven months of 15-minute LiDAR observations from eight monitored corridor locations in Nashville, Tennessee. Its lightweight encoder combines recent turning-movement history, weekly time-slot embeddings, and location embeddings. The Compositional Turning Decomposition framework separately predicts nonnegative approach totals and within-approach turning proportions, then reconstructs movement forecasts from these components. Among the evaluated predefined configurations, LTV-CTDNet achieved a movement-level MAE of 1.8189 and RMSE of 3.8072. Its accuracy gains over the strongest sequence models were modest, but it produced no negative forecasts, while unconstrained learned models generated negative values in approximately 10.6% to 29.2% of raw forecast cells. The framework enforces nonnegative outputs and exact agreement between each model-predicted approach total and the sum of its component movements by construction, providing directly interpretable forecasts without clipping or coherence correction.
|
| 1372 |
SR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-Making
2609.34019
|
cs.LG
|
Shyam Sundar Murali Krishnan, Dean Frederick Hougen |
In many high-stakes applications, machine learning is dominated by black-box models that require post hoc explanations to justify their predictions. These explanations are often unreliable because they do not reflect the model's actual computations, limiting a...In many high-stakes applications, machine learning is dominated by black-box models that require post hoc explanations to justify their predictions. These explanations are often unreliable because they do not reflect the model's actual computations, limiting accountability and trust. A natural alternative is to use models that are interpretable by design. However, existing rule-based approaches, such as RuleFit and decision trees, while transparent, often lack stability and predictive strength, reinforcing a perceived trade-off between traditional performance measures and model understandability. To address this, we propose Sparse Relaxed Regularized Regression Rule-Fit (SR4-Fit), an intrinsically interpretable algorithm for both classification and regression that produces compact and stable rule sets without sacrificing performance. Using demographic data from the U.S. Census Bureau's American Community Survey, SR4-Fit predicts U.S. House election outcomes with high accuracy and interpretability while uncovering demographic interactions missed by black-box models. We further validate SR4-Fit across fourteen benchmark datasets (six classification and eight regression), where it outperforms existing rule-based methods, including RuleFit and decision trees in terms of accuracy, stability, and compactness while remaining competitive with black-box models in predictivity. These results demonstrate that interpretability and predictive reliability need not be mutually exclusive, offering a practical and transparent alternative for high-stakes decision-making.
|
| 1373 |
Structure-Adaptive Tree Field Integrators
2609.34025
|
cs.LG
|
Millend Roy, Soham Samal, Ivan Zelich, Krzysztof Marcin Choromanski |
We present a new class of near-linear algorithms for efficiently integrating general tensor fields defined on trees with distance dependent kernels, the Structure-Adaptive Tree Field Integrators (STAD-TFIs). STAD-TFIs exploit the tree's underlying structure th...We present a new class of near-linear algorithms for efficiently integrating general tensor fields defined on trees with distance dependent kernels, the Structure-Adaptive Tree Field Integrators (STAD-TFIs). STAD-TFIs exploit the tree's underlying structure through decompositions built around path backbones and single vertex separators, and use two-dimensional fast Fourier transforms to compute interactions jointly. By exploiting this structural information, STAD-TFIs achieve more computationally efficient integration than their regular efficient tree field integrators (TFI) counterparts. We provide a detailed theoretical analysis of our proposed approach and complement it with an exhaustive empirical evaluation, ranging from speed tests on synthetic trees, through accelerated Sinkhorn-based relaxations of the Optimal Transport algorithms on real meshes, to Topological Attention Transformers for vision tasks. To the best of our knowledge, we provide some of the first results showing that efficient to compute and accurate relaxations of the geodesic Sinkhorn-based solutions of the Optimal Transport problem can be derived by applying fast TFI methods.
|
| 1374 |
Posterior Regimes and Latent Deception: Variational Bayesian Inference in Hidden Markov Models for Sequential Fraud Detection in Financial Transactions
2609.34031
|
cs.LG
|
Joseph Uririoghene Obukofe, Anthony O'Hare, Chioma Sandra Dike |
We present a three-tier progression of Hidden Markov Models: maximum-likelihood (Baum-Welch), variational Bayesian (VBEM), and a neural variational extension (Neural VBEM), that model each customer's transaction history as a trajectory through a small number o...We present a three-tier progression of Hidden Markov Models: maximum-likelihood (Baum-Welch), variational Bayesian (VBEM), and a neural variational extension (Neural VBEM), that model each customer's transaction history as a trajectory through a small number of latent behavioural regimes, one of which is empirically identified as fraud-associated. The Neural VBEM HMM replaces the fixed Gaussian-multinomial emission family with a learned encoder, compressing a 741-dimensional transaction representation into a 64-dimensional latent space in which the VBEM HMM's posterior operates; a UMAP projection of this space reveals that the discovered regimes are not discrete clusters but ordered segments of a single continuous behavioural manifold, with confirmed fraud concentrated at its extreme. We show that the model's natural output, that is, the posterior probability of regime membership, is routinely mistaken for a fraud probability, and quantify the resulting miscalibration (the regime-membership interpretation error, MRIE); a corrected posterior-predictive score, closes most of this gap. We further distinguish batch (smoothed) inference, which uses look-ahead unavailable at deployment time, from filtered (forward-only) inference, and report both. On IEEE-CIS transaction data, the neural tier achieves a 14.4$\times$ fraud enrichment in its identified regime; while its AUPRC trails a discriminative XGBoost baseline, we show this gap is structural and not incidental, and argue the model is best positioned as a calibrated triage and interpretability layer rather than a drop-in ranking replacement.
|
| 1375 |
UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents
2609.34036
|
cs.LG
|
Wenbo Zhang, Pengcheng Xu, Weizhi Du, Jing Zhang, Hengrui Cai |
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence ...On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to $15.8\%$ relative to standard OPD.
|
| 1376 |
Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation
2609.34041
|
cs.LG
|
Samuel Tetteh, Cody Fleming |
Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard...Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can provide dense semantic feedback, yet it remains unclear whether their scores anticipate collisions and which component drives an observed safety improvement. Episodic cost can also favour policies that make little task progress. To address these gaps, we propose VLM-Safe-RL, a framework that integrates frozen CLIP signals into PPO-Lagrangian through reward shaping and an augmented multiplier update. On MetaDrive Hard, which combines the densest traffic with the largest map, the catastrophe rate falls from 31.6\% to 19.4\%. FormulaOne-L2 analysis finds no evidence that the CLIP signals anticipate collisions and shows that the VLM term has a negligible effect on the Lagrange multiplier. These findings show a conditional reduction in observed catastrophe rate without evidence of collision anticipation.
|
| 1377 |
PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction
2609.34054
|
cs.LG
|
Hyesung Jeon, Hyeongju Ha, Jae-Joon Kim |
Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy ...Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache sharing methods reduce this repeated prefill, but they either require additional training or architectural constraints or retain substantial model computation. Moreover, direct cache reuse causes the current agent to rely on cache states generated by the previous agent's adapter, weakening the role-specific behavior encoded by its own LoRA. We present PReCache, a training-free KV cache sharing framework with two designs, namely PreLRShared and ReBaseShared, that share the base cache computed using the pretrained weights and precompute a compact agent-specific low-rank (LR) cache. To remove repeated prefill, PreLRShared precomputes each agent's LR cache when the shared context is first processed, allowing the current agent to use its own LR cache without reprocessing context processed by previous agents. To improve sharing accuracy, ReBaseShared reconstructs the shared base cache from adapter-free hidden states, reducing the remaining error caused by the previous agent's adapted representation. To minimize its reconstruction cost, we propose two inference schemes tailored to single-stream inference and concurrent serving, performing the same reconstruction after each agent's turn or alongside its execution, respectively. Across multiple models and agent benchmarks, PreLRShared achieves up to a 3.1x TTFT speedup and a 2.3x improvement in per-request throughput over inference without KV cache sharing. ReBaseShared best preserves accuracy overall among the evaluated cache-sharing methods, with an average drop of only 1.1 points relative to inference without cache sharing.
|
| 1378 |
Do World Models Learn Global Understanding?
2609.34058
|
cs.LG
|
Alexander Detkov, Matt Thomson |
AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to...AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundamental insight, we frame "understanding" as learning constraints and propagating their consequences. We construct learning tasks on monoid worlds, sets of states connected by action transitions, where observed training transitions and an unseen constraint jointly determine held-out transitions. Measuring generalization tests whether models can learn global constraints from local transitions and propagate their consequences. We consider inverse, commutativity, composition, and periodicity constraints relevant to spatial and semantic structure. Across attention, recurrent, and state-space architectures, next-state training fits the data but fails to propagate non-trivial constraints. Compositional training, which uses identical paths but hides intermediate states from the input, achieves 96% accuracy on inverse, commutativity, and composition constraints across architectures, yields corresponding improvements in geometric generalization of world models trained on embodied environments and relational generalization in Wikidata-finetuned LLMs. How far do models propagate constraints when inferring an unseen fact may depend on first inferring others? We define proof depth d of a held-out transition, measuring the minimum number of inference rounds to infer the transition, and find that model generalization decreases sharply with proof depth. Increasing compositional path length T improves generalization. These results provide a formal way to investigate global understanding in language and world models and demonstrate that compositional training promotes information propagation and integration.
|
| 1379 |
KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems
2609.34060
|
cs.LG
|
Hyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang, Jae-Joon Kim |
Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeat...Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeatedly prefill the growing context and construct a separate cache with high computation and memory overhead. Selective recomputation reduces this redundancy but still retains substantial model execution, while existing delta correction methods either support only recurring context relations or maintain memory-intensive online correction states for dynamically changing context. For first seen shared context, these methods also construct a reference cache outside the agent workflow, and an approximate correction at the first agent affects the outputs passed to subsequent agents. We present KVCMAS, an online KV cache correction framework that represents cross-agent cache deviations using compact low-rank states and seamlessly chains corrections along the agent workflow without an additional reference prefill. This design supports dynamically changing shared context while preserving an exact first-agent cache. Across multiple language and vision-language workloads, KVCMAS matches or improves the accuracy of prior KV cache sharing methods while achieving the lowest TTFT under highly concurrent serving. Under controlled serving traces, it provides a 2.0x TTFT speedup over inference without KV cache sharing and reduces peak GPU memory by up to 3.7x relative to a prior KV cache correction method. These results establish KVCMAS as an accurate and scalable KV cache sharing approach for prompt-specialized multi-agent serving.
|
| 1380 |
Learning Perturbation Robust Policies for LLM Agents with Stable Optimization
2609.34064
|
cs.LG
|
Pengxin Wang, Yuanzhe LI, Yuxin Ren, Huanrui Yang, Jingdi Chen |
Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, a...Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study how to improve perturbation robustness during policy optimization. We first introduce the notion of a perturbation robust policy and analyze conditions under which perturbed policy updates preserve stable monotonic improvement. Based on this analysis, we introduce Stable Perturbation-Robust Policy Optimization (SPrPO), which applies adaptive and sensitivity-aware perturbations during RL training. We evaluate SPrPO on ALFWorld and WebShop and conduct systematic experiments across multiple perturbation types and scales, showing improved perturbation robustness while maintaining stable policy optimization.
|
| 1381 |
FLARE: Flow Matching with Local Axis-Angle Representations for Stochastic Micromagnetic Evolution
2609.34070
|
cs.LG
|
Pengyu Li, Renjie Tong, Xuanlue Jiang, Jianmin Li, Yuanyuan Zhou |
Long-horizon micromagnetic simulation remains expensive because conventional and learned solvers typically propagate Landau--Lifshitz--Gilbert (LLG) dynamics step by step. Existing learned approaches generally retain stepwise integration or model deterministic...Long-horizon micromagnetic simulation remains expensive because conventional and learned solvers typically propagate Landau--Lifshitz--Gilbert (LLG) dynamics step by step. Existing learned approaches generally retain stepwise integration or model deterministic evolution, leaving full-field, direct-horizon stochastic prediction largely unexplored. We propose FLARE, a flow-matching framework that recasts stochastic finite-time magnetization prediction as conditional transport over anchor-relative local axis-angle rotations. This rotation-space formulation respects the intrinsic geometry of magnetization dynamics and preserves pointwise unit norm by construction. By explicitly conditioning on the physical prediction horizon, FLARE directly generates full-field stochastic endpoints across multiple target times without stepwise integration. Against the strongest single-checkpoint external baseline on each metric, FLARE achieves 29.9% lower angular energy distance ($15.30^\circ$), and a 37.3% lower fair energy score (0.393). On a representative composed 5-ns two-segment protocol, FLARE achieves a $3{,}062\times$ best-batch speedup over the widely used GPU micromagnetic solver MuMax$^3$ on a single GPU.
|
| 1382 |
MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference
2609.34077
|
cs.LG
|
Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu |
Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce ...Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
|
| 1383 |
Beyond One Epoch: Uncertainty-Weighted Sensitivity Regularization for Recommendation Models
2609.34083
|
cs.LG
|
Richard Lettich, Shagun Gupta |
Recommendation models with sparse embeddings and a shared consumer often exhibit the one-epoch phenomenon: a second epoch lowers training loss while sharply degrading generalization. We present a view based on the violation of the prequential principle. On the...Recommendation models with sparse embeddings and a shared consumer often exhibit the one-epoch phenomenon: a second epoch lowers training loss while sharply degrading generalization. We present a view based on the violation of the prequential principle. On the first epoch, an example's label has not affected the embedding rows used to score it. On later epochs, those rows contain a displacement induced by the labels earlier update. This creates an incentive for the shared consumer to exploit this displacement in subsequent epochs, which fails to generalize. We call this self-influence asymmetry. Using an exact scalar model and local influence analysis, we connect this mismatch to the uncertainty in the embeddings and the consumers incentive to exploit it in subsequent epochs. We verify this hypothesis using an embedding-consumer-update interventions in deep recommendation models and propose uncertainty-weighted sensitivity regularization (UWSR) which counteracts this mismatch by augmenting the loss function to penalize the consumer for relying on uncertain embeddings. Unlike existing remedies, UWSR preserves the learned embeddings and across three benchmarks, four-epoch UWSR reduces test cross-entropy by 1.38%-6.78% and improves AUC by 0.0058-0.0231 relative to one-epoch training.
|
| 1384 |
TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care
2609.34088
|
cs.LG
|
Lovely Yeswanth Panchumarthi, Andrew Lu, Saurabh Kataria, Delgersuren Bold, Minxiao Wang |
TRACE (Text-Reinforced Analysis of Cardio ECGs) is a multimodal electrocardiogram (ECG) representation model that learns clinically grounded signal embeddings for downstream cardiac classification. It is designed to address the limitations of existing CLIP-sty...TRACE (Text-Reinforced Analysis of Cardio ECGs) is a multimodal electrocardiogram (ECG) representation model that learns clinically grounded signal embeddings for downstream cardiac classification. It is designed to address the limitations of existing CLIP-style training, which often struggles with noisy clinical text and fails to leverage the complementary strengths of unimodal (from ECG) and cross-modal (between ECG and matched cardiologist reports) learning. To bridge this gap, we propose a hybrid architecture that jointly learns unimodal and cross-modal representations via uncertainty-weighted multi-task learning while utilizing an LLM-based pipeline to extract high-fidelity findings from cardiologist reports. We evaluate TRACE across a spectrum of clinical urgency, establishing robust performance on public benchmarks for arrhythmia classification and structural abnormalities relative to existing unimodal and multimodal ECG models. To demonstrate real-world utility, we further validate the model on acute coronary occlusion (ACO), where the prevailing ST-elevation criteria miss 25-34% of true occlusions. Utilizing a large private ACO dataset with expert-annotated ground truth, TRACE significantly outperforms real-world clinical practice, yielding a 19.0% increase in sensitivity or a 62.6% reduction in false positive rates at the clinical baseline. This extensive evaluation confirms that TRACE delivers both strong performance on benchmark tasks and tangible clinical impact in the most acute, high-risk cardiac scenarios.
|
| 1385 |
When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models
2609.34091
|
cs.LG
|
Yucong Cao, Chenqi Li, Tingting Zhu |
Sparse autoencoders (SAEs) decompose dense model activations into discrete latents, making individual features easy to interpret--and easy to misinterpret. In EEG foundation models, this creates a tempting inference: if removing alpha-band activity strongly ch...Sparse autoencoders (SAEs) decompose dense model activations into discrete latents, making individual features easy to interpret--and easy to misinterpret. In EEG foundation models, this creates a tempting inference: if removing alpha-band activity strongly changes a latent's activation, one might conclude that the latent represents alpha activity. Across 27 settings spanning three backbones, three EEG datasets, and three network depths, this interpretation initially appears compelling: alpha removal changes latent firing 7.3 times more than an equal-width sham notch (95% CI [6.2, 8.7], bootstrapped over settings). However, the alpha filter also deletes far more signal than the sham. After normalizing by removed spectral energy, the ratio falls to 0.28 (95% CI [0.22, 0.36]) and exceeds one in none of the 27 settings. Latents selected for their response to alpha removal are, on clean EEG, slightly anti-correlated with relative alpha power (mean r = -0.073), giving no support for a simple alpha-detector reading. Motivated by this failure case, we propose a validation ladder for semantic interpretations of SAE latents: it asks in turn whether a latent responds, whether that response survives controlling for how much signal the intervention removes, whether it is specific rather than broadly fragile, and whether the proposed property is visible on unperturbed data--while separately testing the stronger claim that the latent matters to a task classifier. Perturbation sensitivity alone does not establish what an SAE latent represents.
|
| 1386 |
Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer
2609.34098
|
cs.LG
|
Seokyong Sheem, Hochang Lee, Suyeong Lee, Daekyum Kim |
Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppr...Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We show that these ablations can lead to misleading conclusions about predictive benefit: poor performance under altered semantics may be taken as evidence that the intended knowledge is beneficial. Across real and controlled experiments, altering semantic content can produce large performance differences even when the model gains little predictive benefit from having that semantic knowledge in the first place. To separate these effects, we distinguish two quantities: content sensitivity and predictive utility. Content sensitivity measures the change in performance when semantic content is altered, whereas predictive utility measures the benefit of the intended semantic knowledge relative to a suitable reference without that knowledge. This distinction motivates an evaluation framework in which the control is chosen according to the question being asked: altered controls assess sensitivity to semantic content, whereas claims that semantic knowledge improves prediction require a suitable reference. Even then, predictive utility is not fixed; it varies across suitable references and decreases when the reference can more easily recover the tested knowledge from other inputs or labeled examples. In a bounded audit of 25 semantic-ablation comparisons across nine studies, only one of 18 explicit predictive-utility claims is paired with a control that clearly isolates the tested semantic contribution. Together, these findings motivate a simple evaluation principle: semantic-ablation controls should be chosen and interpreted according to the question they are intended to answer.
|
| 1387 |
Evolution of fairness in multi-objective reinforcement learning framework
2609.34114
|
cs.LG
|
Jingyi Zhang, Xin Ou, Guozhong Zheng, Shengfeng Deng, Jiqiang Zhang |
Fairness, as a fundamental social norm, continues to pose a longstanding puzzle regarding its emergence. Traditional game-theoretic models largely rely on the assumption of \emph{Homo economicus}, wherein individuals are purely rational and self-interested, ac...Fairness, as a fundamental social norm, continues to pose a longstanding puzzle regarding its emergence. Traditional game-theoretic models largely rely on the assumption of \emph{Homo economicus}, wherein individuals are purely rational and self-interested, acting solely to maximize material payoffs. Such accounts, however, overlook the multidimensional nature of human decision-making, which is often shaped also by other considerations beyond economic incentives. To address this gap, we propose a multi-objective reinforcement learning framework that models the evolution of fairness as a dynamic trade-off between material payoff maximization and fairness-driven moral behavior, regulated by a fairness pressure coefficient. Using simulations of a two-objective Q-learning ultimatum game, we find that increased fairness pressure promotes fair outcomes, as expected. Strikingly, however, under moderate pressure, responder behavior reverses: responders become ``forgiving" by accepting low offers -- a pattern in line with our daily experience. Microscopic analyses reveal that this strategy reversal stems from competition between payoff-maximizing and fairness-oriented preferences. We further extend our framework to an asymmetric setting, where proposers and responders assign different weights to the two objectives. Overall, our work expands the reinforcement learning paradigm from a single-objective to a multi-objective formulation, offering a versatile tool for elucidating a broader range of human social behaviors.
|
| 1388 |
SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving
2609.34117
|
cs.LG
|
Gunho Park, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin, Baeseong Park |
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-...Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.
|
| 1389 |
Probabilistic electrical power demand forecasting with uncertainty quantification
2609.34120
|
cs.LG
|
Mahesh Neupane, Pragya Dhungana, Pradip Khatri, Swechhya Baskota, Hariom Dhungana |
The majority of research on electricity consumption forecasting has focused on deterministic approaches, which generate a single point estimate for each time step in the forecasting horizon. However, the increasing penetration of renewable energy sources and t...The majority of research on electricity consumption forecasting has focused on deterministic approaches, which generate a single point estimate for each time step in the forecasting horizon. However, the increasing penetration of renewable energy sources and the growing complexity of modern smart grids have introduced greater variability and uncertainty into power-system demand and operation. Consequently, probabilistic forecasting, which quantifies the uncertainty and variability associated with future electricity demand, is becoming increasingly important for reliable power-system planning and operation. This study presents an empirical comparison of four contemporary probabilistic forecasting models for electricity consumption, highlighting their respective strengths and limitations. We have performed comparision on real-world power systems related datasets. Across all power-consumption zones, NGBoost demonstrates superior probabilistic forecasting performance, achieving the lowest MAE and RMSE while providing well-calibrated uncertainty estimates with high prediction-interval coverage and reasonably narrow intervals. These results indicate that NGBoost offers a more accurate and reliable forecasting framework than Bayesian, Monte Carlo (MC) Dropout, and Gaussian Process Regression (GPR) models for the considered electricity consumption data.
|
| 1390 |
What Does a Stream Model Buy You in Flow Matching?
2609.34123
|
cs.LG
|
Jian Xu |
Stream-level flow matching replaces the linear interpolant of conditional flow matching (CFM) by a Gaussian-process (GP) stream connecting each source--target pair, and reports lower sample error than \icfm{} on 2-Gaussian, MNIST and CIFAR-10 benchmarks. We as...Stream-level flow matching replaces the linear interpolant of conditional flow matching (CFM) by a Gaussian-process (GP) stream connecting each source--target pair, and reports lower sample error than \icfm{} on 2-Gaussian, MNIST and CIFAR-10 benchmarks. We ask what such a stream model actually contributes. Three results answer the question. (i)~\emph{Reduction.} The stream-level CFM objective depends on the stream law only through the per-time joint law of $(s_t,\sdot_t)$, so the conditional paths a Gaussian stream can reach are exactly the Gaussian conditional paths CFM already parametrises; in the coordinate-wise, shared-scalar-kernel construction gpcfm actually uses, the entire design space collapses to two scalar curves $(m_t,v_t)$, and cross-time covariance affects only estimator variance. (ii)~\emph{The GP is a constrained chart of that space.} One kernel sets both $m_t$ and $v_t$, so the paper's own recipe for widening coverage-shrinking the SE length-scale---destroys the interpolant (the midpoint mean weight falls from $1.03$ to $0.00$). On the 2-Gaussian benchmark this makes the GP chart diverge on $15/200$ runs at high coverage against $0/200$ for a decoupled $(m_t,v_t)$ chart ($p=6.6\times10^{-5}$), and crossing the two curves shows the divergence tracks the mean, not the variance. On MNIST the same sweep does not diverge and the ordering reverses, so whether the coupling is harmful is benchmark-dependent; what holds on both is that the recipe buys nothing---no coverage level beats the paper's own, and past $\max_t\sqrt{v_t}\approx0.6$ both charts degrade. (iii)~\emph{Audit.} The released code does not implement the mechanism it describes: state and velocity are drawn independently ($\mathrm{corr}=0.00\pm0.01$ against an intended $\pm0.83$--$0.99$).
|
| 1391 |
ExpertoRhythm: Morphology-Aware Learning for Waveform Reconstruction and Cuffless Blood Pressure Estimation from Single-Channel PPG
2609.34146
|
cs.LG
|
Amir Arjomand, Kenneth B. Kent, Georgiy Krylov |
Continuous cuffless blood pressure (BP) monitoring from photoplethysmography (PPG) has strong potential for wearable health and telemonitoring, but accurate estimation remains difficult because PPG-to-BP mapping must preserve subtle waveform morphology and pre...Continuous cuffless blood pressure (BP) monitoring from photoplethysmography (PPG) has strong potential for wearable health and telemonitoring, but accurate estimation remains difficult because PPG-to-BP mapping must preserve subtle waveform morphology and pressure-range-dependent dynamics. We introduce ExpertoRhythm, an attention-enhanced 1D U-Net that reconstructs the arterial blood pressure (ABP) waveform from a single-channel PPG signal and derives systolic and diastolic BP directly from the reconstructed waveform. The central contribution is a composite morphology-aware learning objective that integrates range-weighted SmoothL1 reconstruction with a window-range regularizer to emphasize high-dynamic BP segments and reduce amplitude under/over-shoot. On the UCI cuff-less BP dataset with 942 subjects, ExpertoRhythm achieves 2.46/1.46 mmHg MAE for systolic/diastolic BP (SBP/DBP), while obtaining a 30.4% average relative error reduction over pure MSE across waveform reconstruction and BP estimation metrics. Clinical-style evaluation further demonstrates low bias and strong agreement across the BP range, including high-pressure windows up to 200 mmHg, satisfying AAMI criteria and achieving BHS Grade A. These results suggest that morphology-aware waveform reconstruction from a single PPG channel can provide an accurate and practical pathway toward continuous cuffless BP monitoring in wearable and remote-care settings.
|
| 1392 |
SPINET: Sheaf Protein Inverse Folding Network
2609.34153
|
cs.LG
|
Jens Lundsgaard, Colin Mikulski, Zhixuan Yan, Dhananjay Bhaskar |
Proteins change shape as they function, yet most inverse folding models predict amino acid sequences from a single, fixed backbone. A central challenge in protein engineering is to design proteins that undergo specific motions, which requires accounting for ho...Proteins change shape as they function, yet most inverse folding models predict amino acid sequences from a single, fixed backbone. A central challenge in protein engineering is to design proteins that undergo specific motions, which requires accounting for how their structures change over time. This motivates inverse protein folding conditioned on protein motion. We introduce SPINET, which predicts sequences from molecular dynamics trajectories. It uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass. We evaluate SPINET on mdCATH and ATLAS, where it outperforms all evaluated static and ensemble baselines in sequence recovery. On mdCATH, it achieves 56.7% top-1 recovery, compared with 44.5% for the strongest static baseline and 40.7% for the strongest ensemble baseline. We also evaluate whether the predicted sequences are compatible with conformations sampled along the target trajectory. On mdCATH, they achieve a median TM-score of 0.760, and structural recovery favors target conformations over unrelated decoys for 99.5% of test domains.
|
| 1393 |
Transfer Calibrated Prediction Powered Inference
2609.34156
|
cs.LG
|
Aditya T. Vadlamani, Jae Ho Chang, Srinivasan Parthasarathy, Subhadeep Paul |
Prediction-powered inference (PPI) and its power-tuned extension (PPI++) improve confidence intervals by combining a small gold-standard labeled sample with a large AI model's predictions. Its efficiency gain relies on low residual variance, which may not hold...Prediction-powered inference (PPI) and its power-tuned extension (PPI++) improve confidence intervals by combining a small gold-standard labeled sample with a large AI model's predictions. Its efficiency gain relies on low residual variance, which may not hold if the predictor is pre-trained on a different source domain. We propose Transfer Calibrated Prediction-Powered Inference (TC-PPI), adapting the source-domain predictor to the target domain using gold-standard samples through cross-fitting. This approach supports various adaptation methods, such as sparse linear calibration, LoRA, and fine-tuning. Our jointly tuned cross-fit estimator, Joint-TC-Cross-PPI++, maintains unbiasedness and is simultaneously at least as efficient as classical inference, PPI, and PPI++, thereby protecting against negative transfer. We provide high-dimensional MSE bounds for calibration and show empirical improvements over baseline methods across various real-world applications.
|
| 1394 |
WorldGraph: Graph-Native World Modeling
2609.34159
|
cs.LG
|
Zezhong Ding, Yipeng Li, Xike Xie |
World models infer latent states of an environment to capture its underlying dynamics and predict future evolution. Many real-world environments, however, are inherently relational and observed as evolving graphs, where entities, relations, and their propertie...World models infer latent states of an environment to capture its underlying dynamics and predict future evolution. Many real-world environments, however, are inherently relational and observed as evolving graphs, where entities, relations, and their properties change over time. Prior graph-related world models use graph structures to organize internal states or support task-specific reasoning, rather than treating an evolving graph itself as the modeled world. We instead study graph world modeling (GWM), where graph evolution itself constitutes the world dynamics. We formulate graph world modeling over observed graph evolution, latent graph states, and heterogeneous graph-transition predictions. Based on this formulation, we construct GWM-Zero, a benchmark covering node-, edge-, and graph-level transitions over eight temporal graph datasets. We propose WorldGraph, which combines a state-aware graph transformer for multi-granularity structural and transition-conditioned evolution modeling with transition-aware GRPO using dynamic grouping and structure-aware verifiable rewards. Extensive experiments on GWM-Zero show that WorldGraph consistently outperforms representative graph representation, temporal graph learning, graph pretraining, and graph world-model baselines across all three transition granularities.
|
| 1395 |
Hidden Activations are not Enough I: Knowledge Matrices as Higher Representations
2609.34166
|
cs.LG
|
Marco Armenta |
We study the knowledge matrix of a trained feedforward network as a higher representation of its inputs. A network is a pair $(W,f)$, a thin representation $W$ of its quiver and an activation $f$; its function factorizes through the space of quiver representat...We study the knowledge matrix of a trained feedforward network as a higher representation of its inputs. A network is a pair $(W,f)$, a thin representation $W$ of its quiver and an activation $f$; its function factorizes through the space of quiver representations, each input $x$ inducing a representation, and the knowledge matrix $M(x)\in\mathbb{R}^{C\times(d+1)}$ is the contraction of that representation to one matrix whose rows sum exactly to the logits. At one trained network we ask what determines it, what it is invariant to, what it determines, and what its geometry measures. Under (LCS), a locally constant slope diagonal, as for ReLU, the matrix at a regular input is a function of the realized germ; its stabilizer among encodings regular there is exactly the germ stabilizer at inputs with no vanishing coordinate, neuron permutation a special case; and it recovers the germ, whereas hidden activations, gauge-covariant and germ-incomplete, are not enough. Under (LCS) it equals per-class gradient$\times$input plus an exact aggregate bias attribution, grounding it in attribution theory and computing it by $C$ vector-Jacobian products instead of probing. The fixed shape gives an alignment-free per-sample distance between ResNet-152, DenseNet-121 and GoogLeNet; the row-sum identity gives an exact visible/invisible displacement decomposition whose unit-free coherence $A=(d_\Psi/d_M)^2$ puts adversarial germ motion at median $A\le 0.23$, with an attack-family ordering concordant across six architectures (Kendall $W=0.921$; $0.97$ on the three networks at full scale). Two honest negatives: on AlexNet/CIFAR-10 penultimate features win 5 of 6 detectors and all 16 attacks, and a matrix-direction counterfactual fails 0/54.
|
| 1396 |
GPARA: Graph-Posterior-Aligned Refinement and Active Acquisition for Grounding Diffusion Priors
2609.34172
|
cs.LG
|
Wangqian Chen, Hao Wang, Yumeng Zhang, Jiajia Guo, Junting Chen |
Active grounding of a frozen diffusion prior requires jointly determining where new measurements should be taken and how they should be used to refine the current reconstruction. Posterior-ensemble-based methods can estimate acquisition utility from generated ...Active grounding of a frozen diffusion prior requires jointly determining where new measurements should be taken and how they should be used to refine the current reconstruction. Posterior-ensemble-based methods can estimate acquisition utility from generated samples, but require repeated ensemble generation as observations accumulate and capture posterior geometry only through empirical statistics. This paper proposes GPARA, which learns a context-dependent graph surrogate over diffusion prediction residuals, inducing an explicitly reusable posterior response operator that propagates measurement innovations to unobserved variables and evaluates candidate measurements through weighted posterior-risk reduction. Under the matched surrogate, we show that the same response operator also determines expected one-step acquisition benefit and yields an analytic ranking consistent with expected reconstruction improvement. A bounded learned residual calibrates the analytic utility to account for surrogate mismatch, while a small prior ensemble is generated once and reconditioned to update risk weights without repeated diffusion posterior sampling during acquisition. Experiments on two reconstruction tasks spanning physical field and computer vision show consistent improvements in refinement and active acquisition over the evaluated baselines. Ablations further support the complementary roles of step-wise graph refinement, adaptive risk weighting, and analytically anchored calibration.
|
| 1397 |
GradLev: Token-Parallel Test-Time Training Via Costate Prediction
2609.34174
|
cs.LGcs.AI
|
Bo Liu, Qiang Liu |
Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation grad...Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation. GradLev lever- ages this duality: a causal auxiliary network predicts costates across all tokens in parallel; associative scans compute the adapted weights and forward activations and propagate gradients backward; and the resulting gradient targets supervise the predictor via a consistency loss. Exact consistency guarantees exact recovery of the sequential online learner. At deployment, the auxiliary predictor is discarded, and the model updates natively via token-by-token forward and backward passes.
|
| 1398 |
EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates
2609.34185
|
cs.LG
|
Hong Zhang, Zhongjie Duan, Yingda Chen |
Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding o...Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding overhead. Exploiting this flexibility requires accurate rate selection and efficient weight reconstruction for inference. We present EntroPack, an entropy-coded weight compressor that supports arbitrary target bitrates without activation calibration or fine-tuning. It combines row-normalized $E_8$ lattice quantization with a conditional probability model of lattice coordinates. Sampled storage estimates select the quantization resolution without repeated full-stream encoding. The final coordinates are entropy-coded in independently decodable tiles, enabling fast, fused symbol decoding and numerical weight reconstruction on the GPU. EntroPack supports floating-point and integer weight containers, such as BF16, FP16, FP8, and INT8, with storage bitrate controlled independently of numerical precision. Online decoding adds latency that grows with weight count, making the method well suited to compute-intensive workloads such as diffusion denoising and Transformer prefill. Experiments demonstrate fast encoding and modest inference overhead in these settings. When compressing the linear-layer weights of the image generator Z-Image-Turbo, EntroPack achieves substantially lower weight and denoiser output errors than fixed-width formats at comparable storage rates, with modest denoising-step overhead. Targeting 4 bits per parameter, it achieves lower weight and denoiser output errors than NF4, including about 24% lower relative $L_2$ weight error, with less storage. Source code is available at https://github.com/modelscope/entropack.
|
| 1399 |
Cardinality-Stratified Interaction Decomposition for Interpretable Pairwise and Higher-Order Structure in Transactional Basket Data
2609.34191
|
cs.LG
|
Hidetoshi Kawase, Toshihiro Ota |
Transactional basket data can reveal associations among items, but observed co-occurrence conflates item-specific relations with basket-size structure and unmodeled higher-order dependence. We introduce Cardinality-Stratified Interaction Decomposition (CSID), ...Transactional basket data can reveal associations among items, but observed co-occurrence conflates item-specific relations with basket-size structure and unmodeled higher-order dependence. We introduce Cardinality-Stratified Interaction Decomposition (CSID), an interpretable framework that decomposes log-odds contrasts stratified by the number of remaining items into item-set-specific and cardinality-common components, without fitting a global joint distribution. CSID uses an information-weighted, gauge-constrained ridge projection to estimate pair and triple components and to diagnose higher-order contributions to pairwise structure. CSID is designed primarily for interpretable decomposition of association structure rather than for full-distribution prediction. In a simulation with zero pair effects, increasingly strong small-basket cardinality potentials drive ordinary Ising couplings spuriously negative, whereas CSID pair estimates remain centered near zero. Detection power rises with the magnitude of planted triple effects, and local deprojection reduces pair-coefficient RMSE from 0.244 to 0.073. Across three grocery datasets, high-information triple components are reproducible over time. In the matched cross-period partial-transfer evaluation, transferred CSID triple components show closer agreement with later-period stratified contrasts than the nodewise-symmetrized cardinality-aware higher-order pseudolikelihood comparator, with gains in weighted Lin's concordance correlation of 0.038--0.122. These results support CSID as an exploratory and interpretable decomposition framework for pairwise and higher-order association structure in transactional data.
|
| 1400 |
Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
2609.34198
|
cs.LG
|
Jiapeng Li |
Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service ag...Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare upgrades that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and execution outcome, while a task-solvability prediction reverses sign across domains. A separately fixed post-submission OpenHands follow-up on the same 250 issues (eight configurations, 1,981 three-judge cells) reproduces the capability/false-acceptance association (mean Spearman +0.944, exact p=0.000099) and decreasing Youden contrast (mean -0.937, p=0.000397); this is observational, not a new-task replication. Transporting old-version calibration raises SWE-bench comparison error from 3.8 to 19.5 points, with 24.6% undefined bootstrap ratios. A paired audit saves only 5% in interval width at 80 labeled tasks. A randomized self-report test is negative (three adjusted p-values=1.0). These results favor paired audits of current outputs over judge-only release decisions or transported calibration; independent human patch review remains pending.
|
| 1401 |
Learning to Optimize through Solver-Grounded Self-Play
2609.34205
|
cs.LG
|
Xia Jiang, Yaoxin Wu, Chenyu Zhou, Mengzhu Xu, Wim P. M. Nuijten |
Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotat...Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotated or teacher-generated datasets. This dependence introduces a Generalization Ceiling, where models overfit to narrow data distributions, and Capability Anchoring, where models' reasoning is bounded by annotator proficiency and teacher model capability. In response, we propose OPT-Zero, the first fully self-play training framework for optimization modeling that requires zero external training data. OPT-Zero employs a single LLM in a dual-role closed loop: a Proposer that synthesizes increasingly challenging optimization problems alongside their mathematical formulations and solving code, and a Solver that attempts to resolve the problems given only natural-language problem descriptions. Grounded in execution feedback from external optimization solvers, we alternately train both roles using reinforcement learning. This process fosters an auto-curriculum in which the Proposer and Solver co-evolve: generating harder valid problems by the Proposer seamlessly enhances the structural reasoning ability of the Solver. Extensive results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.
|
| 1402 |
SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals
2609.34228
|
cs.LG
|
Jingyun Jia, Antoine Remond-Tiedrez, Aaron Alvarez, Joshua Shunk, Rich Caruana |
Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introd...Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addresses both problems by injecting controlled data-quality problems and feature effects into public tabular datasets: the injected pattern determines the answer, so reference answers are computed automatically and memorized knowledge of the original table is insufficient, while the table keeps its background structure. The injected patterns are modeled on phenomena reported in real data analyses. The benchmark defines 17 question templates in two families: data-quality questions and feature-contribution questions. We evaluate six state-of-the-art LLMs that analyze the data using a Python coding tool, on data-science and business phrasings of 70 validated dataset-template combinations, yielding 1680 graded responses in total. The models detect data-quality problems reliably (83.8% accuracy) but recover feature contributions poorly (41.9%). Finding how features shape the target requires searching over both candidate variables and analytical procedures. To address this issue, we propose the Empirical Layer, a set of precomputed statistical artifacts comprising summaries, fitted feature and interaction effects, and dataset descriptions, which exposes candidate patterns for direct inspection. Access to these artifacts raises feature-contribution accuracy from 41.9% to 68.0%.
|
| 1403 |
CasEm: A Cascade Architecture for Long-Horizon Neural Emulation
2609.34246
|
cs.LG
|
Zhaoyi Li, Jingtao Ding, Shihua Li |
Autoregressive neural emulators can drift or diverge over long rollouts despite accurate short-term predictions. We introduce Cascaded Emulation (CasEm), a one-way rollout architecture that augments an existing full-state backbone with an independently evolvin...Autoregressive neural emulators can drift or diverge over long rollouts despite accurate short-term predictions. We introduce Cascaded Emulation (CasEm), a one-way rollout architecture that augments an existing full-state backbone with an independently evolving model of physically specified aggregates. Its forecasts guide corrections to full-state predictions, without feedback from the backbone to the aggregate model. Effective guidance requires aggregates that cover substantial backbone error, remain accurately predictable, and support useful full-state corrections. We derive a finite-horizon error bound that clarifies these three factors and use empirical diagnostics to guide subsystem selection. Across four ODE/PDE benchmarks, CasEm reduces long-horizon rollout errors across diverse backbones and suppresses the trend toward error divergence in both diffusion tasks using Fourier neural operator backbones. In global climate emulation, CasEm with a regional total-water subsystem reduces 10-year full-state time-mean error by 66.6% and 46.3% for frozen ACE and Spherical DYffusion backbones, respectively, while adding less than 3% to inference time.
|
| 1404 |
Rotated Manifold Optimization for Low-Rank Adaptation
2609.34264
|
cs.LG
|
Yuhui Ding, Javier Zazo, James Hensman |
We propose a novel optimizer for low-rank adaptation (LoRA) that explicitly incorporates the gauge symmetry of low-rank factorization. Our optimizer extends recent matrix optimizers for full-parameter training to the manifold of fixed-rank matrices by interpre...We propose a novel optimizer for low-rank adaptation (LoRA) that explicitly incorporates the gauge symmetry of low-rank factorization. Our optimizer extends recent matrix optimizers for full-parameter training to the manifold of fixed-rank matrices by interpreting them as normalization under a rotated basis. We show how rotation and normalization can be integrated with the fixed-rank manifold efficiently. Our optimizer converges faster to lower held-out loss and achieves better or comparable downstream performance on both supervised finetuning and reinforcement learning tasks.
|
| 1405 |
ZeroCode: On-demand Error-Correcting Code Construction from the Zero Matrix via Reinforcement Learning
2609.34265
|
cs.LG
|
Ju-Hyeong Lee, Yongjune Kim, Sang-Hyo Kim, Dae-Young Yun, Hee-Youl Kwak |
Error-correcting codes (ECCs) are essential across diverse applications, from wireless communications and storage to quantum computing, yet each application imposes distinct design requirements on the parity-check matrix (PCM). To address these on-demand requi...Error-correcting codes (ECCs) are essential across diverse applications, from wireless communications and storage to quantum computing, yet each application imposes distinct design requirements on the parity-check matrix (PCM). To address these on-demand requirements in a unified framework, we propose ZeroCode, a reinforcement learning (RL)-based approach that constructs PCMs sequentially from the all-zero matrix. ZeroCode formulates construction as a discrete sequential decision-making problem and uses proximal policy optimization with action masking to select valid edges. ZeroCode achieves a gain of approximately 1 dB over the prior RL-based construction method at a bit error rate (BER) of $10^{-4}$ for the (32,16) code and outperforms existing genetic, differentiable, and classical code-design methods in our experiments. Beyond optimizing decoding performance, the masking mechanism allows on-demand structural constraints, such as a maximum degree, 4-cycle-free structure, and quasi-cyclic structure, to be flexibly incorporated. Moreover, a single policy rollout yields a library of PCMs with varying edge counts, offering trade-offs between decoding performance and complexity without retraining. Overall, ZeroCode addresses diverse code-design requirements within a unified framework, providing solutions with optimized decoding performance under given constraints.
|
| 1406 |
Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training
2609.34272
|
cs.LG
|
Junlin Chen, Daize Dong, Huanwei Di, Haolong Jia, Jiawei Wu |
BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. When we used FlashAttention-3 to pretrain a 450M-parameter transformer on 50B tokens, however, we ran into a pro...BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. When we used FlashAttention-3 to pretrain a 450M-parameter transformer on 50B tokens, however, we ran into a problem: training was healthy for 25B tokens, then the gradient norm grew a thousandfold and the loss ended 0.2 nats above FP32 attention, without a single NaN. Recomputing the attention backward of just two layers in FP32 removes almost all of the excess gradient. Part of the cause is known: a fused multiply-add in the forward softmax, so far treated as an extreme-input NaN case and never fixed in FlashAttention-3. Repairing it stops the blow-up, but the query gradient is still wrong by more than its own size, and training still drives attention logits to thousands of times their size under accurate gradients. The remaining error comes from a broken conservation law. The softmax score gradient sums to zero along every row, which makes the query gradient blind to where the keys sit as a group; rounding it to BF16 leaves a small nonzero sum that leaks the mean key into the gradient, and the leak grows exactly as late training makes keys large and attention sharp. We introduce GProj (gauge projection), which restores the zero sum after the cast with two rank-one corrections per row. It cuts the remaining median query/key gradient errors from 219%/13% to 0.34%/0.37%, on par with FP32 attention, for 4.7% more time per training step. In matched from-scratch runs it trains to the same loss as FP32 attention, while FlashAttention-3 and key smoothing both destabilize.
|
| 1407 |
Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training
2609.34279
|
cs.LG
|
Yuyang Deng, Yu Wang, Jiayun Wang |
Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training ...Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}volving \textbf{O}ptimization (DEO), which replaces challenger parameter updates with solver-guided task sampling. The KL-regularized challenger objective defines an exponential tilt of a fixed base task distribution. DEO uses this distribution as a sampling target: a frozen LLM generates and mutates tasks, the solver scores them, and an approximate Metropolis selection rule refines the training pool. Only the solver is trained. Theoretically, for an idealized variant that samples exactly from the tilted distribution, and under regularity, local gradient-dominance, and initialization conditions, we show that DEO learns distributionally robust reasoning ability. In experiments, DEO achieves reasoning performance competitive with R-Zero while using over $50\%$ less wall-clock training time, and improves reasoning accuracy over a no-walk ablation. Replacing the task generator with a frozen API-only LLM further improves the local solver, illustrating a capability enabled by removing challenger training.
|
| 1408 |
Agentic High-Dimensional Bayesian Optimization with Hypothesis- and Evidence-Guided Search
2609.34281
|
cs.LG
|
Zhixuan Gao, Ke Xue, Rongxi Tan, Ming Chen, Chao Qian |
High-dimensional Bayesian optimization (HDBO) seeks sample-efficient optimization when the number of variables is large relative to the evaluation budget. Recent LLM-based and agentic BO methods incorporate task knowledge and adapt search decisions during a ru...High-dimensional Bayesian optimization (HDBO) seeks sample-efficient optimization when the number of variables is large relative to the evaluation budget. Recent LLM-based and agentic BO methods incorporate task knowledge and adapt search decisions during a run, but have primarily been evaluated on low- and moderate-dimensional problems. We ask whether this paradigm can transfer to the higher-dimensional regime. Our experiments show that these methods do not remain reliable in the high-dimensional regime, where the challenge is not only where to evaluate, but also which modeling assumption and search geometry to use when the objective's useful structure is unknown. We therefore introduce HERA, a Hypothesis- and Evidence-guided Research Agent that uses task context, optimization feedback, and structural diagnostics to revise search hypotheses, select and configure HDBO strategies, and determine their execution length. PRISM, its numerical optimization engine, generates and evaluates candidates sequentially within each search block, updating numerical models after each observation. HERA remains competitive with strong numerical HDBO baselines and outperforms the evaluated LLM-based and agentic methods on four metadata-free synthetic functions. Across eight real-world tasks, HERA achieves the best mean final objective among all evaluated systems on most benchmarks. Further analyses show that structural diagnostics change strategy use, metadata effects vary across tasks, and adaptive search blocks reduce inference cost.
|
| 1409 |
Epistemic Learning from Imprecise Annotation
2609.34285
|
cs.LG
|
Kaizheng Wang, Siu Lun Chau |
Imprecise annotations may support several plausible labelling distributions, yet learning methods often resolve this ambiguity into a single predictive distribution. This can obscure what the annotation evidence leaves unresolved. We introduce epistemic learni...Imprecise annotations may support several plausible labelling distributions, yet learning methods often resolve this ambiguity into a single predictive distribution. This can obscure what the annotation evidence leaves unresolved. We introduce epistemic learning from credal supervision, a framework that uses convex sets of plausible labelling distributions, called credal sets, as supervision and learns sets of predictive distributions. We instantiate the framework with the pessimistic--optimistic credal classifier (POCC), which combines a shared backbone with two classification heads trained to minimise worst-case and best-case losses over the supervision sets. Their outputs define a predictive credal set whose spread provides an uncertainty score. We also show how credal labels can be obtained through a simple relaxation of existing probabilistic labels, reducing commitment to their precise probability assignments. This construction admits closed-form inner optimisation under cross-entropy loss, enabling efficient training. Assuming the supervision sets contain the true conditional label distributions, and other regularity assumptions, we establish a finite-sample generalisation bound for the averaged predictor with an explicit penalty for supervision imprecision. We evaluate POCC using human annotator disagreement and teacher predictions, alongside label smoothing as a controlled proxy for annotation imprecision. Across these settings, POCC achieves a favourable balance of predictive accuracy, calibration, and uncertainty-based selective classification versus competitive baselines.
|
| 1410 |
One Sequence, Many Decodings: CAGenMol-2 Recasts Drug Design as Masked Molecular Inference
2609.34301
|
cs.LG
|
Yanting Li, Enyan Dai, Lei Wang, Wen-Cai Ye, Li Liu |
Drug design couples property evaluation, conditional generation, structure-based design, and local optimization, yet machine learning systems typically address these capabilities with separate task-specific models. We introduce CAGenMol-2, a masked diffusion m...Drug design couples property evaluation, conditional generation, structure-based design, and local optimization, yet machine learning systems typically address these capabilities with separate task-specific models. We introduce CAGenMol-2, a masked diffusion molecular language model that represents molecules, continuous scalar properties, and 3D protein pockets within a single wrapped sequence. Within this pretrained interface, downstream operations are selected by which sequence regions are observed or masked at inference, allowing one checkpoint to perform property prediction, property- and pocket-conditioned generation, and partial-constraint design without task-specific architectures or backbone fine-tuning. We further propose Adaptive Fragment Optimization (AdaFO), a gradient-free mask-and-refill search that turns the masked decoder into an iterative local molecular optimizer. On CrossDocked2020, AdaFO increases Success Rate from 30.2\% to 70.8\%, the best reported under this protocol, while largely preserving drug-likeness and diversity. Finally, scaffold-preserving directional editing and CRBN/VHL case studies demonstrate its use in compound design workflows spanning local molecular editing, structure-based prioritization, and downstream simulation-based screening.
|
| 1411 |
Riemannian Difference-of-Convex Optimization for K-Means Clustering
2609.34310
|
cs.LG
|
Meng Xu, Bo Jiang, Hanfu Zhang, Ya-Feng Liu, Anthony Man-Cho So |
K-means is a widely adopted clustering approach in signal processing and machine learning. In this paper, we study K-means clustering through a cardinality-constrained formulation on a compact embedded submanifold. We replace the cardinality constraint with a ...K-means is a widely adopted clustering approach in signal processing and machine learning. In this paper, we study K-means clustering through a cardinality-constrained formulation on a compact embedded submanifold. We replace the cardinality constraint with a difference-of-convex (DC) penalty and establish a global error bound to prove that the penalized and constrained formulations share the same global minimizers whenever the penalty parameter exceeds a finite threshold. To solve the resulting nonsmooth Riemannian DC problem, we reformulate it as a minimax problem and propose RADA-DC, a Riemannian alternating descent ascent method combining dual regularization with DC linearization. Under standard assumptions and suitable parameter choices, RADA-DC finds an $\epsilon$-Riemannian critical point within $O(\epsilon^{-3})$ iterations. We conduct experiments on synthetic and real-world datasets to demonstrate that the proposed method outperforms the tested baselines, including K-means++, in solution quality at competitive computational cost when the number of clusters is large.
|
| 1412 |
One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents
2609.34321
|
cs.LG
|
Ziqiang Wang, Li Gu, Zhixiang Chi, Linlian Jiang, Zihuan Jiang |
GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs),...GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.
|
| 1413 |
Routing Without Embeddings: Fast And Interpretable Routing With Regular Expressions
2609.34326
|
cs.LG
|
Yifan Lu, Qiyue Zhang, Haotian Shan, Hanjie Chen, Jiarong Xing |
Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Fi...Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders may already capture the query properties needed for routing. We therefore investigate which properties matter and whether they can be extracted directly from text without a neural encoder. We introduce REGEXROUTE, a pipeline that uses sparse autoencoders (SAEs) to discover interpretable regular-expression (regex) features. Using unlabeled text, an LLM turns descriptions of grouped SAE latents into regex extractors and refines them to match latent activation patterns. These extractors supply numerical features to a lightweight routing head, eliminating neural encoding at inference (Figure 1a). Across four benchmarks, one fixed set of 128 features achieves 76.43% average routing accuracy, comparable to 76.41% for the strongest neural text encoder baseline, with much smaller latency and strong robustness. These findings establish explicit, interpretable text features as a practical basis for designing and understanding LLM routers.
|
| 1414 |
The Composition Gap in Dataset Distillation
2609.34343
|
cs.LG
|
Guang Li, Takahiro Ogawa, Miki Haseyama |
Dataset distillation compresses a training set into a small synthetic set, usually evaluated one at a time. In federated and data-governance settings, several parties distill their own data and a user trains on their union. We ask whether the union of separate...Dataset distillation compresses a training set into a small synthetic set, usually evaluated one at a time. In federated and data-governance settings, several parties distill their own data and a user trains on their union. We ask whether the union of separately distilled sets reproduces training on the union of the real data composability and show that it can fail even when every source is distilled exactly and the total budget admits an exact joint distillate. Compressing a training trajectory into fewer steps transforms the source statistics nonlinearly, so averaging compressed sources differs from compressing their average. For quadratic objectives we derive the exact composition error for two-to-one step compression in terms of the source-Hessian variance and the linear terms of the losses, and on a smooth network at small step sizes this prediction captures the local endpoint discrepancy in magnitude and direction. For learned synthetic sets, however, the composed error decomposes exactly into this local discrepancy and an aggregate source residual. Under endpoint matching the residual exceeds the structural term by more than an order of magnitude, and under distribution matching the two terms partly cancel. Joint distillation also retains an accuracy advantage when both sets are distilled from the same dataset, where the local discrepancy is exactly zero. Training fidelity and downstream accuracy are therefore distinct requirements, neither established by evaluating each set on its own.
|
| 1415 |
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
2609.34344
|
cs.LG
|
Yuchen Cai, Ding Cao, Qixiang Yin, Xin Xu, Kai Yang |
Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (R...Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional effective manifold in activation space associated with RL-induced gains. We uncover two geometric properties. (1) Effective Manifold Capacity: the capacity needed to reproduce RL gains can be very small but is not infinitely compressible; at extremely low capacity, intervention dimensionality and input-dependent expressiveness become key constraints, and this requirement varies with injection depth. (2) Control Manifold Separation: effective control directions lie mainly in the low-variance complement of the activation principal subspace. Within a task and base model, the learned geometry stays largely consistent across training configurations, and across tasks geometric alignment correlates with capability transfer. Experiments on 5 LLMs and 6 verifiable-reward tasks support these findings. We then propose Alpha-Stabler, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement. Alpha-Stabler stabilizes training for 2,000 steps and consistently improves RL gains, offering practical insights for robust post-training. Code: https://github.com/caiyuchen-ustc/On_Policy_Vector_Training
|
| 1416 |
FAST-Brain: A Flow-Aligned Spatio-Temporal Surrogate Brain Model
2609.34354
|
cs.LG
|
Shucheng Liu, Changchun Shi, Kai Zhang, Hongtu Zhu |
Modeling resting-state functional magnetic resonance imaging (rs-fMRI) data is crucial for understanding brain-wide neural activity. However, traditional methods struggle to capture complex temporal dynamics over long horizons, to account for the brain's anato...Modeling resting-state functional magnetic resonance imaging (rs-fMRI) data is crucial for understanding brain-wide neural activity. However, traditional methods struggle to capture complex temporal dynamics over long horizons, to account for the brain's anatomical spatial structure, and to model high-dimensional ambient signals that lie on a low-dimensional intrinsic subspace. We propose FAST-Brain, a unified flow-aligned spatio-temporal surrogate brain model that addresses all three challenges. At its core is a flow-aligned generative framework that directly predicts the clean blood-oxygen-level-dependent (BOLD) signal, paired with a graph convolutional network that captures spatial structural constraints and a Transformer that models long-range temporal dependencies. Theoretically, we show that under a low-dimensional subspace assumption, the approximation error of our model scales with the intrinsic dimension rather than the ambient dimension, which justifies our direct modeling of the BOLD signal. Extensive experiments on synthetic and Human Connectome Project datasets demonstrate that FAST-Brain achieves state-of-the-art performance in recovering functional connectivity, effective connectivity, and the implicit low-dimensional signal subspace.
|
| 1417 |
Spexis: Speculative Lookahead Scheduling for LLM Inference
2609.34370
|
cs.LG
|
Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi, Sungkyun Kim, Jinho Lee |
Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with ...Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.
|
| 1418 |
Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication
2609.34373
|
cs.LG
|
Mihir Chauhan, Aniket Bera |
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate ...Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.
|
| 1419 |
ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
2609.34374
|
cs.LG
|
Eric Frankel, Banghua Zhu, Sewoong Oh, Lillian J. Ratliff |
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant altern...Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF, DPO, and GRPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines in a series of experiments on an increasing scale. Our code is available at https://github.com/SewoongLab/abc-align .
|
| 1420 |
LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models
2609.34375
|
cs.LG
|
Luzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao |
Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controlla...Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent $\mathbf{z}$ and learned-query residual-context embeddings $\mathbf{u}$. Only $\mathbf{z}$ is propagated by the dynamics model and used for planning, while $\mathbf{u}$ captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.
|
| 1421 |
GeoCFM: Positive-Only Conditional Flow Matching for Mineral Occurrence Sampling
2609.34398
|
cs.LG
|
Moshe Eliasof, Eldad Haber |
Critical mineral discovery is a positive-only problem: deposits are observed as sparse locations, while unlabeled regions are not reliable negatives, and similar geophysical signatures can arise from different subsurface states. We therefore model mineral targ...Critical mineral discovery is a positive-only problem: deposits are observed as sparse locations, while unlabeled regions are not reliable negatives, and similar geophysical signatures can arise from different subsurface states. We therefore model mineral targeting as learning a conditional spatial distribution over occurrence locations, $\pi(p\mid d)$, given geo-images $d$, rather than predicting a deterministic per-pixel score map. We introduce GeoCFM, a conditional flow-matching model that generates mineral occurrence point sets conditioned on multi-channel geo-images; GeoCFM learns a point-wise transport field in $\mathbb{R}^2$, using UNet features with point-conditioned velocity prediction to bridge dense rasters and sparse supervision without pseudo-negatives. On a synthetic magnetics--geochemistry benchmark with latent activation and on USGS Earth MRI data with a spatially disjoint tile split, GeoCFM improves geometric agreement with observed occurrences over score-map and non-conditional baselines, while representing epistemic uncertainty through conditional sampling.
|
| 1422 |
Unlocking Few-Step Diffusion for Faithful Previews
2609.34406
|
cs.LG
|
Jing Jia, Sifan Liu, Guanyang Wang |
Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initia...Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.
|
| 1423 |
MASCIT: A Mask-Aware State Space Classifier for Naturally Irregular Time Series
2609.34409
|
cs.LG
|
Yoo-Min Jung, Hyeon-Gi Kim, Jonghun Park |
Naturally irregular time series combine asynchronous observations, missing values, unequal lengths, and nonuniform sampling, while dense adapters can discard temporal structure. We propose a mask-aware state space classifier for irregular time series (MASCIT),...Naturally irregular time series combine asynchronous observations, missing values, unequal lengths, and nonuniform sampling, while dense adapters can discard temporal structure. We propose a mask-aware state space classifier for irregular time series (MASCIT), which supplies observation masks to the encoder and excludes invalid steps from gated temporal aggregation. Across 34 irregular time series datasets, MASCIT yielded the strongest aggregate point estimate and was the only evaluated neural model with three-seed results on every dataset. MASCIT retained the lowest point rank across six overlapping irregularity indicators, while factorial ablations favored partial over full selectivity. These results support selective state space models as effective, executable backbones for naturally irregular time series classification.
|
| 1424 |
PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety
2609.34415
|
cs.LG
|
Ding Jia, Wei Liu, Xianglong Du, Yingjie Li, Yingqing Yang |
The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, ...The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, a bilingual safety benchmark with 155,780 states labeled through multi-model adjudication. Evaluating updated context before the next LLM inference, the trained guard achieves 91.46% unsafe-class F1 and 90.63% exact-boundary detection under complete source holdout. In AgentDojo, it reduces non-DoS targeted attack success from 20.82% to 0.40%.
|
| 1425 |
MegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid Parallelism
2609.34420
|
cs.LG
|
Tong Qiao, Ao Zhou, Yingjie Qi, Chunming Hu, Jianlei Yang |
Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the at...Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated topology-aware bias matrix jointly incur significant per-layer memory overhead, and heavy graph embedding layers result in severe workload imbalances. These characteristics are unique to GT training and are not addressed by parallelism techniques designed for either conventional GNNs or Transformers, making a dedicated solution necessary. This paper introduces MegaGraph, the first automated hybrid parallelism framework designed for efficient GT training. MegaGraph designs three specialized strategies, namely graph-aware context parallelism, heterogeneous pipeline parallelism, and hybrid data parallelism, to support efficient training on large-scale graphs. However, coordinating these three parallelism strategies yields an exponentially large configuration space. To address this complexity, an automatic search engine leverages precise cost models via a Profile - Model - Search workflow to identify the optimal parallelism configuration. Evaluations demonstrate that MegaGraph enables training on large-scale graphs where state-of-the-art baselines fail due to out-of-memory (OOM) errors. The framework reduces per-device peak memory by up to 77.8\% and achieves up to 4.51$\times$ training speedup while maintaining model accuracy.
|
| 1426 |
On the Relation Between Interval Regret and Dynamic Regret
2609.34423
|
cs.LG
|
Yi-Han Wang, Peng Zhao, Zhi-Hua Zhou |
Non-stationary online learning has attracted much attention in recent years, as static regret is insufficient to guide algorithm design in changing environments. To address this limitation, interval regret and dynamic regret have been introduced as two represe...Non-stationary online learning has attracted much attention in recent years, as static regret is insufficient to guide algorithm design in changing environments. To address this limitation, interval regret and dynamic regret have been introduced as two representative performance metrics that strengthen static regret in complementary directions. Interval regret requires an online algorithm to achieve competitive static regret over every local time interval, whereas dynamic regret evaluates performance against an arbitrary sequence of time-varying comparators. Despite their importance, the relation between these metrics has long remained unclear. Prior work has often regarded interval regret as the stronger notion, based on the intuition that local guarantees should naturally induce global guarantees. Consequently, it is widely conjectured that an algorithm with optimal interval regret should automatically attain optimal dynamic regret. In this paper, we first establish a negative result that refutes this intuition of a metric-level implication. Specifically, for both convex and curved functions (including exp-concave and strongly convex functions), we show that there exist instances in which an algorithm with optimal interval regret nevertheless fails to achieve optimal dynamic regret. We then show how to leverage local adaptivity to obtain optimal dynamic regret. In particular, optimal dynamic regret can be attained by invoking an interval regret minimization process over an enlarged Euclidean ball containing the original convex feasible domain and using a suitable domain-converted surrogate loss. This reduction applies to both convex and curved functions. As a byproduct, we obtain the first proper and efficient algorithm with optimal dynamic regret for exp-concave functions, improving prior results while significantly simplifying the analysis.
|
| 1427 |
Q-learning Penalized Transformer for Safe Offline Reinforcement Learning
2609.34426
|
cs.LG
|
Shengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou, Anning Hu |
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and co...This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizing rewards, and adhering to the behavior regularization imposed by the offline dataset. To tackle this trilogy challenge, we propose Q-learning Penalized Transformer policy (QPT), a \emph{training--inference consistent} framework that bridges conditional sequence modeling with constraint-aware value estimation. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost, retaining strong behavior regularization. To inject explicit safety semantics during learning, we augment sequence-model training with a Q-shaped penalty using learned reward and cost Q-functions to favor high return under low constraint violation. At inference, the same Q-functions enforce the cost threshold and choose the highest-reward feasible action, closing the loop between training and deployment. We provide a principled analysis under stylized near-deterministic CMDPs, characterizing how Q-penalized conditional generation improve safety and performance. Empirically, QPT consistently outperforms strong safe offline RL baselines across 38 tasks on the DSRL benchmark, and exhibits robust zero-shot adaptation to different constraint thresholds.
|
| 1428 |
Harmonizing Spectral Evolution in Conditional Flow Matching for TTS
2609.34431
|
cs.LG
|
Isha Pandey Varad Deshpande Abhijat Bharadwaj Ganesh Ramakrishnan |
Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to th...Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our method dynamically modulates mel-spectrogram sub-bands during ODE integration, synchronizing spectral development by penalizing aggressive low-frequency growth and boosting lagging high-frequency details. Validated across diverse architectures (Matcha-TTS, F5-TTS, IndicF5), our approach reduces the required Number of Function Evaluations (NFE) from 32 to 26 and improves Frechet Audio Distance (FAD) by up to 61%, all without compromising mean opinion scores, speaker similarity, and speech intelligibility.
|
| 1429 |
Admissible Diffusion for Multimodal Interventional Trajectories
2609.34433
|
cs.LG
|
Xing Han, Shravan Chaudhari, Jiarui Shao, Paul Pu Liang, Suchi Saria |
Generating a plausible clinical trajectory does not establish what would happen under a different treatment. We present ADMIT, a framework combining irregular multimodal representations, treatment-conditioned latent diffusion and explicit constraints on genera...Generating a plausible clinical trajectory does not establish what would happen under a different treatment. We present ADMIT, a framework combining irregular multimodal representations, treatment-conditioned latent diffusion and explicit constraints on generated states or actions. We formulate its interventional target through sequential g-computation and distinguish causal assumptions from constraint satisfaction. Its admissibility mechanism translates physiological prior knowledge into explicit constraints on generated states and proposed actions. Treatment-exposure dynamics condition latent transitions, while state projection or action gating applies the constraints during rollout so that they influence subsequent trajectory generation. In our preliminary experiments, multimodal inputs improved supervised hidden-state recovery and reduced treatment-contrast error. In a simulated dosing-schedule experiment with leak-free history encoding, ADMIT predicted most of the tumor-volume change caused by redistributing a fixed total dose. An exposure input improved these predictions around a temporary dose reduction whether or not the assumed clearance rate was correct, but reduced the predicted size of a dose effect, and a deterministic recurrent baseline matched ADMIT's average predictions. Exposure projection reduced constraint violations, although enforcement remained incomplete. Semi-synthetic experiments using eICU context illustrated treatment-response generation under fixed and adaptive policies. Observational examples further characterize model treatment sensitivity. ADMIT provides a framework for testing whether complementary observations and physiological restrictions improve intervention trajectories, with representation recovery, effect accuracy and rule enforcement assessed separately.
|
| 1430 |
Making LLMs Truly Forget: Deep Unlearning by Searching, Selecting, and Severing Knowledge Paths
2609.34442
|
cs.LG
|
Jialu Wang, Peizhi Niu, Haoteng Yin, Hans Hao-Hsun Hsu, Pan Li |
While an unlearned language model may no longer recall a fact directly, the fact often remains recoverable through multi-hop reasoning over related knowledge. Most existing unlearning techniques overlook this vulnerability, targeting facts in isolation while l...While an unlearned language model may no longer recall a fact directly, the fact often remains recoverable through multi-hop reasoning over related knowledge. Most existing unlearning techniques overlook this vulnerability, targeting facts in isolation while leaving their supporting knowledge intact. To achieve true forgetting, we propose a general deep unlearning framework compatible with existing unlearning algorithms. Our approach adaptively explores both explicit responses and latent internal representations to discover valid reasoning paths, compiles them into a confidence-aware supporting subgraph, and we apply a graph minimum cut to sever all recovery paths while preserving unrelated knowledge. To rigorously evaluate deep unlearning, we introduce a model-specific pipeline that extracts and completes knowledge graphs from raw text, filtering them by calibrated model confidence to reflect what the model genuinely retains. Comprehensive experiments demonstrate that selectively unlearning supporting knowledge yields substantially deeper forgetting than superficial methods while preserving model utility, highlighting that genuine unlearning requires breaking the relational structures that enable factual reconstruction.
|
| 1431 |
Livin' on a Prior: Likelihood Score Approximation for Inverse Problems
2609.34446
|
cs.LG
|
Rostislav Makarov, Tal Peer, Danilo de Oliveira, Timo Gerkmann |
Generative models have found great success as data-driven methods of solving inverse problems. Two popular approaches work either by combining a pretrained generative prior with a known degradation model, or by training a conditional generative model directly ...Generative models have found great success as data-driven methods of solving inverse problems. Two popular approaches work either by combining a pretrained generative prior with a known degradation model, or by training a conditional generative model directly from paired data. We target a setting that spans both regimes: unknown degradations can be learned from few paired examples, while known degradations can be learned from self-generated samples. We introduce Likelihood Score Approximation (LSA), a generative framework that keeps a pretrained unconditional model fixed and learns an observation-conditioned model that approximates the likelihood score from paired samples. Within a conditional stochastic-interpolant framework, LSA can be trained in either score or velocity coordinates, independently of the unconditional model's native parameterization, and supports both deterministic and stochastic sampling. We further show empirically that the prior model can be swapped post-training while keeping the same LSA model. Across speech and image inverse problems, LSA operates effectively even at roughly 0.01% of the full training dataset. On the ImageNet-256 benchmark it achieves competitive or better restoration quality than strong posterior-sampling baselines while requiring up to several orders of magnitude fewer network evaluations.
|
| 1432 |
SPACE-LoRA: Allocating Activation-Subspace Protection for Continual Learning
2609.34453
|
cs.LG
|
Seunghyun Yoo, Kiseok Kim, Hyeontae Joo, Junyeop Bang, Hwangnam Kim |
This study addresses the catastrophic forgetting problem that occurs when sequentially learning successive tasks using Low-Rank Adaptation (LoRA) from a lifelong learning perspective. While existing approaches have primarily constrained parameter updates or le...This study addresses the catastrophic forgetting problem that occurs when sequentially learning successive tasks using Low-Rank Adaptation (LoRA) from a lifelong learning perspective. While existing approaches have primarily constrained parameter updates or learning subspaces to reduce interference with past knowledge, they have not fully considered additive interference. This occurs when a newly added residual adapter on top of a fixed past model generates non-zero responses along input directions important for old tasks, thereby altering previous predictions. To this end, we propose Subspace Protection with Allocated Capacity for Efficient Continual Adaptation (SPACE-LoRA). SPACE-LoRA directly suppresses the responses of the new residual branch along input activation directions that are important for old tasks and adaptively determines the protection coverage for each module based on past-task sensitivity estimated via a common Fisher sensitivity-based coverage target. Under a fixed LoRA rank, this approach adaptively adjusts module-specific protection coverage while suppressing interference along input directions sensitive to old tasks. We assess the effectiveness of activation-subspace protection in mitigating catastrophic forgetting and examine the role of sensitivity-guided protection in continual learning across diverse tasks. Code is available at https://anonymous.4open.science/r/SPACE-LoRA-7864.
|
| 1433 |
ZonoGPT: Towards An Abstract Domain for Verifying Large GPT Models
2609.34457
|
cs.LG
|
Hai Duong, Thanh Le, ThanhVu Nguyen |
Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, such as robustness, safety, and fairness, neural network verification techniques prove required properties and pro...Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, such as robustness, safety, and fairness, neural network verification techniques prove required properties and provide auditable guarantees before deployment. However, prior work remains limited to small or restricted Transformers, and maintaining precision across deep models remains challenging. In this work, we introduce ZonoGPT, an abstract domain for verifying large transformers that maintains a space complexity independent of network depth. ZonoGPT uses a structured zonotope and a generator reduction mechanism to efficiently preserve correlations. To maintain precision, it introduces block-specific fused transformations for Attention and LayerNorm that retain feature relations, along with an affine transform for GELU that preserves generator relations. These mechanisms enable \tool{} to be the first approach to verify standard architectures, scaling to official HuggingFace models up to GPT-2 Medium (24 blocks, 300M+ parameters) and successfully verifying 1,339 instances across text and vision tasks.
|
| 1434 |
PhysioTRACE: Provenance-Aware Stress Tests for Physiological Foundation Models
2609.34466
|
cs.LG
|
Ayana Mussabayeva, Anuar Aimoldin, Olivier Oullier, Xue Liu, Kun Zhang |
Physiological foundation models encode how a signal was recorded alongside the physiology it reflects. When recording conditions are associated with diagnosis, this acquisition provenance can become a shortcut, yet the usual evidence, shifted transfer and prov...Physiological foundation models encode how a signal was recorded alongside the physiology it reflects. When recording conditions are associated with diagnosis, this acquisition provenance can become a shortcut, yet the usual evidence, shifted transfer and provenance decodability, does not show whether a predictor uses it. We introduce PhysioTRACE, a four-axis behavioral audit for frozen encoders that separates what a probe can decode from what a fixed task head relies on. Recover scores how decodable provenance is; Stress reverses only the provenance-target association on the same held-out records; Intervene removes a train-localized provenance component; and Verify certifies that removal only if it beats matched random projections within a declared utility margin. Each audit thus ends in one of three verdicts: no reliance, or reliance with the remedy certified or refused. Across EEG and ECG, five training objectives, and five frozen foundation models, the relation between Recover's calibrated score and out-of-distribution utility changes sign between datasets, so neither can stand in for a reliance test. On paired EEG views where the shortcut is known by construction, the audit detects it (the exposed head loses about 0.2 AUROC when the association is reversed, while a control head is unaffected) and certifies removal of a rank-two component that restores control-level behavior without measurable utility loss, for both encoder objectives tested. On real ECG device metadata it returns all three verdicts: it certifies a remedy that removes 91% of one model's excess vulnerability, finds no reliance where device and diagnosis are barely associated, and refuses the remedy for a second model whose localized direction also carries task signal. Robustness to how inputs were recorded therefore needs a behavioral test, and PhysioTRACE provides one that can pass, fail, or refuse a remedy.
|
| 1435 |
Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning
2609.34467
|
cs.LG
|
Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang |
Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} ...Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, which weakens action grounding and hurts generalization and fine-tuning efficiency. In this work, we present Alignment-Guided Flow Transformer (AGFT), a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation. While prior research has predominantly emphasized bi-modal vision--language alignment, we systematically formalize and study tri-modal alignment in VLA models, and provide both ablations and analysis to isolate its role in improving adaptation and robustness. To further accelerate deployment, we adopt a flow-matching objective, enabling substantially fewer inference steps than diffusion-based policies while maintaining accuracy. Theoretically, we establish a quantitative connection between the tri-modal alignment gap and the optimization tightness of flow matching; empirically, experiments on the extensive benchmark show that AGFT achieves superior success rates and lower inference latency compared to SOTA baselines, underscoring tri-modal alignment as a key ingredient for scaling robust VLA manipulation.
|
| 1436 |
Deep kernel hedging
2609.34474
|
cs.LG
|
Jean-Loup Dupret, Donatien Hainaut, Edouard Motte |
We introduce a deep kernel hedging framework that combines the flexibility of deep learning with the structural inductive bias of kernel methods. The hedging functional is restricted to a reproducing kernel Hilbert space whose kernel is parameterized through a...We introduce a deep kernel hedging framework that combines the flexibility of deep learning with the structural inductive bias of kernel methods. The hedging functional is restricted to a reproducing kernel Hilbert space whose kernel is parameterized through a neural network embedding of the input features. The framework minimizes a regularized empirical risk under convex loss functions and can accommodate path-dependent information through truncated time-augmented signature features. We derive a generalized representer theorem for the joint hedging problem, reducing the empirical optimization to a finite-dimensional problem. To further reduce the computational cost associated with large kernel matrices, we develop a scalable random Fourier feature approximation and establish convergence guarantees. The random Fourier parameters are sampled once and remain fixed throughout training, while the deep kernel adapts to market data through the learned neural representation. We evaluate the performance of the proposed deep kernel approach on both synthetic and real data and compare it with standard kernel methods and classical deep hedging architectures. Numerical results indicate competitive and robust hedging performance, particularly in low-data regimes, which highlights the benefits of combining expressive neural representations with the inductive bias of kernel methods.
|
| 1437 |
Causal Routing for Unlearning
2609.34475
|
cs.LG
|
Bardh Prenkaj, Andrea D'Angelo, Davide Mottin, Federico Fontana, Davide Gabrielli |
LLMs cannot forget the way we delete a file. Strangely, we are asked to remove something that was never put anywhere in particular. What the model took from a piece of text is now smeared across billions of weights. Existing methods rewrite all of them to chan...LLMs cannot forget the way we delete a file. Strangely, we are asked to remove something that was never put anywhere in particular. What the model took from a piece of text is now smeared across billions of weights. Existing methods rewrite all of them to change one thing, and none of them say which part produced that change. To address this, we introduce Causal Routing for Unlearning (CRU) by asking where the concept is expressed in the model and suppressing only that part. One untrained forward pass over the forget set ranks neurons by how their activations vary. Then, small routing modules on those neurons gate and suppress only the concepts that need to be forgotten. In CRU, the base model is frozen, and any change in behavior is caused only by the gated neurons; hence, why the routing is causal. Due to our parameter efficiency (only ~0.01% as many parameters as the base model), unlearning a concept costs 14 GiB, whereas the baselines require 71 GiB. On TOFU, CRU is indistinguishable from the retained model (p > 0.05, KS test) and is never Pareto-dominated, whereas every compared baseline matches its forgetting on the larger-forget batches only by collapsing utility. On RWKU, it achieves an adversarial-probe recall of 0.052, compared to 0.250 for the strongest baseline, meaning the knowledge is gone, not merely harder to reach. Thus, deciding on the intervention at query time, rather than fixing it beforehand, is the axis along which we argue that unlearning should proceed.
|
| 1438 |
Learn Here, Move Less Elsewhere: Input-Conditioned Plasticity from Retained-Domain Activation Atlases
2609.34478
|
cs.LG
|
Jiangtao Lin, Bangyang Wei, Yihang Ding, Siyi Liu, Yuhan Dong |
Task-specific fine-tuning can rewrite a language model's answers beyond the training task, complicating updates that must preserve existing behavior. We introduce ATLAS, which turns retained-domain representations into an input-dependent rule for task adaptati...Task-specific fine-tuning can rewrite a language model's answers beyond the training task, complicating updates that must preserve existing behavior. We introduce ATLAS, which turns retained-domain representations into an input-dependent rule for task adaptation. An activation atlas supplies local reference centers and directional filters to a shared low-rank residual. Target supervision learns the residual, while retained geometry shapes its action throughout training and inference. On Qwen3-8B, ATLAS achieves lower mean retained-output Kullback-Leibler (KL) divergence than all seven published baselines at shared coding-performance requirements, with consistent advantages across multiple training seeds. Structural comparisons identify the contributions of retained reference states and directional conditioning, and answer-level analyses show fewer rewritten mathematical answers and more stable commonsense choices. Experiments spanning five backbones and two retained domains further demonstrate coding gains with reduced retained-output movement. With compact storage and modest decoding overhead, ATLAS provides a practical mechanism for acquiring specialized skills while maintaining continuity in existing responses.
|
| 1439 |
FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL
2609.34488
|
cs.LG
|
Xun Wang, Ruishuo Chen, Yu Chen, Zhuoran Li, Longbo Huang |
Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation,...Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to naturally support elastic inference across recurrent depths. Surprisingly, we find that pretrained looped policies exhibit severe recurrent-depth specialization: reliable decisions are concentrated near the full trained depth, tying deployment computation to this depth even when less computation may suffice. Achieving depth elasticity, i.e., reliable decisions across recurrent depths with adaptive computation at deployment, therefore remains a key challenge. To address this, we propose FlexLoop, a novel post-training framework that converts pretrained fixed-depth looped policies into depth-elastic policies. FlexLoop keeps training on the original RL objective to preserve full-depth capability while performing adjacent-depth policy distillation to progressively transfer decision quality from deeper to shallower recurrent steps. The resulting policy supports reliable inference across recurrent depths and enables state-wise adaptive inference through recurrent-depth consistency. Experiments on $30$ online and offline long-horizon goal-conditioned environments show that FlexLoop preserves full-depth performance while making shallower depths effective. Keeping competitive performance, FlexLoop reduces average recurrent depth by up to $\bf{43\%}$ and achieves up to $\bf{1.34\times}$ wall-clock speedup in a stress test.
|
| 1440 |
When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation
2609.34489
|
cs.LG
|
Minjae Park, Taesun Yeom, Jaeho Lee |
Data pruning reduces the training cost of knowledge distillation (KD). However, the preferred teacher capacity changes with the data budget: smaller teachers can outperform larger ones when limited training data are available. Understanding what drives this sh...Data pruning reduces the training cost of knowledge distillation (KD). However, the preferred teacher capacity changes with the data budget: smaller teachers can outperform larger ones when limited training data are available. Understanding what drives this shift is important not only for teacher choice but also for identifying which samples are useful for distillation. We analyze teacher supervision by decomposing it into relational ordering---the ranking of classes---and score geometry---the magnitudes and margins of class probabilities---and show that the small-teacher advantage in the low-data regime arises not only from score geometry but also from relational ordering. Beyond understanding teacher capacity, our analysis reveals two properties of effective subsets: samples should match the difficulty appropriate for the available budget, and their relational signals should be diverse rather than redundant. Based on these findings, we propose DVA (Difficulty- and Volume-Aware data selection for KD), a training-dynamics-free method, which uses a small teacher as a proxy for budget-aware difficulty filtering and class-conditional relational volume maximization. Despite requiring no training dynamics statistics, our method remains competitive with training-dynamics-based methods while consistently outperforming training-dynamics-free baselines.
|
| 1441 |
M3OS: A Monte Carlo Graph Search-Orchestrated Multi-Agent LLM System for Evidence-Traced Molecular Optimization
2609.34491
|
cs.LG
|
Junjie Wang, Yaowei Jin, Ruohui Tang, Guonan Cui, Haojie Wang |
Small-molecule optimization integrates medicinal-chemistry reasoning and computational evidence through iterative, multi-objective decisions. When large language models (LLMs) reason over optimization histories stored primarily in conversational context, they ...Small-molecule optimization integrates medicinal-chemistry reasoning and computational evidence through iterative, multi-objective decisions. When large language models (LLMs) reason over optimization histories stored primarily in conversational context, they must recover candidate identities, prior evaluations, and task constraints to guide subsequent decisions. We present M3OS, a multi-agent LLM system that decouples molecular-design reasoning from optimization-state management through Monte Carlo graph search. A persistent graph links evaluated candidates, parent-child transformations and evaluation evidence, while rewards and visit statistics guide LLM-assisted parent selection. Two branches combine tool-driven candidate generation with knowledge- and case-guided medicinal-chemistry editing. An execution harness controls graph updates through structured output extraction, molecular validation and task-bound evaluation. Agents receive role-specific contexts, while the graph preserves optimization trajectories beyond their active contexts. Across three molecular optimization benchmarks, M3OS achieves higher success rates than baselines, supporting the integration of persistent search state, specialized agents and controlled execution for multi-constraint optimization.
|
| 1442 |
LLN: Learnable Lens Networks for Parameter-Efficient Long-Horizon Dynamical Prediction
2609.34493
|
cs.LG
|
Binbin Yong, Zhao Su, Lan Guo, Haoran Li, Jun Shen |
Explicit residual connections of the form (x+f(x)), often combined with normalization layers, have become a standard strategy for training very deep neural networks. However, residual addition primarily provides an algebraic shortcut for gradient propagation, ...Explicit residual connections of the form (x+f(x)), often combined with normalization layers, have become a standard strategy for training very deep neural networks. However, residual addition primarily provides an algebraic shortcut for gradient propagation, while leaving the evolution of feature geometry across layers largely unconstrained. We introduce Learnable Lens Networks (LLN), a physics-inspired architecture that replaces direct feature-space residual accumulation with learnable optical transport in an augmented position-angle phase space. Each layer alternates between free propagation, which provides an implicit transport path, and a learnable lens field that performs nonlinear trajectory transformation and focusing. Theoretically, we establish that LLN transport is globally invertible and volume-preserving for any differentiable lens field, with the implemented coordinate-wise Gaussian transport further satisfying symplecticity. Importantly, these structural constraints do not limit expressivity: with unrestricted embeddings and readouts, LLN retain universal approximation of continuous end-to-end maps. Experiments across diverse dynamical systems demonstrate that LLN improves long-horizon prediction while using substantially fewer parameters than same-depth comparators. Further analysis reveals stable depth-wise gradient transport and interpretable learned dynamics under the coupled propagation and refraction design.
|
| 1443 |
QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning
2609.34497
|
cs.LG
|
Yuehu Gong, Shutong Ding, Mokai Pan, Yimiao Zhou, Jiashu Hou |
Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is d...Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.
|
| 1444 |
Scalable GNN-based Knowledge Graph Representation Learning with Efficient Message Passing
2609.34499
|
cs.LG
|
Huu Tan Mai, Cuong Xuan Chu, Heiko Paulheim, Daria Stepanova |
Graph neural networks (GNNs) excel at representation learning on Knowledge Graphs (KGs), achieving stateof-the-art performance on tasks like link prediction or entity classification. However, their high computational complexity, inherent to their user-defined ...Graph neural networks (GNNs) excel at representation learning on Knowledge Graphs (KGs), achieving stateof-the-art performance on tasks like link prediction or entity classification. However, their high computational complexity, inherent to their user-defined message passing (MP) algorithm, still prohibits their widespread adoption, especially for large KGs. Current efforts to mitigate the scalability bottlenecks of GNNs on KGs, such as subgraph sampling, are often task- and model-specific, and do not reliably guarantee lossless (if applicable) runtime/space reductions. To address this, we extend Relational Sparse Matrix Multiplication (RSPMM), originally designed to losslessly lower the space complexity of composition-based MP with pointwise composition functions, to support more expressive functions (e.g., 2x2 block-diagonal matrix multiplication, Givens rotation, circular correlation). Our method delivers significant task-independent reductions in runtime and space for current GNNs on KGs and facilitates efficient re-implementations of GNNs that maintain near state-of-the-art performance on challenging KG tasks, for a fraction of computational costs.
|
| 1445 |
Distribution-Conditioned Task Routing for Class-Incremental Learning
2609.34503
|
cs.LG
|
Longhuan Xu, Zhipeng Zhou, Wei Ji, Chunyan Miao, Peilin Zhao |
Parameter-efficient adaptation enables continual learners to acquire task-specific knowledge through compact model updates while maintaining strong within-task performance. However, class-incremental inference requires each input to be classified among all cla...Parameter-efficient adaptation enables continual learners to acquire task-specific knowledge through compact model updates while maintaining strong within-task performance. However, class-incremental inference requires each input to be classified among all classes seen so far without access to its task identity. For learners equipped with task-specific parameter-efficient modules, this introduces a critical task-routing challenge beyond catastrophic forgetting. We study post-hoc task routing without retraining the learner or introducing a separately trained router. Such training-free inference-time calibration remains comparatively underexplored in parameter-efficient class-incremental learning. We identify three sources of routing error (feature-level, task-level, and class-level misalignment) and propose Feature Distribution Calibration (FDC). Its three components address these misalignments: Task Subspace Filtering (TSF) suppresses feature components outside each task's principal subspace, Residual Likelihood Calibration (RLC) evaluates the typicality of its subspace residual, and Prototype Affinity Calibration (PAC) measures compatibility with the task's class prototypes. Experiments demonstrate plug-and-play applicability to eight parameter-efficient class-incremental methods using a shared encoder. With one component configuration selected per method across all five benchmarks, FDC improves final accuracy in all 40 method-dataset pairs by 4.39 percentage points on average. Enabling all components improves 35 of the 40 pairs, with an average gain of 4.45 points. When applied to a simple baseline, FDC achieves strong overall performance.
|
| 1446 |
KiT: A Foundation Model for Financial Time-Series Forecasting using DiffusionTransformers
2609.34507
|
cs.LG
|
Boyu Zhang, Haorui Li |
Financial candlestick forecasting is fundamental to quantitative investment, yet it remains exceptionally challenging due to extremely low signal-to-noise ratios and vast heterogeneity across markets and instruments. Existing approaches have largely attempted ...Financial candlestick forecasting is fundamental to quantitative investment, yet it remains exceptionally challenging due to extremely low signal-to-noise ratios and vast heterogeneity across markets and instruments. Existing approaches have largely attempted to introduce deep learning to capture hidden temporal features, but most adopt an auto-regressive formulation, which leads to error accumulation during inference. Meanwhile, general-purpose time-series foundation models are not tailored to the unique structure of k-line data and yield unsatisfactory performance on downstream candlestick forecasting tasks. To tackle these problems, we introduce KiT, a K-line Diffusion Transformer foundation model, and reformulate future prediction as conditional path generation via flow matching: given a historical context window, the model generates an ensemble of plausible future OHLCV trajectories. We pre-train KiT at multiple parameter scales on billions of candlestick bars spanning multiple markets and timescales. Across three markets and seven resolutions, KiT attains a mean return RankIC of 0.057 and a mean volatility RankIC of 0.66, leading at every timescale and outperforming both task-specific financial forecasters and general time-series foundation models. Code will be available at: https://github.com/Luciferbobo/KiT.
|
| 1447 |
HALO: Enhancing Time Series Generation via Hyperspherical Latents and Masked AutoregRessive Modeling
2609.34511
|
cs.LG
|
Chunyi Hou, Xiangfei Qiu, Hanyin Cheng, Yutong Li, Bin Yang |
Most existing time series generators rely on a two-stage modeling paradigm: the first stage learns discrete latent representations of time series; the second stage performs autoregressive modeling on these discrete latents through next token prediction. Howeve...Most existing time series generators rely on a two-stage modeling paradigm: the first stage learns discrete latent representations of time series; the second stage performs autoregressive modeling on these discrete latents through next token prediction. However, this paradigm suffers from two stage-specific limitations: the first stage can lead to information loss when discretizing continuous time series, while the second stage is prone to error accumulation during autoregressive generation. To address these limitations, our core idea is to perform generative modeling in a continuous latent space with a more efficient autoregressive framework. We propose HALO, which enhances time series generation via Hyperspherical Latents and Masked Autoregressive modeling to achieve this goal by tackling two key bottlenecks: (1) variance and scale heterogeneity of continuous latent representations; (2) the difficulty of balancing generation efficiency with temporal correlation modeling. HALO first introduces a hyperspherical VAE that constrains continuous latents to a fixed-radius hyperspherical shell, effectively stabilizing the numerical fluctuations of continuous latent representations. Secondly, we develop a masked autoregressive model that balances parallel decoding and temporal correlation learning, substantially reducing the number of inference steps required for generation and improving generation stability. Our extensive experiments demonstrate that HALO achieves state-of-the-art generation performance while offering significantly improved inference efficiency over existing advanced baselines.
|
| 1448 |
Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers
2609.34538
|
cs.LG
|
Guanghao Li, Zihan Su, Hao Yu, Jinyang Jiang, Tao Ren |
Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and ver...Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation to prefix representations from the same recurrent depth. We find that queries and keys approach their final-depth representations earlier than values, and controlled prefix-channel interventions show that mature values substantially improve shallow draft predictions. Motivated by this asymmetry, we introduce Depth-Asynchronous Self-Speculation (DAS), which decouples the depth of draft computation from the depth of verified-prefix representations it reads. Its Mature-V primitive lets shallow queries retrieve full-depth prefix values without additional recurrent computation. We further develop DAS-Wave, which combines depth-asynchronous prefix reads with carried parallel refinement, progressive block growth, and an independent full-depth verifier. Across four recurrent-model checkpoints and mathematics and code workloads, DAS-Wave achieves 4.00--6.96$\times$ mean throughput speedup over paired full-depth autoregressive decoding in the same inference stack. These results identify prefix-information depth as an effective design axis for recurrent self-speculation.
|
| 1449 |
On Parameter Symmetries and Conservation Laws in Gradient Flow
2609.34549
|
cs.LG
|
Khang Nguyen, Guido Mont\'ufar |
Parameter space symmetries and conservation laws play an important role in understanding the loss landscapes and implicit biases of neural networks. Inspired by Noether's theorem in physics, prior works have sought to derive conservation laws under gradient fl...Parameter space symmetries and conservation laws play an important role in understanding the loss landscapes and implicit biases of neural networks. Inspired by Noether's theorem in physics, prior works have sought to derive conservation laws under gradient flow from parameter symmetries, but the scope and limitations of this connection remain unclear. We develop a unified geometric framework that clarifies the precise relationship between the two notions, including the conditions under which symmetries correspond to conservation laws. We introduce a notion of compositional identifiability and use it to establish a general inheritance principle for complete characterizations of symmetries and conservation laws in multilayer networks. We apply the framework to multi-head and grouped-query attention, polynomial neural networks, and square deep linear networks.
|
| 1450 |
Verifying Neural Networks with Reinforcement Learning
2609.34553
|
cs.LG
|
Hai Duong, Thanh Le, ThanhVu Nguyen |
Formal verification can play a key role in ensuring the reliability of Deep Neural Networks (DNNs) deployed in safety-critical systems. Modern DNN verifiers employ a branch-and-bound framework, which alternates between branching (splitting into smaller subprob...Formal verification can play a key role in ensuring the reliability of Deep Neural Networks (DNNs) deployed in safety-critical systems. Modern DNN verifiers employ a branch-and-bound framework, which alternates between branching (splitting into smaller subproblems) and bounding (pruning subproblems) to efficiently explore the verification space. However, existing branching heuristics make greedy decisions based on static scoring functions. They do not anticipate long-term efficiency or leverage the growing availability of verification data to improve performance. This work introduces RSB, a reinforcement learning framework that learns to refine baseline branching heuristics. It trains an actor-critic architecture to maximize cumulative future rewards rather than immediate scores. The actor generates attention weights from observations of raw neuron features and learned graph embeddings, which rescale baseline heuristic scores to guide neuron branching. Evaluation on 600 challenging instances demonstrates that RSB consistently outperforms state-of-the-art branching heuristics, solving 11% more instances while reducing branch exploration by 50%.
|
| 1451 |
PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding
2609.34555
|
cs.LG
|
Qiuyang Zhang, Kai Zhou, Kai Lu, Haocheng Lu, Jian Zhou |
Long-context LLM serving is increasingly bottlenecked by decode, where large KV caches limit batch size and underutilize GPUs. Sparse KV cache offloading expands effective capacity by storing most historical KV blocks in CPU DRAM and recalling only selected bl...Long-context LLM serving is increasingly bottlenecked by decode, where large KV caches limit batch size and underutilize GPUs. Sparse KV cache offloading expands effective capacity by storing most historical KV blocks in CPU DRAM and recalling only selected blocks on demand. However, we find that existing offloading systems shift the bottleneck to CPU-GPU recall I/O: recall volume varies widely across layers, decode steps and requests, while headwise sparse selection fragments recalls into many small PCIe transfers. This paper presents PulseInfer, an I/O-centric sparse KV cache offloading system. PulseInfer hides variable recall latency with interruptible layer-wise scheduling, adapts offloading decisions with IO-Adaptive Offloading Admission, and coalesces fragmented transfers using SoloHead sparse selection and a gather-scatter I/O engine. Implemented on SGLang, PulseInfer improves decode throughput by up to 4.7x over SGLang and 2.6x over the best existing offloading baseline, while reducing TPOT by up to 76% and preserving near-lossless accuracy.
|
| 1452 |
Brain-Conditioned Action Policies for Neural Motor Decoding
2609.34561
|
cs.LG
|
Luyao Jin, Running Zhao, Huan Zhao, Vincent C. K. Cheung, Wei-Hsin Liao |
Motor brain-computer interfaces (BCIs) aim to decode motor intention, enabling people with paralysis to control external devices. Neural motor decoding typically learns task-specific mappings from neural activity to kinematics, yet remains constrained by scarc...Motor brain-computer interfaces (BCIs) aim to decode motor intention, enabling people with paralysis to control external devices. Neural motor decoding typically learns task-specific mappings from neural activity to kinematics, yet remains constrained by scarce paired neural-action data. We propose BrainVLA, a framework that enables neural motor decoding by drawing on a pretrained vision-language-action (VLA) model through language-mediated alignment. BrainVLA mitigates reliance on scarce paired neural-action data by leveraging VLA policies. We first construct VLA-compatible datasets including paired neural activity, action signals, language instructions, and rendered visual observations. Then, we adapt the OpenVLA-OFT policy to the target action spaces through LoRA fine-tuning. To establish an effective interface through which neural activity can convey motor intention to adapted VLA policies and guide action generation, we train a neural encoder via neural-language alignment, using language representations as semantic targets to capture latent motor intent from neural activity. The resulting neural representations serve as an endogenous intention signal to guide VLA policies to generate executable actions, while visual observations provide complementary information about the evolving task state. BrainVLA is evaluated on two neural motor datasets with different action dimensionalities using causal rollout decoding. It outperforms the evaluated baselines in cross-session decoding $R^2$ and task success rate, while demonstrating high training data efficiency. These results establish a route for neural motor decoding to draw on large-scale robotic priors through brain-conditioned VLA policies.
|
| 1453 |
Single-Layer MeMo as a Randomized Hamming-Kernel Classifier
2609.34562
|
cs.LG
|
Alessandro Straziota |
MeMo (Zanzotto et al., 2025) is a recent language-model architecture that stores associations between token contexts and next tokens in a correlation matrix memory. In this work, we study its single-layer form and show that its ideal retrieval rule is a multic...MeMo (Zanzotto et al., 2025) is a recent language-model architecture that stores associations between token contexts and next tokens in a correlation matrix memory. In this work, we study its single-layer form and show that its ideal retrieval rule is a multiclass classifier based on the positional Hamming kernel. The MeMo architecture represents both the sequence features and the output labels with Gaussian random codes. Its score is therefore a doubly randomized sketch of the ideal classifier. Under independent input and output codebooks, we bound the errors introduced by context sketching and output decoding, characterize their dependence on model and data parameters, and give a margin-based guarantee for recovering the ideal prediction. Controlled simulations support the trends predicted by the analysis. On a restricted WikiText-2 next-token task, we compare single-layer MeMo with classical baselines and show that it can offer a useful trade-off among predictive accuracy, memory, and throughput, particularly on a GPU, where its matrix operations can be parallelized.
|
| 1454 |
Retracing Hodgkin and Huxley: State Recovery Does Not Certify Mechanism
2609.34566
|
cs.LG
|
Peiyu Zang, Jiayi Hao, Yongqiang Cai |
Predicting observed dynamics does not establish recovery of the underlying physical mechanism. Can machine learning retrace the hidden-state reasoning behind the Hodgkin-Huxley (HH) model? We train structured latent models on simulated current and voltage, wit...Predicting observed dynamics does not establish recovery of the underlying physical mechanism. Can machine learning retrace the hidden-state reasoning behind the Hodgkin-Huxley (HH) model? We train structured latent models on simulated current and voltage, withholding gate identities and trajectories from training and model selection. We then test response prediction, state recovery, protocol transfer, and agreement with HH dynamics. Prediction error and its cross-seed spread both drop sharply at three latent dimensions under the tested protocols, while gate recovery under new protocols improves through five to six coordinates. State recovery depends on which observations the chart uses. Observed voltage improves current-clamp decoding relative to freely predicted voltage. Under voltage clamp, adding latent state to command voltage raises m-state $R^2$ from 0.976 to above 0.99, yet the transported field disagrees with HH on identical smooth samples. Known invertible HH coordinates achieve high fast-m field agreement under the same audit procedure. An exact HH identity decomposes the discrepancy into time-scale-weighted state error and a residual in the transported field; these terms can cancel or reinforce. These findings concern the tested models and charts. They support evaluating state and dynamics recovery separately, including chart inputs and transported-field agreement across interventions.
|
| 1455 |
Compute Time Scaling with Recursive Models for Combinatorial Optimization
2609.34585
|
cs.LG
|
Zhengxi Zhang, Paul Swoboda |
We propose Tiny Recursive Models for Combinatorial Optimization (\ours{}), a general neural method for combinatorial optimization that scales both depth (how often we recursively invoke our network) and width (how much we sample in parallel). Both are fundamen...We propose Tiny Recursive Models for Combinatorial Optimization (\ours{}), a general neural method for combinatorial optimization that scales both depth (how often we recursively invoke our network) and width (how much we sample in parallel). Both are fundamental for combinatorial optimization: hard instances demand a large amount of compute, while a small network is essential to avoid overfitting and capture the algorithmic essence of optimization. In particular, our method consists of a graph-aware tiny recursive model that iterates on a latent state with adaptive halting and needs only a lightweight problem-specific decoder. Compared with previous heatmap-based general neural solvers, it achieves a better balance between solution quality and inference speed on both the Traveling Salesman Problem~(TSP) and the Maximum Independent Set~(MIS) problem, and remains competitive with hybrid methods that combine neural components with heuristics specific to each problem. With the same backbone architecture for both tasks, \ours{} outperforms every diffusion-based solver on TSP from 500 to 10,000 cities at a lower inference cost, and on the standard Erd\H{o}s--R\'enyi-[700-800] MIS benchmark it surpasses all neural solvers except those that only work well on MIS. We then explore self-relabeling for self-supervised training. We periodically replace the current set of training labels with the model's own better solutions, as an alternative training signal. Self-relabeling can, while forgoing supervision from near-optimal solutions, still result in on-par quality.
|
| 1456 |
Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation
2609.34593
|
cs.LG
|
Lican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong |
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q...Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample utilization efficiency. To address these issues, this paper introduces a weighted Bellman residual minimization framework that incorporates density ratio weighting by effectively integrating expert demonstrations with behavioral data. The proposed weighting scheme departs from the conventional completeness assumption commonly imposed in the theoretical analysis of deep reinforcement learning. We establish a sharp convergence rate for density ratio estimation and derive the convergence rate for the excess risk of resulting deep $Q^*$ estimator. Extensive empirical evaluations demonstrate that, compared to existing methods, our method achieves significant improvements in numerical performance and policy generalization, providing specific guidance for the rational utilization of expert demonstrations.
|
| 1457 |
The Low-Rank Structure of VLA Reinforcement Learning
2609.34599
|
cs.LG
|
Minjae Oh, Yoonah Park, Jongwon Lim, Yohan Jo |
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $\pi_{0.5}$ and GR00T~N1.5/N1....Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $\pi_{0.5}$ and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CALVIN induces substantially lower-rank parameter updates that are highly concentrated in the action expert's Timestep Modules, a small and previously overlooked component. Through systematic module-replacement experiments, we further show that these modules capture a disproportionate share of the performance gains from RL. We then characterize what is encoded in these Timestep Modules. First, we show that RL specializes them to the discrete denoising timesteps used during rollouts, and that this discrete-timestep training underlies the low-rank updates. Second, we find that among their outputs, the shift vector changes most distinctly under RL, and through probing, we show that shift update directions strongly predict task success (ROC-AUC up to $99.6\%$). Third, we find that the geometry of shift updates reflects task relationships, as their pairwise similarity correlates with cross-task transfer patterns. Building on these findings, we show that steering along shift update directions further improves RL-trained policies without additional RL training. Overall, we provide a systematic understanding of how RL reshapes VLA policies by studying how learned signals are encoded in parameter space, offering insights into more efficient and interpretable VLA post-training.
|
| 1458 |
When local gains fail to transfer: Frozen Earth-observation embeddings across wildfires
2609.34602
|
cs.LG
|
Philipp Stark, Alexandros Sopasakis, Ola Hall |
Frozen Earth-observation embeddings are judged almost entirely by spatially blocked cross-validation inside one study region. We show that this number does not predict accuracy in a new region; we show why; and we show the one setting in which such a model doe...Frozen Earth-observation embeddings are judged almost entirely by spatially blocked cross-validation inside one study region. We show that this number does not predict accuracy in a new region; we show why; and we show the one setting in which such a model does keep working, using a protocol that needs only a linear probe and labels one already has. The testbed is wildfire, with Copernicus burned-area maps of six fires in Greece and Spain and descriptors from the year before each fire, comparing TESSERA and AlphaEarth with ESA WorldCover classes and annual Sentinel-2 index summaries. Inside a fire, the embeddings identify the burned land 0.05 to 0.13 ROC AUC better than the index summaries, and repeated fold allocations, spatial buffers, a block bootstrap, and gradient-boosted trees leave that margin unchanged. On a fire in another region, they lose 0.15 to 0.18 AUC, and the index summaries lose 0.06, so the three end within a few hundredths of each other. The representation is not the cause. Eight labelled blocks from the new region restore the embedding advantage and give a higher AUC than 59,000 labelled pixels from other regions, and the weight vector fitted in one region is nearly orthogonal to the vector fitted in the others, so the part that carries across regions is small and low-dimensional. Forecasting within a region is a different matter. Fitted on a fire that burned in 2023 and applied to a fire twelve kilometres away that burned in 2024, where nothing used postdates the target fire, TESSERA reaches 0.772 AUC and loses 0.04 against a classifier fitted inside the 2024 fire, while classifiers fitted in other regions lose 0.09 to 0.18. A region with one mapped fire can therefore forecast susceptibility for later fires there; a region without one cannot borrow a model from elsewhere, and every evaluation of a frozen embedding should report a held-out region.
|
| 1459 |
Shaping Persistent Representations from Independent Interactions
2609.34604
|
cs.LG
|
Ji Dai, Quan Fang, Junyu Gao, Rongfeng Guo, Haoyan Rong |
World models learn environment dynamics from interaction experience. These dynamics depend on the current state and actions, as well as on properties that persist across interactions. Yet standard predictive training can reduce error using local evidence alone...World models learn environment dynamics from interaction experience. These dynamics depend on the current state and actions, as well as on properties that persist across interactions. Yet standard predictive training can reduce error using local evidence alone, without organizing persistent information into reusable context. We introduce SPRII, a training principle that uses relations between interactions as weak supervision for persistent context while retaining the learner's native objective. For example, different trajectories of the same system share persistent properties even when their states and actions differ. SPRII uses such relations to guide context learning without numerical property labels. Two composable components encourage contexts from related interactions to agree (Align) and use one interaction's context to predict another's future (Cross). Our analysis distinguishes three linked questions: what persistent information is accessible in the learned context (Formation), how that context influences a fixed predictor (Use), and whether it reduces task error (Value). Success at one stage does not guarantee success at the next. Controlled experiments show that more reliable relations improve representation organization, but adding a shared-property constraint can reduce access to a property that remains shared. Context substitutions change predictions at fixed model weights, while the benefit from history depends on prediction horizon and readout. Evaluations span thirteen settings, including controlled physical systems, public dynamics tasks, robotic and tactile data, and partner interaction, across multiple learner families. Relative to the corresponding baselines, SPRII yields average gains of over 10% in downstream task performance and over 15% in persistent-property readout. The project page is available at https://persistent-learning-review.netlify.app/interactive.html.
|
| 1460 |
PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
2609.34605
|
cs.LG
|
Youzhi Liu, Ruobing Zheng, Boyuan Tong, Tianqi Li, Pingqi Li |
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objecti...Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
|
| 1461 |
Beyond Site Agreement: Re-estimation for Brain Network Generalization
2609.34611
|
cs.LG
|
Yingxu Wang, Kunyu Zhang, Yanwu Yang3, Thomas Wolfers, Yujie Wu |
Cross-site out-of-distribution (OOD) generalization in resting-state functional magnetic resonance imaging (rs-fMRI) often relies on learning task-discriminative representations from full-scan functional connectivity (FC) graphs and promoting invariance across...Cross-site out-of-distribution (OOD) generalization in resting-state functional magnetic resonance imaging (rs-fMRI) often relies on learning task-discriminative representations from full-scan functional connectivity (FC) graphs and promoting invariance across source sites. However, FC graphs are estimated from finite, temporally correlated blood-oxygen-level-dependent (BOLD) sequences. Cross-site agreement therefore does not necessarily imply that predictive evidence remains supported under FC re-estimation within the same scan. In this paper, we propose Brain Network Re-estimation-Informed OOD Learning (BRIO), a framework that uses within-scan FC re-estimation to guide cross-site alignment. BRIO maps fullscan graphs and their re-estimates into consistently indexed connectome factors, enabling comparisons of their predictive contributions. It assesses re-estimation support from changes in these contributions relative to within-class subject variability and class separation. For each source-site pair and class, this task-calibrated support from both sites is combined with predictive relevance to form pairwise qualifications, which determine relative factor weights and overall alignment strength. Leave-one-site-out experiments on four real-world datasets (ABIDE, REST-metaMDD, SRPBS, and ABCD) show that BRIO consistently outperforms competitive baselines, with relative improvements of up to 3.8% in accuracy. These gains also persist under an alternative brain parcellation on ABIDE.
|
| 1462 |
Learning Regional Snow Water Equivalent and Snow Height Variations from Sentinel-1 InSAR Acquisitions
2609.34614
|
cs.LG
|
Luca Barco, Lorenzo Innocenti, Bianca Bartoli, Claudio Rossi, Edoardo Arnaudo |
Managing water resources in mountainous regions depends heavily on reliable Snow Water Equivalent (SWE) and Snow Height (HS) data, yet these variables remain difficult to track at scale. This study evaluates three machine learning architectures (XGBoost, U-Net...Managing water resources in mountainous regions depends heavily on reliable Snow Water Equivalent (SWE) and Snow Height (HS) data, yet these variables remain difficult to track at scale. This study evaluates three machine learning architectures (XGBoost, U-Net and SegFormer) for the joint estimation of SWE and HS variations from Sentinel-1 InSAR data over the Italian Alps, using the IT-SNOW reanalysis as reference. SegFormer achieves the best results on both targets, with an MAE of 10.391 cm for HS and 27.113 mm w.e. for SWE and the lowest variability across initializations. A feature sensitivity analysis shows that including all available features does not guarantee the lowest error, with model- and task-specific sensitivities. Spatial metrics (R2, Pearson's r) separate the three architectures far more clearly than mean error (MAE, RMSE) does, and decomposing the error per window attributes most of it to a systematic offset in the estimated mean variation rather than to the spatial pattern.
|
| 1463 |
Evaluating Dynamical Fidelity through Predictive Structure in Physical Representations
2609.34627
|
cs.LG
|
Oskar Bohn Lassen, Joao Paulo de Souza Boger, Simon Driscoll, Stephen I. Thomson, Sebastian Schemm |
Machine-learning models for physical systems are currently evaluated primarily through errors between predicted and reference states and, increasingly, through tests of physical consistency. These metrics assess whether predictions are accurate and satisfy sel...Machine-learning models for physical systems are currently evaluated primarily through errors between predicted and reference states and, increasingly, through tests of physical consistency. These metrics assess whether predictions are accurate and satisfy selected physical requirements, but provide limited insight into whether learned trajectories reproduce the underlying dynamics. Domain experts examine such relationships through physical representations that expose relevant processes, interactions, and responses, but these analyses are often separated from typical machine-learning evaluation. We introduce a practical framework for evaluating dynamical fidelity through predictive structure in physical representation spaces. Experts define the representations, while reference trajectories determine which relationships are predictive and retained as evaluation tests. We demonstrate the approach in atmospheric forecasting using ERA5 representations of planetary-wave activity and Northern Annular Mode evolution, and evaluate Pangu-Weather, GraphCast, and FengWu. The models exhibit distinct departures from reference predictive structure that are not reflected by conventional forecast errors. The framework thereby turns domain-expert representations into systematic tests of learned physical dynamics without prescribing the relationships in advance.
|
| 1464 |
DisKO: Deep Koopman Learning in Distribution Space from Unpaired Snapshots
2609.34629
|
cs.LG
|
He Ma, Xiaochen Liu, Wanfeng Lu, Ying Wang, Wei Lin |
Many complex systems are observed only through temporally unpaired distribution snapshots, making trajectory-based dynamical learning difficult without additional assumptions. We therefore formulate the problem directly in distribution space, treating the dist...Many complex systems are observed only through temporally unpaired distribution snapshots, making trajectory-based dynamical learning difficult without additional assumptions. We therefore formulate the problem directly in distribution space, treating the distribution itself as the dynamical state. The challenge is that distribution space is infinite-dimensional, making compact and approximately closed representations difficult to learn from finite snapshots. We introduce DisKO, which extends deep Koopman learning to distribution dynamics by jointly learning predictive distributional observables, a finite-dimensional Koopman representation, and a generative map back to the full distribution. Across seven diverse benchmarks, DisKO achieves state-of-the-art extrapolation performance, with substantially slower error accumulation on long-horizon prediction tasks. DisKO further recovers leading Koopman eigenvalues and eigenfunctions on systems with analytic spectra, revealing meaningful dynamical structure in the learned representation.
|
| 1465 |
GenMem: Generative Symbolic Memory for Self-Evolving Harness
2609.34633
|
cs.LG
|
Xinke Jiang, Tao Feng, Weixuan Xu, Zhixin Zhang, Zhibang Yang |
Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to...Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to discriminative retrieval and to address the sparse, hierarchical, and highly redundant structure of reusable experience: only a small, task-dependent subset of trajectories and memories warrants retention, retrieval, or revision. Learning these operations is further complicated by sparse, delayed, and indirect task-level feedback, with weak supervision across the memory lifecycle. Moreover, continual memory evolution introduces an architectural tension as addressing invariance: stored experience is perpetually revised, yet the addressing interface consumed by learned retrieval policies must remain stable. To address, we present GenMem, which reformulates memory management as generative symbolic addressing. Its core mechanism is the Symbolic Identifier (SID), a multi-level discrete token tuple drawn from a Cartesian-product address space that factorizes a million-scale sparse memory space using fewer than one hundred discrete symbols. Instead of generating ever-changing raw content, the memory agent learns to generate SIDs, while memory evolution rewrites the payload at a fixed address without shifting the address itself. Architecturally, GenMem couples a MemRetriever and a MemEvolver within a multi-agent harness, trained via GRPO with dense process and outcome rewards with two-channels optimization. Under offline memory evolution, experiments spanning ALFWorld, WebShop, multi-hop QA, medical reasoning, and deep research evaluate GenMem against strong memory-augmented baselines...
|
| 1466 |
Correction-space Cross-variate Interaction for Test-time Adaptation in Time Series Forecasting
2609.34638
|
cs.LG
|
Yuanyuan Deng, Mykola Pechenizkiy, Songgaojun Deng |
Test-time adaptation (TTA) is a promising paradigm for handling distribution shift in time-series forecasting (TSF), where models adapt at inference time, often leveraging delayed observed data to refine predictions. In the multivariate setting, distribution s...Test-time adaptation (TTA) is a promising paradigm for handling distribution shift in time-series forecasting (TSF), where models adapt at inference time, often leveraging delayed observed data to refine predictions. In the multivariate setting, distribution shifts often exhibit cross-variate dependencies, yet existing TSF-TTA methods adapt each variate independently and ignore this cross-variate structure. Exploiting such structure motivates cross-variate interaction, but coupling variates through backbone predictions introduces direct pathways for mixing uncorrected errors across variates, a concern under the delayed supervision of TSF-TTA. We identify the \emph{interaction space} as a key design choice, and show that acting on adapter corrections that refine backbone outputs, the \emph{correction space}, rather than on the predictions themselves, avoids directly propagating backbone errors across variates. We build on this to propose \textsc{CoRe} (\textsc{Co}rrection-space Interaction \textsc{Re}finement), realizing correction-space interaction through (i) Shared-anchor Correction Refinement (SCR), which combines each variate's correction with a shared anchor through a parameter-efficient bottleneck, and (ii) input-conditioned spectral gating, which adaptively modulates the refinement from the current input window. Across seven backbones, six datasets, and four prediction horizons, \textsc{CoRe} reduces MSE by 25.82\% on average over backbones and 10.57\% over the state-of-the-art TSF-TTA method, with stronger gains at medium-to-long horizons and modest computational overhead. Data and code are available at: https://github.com/yyddou/CoReTTA
|
| 1467 |
Uniform Race: Parameter-Free Approximate Rejection Sampling
2609.34639
|
cs.LG
|
Seiyun Shin, Juhyeong Pang, Kwang-Sung Jun |
We study approximate sampling: given $N$ independent samples from a proposal distribution $\mu$, the goal is to select one whose distribution is close to a target $\pi$ specified only up to a normalizing constant. Block and Polyanskiy (2023) provide finite bud...We study approximate sampling: given $N$ independent samples from a proposal distribution $\mu$, the goal is to select one whose distribution is close to a target $\pi$ specified only up to a normalizing constant. Block and Polyanskiy (2023) provide finite budget error bounds for approximate rejection sampling (RS) as a function of the acceptance threshold $M$. The threshold $M$ giving the smallest bound, however, depends on properties of $(\pi,\mu)$ that are typically unavailable from the observed sample. This raises a natural question: Can one attain the best RS guarantee without taking $M$ as input? We answer affirmatively by proposing a parameter-free sampling algorithm called uniform race (UR), based on importance weights, which are ratios of target to proposal probabilities (or densities). It divides each observed weight by an independent uniform random variable to form a score and returns the candidate with the largest score. For every budget $N$, its total variation error satisfies the RS upper bound for every fixed threshold $M$ simultaneously, thereby achieving the best such bound in hindsight. We also characterize its output distribution conditional on the largest score, identifying when it is exactly the target $\pi$. Uniform race has no larger total variation error than a natural budget-calibrated RS derived from Rohatgi et al. (2025) and sampling importance resampling (SIR). In particular, we exhibit instances where UR's error is exponentially smaller in $N$ than that of either baseline. Furthermore, we establish conditions under which attaining this RS guarantee for every $(\pi,\mu)$ uniquely determines the selection probabilities as those of UR. Finally, test-time scaling experiments on LLM math-reasoning tasks corroborate the theoretical comparisons and demonstrate that UR remains competitive in ground-truth accuracy without requiring threshold selection.
|
| 1468 |
Tilted Schr\"odinger Bridge Matching
2609.34642
|
cs.LG
|
Sergei Kholkin, Evgeny Burnaev, Alexander Korotin |
Schr\"odinger bridges provide an entropy-regularized framework and a principled solution for unpaired domain translation. In practice, a pretrained bridge may need to be adapted to human preferences or physical constraints through a reward a problem closely re...Schr\"odinger bridges provide an entropy-regularized framework and a principled solution for unpaired domain translation. In practice, a pretrained bridge may need to be adapted to human preferences or physical constraints through a reward a problem closely related to reward tilting in diffusion models but underexplored for Schr\"odinger bridges. We introduce Tilted Schr\"odinger Bridge Matching (TSBM), a post-training method for fine-tuning a learned bridge $P$ between source $p_0$ and target $p_1$ toward a reward-tilted target $p_1^r\propto p_1e^r$, while preserving source $p_0$. We formulate this adaptation as alternating optimization initialized from $P$, provide theoretical justification, and derive a practical algorithm based on Adjoint Matching. We evaluate TSBM on unpaired image-to-image translation targeting digit properties in MNIST and facial attributes in CelebA.
|
| 1469 |
Universal Dynamic Portfolios
2609.34643
|
cs.LG
|
Yu-Jie Zhang, Yu-Xiang Wang, Peng Zhao, Kevin Jamieson |
Cover's Universal Portfolio (Cover, 1991) matches the performance of the best constant rebalanced portfolio in hindsight. We generalize this framework to compete with an arbitrary comparator sequence $\mathbf{u}_1,\ldots,\mathbf{u}_T$, leading to a dynamic reg...Cover's Universal Portfolio (Cover, 1991) matches the performance of the best constant rebalanced portfolio in hindsight. We generalize this framework to compete with an arbitrary comparator sequence $\mathbf{u}_1,\ldots,\mathbf{u}_T$, leading to a dynamic regret minimization problem for the log loss where existing methods break down due to potentially unbounded gradients. The log loss is exp-concave, a curvature property that classically yields fast rates for static regret, yet we show that this advantage generally disappears in the dynamic setting. In particular, a linear-loss-type $\sqrt{TP_T}$ dependence is unavoidable, where $P_T=\sum_{t=2}^T\lVert\mathbf{u}_t-\mathbf{u}_{t-1}\rVert_1$ is the standard path length. This limitation stems from the coarse nature of $P_T$, which obscures finer spatial and temporal structure of the comparator sequence. We therefore introduce two structure-aware measures---the Jensen-Shannon distance for spatial structure and the JS$^q$-path length for temporal structure---under which faster rates are attainable when the comparator sequence has favorable structure. To achieve sharp bounds for both measures simultaneously, we develop Universal Dynamic Portfolio, a parameter-free method that combines a new Dirichlet Hedge algorithm with a fixed-share update, while retaining a near-optimal $P_T$ guarantee in the worst case. Finally, under an additional bounded-gradient assumption, we show that OPS admits the faster $T^{1/3}P_T^{2/3}$ dynamic regret rate over all comparator sequences. We attain this rate with a tractable proper algorithm that applies more broadly to general online exp-concave optimization over arbitrary compact convex domains.
|
| 1470 |
Minimax Last-Iterate Convergence in Matrix Games with Observed Actions
2609.34656
|
cs.LG
|
Yuheng Zhang |
We study last-iterate convergence in unknown two-player zero-sum matrix games with bandit payoff feedback and observed opponent actions. For games with $d$ actions per player, we develop an algorithm achieving a duality gap of $\widetilde{\mathcal{O}}(\sqrt{d/...We study last-iterate convergence in unknown two-player zero-sum matrix games with bandit payoff feedback and observed opponent actions. For games with $d$ actions per player, we develop an algorithm achieving a duality gap of $\widetilde{\mathcal{O}}(\sqrt{d/t})$ with high probability, simultaneously at every round $t$. This improves the dimension dependence of the best previously known guarantee by a factor of $d^{3/2}$. The rate matches a standard bandit lower bound, establishing minimax optimality in both the number of actions and the number of rounds, up to logarithmic factors. The algorithm is computationally efficient, requiring only $\mathcal{O}(d)$ time and memory per round. Our technical contribution is a joint design of adaptive averaging and corrected exponential weights that absorbs estimation variance, together with a potential argument that bounds phase durations.
|
| 1471 |
FestDPO: Few-step Generator Alignment with Direct Preference Optimization
2609.34673
|
cs.LG
|
Jaewoo Lee, Kyuil Sim, Hyeongyu Kang, Kanghoon Lee, Woocheol Shin |
Few-step generative models can generate high-fidelity samples within a few function evaluations. Despite this efficiency, generated samples may not exhibit desirable properties. When these properties are difficult to encode as an explicit reward function, dire...Few-step generative models can generate high-fidelity samples within a few function evaluations. Despite this efficiency, generated samples may not exhibit desirable properties. When these properties are difficult to encode as an explicit reward function, direct preference optimization (DPO) can align generative models using pairwise preference feedback without training a separate reward model. However, extending DPO to few-step generative models is challenging because few-step generative models are generally implicit, making the likelihood evaluation required by DPO intractable. To address this challenge, we introduce Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples. By exploiting the fast sampling capabilities of few-step generative models, our approach makes sample-based approximation of DPO loss computationally feasible. Furthermore, the sample-based formulation makes FestDPO agnostic to the model family and sampling procedure. Our toy experiment demonstrates that FestDPO matches the reward-tilted target distribution across four few-step generators. For real-world tasks, we evaluate FestDPO in two domains: text-to-image generation and protein backbone generation. In text-to-image generation, FestDPO outperforms preference optimization baselines in both win rates against the base models and human evaluation scores. In protein backbone generation, it achieves a higher $\beta$-sheet fraction and better structural designability than the baselines.
|
| 1472 |
Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
2609.34677
|
cs.LG
|
Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen, Amir Bar, Jong Chul Ye |
World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental ...World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for the current prediction, and which available retrieval cues should be trusted to find them? This is challenging because fixed criteria based on recency, pose overlap, or visual similarity can be unreliable across environments and queries. We propose Future-Aware Recall (FAR), a framework that learns episodic recall from future-aware predictive supervision and adaptive multi-cue scoring. During training, FAR measures predictive utility by the conditional log-likelihood of the realized future given recalled context, approximated by negative diffusion prediction loss, and uses it to train a retriever that remains future-blind at inference. The retriever learns cue-specific relevance and automatically determines which available retrieval cues, such as time, pose, vision, and audio, to trust for each query when selecting memories. Across three complementary settings, FAR outperforms hand-designed recall even with the same retrieval cues, automatically adapts which available cues to trust, and recalls the right history as the world changes. Together, these results establish FAR as a flexible, principled approach to episodic memory access in world models.
|
| 1473 |
From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling
2609.34679
|
cs.LG
|
Wangxuan Fan, Xiaoyu Nie, Zhoutian Shi, Xiangcheng Meng, Shipei Zeng |
Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and sh...Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential interaction under limited information. We propose a dynamic bipartite matching framework that combines large language model (LLM) agents with contextual bandits. In a simulated Chinese marriage market, economically grounded LLM agents evaluate locally encountered candidates, while agent-specific Logistic-UCB models learn reciprocal acceptance from realized proposal outcomes. The mechanism therefore separates two decisions---\emph{whom do I like?} and \emph{who is likely to like me back?}---without requiring ex ante market-wide preference rankings. We first validate LLM-induced mate preferences against the empirical conditional-logit reference across multiple LLM backbones. In the $50\times50$ matching experiment, Bandit-UCB achieves the highest mean mutual welfare (56.01 versus 54.87 for Gale--Shapley), a smaller gender rank gap than the classical baselines, and the fewest blocking pairs among the LLM-ABM policies. Learned acceptance models show economically interpretable gender-differentiated associations, while counterfactual setups reveal no systematic unilateral advantage from prior search knowledge. Overall, these results support the advantages of decentralized matching with LLM-based behavioral modeling and online learning under incomplete information for economic simulation and computational social science research.
|
| 1474 |
QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization
2609.34680
|
cs.LG
|
Qiulin Shang, Zhoutong Wu, Jie Hu, Kun Yuan |
Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect t...Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the network. The useful algorithmic decomposition is therefore not fully known before search. LLM-driven program evolution offers a way to explore these choices, but performance scores alone do not explain which design should change next. We introduce QuantForge, a PTQ discovery system that records competing explanations, selects controls that distinguish them, and checks that successor code implements the resulting conclusions. This residual compilation guides program revisions while retaining useful programs even when their original explanations are rejected. Remeasuring the revised program reveals the next error to address. This process discovers HiRes, a fixed MXFP4 quantizer that shapes coordinates, refines legal code assignments, and recovers errors along attention and MLP paths. Each stage acts on residuals measured after the preceding stage has executed. Across seven tasks, HiRes achieves the lowest seven-model Robust Fit (0.09300) and the lowest quantized Fit-7 at 32B. In matched-budget comparisons of LLM-driven program evolution, each with 240 evaluator calls, QuantForge reaches a held-out transfer target in six of eight runs, compared with three each for textual memory and reflection memory, and one for score-only evolution, despite evaluating fewer new programs. These results show that QuantForge improves the discovery of transferable PTQ algorithms by turning controlled evidence into subsequent program changes.
|
| 1475 |
SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining
2609.34681
|
cs.LG
|
Qiulin Shang, Binyu Wang, Yongqi Qiao, Songde Rao, Zhoutong Wu |
Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance...Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimization dynamics. Online learned scheduling within the Learning to Optimize (L2O) framework offers a dynamic alternative, but remains brittle at LLM scale due to noisy signals, delayed feedback, and the risk of catastrophic divergence. We propose SOLAR (State-driven Online Learning rAte scheduleR), a stabilized framework for reliable online LR adaptation. SOLAR uses a base schedule as a reference and learns bounded, state-dependent residual corrections for individual parameter groups. Each correction re-anchors to the base at every step, allowing the policy to adapt the LR without relearning the warmup-decay profile. A lightweight state representation and progress-aware reward guide online learning, while a Circuit-Breaker restores training after rare unsafe actions. Across autoregressive language-model pretraining, SOLAR improves final perplexity over tuned static schedules and automatic LR tuners for dense models from 60M to 1B, AdamW and Muon, and two MoE settings up to 3B. Matched 130M controls show that adding base anchoring and action bounds improves a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87 on the same two seeds. A residual policy trained on a 60M proxy can also be frozen and reused at larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range. These results establish SOLAR as a practical learned LR controller for LLM pretraining.
|
| 1476 |
AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
2609.34683
|
cs.LG
|
Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin, Yao Lai, Haoran Wu |
The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in th...The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Existing benchmarks primarily focus on simple single-turn chatbot workloads. LLM applications are increasingly agentic: coding agents, terminal execution systems, and tool-use agents issue multi-turn requests with growing context lengths. We introduce AgentPerfBench, a benchmark suite for agentic inference. It uses real traces from agentic benchmarks, such as SWE-Bench and TerminalBench, alongside standard chat baselines. This enables benchmarking of models on multi-turn tasks involving tool calling, skill utilization, and increasing context lengths. AgentPerfBench also samples from empirical distributions of input length, output length, and turn count derived from the real traces, generating representative synthetic profiles for cheap and accurate measurements on new hardware. In addition, we further find that several existing benchmarks fail to accurately reflect real hardware performance for two key reasons: 1) they do not account for realistic context-length growth, and 2) they measure inference performance without operating at hardware saturation. We discuss these issues in detail and provide rich kernel-level Nsight Compute (NCU) traces to construct a new multi-dimensional roofline model that captures hardware-system limitations in both memory bandwidth and memory capacity footprint. The benchmarking suite then includes automated scripts to identify potential bottleneck conditions on emerging hardware when evaluated with diverse agentic traces. Together, these contributions quantify the chat-to-agentic gap in current inference benchmarks and characterise per-kernel GPU resource utilisation via roofline analysis.
|
| 1477 |
Predicting Delayed Train Trajectories on the Dutch Railway Network: Explainable AI Evaluation of Topological, Operational and Weather Features with Tree Based Ensemble Methods
2609.34692
|
cs.LG
|
Jia Long Bao, Ali Mohammed Mansoor Alsahag, Seyed Sahand Mohammadi Ziabari |
The reliable prediction of passenger train delays is a critical component of railway management. While contemporary research frequently attempts to maximize absolute accuracy by deploying opaque deep learning architectures, the underlying data mechanics drivin...The reliable prediction of passenger train delays is a critical component of railway management. While contemporary research frequently attempts to maximize absolute accuracy by deploying opaque deep learning architectures, the underlying data mechanics driving longitudinal predictive decay remain underexplored. Consequently, this study provides an explainable temporal robustness analysis of network-wide railway delay prediction. Focusing on the Dutch railway network, this research utilizes interpretable tree-based ensembles to integrate granular topological, environmental, and operational features. The overarching finding establishes that while feature-rich tree-based models improve simultaneous (within-month) prediction, predictive performance systematically degrades when evaluated across non-simultaneous (future) months. Furthermore, multi-horizon SHAP and dispersion analyses explicitly link this degradation to environmental feature volatility and instability within the statistical target definition. Ultimately, this thesis demonstrates that richer feature sets alone are insufficient to resolve long-term forecasting constraints, underscoring the necessity to transition toward dynamic, season-aware architectures anchored by absolute operational boundaries.
|
| 1478 |
Learning Propagation Geometry from Message-Passing Feedback
2609.34711
|
cs.LG
|
Yingxu Wang, Kunyu Zhang, Xinwang Liu, Mengzhu Wang, Siyang Gao |
Learning local geometry enables graph neural networks (GNNs) to adapt how they compare and integrate neighborhood information. However, estimating geometry from aggregated representations can overlook variation among individual messages and dependencies across...Learning local geometry enables graph neural networks (GNNs) to adapt how they compare and integrate neighborhood information. However, estimating geometry from aggregated representations can overlook variation among individual messages and dependencies across feature dimensions. We propose GeoF, a recurrent framework that jointly evolves node features and propagation geometry through message-passing feedback. Each node maintains a local symmetric positive-definite geometry, initialized from a structure-aware prototype atlas and parameterized in block log-triangular coordinates. At each step, the geometry determines neighborhood weights, while triangular frame transport maps transformed source messages into the target node's local coordinates before aggregation. Weighted second-order statistics of residuals between aligned messages and the transformed target state capture directional variation and within-block dependencies, yielding a geometric update target. A shared controller learns complementary corrections through task supervision. A bounded log-triangular update combines these corrections, the target, and the previous geometric state while preserving positive definiteness. The geometry governs subsequent propagation, closing the feedback loop. With parameters shared across recurrent steps, task-specific readouts support node classification, link prediction, and graph classification. Experiments on benchmark datasets show that GeoF consistently outperforms state-of-the-art GNN baselines.
|
| 1479 |
Edge-Level Automorphism in GNNs: A Quantitative Framework and Effective Designs For Link Prediction
2609.34729
|
cs.LG
|
Chen Shao, Donald Loveland, Tobias K\"afer, Danai Koutra |
Graph Neural Networks (GNNs) are effective for learning node and link embeddings through permutation-equivariant aggregation. However, standard GNNs collapse automorphic nodes, i.e., those with identical structural roles (or orbits) into indistinguishable repr...Graph Neural Networks (GNNs) are effective for learning node and link embeddings through permutation-equivariant aggregation. However, standard GNNs collapse automorphic nodes, i.e., those with identical structural roles (or orbits) into indistinguishable representations, leading to the node automorphism problem. This collapse limits their expressive power and degrades link prediction performance. Existing approaches to characterize GNN expressiveness rely primarily on Weisfeiler-Lehman (WL) analyses, but these methods are typically qualitative and often misaligned with empirical results. To address this gap, we begin by introducing a novel quantitative framework to assess GNN expressiveness for link prediction. We first formalize edge-level automorphism through edge orbits, which capture the set of structural role pairs for nodes that share a link. Then, we introduce the edge automorphism ratio (EAR), a scalar metric that quantifies a GNN's ability to distinguish links in a given graph. We empirically demonstrate that EAR correlates strongly with performance, validating its practical benefit. Building on this insight, we design EDGE-ORBIT EQUIVARIANT GRAPH NEURAL NETWORK (EO-GNN), a GNN architecture that addresses automorphism collapse while preserving equivariance and incurring minimal computational overhead. EO-GNN accomplishes this through two core designs combined with WL-based node hashes: (i) automorphism-aware dropouts and (ii) subgraph orbit-biased aggregation. Empirical evaluations on synthetic and real graphs show improvements of up to 42.36% and 28.44%, respectively, in predicting links in scenarios with high automorphism.
|
| 1480 |
Predictive Dual Smoothing for Column Generation
2609.34740
|
cs.LG
|
Senne Berden, Noah Schutte, Andrea Lodi, Tias Guns |
Solving large-scale linear programs efficiently is an important challenge in many optimization settings. A key technique is column generation, which alternates between solving the master problem over a restricted subset of the variables, and using a pricing su...Solving large-scale linear programs efficiently is an important challenge in many optimization settings. A key technique is column generation, which alternates between solving the master problem over a restricted subset of the variables, and using a pricing subproblem to identify new variables to add. The pricing subproblem is guided by the dual solution of the current restricted master problem, but oscillations in these dual solutions can substantially slow convergence. Dual stabilization methods address this issue. Dual smoothing is a common stabilization method, which guides the pricing subproblem using a combination of the current dual solution and duals from previous iterations. However, while past dual solutions can stabilize the dual trajectory, they do not necessarily guide pricing towards useful new variables. We therefore introduce predictive dual smoothing, which instead combines the current dual solution with a learned prediction of future duals to steer pricing towards variables that are more useful in subsequent iterations. The predictor is trained offline using supervision extracted from standard column generation trajectories and is used only to modify the pricing subproblem's objective function, while exact reduced-cost checks and fallback pricing with the unsmoothed duals preserve correctness. Experiments on cutting stock and generalized assignment problems show that predictive dual smoothing substantially reduces generated columns and wall-clock time relative to standard column generation and existing classical and learned stabilization methods. These gains extend to out-of-distribution instance sizes, and predictive smoothing provides further improvements when combined with strong classical stabilization.
|
| 1481 |
No Pain, More Gain: Iterative Merging for Effective Multi-Teacher On-Policy Distillation
2609.34745
|
cs.LG
|
Seonghyeon Kim, Chaeyun Jang, Noah Lee, Boseop Kim, Juho Lee |
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different pos...Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOPD can struggle to recover some teacher capabilities. Because distillation occurs on student-generated prefixes, the student initialization can strongly affect subsequent recovery. However, initial benchmark performance is not a reliable predictor of a good MOPD initialization. For example, merge initialization can start below SFT warm-up yet finish higher after MOPD. We further find that effective merging depends on both the relative teacher contributions and the overall merge scale, with some strong configurations lying outside the simplex of convex parameter averaging. Thus, selecting a good merge initialization requires evaluating not only its immediate performance but also the learning it enables under MOPD, making one-shot coefficient search difficult. We propose Iterative Merging for MOPD (IM-MOPD), which starts from a uniform merge and progressively adds task-vector increments for under-recovered domains during distillation. In a 5-domain setting, IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up, showing that effective teacher contributions can be determined progressively during training.
|
| 1482 |
A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees
2609.34750
|
cs.LG
|
Vojt\v{e}ch K\r{u}r, Adam Kuku\v{c}ka, Tom\'a\v{s} Br\'azdil, V\'it Musil |
Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We i...Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.
|
| 1483 |
Instance-Adaptive Prompts as Context for Time-Series Foundation Models
2609.34786
|
cs.LG
|
Zehao Xiao, Shifeng Xie, Lei Zan, Jianfeng Zhang, Lujia Pan |
Longer histories can improve time-series foundation models (TSFMs), but require substantially higher inference cost. We therefore ask whether contextual information can be provided more efficiently through a compact set of learned token embeddings. We introduc...Longer histories can improve time-series foundation models (TSFMs), but require substantially higher inference cost. We therefore ask whether contextual information can be provided more efficiently through a compact set of learned token embeddings. We introduce PaCTS, which generates a small set of instance-adaptive latent prompts in the form of continuous embedding tokens conditioned on the visible context. These prompts serve as compact context surrogates for frozen TSFMs. PaCTS constructs them from instance-specific global statistics and further refines them with segment-level temporal information, capturing both global characteristics and local temporal variations. The prompt module is jointly trained and deployed across heterogeneous time series with the frozen backbone. Extensive experiments demonstrate the effectiveness of prompts as context, consistently improving forecasting across context lengths and model architectures. With a shorter input context, PaCTS can outperform the same frozen backbone using double context while requiring substantially less inference computation. Compared with weight-space adaptation methods, PaCTS achieves stronger improvements and better out-of-distribution generalization.
|
| 1484 |
Separating personal from population gains when calibrating EEG foundation models for new users
2609.34801
|
cs.LG
|
Xilin Tao, Kani Chen |
Foundation models are increasingly adapted to individual users, but an apparent personalization gain can simply reflect a stronger population model. This distinction matters for brain-computer interfaces, where every new user must be calibrated. We evaluated p...Foundation models are increasingly adapted to individual users, but an apparent personalization gain can simply reflect a stronger population model. This distinction matters for brain-computer interfaces, where every new user must be calibrated. We evaluated personal adaptation of three frozen EEG foundation models (CBraMod, REVE and LaBraM) in 235 held-out subjects from three motor-imagery datasets, comparing each subject's adapter with the population model and with adapters fitted to other subjects. Using all first-half session labels, personal adapters improved mean balanced accuracy over the population model by 1.5-5.4 percentage points and outperformed exchanged adapters by 2.3-7.3 points in all nine model-dataset combinations. The size of this benefit depended on population training: with four times the original budget, median gains remained positive (1.0-2.0 points) but were smaller for every model, and no population model reached a confirmed plateau. Acquiring the benefit cheaply was unreliable: few-label calibration was consistently non-negative on only one dataset, and in CBraMod neither unlabeled context nor meta-learned initialization outperformed matched controls. Personalization should therefore be evaluated against both a population reference and exchanged parameters, across population-training budgets.
|
| 1485 |
Polylogarithmic Nash Regret in Matrix Games with Bandit Feedback
2609.34812
|
cs.LG
|
Yuheng Zhang |
We study Nash regret minimization in unknown finite matrix games with bandit payoff feedback and observed opponent actions. We develop Optimistic Payoff Balancing (OPB), which achieves instance-dependent $\mathcal{O}(\log^2 T)$ Nash regret against arbitrary ad...We study Nash regret minimization in unknown finite matrix games with bandit payoff feedback and observed opponent actions. We develop Optimistic Payoff Balancing (OPB), which achieves instance-dependent $\mathcal{O}(\log^2 T)$ Nash regret against arbitrary adaptive opponents, including games with nonunique equilibria. This resolves the open problem posed by Maiti et al. (2025), extending their polylogarithmic guarantee under bandit feedback from $2\times2$ games to arbitrary finite dimensions. To handle nonunique equilibria, we construct a reference strategy that leaves room for local adjustments. We order independent payoff differences by estimation accuracy and scale these adjustments by uncertainty, allowing the learner to exploit the opponent's imbalance to offset estimation costs. Our result thus shows that observing opponent actions suffices for polylogarithmic Nash regret in general finite matrix games.
|
| 1486 |
Gaussian Neural Networks
2609.34825
|
cs.LG
|
Peter Kuhn, Victoria Heusinger-He{\ss} |
Gaussian neural networks (GaNNs) are proposed as a novel regularization mechanism for neural networks. From a Bayesian perspective standard regularization techniques can be viewed as imposing priors over weight-space. Assuming priors over activation-space rema...Gaussian neural networks (GaNNs) are proposed as a novel regularization mechanism for neural networks. From a Bayesian perspective standard regularization techniques can be viewed as imposing priors over weight-space. Assuming priors over activation-space remains a largely unexplored possibility. GaNNs assume such priors. They do this by treating activities from earlier layers like signals with Gaussian noise and predicting the properties of the noise distribution using an additional unsupervised loss. While training, the unsupervised loss acts as a penalty on unexpected activities, allowing greater weight updates in less surprising directions. The paper demonstrates the superiority of Gaussian neural networks over standard neural networks on a variety of classification and regression tasks. We also investigate the ability of GaNNs to quantify uncertainty.
|
| 1487 |
Structured Neural SDEs for Functional Calibration
2609.34831
|
cs.LG
|
Francesco Piatti, Andrea Iannucci, Thomas Cass |
Neural Stochastic Differential Equations (Neural SDEs) provide flexible continuous-time generative models, but generic neural drift and diffusion networks are costly to simulate on long horizons and can give unstable gradients when the training signal is a pat...Neural Stochastic Differential Equations (Neural SDEs) provide flexible continuous-time generative models, but generic neural drift and diffusion networks are costly to simulate on long horizons and can give unstable gradients when the training signal is a path functional rather than a pointwise observation. We introduce SLiSDE, a family of Neural SDE models built from structured linear stochastic layers. Parallel-in-time simulation is obtained at the layer level, while expressivity is recovered by gated in-flow stacking: previous-layer paths modulate the next layer's latent flow through learned gates. For functional calibration tasks in which rare paths dominate the loss, we add an optional Girsanov tilt that acts as a learned importance sampler with an exact likelihood-ratio correction. We prove well-posedness, a discretisation error bound, validity of the change of measure, and a universality result: the terminal laws of the gated stack are dense in the space of square-integrable laws. Experiments on functional calibration benchmarks show that the structured model outperforms fully neural SDE baselines while retaining parallel-time simulation and stable importance weights.
|
| 1488 |
QiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities
2609.34842
|
cs.LG
|
Hanyin Cheng, Linfeng Wang, Zhengbo Qu, Yang Shu, Zhongwen Rao |
Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aw...Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aware multimodal TSFM that models the two types of modalities separately. For endogenous modalities, to capture how they evolve along with the underlying temporal dynamics, we introduce an Endo-Multimodal Predictor and Endo-Multimodal Supervision to explicitly learn their evolution from history to the future. For exogenous modalities, to generalize across domains and across various modality types and numbers under the scarcity of exo-multimodal pretraining data, we propose an Exo-Multimodal Retrieval Enhancer that enables rapid downstream adaptation without updating the TSFM parameters. We further introduce Endo-Modality Proxy Training to train this retrieval module without exogenous multimodal pretraining data. Extensive experiments across unimodal and multimodal benchmarks demonstrate strong forecasting performance in scenarios both with and without exogenous modalities.
|
| 1489 |
Context-dependent time-series prediction via HyperReservoirs
2609.34847
|
cs.LG
|
Kohei Tsuchiyama, Takatomo Mihana, Ryoichi Horisaki, Andr\'{e} R\"{o}hm |
Time series prediction is a common application of reservoir computing. When the training and testing time series data contains multiple dynamical regimes, because an underlying parameter is changing, or the data in fact consists of multiple distinct systems, s...Time series prediction is a common application of reservoir computing. When the training and testing time series data contains multiple dynamical regimes, because an underlying parameter is changing, or the data in fact consists of multiple distinct systems, simple application of the reservoir computing principle produces high prediction errors. Here, we propose a HyperReservoir as an extended model of reservoir computing especially designed for such cases. The HyperReservoir combines a main reservoir with a smaller context reservoir, where the latter modulates the output weights of the former. This structure resembles the hypernetworks from deep neural network literature. However, in contrast, HyperReservoirs retain the simple training via linear regression of standard reservoir computing. We compare the proposed architecture with a conventional ESN, in which context acts at the input, and a full-matrix Conceptor, in which context modulates the reservoir state space. We evaluate all three models on time-series prediction tasks based on Lorenz and R\"ossler systems, including for varying bifurcation parameters and time sampling scales. We find that the HyperReservoir achieves the lowest mean test error in all three tasks, and particularly outperforms conceptors on data that is sampled from the same attractor but at different time scales.
|
| 1490 |
When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation
2609.34849
|
cs.LG
|
Xinke Jiang, Tao Feng, Zhibang Yang, Zhixin Zhang, Weixuan Xu |
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement thi...Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing dense token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward--distillation training through a neural tangent kernel (NTK) analysis. We introduce the cross-signal NTK $K_{DR}(n)$, a token-level statistic that measures the alignment between reward and distillation gradients at position n. Through this analysis, we identify two failure modes: 1 Magnitude drowning, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to rise despite its explicit inclusion in the training objective; and 2 Localized directional conflict, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token ($K_{DR}(n)\!<\!0$). The severity of these effects depends on the optimization regime: the gradient-norm ratio $\kappa\!=\!\|\nabla\mathcal{L}_R\|/\|\nabla\mathcal{L}_D\|$ varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the M3 family, which combines magnitude normalization with three strategies...
|
| 1491 |
Learning High-Risk High-Precision Motion Control
2609.34851
|
cs.LG
|
Nam Hee Kim, Markus Kirjonen, Perttu H\"am\"al\"ainen |
Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we f...Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we focus on an under-researched class of high-risk, high-precision motion control problems where actions carry irreversible outcomes, driving sharp peaks and ridges to plague the state-action reward landscape. Using computational pool as a representative example of such problems, we propose and evaluate State-Conditioned Shooting (SCOOT), a novel DRL algorithm that builds on advantage-weighted regression (AWR) with three key modifications: 1) Performing policy optimization only using elite samples, allowing the policy to better latch on to the rare high-reward action samples; 2) Utilizing a mixture-of-experts (MoE) policy, to allow switching between reward landscape modes depending on the state; 3) Adding a distance regularization term and a learning curriculum to encourage exploring diverse strategies before adapting to the most advantageous samples. We showcase our features' performance in learning physically-based billiard shots demonstrating high action precision and discovering multiple shot strategies for a given ball configuration.
|
| 1492 |
Attention-based Hierarchical Variational Information Bottleneck for Robust Multi-Agent Communication under Variable Bandwidth
2609.34860
|
cs.LG
|
Lukas Koch Vindbjerg, Qi Zhang, Yury Brodskiy, Lukas Esterle |
Learning-based multi-agent communication under limited bandwidth does not only require deciding what to communicate, but also structuring messages so that partial transmissions remain useful. We study this problem under prefix truncation, where only the first ...Learning-based multi-agent communication under limited bandwidth does not only require deciding what to communicate, but also structuring messages so that partial transmissions remain useful. We study this problem under prefix truncation, where only the first part of each message is received. To address it, we propose \textbf{AH-VIB}, an attention-based autoregressive variational communication model that combines a variational information bottleneck (VIB) with sequential message generation and a hierarchical robustness loss. We evaluate AH-VIB on a custom cooperative object-inspection and occupancy-mapping task, where agents equipped with a limited field-of-view sensor coordinate to scan inspection objects in an occupancy-grid world, under variable and fixed bandwidth conditions, and compare it against MADDPG, CommNet, a flat VIB baseline, and an autoregressive MLP ablation. AH-VIB achieves competitive mean return while improving performance reliability under the most constrained bandwidth conditions. These results indicate that AH-VIB improves the reliability and graceful degradation of learned communication under bandwidth constraints.
|
| 1493 |
From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers
2609.34866
|
cs.LG
|
Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei, Milad Hosseini, Adrian Weller |
Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, ...Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention output instead, over all three projections at once, and reuse it throughout the pipeline. JAB defines one scalar loss over the joint Q, K, V weights of a block, evaluated against the block's real causally-masked attention output, and uses it twice: to fit the quantized weights (GPTQ warm start, then STE with learnable scales), and to score the block for a multiple-choice knapsack allocation. On attention-only quantization of Mistral-7B this works. At 3 bits JAB recovers 77-90% of the gap between uniform GPTQ and full precision, and its sensitivity estimate tracks an oracle costing 73 forward passes to within a fraction of a point. It stops working once MLP layers enter the allocation. A role-aware offset rule needing no sensitivity estimate at all beats JAB on GPT-2's MLP and on the full Mistral-7B model: with a 3-bit floor it quantizes 96.4% of the weights to 4.5 bits per parameter at 6.933 perplexity, within 4.4% of full precision (6.643) at 3.56x compression, against 7.158 for JAB at the same budget. Which matrix a weight sits in matters more than any sensitivity estimate we computed. Two things came out sideways. Block-local reconstruction is an unreliable proxy for end-to-end perplexity: one run improved a block's own objective 4.6x while perplexity rose 32x, which is why every allocation here is validated end-to-end. And on attention-only quantization, fine-tuning moved weights farther from their pretrained values while pulling attention outputs closer, with net gains. Post-training seems to recover attention behavior, not weights.
|
| 1494 |
Reference-Tail Trust:Certified Probability Floors for Learned Updates Inside a Deployed Network
2609.34904
|
cs.LG
|
Abdolvahab Khalili Sadaghiani, Jose Nunez-Yanez |
Graph neural networks (GNNs) need to exploit improved message passing without surrendering control over predictions already trusted in deployment. We introduce Reference-Tail Trust (RTT), a framework that admits learned updates inside a frozen GNN and certifie...Graph neural networks (GNNs) need to exploit improved message passing without surrendering control over predictions already trusted in deployment. We introduce Reference-Tail Trust (RTT), a framework that admits learned updates inside a frozen GNN and certifies the prediction actually served. RTT couples graph-based proposal states with a constrained internal optimizer: each displacement is charged for its worst-case terminal cross-entropy increase through the incumbent's remaining message-passing layers. A trajectory-validated tube and an independent checker enforce per-node probability floors, $p^{\mathrm{s}}_{ic} \ge e^{-H_{\mathrm{row}}} p^{\mathrm{r}}_{ic}$, and a call-level budget, $\sum_i w_i D_\infty(p^{\mathrm{r}}_i \| p^{\mathrm{s}}_i) \le H^+$, uniformly over labels. Calls whose adapted outputs pass certification require no separate full incumbent rollout; failed certificates trigger whole-call fallback. We derive the exact probability-floor frontier by water-filling, characterize architecture-constrained efficiency, and establish conditions under which internal propagation exploits evidence unavailable to restricted output correctors. In the reported ogbn-arxiv audit, RTT achieves $6.5\times 10^{-3}$ nats of mean gain per call, with a one-sided 95% regression-rate upper bound of 0.95% and a 95% negative-flip upper bound of 0.51% on the uninspected part of the reserved node population. Its mean gain is 61% of a cross-fitted posterior-based frontier estimate and exceeds the strongest matched one-pass corrector by $+0.9\times 10^{-3}$ nats. Reported experiments span eight proposals, six graph-incumbent families, structural and temporal graph shifts, and molecular prediction, with additional image and tabular evaluations. RTT makes GNN adaptation a budgeted, certifiable inference decision rather than an unconditional model replacement.
|
| 1495 |
Muon Sublates the Edge of Stability in LLM Pretraining
2609.34915
|
cs.LG
|
Yanzhe Chen, Qifang Zhao, Xiaoxiao Xu, Fanghui Liu |
Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability me...Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary $2\rho_b/\eta$, while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining
|
| 1496 |
Drug-Target Interaction Prediction via Hierarchical Sequential Cross-Attention over Chemical and Protein Language Models
2609.34921
|
cs.LG
|
Khadidja Henni, Hamza Abdelali, Abdelkrim Aries, Neila Mezghani, Brigitte Vannier |
Predicting Drug-Target Interactions~(DTIs) is a central task in computational drug discovery, with direct applications in virtual screening, drug repurposing, and therapeutic candidate prioritization. Although recent deep learning methods have improved DTI pre...Predicting Drug-Target Interactions~(DTIs) is a central task in computational drug discovery, with direct applications in virtual screening, drug repurposing, and therapeutic candidate prioritization. Although recent deep learning methods have improved DTI prediction, many sequence-based models still process drugs and proteins independently and only combine their representations at a late prediction stage. This limits their ability to explicitly model cross-molecular dependencies between chemical substructures and protein sequence regions. In this paper, we propose a sequence-only DTI prediction architecture that combines two pre-trained language models, ChemBERTa for drug SMILES strings and ESM-2 for protein amino acid sequences, with a hierarchical interaction module. The proposed model first extracts contextual representations using pre-trained encoders, then applies 1D convolutional layers to condense local sequence patterns, followed by a sequential bidirectional cross-attention mechanism inspired by the induced-fit view of molecular recognition. Finally, attention-based pooling constructs fixed-size interaction-aware vectors for binary prediction. Experiments on BIOSNAP, Davis, and BindingDB show that the proposed model achieves the best performance on BIOSNAP, matches the best AUROC on Davis, and remains competitive on BindingDB while using only 25.2 million trainable parameters. Ablation results confirm the contribution of both the CNN and cross-attention modules, and cold-start experiments indicate promising generalization to unseen proteins and drugs.
|
| 1497 |
Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
2609.34924
|
cs.LG
|
Sebastian Bobadilla-Suarez, Bob Suh, Ryan Fortin |
An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape on...An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape only if that set expands. Rewriting scaffolding (tools, verifiers, decomposition) expands what an agent reaches without touching a weight, so frozen weights buy an eventual ceiling but no stationarity along the way. The criterion also separates three regimes usually merged: search within a fixed class, test-time training that raises the ceiling itself, and scaffold rewriting between them. Audit the scaffold, not the checkpoint. The same ceiling binds sideways. Best-of-$k$ orchestration realizes the best worker's ceiling exactly: width buys rate, not budget. Re-consulting a fixed pool has a horizon computable in advance, decided by the pool alone, and the one arrangement that would beat it, a weighted vote, needs diversity real workers lack: on 30 same-family workers the failure overlap sits at its maximum, and a majority fails 23/55 (42%) of tasks. We obtain the criterion by reading refinement as gradient boosting on the residual error between draft and target, a patch or git diff, and then measuring where that reading breaks: patches compose instead of standing beside each other to be voted on, and failures overlap. What we measure is saturation. Per-round improvement decays toward zero on SWE-bench, and churn decays geometrically across 401 production sessions, a shape shared with a pre-AI human baseline that establishes the regime without identifying its cause. Both breaks are engineering choices rather than laws about code, so together they specify a harness worth building.
|
| 1498 |
XMatch: Enhancing Covariate-Aware Time Series Forecasting through Tree-Structured Exogenous Matching
2609.34939
|
cs.LG
|
Ziyang Zhang, Hanyin Cheng, Xiangfei Qiu, Yang Shu, Bin Yang |
Future exogenous variables provide valuable information for forecasting endogenous time series. Existing covariate-aware methods primarily learn the direct influence of exogenous variables on endogenous variables. However, these effects can be complex and chan...Future exogenous variables provide valuable information for forecasting endogenous time series. Existing covariate-aware methods primarily learn the direct influence of exogenous variables on endogenous variables. However, these effects can be complex and change with the pattern of the exogenous variables, making them difficult to capture. Beyond this perspective, we observe that a given exogenous pattern often co-occurs with only a small set of endogenous response patterns. These associations motivate a strategy that matches future and historical exogenous patterns and uses the corresponding endogenous patterns to enhance forecasting. However, in real-world forecasting scenarios with multiple exogenous variables, each exogenous variable provides a distinct dimension for matching, creating a dilemma for this strategy between precise matching and sufficient historical support. To bridge this gap, we propose XMatch (EXogenous MATCHing), a covariate-aware forecasting model that realizes the aforementioned strategy through a tree-structured matching process that adaptively adjusts the number of exogenous variables used as matching conditions. Specifically, we first introduce the ProtoTree Creator, which organizes historical correspondences between exogenous and endogenous patterns into a ProtoTree, whose deeper levels incorporate additional exogenous variables for matching. For forecasting, we then design the ProtoTree Matcher, which uses future exogenous variables to query the ProtoTree and adaptively determines how many exogenous variables to use for matching based on exogenous pattern similarity and historical support. Finally, the matched endogenous patterns are used as explicit historical evidence to enhance forecasting. Extensive experiments on 12 real-world datasets demonstrate that XMatch outperforms state-of-the-art baselines.
|
| 1499 |
Adjoint Guidance Flow: Amortized Critic Guidance for VLA Policies
2609.34944
|
cs.LG
|
Jeongsol Kim, Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin |
Flow-based Vision-Language-Action (VLA) policies are typically trained by behavior cloning and thus do not explicitly optimize long-term task return. Critic guidance steers generation toward higher-value actions, but existing methods differentiate the critic t...Flow-based Vision-Language-Action (VLA) policies are typically trained by behavior cloning and thus do not explicitly optimize long-term task return. Critic guidance steers generation toward higher-value actions, but existing methods differentiate the critic through a one-step surrogate of the sampler and back-propagate a critic ensemble at every flow step. In contrast, here we propose Adjoint Guidance Flow (AGF), which amortizes trajectory-aware critic guidance into a lightweight guidance network while preserving the pretrained VLA policy. Specifically, we formulate critic-guided flow generation as a deterministic optimal control problem, whose optimal guidance is a costate that carries the terminal critic gradient back through the remaining flow, and regress the guidance network onto this costate while keeping both the VLA and critic frozen. This design provides favorable memory and throughput scaling during training, and inference needs one guidance-network forward pass per step, without the critic ensemble, back-propagation, or adjoint computation. Across LIBERO, RoboCasa, and LIBERO-Pro, AGF consistently improves pretrained VLAs, remains competitive with critic-guidance and policy-fine-tuning baselines, and is the most robust method when a single guidance strength is deployed across tasks. Compared with QGF, AGF runs $3.6\times$ faster per guidance step with $7.0\times$ fewer parameters, with comparable and even better performance, showing that critic guidance can be trajectory-aware and lightweight.
|
| 1500 |
ALICE: In-context, Zero-shot, Mutual Information Estimation
2609.34962
|
cs.LG
|
Giulio Franzese, Simone Rossi, Pietro Michiardi |
Estimating mutual information (MI) from samples is a central objective in a variety of scientific fields. Modern neural estimators are accurate in the large-data regime, but they fall short when data is scarce, and each must be fit anew for every distribution ...Estimating mutual information (MI) from samples is a central objective in a variety of scientific fields. Modern neural estimators are accurate in the large-data regime, but they fall short when data is scarce, and each must be fit anew for every distribution under study. Current estimators are moreover tied to specific data types. These constraints limit their adoption in many applications where per-distribution training is impractical and sample sizes are small. We present ALICE, a foundation model that removes per-distribution training, while achieving competitive estimation accuracy. Trained exclusively on a broad family of synthetic distributions, ALICE acts as an in-context estimator of rectified-flow velocity fields: conditioned on samples of an unseen distribution, it estimates that distribution's velocity field without any explicit training. MI is then obtained through a fixed identity that integrates the squared difference between the joint and conditional fields. We validate ALICE on a standard, challenging benchmark and apply it in three domains, biology, genetics, and neuroscience, whose data the model has never seen. For the first time, we show that a single model closes the gap with neural estimators trained separately for each distribution, while natively supporting different data dimensionality and sample cardinality, enabling zero-shot MI analysis across scientific domains.
|
| 1501 |
Cyclostationary Phase Conditioning for Medical Time Series Diffusion
2609.34965
|
cs.LGcs.AI
|
Samuel Ruiperez-Campillo, Michele Copetti, Jorge da Silva Goncalves, Sonia Laguna, Thomas Hofmann |
Many physiological time series, such as cardiac and brain recordings, exhibit cyclostationarity: their statistics vary periodically with an underlying cycle phase. Corruption from motion, poor contact, and physiological interference obscures morphology needed ...Many physiological time series, such as cardiac and brain recordings, exhibit cyclostationarity: their statistics vary periodically with an underlying cycle phase. Corruption from motion, poor contact, and physiological interference obscures morphology needed for diagnosis, making signal restoration essential. Existing diffusion approaches condition on corrupted observations alone and must learn cyclic structure implicitly. We instead propose two inductive biases which encode cyclostationarity: a shift-covariant wavelet representation and dense per-sample phase conditioning inferred from the corrupted input. We further introduce a training-free cyclostationarity index that quantifies phase structure and predicts when phase conditioning will help. Finally, we propose antithetic coupling of reverse trajectories to reduce sampling variance while achieving comparable performance with fivefold fewer network evaluations. Across modalities, our results show that explicitly encoding measurable cyclic structure improves physiological time-series restoration.
|
| 1502 |
Teach to Learn: Hint Annealing for Self-improving LLM Reasoning
2609.34975
|
cs.LG
|
Zile Wang, Zijian Li, Haodong Wang, Jian Liu, Qianli Liu |
Group Relative Policy Optimization (GRPO) improves language-model reasoning by comparing verified rewards among multiple solution rollouts for each query. However, difficult training queries can yield only incorrect rollouts, leaving GRPO with no reward contra...Group Relative Policy Optimization (GRPO) improves language-model reasoning by comparing verified rewards among multiple solution rollouts for each query. However, difficult training queries can yield only incorrect rollouts, leaving GRPO with no reward contrast or learning signal. Prior hint-based methods construct auxiliary hints from solution evidence and use them to re-solve failed queries, recovering learning signal. Yet the resulting trajectories are typically treated as ordinary solution trajectories despite being generated under an assisted condition unavailable at evaluation. We discover hinted reward shift: recovered reward contrast can concentrate policy updates on hinted trajectories, limiting improvement without hints. This also creates a trade-off: increasing hinted trajectories can accelerate early learning but intensify reward shift later. To address this problem, we propose HATCH (Hint-Annealed Self-Teaching), an online single-policy framework that learns from both generating and using its own hints to improve reasoning without assistance. To mitigate hinted reward shift, we introduce online weighting to anneal the contribution of hinted trajectories. However, learning to generate hints can conflict with improving query solving. We therefore use gradient projection to remove the opposing component of hint-generation updates. Together, these designs support self-improvement by enabling the policy to create learning opportunities for itself and turn them into stronger reasoning without hints. We evaluate our method on mathematical reasoning benchmarks and outperform state-of-the-art methods by 1.02 pp on Llama-3.2-1B-Instruct, 2.84 pp on Qwen3-1.7B, and 4.32 pp on Qwen3-8B.
|
| 1503 |
Inspector: Conversational and Lightweight Analyzer of Analog Circuit Layouts Using LLM and CNNs
2609.34976
|
cs.LG
|
Abril Cano Castro, Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni |
The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intellig...The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDSII file represents the industry-standard database containing the ultimate and most accurate source of information of the analog circuit, encapsulating the complex physical geometries and parasitic realities that define tape out performance. This paper proposes a novel framework that combines fine-tuned LLMs and CNNs to analyze GDSII files of analog circuits, enabling a conversational interface between the tool and the designers. Experimental results using thousands of analog designs across four realistic tasks demonstrate that the proposed solution outperforms state-of-the-art general-purpose massive VLMs by a significant margin (up to 81%), thus providing a lightweight solution to the problem of GDSII analysis.
|
| 1504 |
MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR
2609.34990
|
cs.LG
|
Yangyang Ren, Haodong Zhu, Sheng Xu, Yanjing Li, Nikolai Yu. Zolotykh |
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posterior...Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts. However, existing methods overlook how reliably learning signals are extracted from sampled responses. In GRPO, a response's advantage depends on both its own outcome and the randomly sampled outcomes of its peers through group normalization. Our theoretical and experimental analyses show that uncertainty in group composition introduces composition noise, a non-vanishing variance component that imposes an irreducible lower bound on gradient estimation error and impairs downstream prompt selection. We propose MaPP (Marginalized Posterior-Predictive), a unified framework for data-efficient RLVR that denoises response-level advantage estimation and improves prompt selection using a shared Beta posterior. For each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage through closed-form Beta-Binomial marginalization. The resulting posterior-predictive estimator has an error that provably diminishes as the posterior concentrates. Using the same posterior, MaPP derives an uncertainty-aware prompt selection score to improve data efficiency without additional rollout cost. Experiments on mathematics, planning, and visual geometry across five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget and setting a new state of the art.
|
| 1505 |
Composable Decoding on the Probability Simplex: Theory and Implementation
2609.34992
|
cs.LG
|
Xiaotong Ji, Ahmed Khaled Khamis, Rasul Tutunov, Matthieu Zimmer, Haitham Bou-Ammar |
Decoding for large language models is typically treated as a collection of isolated sampling strategies, with limited theoretical understanding of the behaviours they induce and how their underlying objectives relate. We formulate decoding as an optimisation p...Decoding for large language models is typically treated as a collection of isolated sampling strategies, with limited theoretical understanding of the behaviours they induce and how their underlying objectives relate. We formulate decoding as an optimisation problem over next-token distributions on the probability simplex, balancing expected model score against regularisation under support constraints. This view recovers familiar decoding methods through choices of regularisers and support constraints; more importantly, it enables new decoders to be constructed by composing distributional preferences within a single optimisation problem without external rewards, learned critics, or model parameter updates. We introduce CompoSimplex, a library with configurable support rules, regularisation primitives, and simplex solvers for constructing and evaluating compositional decoders. We evaluate standard samplers, individual regularisers, and compositions across multiple models and reasoning tasks. Our results show that compositions can realise trade-offs between single-sample quality, multi-sample quality, and diversity that are not attained by individual decoding objectives.
|
| 1506 |
Price Stability in the European Union: A Systemic Approach Using Random Matrix Theory
2609.35011
|
cs.LG
|
Sami Diaf |
Price stability remains a pillar in monetary policy practices and carries a special importance within monetary unions. Mainstream economics tried to leverage price stability using price indices and several metrics to shed light on specific dynamics and optimal...Price stability remains a pillar in monetary policy practices and carries a special importance within monetary unions. Mainstream economics tried to leverage price stability using price indices and several metrics to shed light on specific dynamics and optimal macroeconomic levels. The wide availability of data led researchers to consider the study of systems using Random Matrix Theory, based on inner correlation patterns. This aims to enhance the multivariate analysis by removing noisy patterns from the signal and improve data quality for further inferences. This work considers the collection of monthly inflation indices in the Eurozone as a \textit{system} of prices to analyze its eigenvalues' statistical and asymptotic properties and uncover inner country-level insights. Results confirm the system cannot assumed to be randomly generated, and the data exhibit noise-dominated patterns, due to small and persistent variations at the country-level. The latter make the inter-country correlations more dynamic and the separation of the signal from the noise quiet difficult. Findings identified two countries as distorting inflation dynamics besides three other distinct, regional-based groups of countries. Variability sources might stem from economic episodes fueling inflation spikes in some countries, as well as methodological aspects used to ensure data quality and representativeness in the European Union. Despite being complex, the system demonstrates a certain stability, in terms of self-organization; while large monthly fluctuations cannot be considered as rare events, but part of the data-generating process.
|
| 1507 |
Addressing Spatial Indistinguishability in Spatiotemporal Prediction via Optimal Transport-Guided Masking
2609.35021
|
cs.LG
|
Guangyu Wang, Jiawei Tong |
Spatiotemporal prediction aims to learn discriminative representations from correlated temporal signals over spatial structures for accurate future inference. A central challenge is \emph{spatial indistinguishability}: different nodes may share similar histori...Spatiotemporal prediction aims to learn discriminative representations from correlated temporal signals over spatial structures for accurate future inference. A central challenge is \emph{spatial indistinguishability}: different nodes may share similar historical patterns yet evolve toward divergent futures, severely degrading forecasting performance in real-world sensor networks. Existing embedding-based and graph neural network (GNN)-based approaches can partially detect such ambiguous nodes but rely on historical similarity, struggling to capture \emph{future behavioral divergence}. We propose \textbf{STOT} (\textbf{S}patio\textbf{T}emporal \textbf{O}ptimal \textbf{T}ransport), a self-supervised framework that resolves spatiotemporal ambiguity via structured masking guided by optimal transport. Our key idea treats indistinguishability as a \emph{disambiguation} problem: future states are inferred by exploiting concurrent spatial correlations and their time-varying similarity. We design a similarity-aware metric for dynamic inter-node relationships and an optimal transport-based masking strategy to emphasize ambiguous positions during pre-training. A batch consistency constraint preserves semantic coherence, while a random-walk masking mechanism promotes structured context exploration. Experiments on six real-world datasets show that STOT performs competitively with state-of-the-art baselines on the evaluated benchmarks and improved interpretability through transport-plan visualizations.
|
| 1508 |
Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons
2609.35029
|
cs.LG
|
Mana Sakai, Masaaki Imaizumi |
Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh ...Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh et al., 2026), which formalizes when transfer is effective, we investigate conditions that ensure fast transfer in the growing-horizon regime. Specifically, we study learning-rate transfer in a shallow linear network with a single trainable hidden matrix, trained by full-batch gradient descent. Under additional spectral assumptions, our main results are threefold. (i) We prove fast learning-rate transfer as $n,T\to\infty$ whenever $T=o(\sqrt{n})$. (ii) We characterize the transfer rates through the finite-width perturbation scale, the first-order sensitivities of the loss and its learning-rate derivative to finite-width perturbations, and the local loss curvature. (iii) We derive limiting distributions for the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix. These results clarify how spectral structure and local loss sensitivities govern learning-rate transfer at growing horizons.
|
| 1509 |
THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout
2609.35035
|
cs.LG
|
Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni |
The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intellig...The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDSII file represents the industry-standard database containing the ultimate and most accurate source of information of the analog circuit, encapsulating the complex physical geometries and parasitic realities that define tape out performance. This paper proposes THEIA, a novel dataset containing thousands of layout images paired with question-answer conversations, along with a benchmark that employs a fine-tuned vision-language model (VLM) to analyze GDSII files of analog circuits, enabling designers to interact with and query physical layouts as intuitive, meaningful entities. Experimental results using thousands of analog designs across five realistic tasks demonstrate that the proposed fine-tuned VLM outperforms state-of-the-art general-purpose VLMs by a significant margin (up to 73%), highlighting a fundamental gap between general-purpose multimodal reasoning and domain-specific layout understanding.
|
| 1510 |
BA-DPO: Bias-Adjusted Direct Preference Optimization for Language Model Alignment
2609.35044
|
cs.LG
|
Antonio Ferrara, Alberto Rumi, Francesco Bonchi |
Preference-based alignment methods such as Direct Preference Optimization (DPO) use pairwise preferences labeled by human annotators to fine-tune language models. However, annotators carry systematic biases toward some attributes: a name that signals a gender ...Preference-based alignment methods such as Direct Preference Optimization (DPO) use pairwise preferences labeled by human annotators to fine-tune language models. However, annotators carry systematic biases toward some attributes: a name that signals a gender or an ethnicity, a persona, a language variety, a formatting convention, or length. If not properly addressed, these systematic biases can be absorbed and amplified during alignment. Existing methods address length bias or annotator disagreement, but fail to eliminate biases toward arbitrary attributes. To address this limitation, we propose Bias-Adjusted DPO (BA-DPO), a generalization of DPO that adds one bias parameter per annotator toward responses carrying a declared attribute. We prove that the objective is convex in the bias parameters and that the votes identify each annotator's bias up to a shared constant. The remaining constant is what fixes the aligned model's attribute rate: by default the rate of the reference model, or a target rate, which we use to bring a biased policy to statistical parity. On a corpus with planted biases, DPO drives the attribute from a balanced start to probability 0.96 and BA-DPO removes 81 to 95\% of that shift; on MultiPref with real annotators it removes about half of DPO's lengthening. Both hold at 0.5B with full fine-tuning and at 8B with LoRA, at no higher KL than DPO and no loss in judged quality.
|
| 1511 |
Universality and Generalization of Causal Transformers Across Context Lengths
2609.35055
|
cs.LG
|
Takashi Furuya, Maarten V. de Hoop, Gabriel Peyr\'e |
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arb...Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by $\alpha$-H\"older sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a $\beta$-smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error $O((\log\log N/\log N)^{\beta/(d+2)})$ from $N$ iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with $d$ the token dimension and no maximum-length factor. Finally, experiments on physical time series support the H\"older-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.
|
| 1512 |
TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL
2609.35058
|
cs.LG
|
Yibin Huang, Xinming Xu, Conghui Zhu |
Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early e...Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.
|
| 1513 |
Beyond Gradient Flow: Identifiability and Recovery from Distribution Snapshots
2609.35060
|
cs.LG
|
Nam D. Nguyen, Valeriya Malysheva |
Inferring dynamics from snapshots of evolving distributions is fundamentally underdetermined: the Fokker-Planck equation constrains the drift $F$ only through its score-weighted divergence $\nabla\cdot F+F\cdot\nabla\log\rho$, leaving a $\rho$-solenoidal gauge...Inferring dynamics from snapshots of evolving distributions is fundamentally underdetermined: the Fokker-Planck equation constrains the drift $F$ only through its score-weighted divergence $\nabla\cdot F+F\cdot\nabla\log\rho$, leaving a $\rho$-solenoidal gauge invisible to any single-time constraint. Time-indexed transport formulations cannot resolve this ambiguity: every admissible marginal path admits a curl-free explanation, minimum-action reconstruction selects it, and marginal fit alone cannot distinguish dynamically inequivalent explanations. Requiring one autonomous field to explain several marginals instead makes part of the hidden circulation visible as $\nabla\log\rho$ changes across marginals. Separating instantaneous Fokker-Planck source constraints from the snapshot experiment, we show that the source constraints identify the field modulo the kernel of a stacked score-weighted divergence operator. For generic Gaussian shape variation, source constraints at $K\ge m$ time points in intrinsic dimension $m$ eliminate every polynomial gauge direction, whereas finitely many density snapshots alone admit aliasing; we give the obstruction explicitly. At a Gaussian anchor, for Sobolev smoothness $s$ and $n$ samples per time point, we derive a conditional lower rate $(nK)^{-2s/(2s+m+1)}$ for the tangent snapshot experiment, with a matching upper rate in a degreewise benchmark. Strong-form fitting is non-orthogonal to score error and cannot be repaired by spectral filtering. Instead, we estimate using smooth test functions while retaining the known diffusion term, and derive a finite-sample bound that separates sampling error from fixed-grid quadrature bias. Planted-circulation experiments confirm the predicted gauge contraction and expose a design tension between cross-slice information and covariance-aware whitening.
|
| 1514 |
Depot-Closed Multi-Component Construction for Neural Vehicle Routing
2609.35066
|
cs.LG
|
Shinichiro Hamada, Hisashi Kashima |
Most neural constructive solvers for the vehicle routing problem (VRP) use route-by-route construction, extending one route until completion before starting the next. This commits route membership early and hinders global coordination across routes. We propose...Most neural constructive solvers for the vehicle routing problem (VRP) use route-by-route construction, extending one route until completion before starting the next. This commits route membership early and hinders global coordination across routes. We propose multi-component construction, which maintains many route components simultaneously and merges them in an arbitrary order. This removes the depot-return cue that route-by-route construction obtains from the remaining capacity; to compensate, we introduce an interpretation in which every component is treated as an implicitly depot-closed route. Under this depot-closed interpretation, every intermediate state of standard CVRP construction is a complete feasible solution, and the exact cost reduction of a merge is the Clarke-Wright saving. The neural policy combines this CW-saving signal with the evolving component state to learn what to connect and when to connect. A policy trained only on CVRP100 outperforms the reported results of representative neural solvers on CVRP100-500 with greedy inference and, reused for ruin-and-reconstruct, performs strongly at all evaluated sizes up to CVRP1000. In a zero-shot Constraint Tightness evaluation with capacities from $C=10$ to $500$, it outperforms the reported neural solvers at every capacity. Controlled analyses show that robustness persists without CW grounding and point to learned route-closing behavior as a plausible contributor to the tight-regime degradation of learned route-by-route solvers.
|
| 1515 |
Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning
2609.35072
|
cs.LG
|
Xuesong Jia, Ziao Yang, Zhanhe Huang, Hongfu Liu |
We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement lea...We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.
|
| 1516 |
Interrelating Fruchterman-Reingold Graph Visualization and Agglomerative Clustering
2609.35073
|
cs.LG
|
Alexandre Benatti, Luciano da F. Costa |
Graph visualization methods and agglomerative clustering have been frequently considered in data analysis and pattern recognition. Because these approaches are interrelated and complementary, it is of particular interest to investigate their associations. In t...Graph visualization methods and agglomerative clustering have been frequently considered in data analysis and pattern recognition. Because these approaches are interrelated and complementary, it is of particular interest to investigate their associations. In this work, we study the possible relationship between the Fruchterman-Reingold graph visualization method and four types of agglomerative clustering adopting single- and complete-linkage, average, and Ward's linkage criteria. Three types of datasets have been considered in 2 and 10 dimensions, as well as the PCA projection of the latter to two dimensions. The results obtained suggest that the relationship between the methods considered did not vary much for the three types of data mentioned above. At the same time, the agglomerative methods tended to yield results that are mostly similar to each other, while presenting moderate similarity with the original data. The Fruchterman-Reingold visualization resulted similar to the original data, but exhibited relatively smaller similarity to the agglomerative methods.
|
| 1517 |
ReCo: When to Relocate Sensor Kits under Deployment Constraints -- A NILM Case Study
2609.35075
|
cs.LG
|
Haokun Chen, Yu Tong, Yehai Chen |
Many sensing tasks obtain training labels only by deploying instruments in the field. With a limited number of sensor kits, a collection deadline, and measurement downtime at every move, the collector must repeatedly decide whether to stay at the current site ...Many sensing tasks obtain training labels only by deploying instruments in the field. With a limited number of sensor kits, a collection deadline, and measurement downtime at every move, the collector must repeatedly decide whether to stay at the current site or relocate. We study this decision in non-intrusive load monitoring (NILM), which estimates the power drawn by individual appliances from a home's main meter and is trained on data from homes temporarily fitted with appliance-level sub-meters. In NILM, appliance usage varies with the appliance, season and climate, and the value of new data depends on how diverse the combinations of target operation and background load are. To address this, we propose a constraint-based relocation framework and instantiate it for NILM as ReCo (Relocation by Coverage gain). ReCo counts new operating regimes in a joint target-background feature space, forecasts each home's future gain from the data collected so far, and each night weighs the gain of staying against the gain of moving elsewhere after the downtime. In replayed deployments on the Plegma dataset under two kit counts and two downtime costs, ReCo outperforms fixed-dwell and count-based schedules and a threshold rule using the same metric in every setting. Its advantage is not explained by collecting more days alone and reflects allocating the days to more valuable homes and periods.
|
| 1518 |
Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers
2609.35080
|
cs.LG
|
Preben Johnsen Bentdal, Nello Blaser, Xue-Cheng Tai |
We study post-hoc refinement of frozen node classifiers: given only the graph $G$ and class distributions $Q$ predicted by a frozen model, can we improve accuracy without access to node features, model parameters, or gradients? APPNP answers this by propagatin...We study post-hoc refinement of frozen node classifiers: given only the graph $G$ and class distributions $Q$ predicted by a frozen model, can we improve accuracy without access to node features, model parameters, or gradients? APPNP answers this by propagating logits with a restart towards the initial predictions, minimizing the anchored Dirichlet energy. Instead, we consider the Potts energy, and decompose it into a Dirichlet term, which penalizes disagreement between neighbouring nodes, and a Gini term, which penalizes indecision within each node. This decomposition motivates Propagate, Then Sharpen (PtS), which alternates between propagation of class probabilities and node-wise, mass-preserving sharpening, with only one additional hyperparameter selected using labelled validation nodes. Across nine homophilic graphs, with a frozen MLP backbone, PtS improves mean test accuracy over independently tuned APPNP by $1.71$ percentage points on clean inputs and $3.90$ under severe Gaussian feature corruption. Gains over APPNP become smaller, but remain positive with frozen GCN and GraphSAGE backbones. Sharpening also removes most of the accuracy loss of deep propagation: on clean inputs without restart, accuracy falls by $2.2$ points between $2$ and $100$ propagation steps under PtS, compared with $33.8$ for APPNP.
|
| 1519 |
Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
2609.35082
|
cs.LG
|
Yangyang Ren, Haodong Zhu, Linlin Yang, Sheng Xu, Peichao Lai |
Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credi...Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence beyond the realized suffixes observed at an anchor while aggregating alternative continuations according to their empirical frequencies. Visit-local averaging pools realized suffix returns at shared anchors and respects observed frequencies, but does not recursively propagate evidence across rollouts, whereas shortest-path estimators have global reach but allow a rarely observed route to dominate an anchor's value. We introduce Cross-Rollout Bellman Closure (CRBC), which merges each rollout group into a finite empirical process with absorbing success and failure boundaries and evaluates its behavior-policy Bellman fixed point with one linear solve. This fixed point uses the same empirical action and transition frequencies to propagate evidence through shared anchors and aggregate alternative continuations. Backing up the resulting state values through observed transitions yields action values, whose gain over the corresponding state value provides step-level credit. A corresponding finite-depth family recovers visit-local return averaging at zero depth and converges to the exact closure as depth increases. The normalized closure credit is combined with the trajectory-level group advantage for policy optimization, without additional environment rollouts. Across ALFWorld, WebShop, and Sokoban benchmarks with multiple model scales, CRBC consistently improves final performance and learning efficiency. For example, CRBC outperforms the strongest evaluated baseline by 5.59 percentage points on ALFWorld with Qwen2.5-1.5B-Instruct.
|
| 1520 |
GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents
2609.35084
|
cs.LG
|
Haodong Zhu, Yangyang Ren, Changbai Li, Sheng Xu, Linlin Yang |
Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an act...Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its retrospective relation to the realized outcome. Hindsight credit assignment (HCA) instead attributes credit through the ratio of hindsight to behavior-policy probabilities, but estimating the hindsight distribution requires an auxiliary model or an extra pass. To address this estimation bottleneck, we propose GraphHCA, a model-free realization of HCA that eliminates explicit hindsight-distribution estimation. For terminal-goal tasks with deterministic transitions, Bayes' rule reduces the hindsight ratio to a ratio of behavior-policy success probabilities at consecutive states. Taking logs yields a state-wise success potential, whose increment across a transition provides step-level credit. GraphHCA estimates this potential from pooled rollouts through a discounted recursion on the induced transition graph, which admits a unique fixed point on any directed graph. The resulting step-level signal is combined with the trajectory-level advantage, requiring neither a learned hindsight model nor an extra forward pass and recovering GRPO when the step-level weight is zero. Among all compared baselines, GraphHCA achieves state-of-the-art results on ALFWorld and WebShop at both LLM scales, and on Sokoban with a vision-language agent. For example, on ALFWorld it improves overall success rate by up to 24.6 points over GRPO and by up to 4.7 points over the strongest step-level baseline.
|
| 1521 |
Retrieval-Augmented Diffusion Modeling for Stochastic Discount Factor Portfolios
2609.35086
|
cs.LG
|
Kelvin J. L. Koa, Xinyang Li, Ke-Wei Huang |
In this work, we study portfolio optimization under the stochastic discount factor (SDF) framework by learning market state representations that capture the underlying risk structures of financial data. This is challenging due to several factors: financial mar...In this work, we study portfolio optimization under the stochastic discount factor (SDF) framework by learning market state representations that capture the underlying risk structures of financial data. This is challenging due to several factors: financial markets exhibit non-stationary dynamics with shifting regimes, multimodal inputs such as price and news data often contain stochastic noise, and existing diffusion-based approaches, while effective for modeling stochastic dynamics, rely on assumptions such as isotropic Gaussian noise that fail to capture the state-dependent nature of financial uncertainty. To address these challenges, we introduce RADAR, a retrieval-augmented diffusion framework that learns market representations by conditioning on similar historical regimes. RADAR leverages retrieval to construct context-dependent noise distributions, applies conditional diffusion to denoise multimodal representations, and initializes the diffusion process using empirical statistics to reflect state-dependent uncertainty. Experiments show that RADAR achieves state-of-the-art performance on key risk-adjusted metrics while producing economically meaningful signals on asset returns and correlations.
|
| 1522 |
SpikeLite: Lightweight Spiking Neural Networks for Time-Series Forecasting
2609.35097
|
cs.LG
|
Bang Hu, Changze Lv, Mingjie Li, Xiaoqing Zheng, Wei cao |
Spiking neural networks (SNNs) offer an energy-efficient paradigm for time-series forecasting through spike-driven computation. However, recent SNN forecasters often pursue higher accuracy through increasingly complex attention mechanisms, or specialized neuro...Spiking neural networks (SNNs) offer an energy-efficient paradigm for time-series forecasting through spike-driven computation. However, recent SNN forecasters often pursue higher accuracy through increasingly complex attention mechanisms, or specialized neuronal dynamics, weakening the lightweight motivation of SNNs. We introduce SpikeLite, a spiking forecasting framework built around two modules: a Frequency-Selective Spiking Encoder (FSSE) for frequency-sensitive temporal encoding and a Sparse Spiking Channel Attention (SSCA) module for selective cross-channel interaction. FSSE exploits the low-pass filtering behavior of LIF dynamics to reorganize each input sequence into frequency-sensitive components while collectively preserving the input at the decomposition stage. SSCA then learns a binary mask from encoded channel representations and uses it to selectively exchange information within spike-driven self-attention, retaining informative cross-channel interactions while suppressing redundant ones. When explicit channel interaction is unnecessary, SpikeLite uses the lighter FSSE-only channel-independent path. Experiments under the SeqSNN and SpikF protocols cover four standard multivariate and eight long-term forecasting benchmarks. SpikeLite achieves the best aggregate performance under both protocols, with an average $R^2$ of 0.790 and RSE of 0.440, and lowest average MSE/MAE of 0.343/0.345 in long-term forecasting. Moreover, evaluation on the ECL dataset shows that SpikeLite achieves the lowest reported energy consumption, further demonstrating its potential for energy-efficient time-series forecasting.
|
| 1523 |
E3J: An Efficient and Open-Source Backend for Euclidean Equivariant Operations on GPU and TPU
2609.35099
|
cs.LG
|
Olivier Peltre, Armand Picard, Adrien Pichard, Miguel Bragan\c{c}a, Luca Giacomoni |
We present e3j, a fast Euclid-equivariance backend for geometric deep learning applications with JAX bindings for GPU and TPU. Leveraging both optimized CUDA and Pallas kernels and algorithmic improvements, the library achieves state-of-the-art throughput and ...We present e3j, a fast Euclid-equivariance backend for geometric deep learning applications with JAX bindings for GPU and TPU. Leveraging both optimized CUDA and Pallas kernels and algorithmic improvements, the library achieves state-of-the-art throughput and runtime on both forward and backward paths. On a machine learning interatomic potential (MLIP) use case, it outperforms established backends, measuring up to 34% speed-up over cuEquivariance on water box NPT simulation using MACE, while remaining fully open source. E3j achieves over 80% efficiency over the H100 maximum memory bandwidth on tensor product operations, and in many cases more than doubles throughput of message passing convolutions forward compared to previously available backends. In addition, with the release of dedicated Pallas TPU kernel, e3j opens the possibility of large scale equivariant deep learning workloads on TPU architectures, which has so far been difficult to achieve. Our benchmarks show that e3j also achieves over 80% of a TPUv6e memory bandwidth, up to one order of magnitude more than e3nn-jax. The library is available on GitHub, PyPI and is released under an open source Apache 2.0 license.
|
| 1524 |
DRIFT: Disentangled Responsive-Invariant Flow Transport for Single-Cell Perturbation Prediction
2609.35106
|
cs.LG
|
Mustapha Bounoua, Giulio Franzese, Pietro Michiardi |
Predicting cellular responses to perturbations is a central problem in cellular biology, with broad applications in systems biology and drug discovery. This task is challenging because cellular responses can be complex and cell-state dependent, intrinsic cell-...Predicting cellular responses to perturbations is a central problem in cellular biology, with broad applications in systems biology and drug discovery. This task is challenging because cellular responses can be complex and cell-state dependent, intrinsic cell-to-cell variability can be confounded with perturbation effects, and destructive single-cell RNA sequencing precludes paired measurements of the same cell before and after treatment. Flow matching transports control cells to perturbed states flexibly, but acting on the full cell state can confound perturbation effects with pre-existing cell-to-cell variability. Disentangled approaches separate responsive from invariant components, but model perturbations through prescribed mechanisms, such as latent shifts or graph edits, limiting their flexibility. We address both limitations in a unified framework. A variational encoder disentangles each cell into an invariant block, capturing state unaffected by the perturbation, and a responsive block, capturing state it changes, through conditional priors and an information-theoretic invariance constraint. Conditional flow matching transports only the responsive block, conditioned on the perturbation and invariant state, yielding a flexible, data-driven model of perturbation effects without confounding pre-existing variability. Across several benchmarks, our method outperforms the strongest published method in settings involving combinatorial and unseen perturbation prediction.
|
| 1525 |
SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression
2609.35113
|
cs.LG
|
Ziwen Zhang, Xiju Wu, Yuheng Jing, Runxiang Wang, Boxiao Wang |
Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is exp...Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is expensive, and compact benchmarks lack systematic evidence of preserved task diversity and algorithm discriminability. SymbolicArena provides a unified infrastructure for benchmark distillation and dynamic evaluation. The framework standardizes 664 heterogeneous tasks with executable ground truth expressions and distills the Full Task Set into Core50, a validated benchmark of 50 tasks. The distillation process preserves task coverage and algorithm discrimination under explicit balance constraints. SymbolicArena applies a unified execution protocol to heterogeneous SR algorithms and produces comparable outputs and search trajectories. Multi Axis Evaluation characterizes numerical quality, symbolic quality, and search behavior. Core50 reduces evaluation workload by 92.5% and maintains agreement with Full Task Set evaluations. Experiments show that SymbolicArena achieves 72.6% to 86.7% lower approximation error than alternative selectors, further supporting its fidelity to the Full Task Set. Evaluation reveals a substantial gap between numerical fitting and symbolic recovery across current SR methods, suggesting that reliable equation recovery remains an open challenge.
|
| 1526 |
A Multimodal Autonomic Sensing Framework for Objective Assessment of Patient Responses to Dental Pulp Stimulation
2609.35121
|
cs.LG
|
Youngsun Kong, Yubin Choi, Dongjin Song, Dong-Guk Shin, I-Ping Chen |
Patient responses to dental pulp testing, ranging from no sensation to intense pain, provide important information for assessing pulp status in endodontic diagnosis. However, pain is a subjective sensory and emotional experience that varies considerably across...Patient responses to dental pulp testing, ranging from no sensation to intense pain, provide important information for assessing pulp status in endodontic diagnosis. However, pain is a subjective sensory and emotional experience that varies considerably across individuals and can be difficult to communicate. We investigated whether complementary autonomic signals could support objective assessment of responses during dental examination. Forty-nine patients underwent cold pulp testing, yielding no-response, mild-response, and intense-response conditions. The framework integrated ECG-derived skin nerve activity (SKNA) and R-R intervals (RRI), together with electrodermal activity (EDA), using temporal convolutional network encoders with attention-based mid-level fusion. Individual baseline signals and subject-level covariates, including anxiety scores and biological sex, were also incorporated. The framework achieved 80.2% balanced accuracy, 75.2% sensitivity, and 85.2% specificity for binary classification of no response versus mild or intense response. For three-class classification, it achieved 60.0% balanced accuracy and a 58.8% macro-averaged F1 score. Ablation and attention-weight analyses indicated that EDA contributed most strongly to model performance, followed by RRI, while SKNA improved balanced accuracy by approximately five percentage points. Age was significantly associated with model performance. These findings support the feasibility of multimodal autonomic sensing for objective, non-invasive assessment of responses to dental pulp stimulation.
|
| 1527 |
Explaining Hyperbolic Neural Networks via Geometry-Aware Relevance Propagation
2609.35128
|
cs.LG
|
Ping Xiong, Shanglin Li, Yi Ding, Thomas Schnake, Shinichi Nakajima |
Hyperbolic neural networks introduce geometric operations that require explicit treatment in relevance propagation. Equivalent geometric realizations can produce different feature attributions, even when local relevance is conserved. We study this problem thro...Hyperbolic neural networks introduce geometric operations that require explicit treatment in relevance propagation. Equivalent geometric realizations can produce different feature attributions, even when local relevance is conserved. We study this problem through Geometric Representation Invariance (GRI), a specialization of Implementation Invariance, and zero-curvature consistency, which requires identity relevance propagation when a geometric module approaches the identity. We propose LRP-radial-all for origin-centered radial modules, treating geometric scaling as modulation and assigning relevance entirely to the signal branch. The rule conserves relevance, is invariant to equivalent radial factorizations, and satisfies zero-curvature consistency, yielding GRI for a specified Poincar\'e-Lorentz logarithmic-map construction. In contrast, a conservative LRP-half baseline can violate both consistency criteria. Experiments on hyperbolic MNIST, sEEG, and CIFAR-10 classifiers assess attribution fidelity, qualitative explanations, and runtime. LRP-radial-all achieves competitive attribution fidelity across datasets with runtime comparable to Gradient$\times$Input and substantially lower than Integrated Gradients. These findings motivate geometry-aware propagation rules that distinguish relevance conservation from consistency across equivalent computations.
|
| 1528 |
CTP-FL: Common-Trajectory Gradient Prediction for Federated Learning
2609.35130
|
cs.LG
|
Junkang Liu |
Communication-efficient federated optimization commonly spends several gradient evaluations between server updates. Existing local-update methods use this computation to advance an independent model on each client. Under heterogeneous data, however, these mode...Communication-efficient federated optimization commonly spends several gradient evaluations between server updates. Existing local-update methods use this computation to advance an independent model on each client. Under heterogeneous data, however, these models evaluate gradients at different locations, making the aggregated update difficult to interpret as a gradient of the global objective. We study an alternative use of the same computation budget: \emph{evaluate the global objective along a shared, predicted path}. We propose Common-Trajectory Predictive Federated Learning (\texttt{CTP-FL}). At each round, all clients construct the same sequence of query points from the current global model and the previous aggregated direction, evaluate $K$ stochastic gradients along this sequence, and upload their average. The server then performs a single global update. Thus, \texttt{CTP-FL} uses $K$ mini-batch gradients per client and one model-sized vector in each communication direction, matching the per-round computation and communication of full-participation FedAvg-M. Shared query points make the aggregated direction an unbiased estimator of the average \emph{global} gradient along the predicted path. The remaining discrepancy from the gradient at the current model is controlled by the path length, without assuming bounded client-gradient dissimilarity or bounded gradients. For smooth non-convex objectives, we establish an $\mathcal{O}\!\left( \sqrt{L\Delta\sigma^2/(NKR)}+L\Delta/R \right)$ average-stationarity bound under full participation. The analysis isolates a testable trade-off: extending the prediction path provides more forward-looking gradient information but increases its displacement bias.
|
| 1529 |
FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales
2609.35138
|
cs.LG
|
Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang |
Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to shor...Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.
|
| 1530 |
CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion
2609.35139
|
cs.LG
|
Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer, Haoxuan Xu, Guoliang Xing |
Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can...Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61$\times$ speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.
|
| 1531 |
Small transformers track Bayesian evidence for latent common causes via a context-invariant mechanism
2609.35161
|
cs.LG
|
Amir Mohammadpour, Michael Franke |
We present an in-depth investigation of how a form of Bayesian reasoning about common causes can emerge as a cross-contextual generalization in small, tractable transformers. Incrementing on recent work, our set-up (i) disentangles causal mechanisms in the mod...We present an in-depth investigation of how a form of Bayesian reasoning about common causes can emerge as a cross-contextual generalization in small, tractable transformers. Incrementing on recent work, our set-up (i) disentangles causal mechanisms in the model from the causal structure of the true data-generating process, (ii) orients more towards natural language prediction by considering inference of latent common causes, and (iii) considers whether and how Bayesian evidence accumulation for latent common causes can be implemented in representations and mechanisms that allow for cross-context generalization to novel test cases.
|
| 1532 |
Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion
2609.35166
|
cs.LG
|
Dmitrii Moor, Federico Tomasi, Paul N. Bennett, Alice Wang, Mounia Lalmas |
Discrete diffusion models offer the ability to re-draft, revisiting and correcting earlier tokens throughout generation. This capability depends on the forward corruption process that defines what the denoiser learns to correct. Masked diffusion models fix tok...Discrete diffusion models offer the ability to re-draft, revisiting and correcting earlier tokens throughout generation. This capability depends on the forward corruption process that defines what the denoiser learns to correct. Masked diffusion models fix tokens once they are unmasked, while uniform diffusion permits revisions but relies on uniformly random token substitutions. We instead learn which substitutions are most useful for training the denoiser to re-draft. We introduce Variational Stackelberg Discrete Diffusion (VSDD), a framework for learning a semantically aware corruption process. VSDD formulates training as a leader-follower game: the leader defines a Markovian corruption process parameterized by the denoiser's token embeddings, while the follower optimizes a variational denoising objective with the corruption process held fixed. The leader rewards corruptions based on how much the denoiser improves after learning from them, rather than on how easily the current denoiser can reconstruct them. We measure this improvement under a fixed reference corruption process, approximate the follower's response with a one-step gradient update, and optimize the leader using a score-function estimator. We evaluate VSDD across molecular, text, and playlist generation. VSDD substantially improves molecular validity over uniform and masked diffusion, reduces text perplexity relative to uniform diffusion while remaining competitive with masked diffusion, and achieves sizable improvements in offline playlist recommendation metrics.
|
| 1533 |
EdgeCraft: Automated Model Crafting for Edge IoT
2609.35167
|
cs.LG
|
Genglin Wang, Kaiwei Liu, Liekang Zeng, Wangsong Yin, Shangcheng Jin |
Machine learning (ML) increasingly powers Internet of Things (IoT) applications at the edge. Yet producing a deployable edge ML artifact for a specific scenario requires navigating a huge search space spanning data representation, model design, training on dom...Machine learning (ML) increasingly powers Internet of Things (IoT) applications at the edge. Yet producing a deployable edge ML artifact for a specific scenario requires navigating a huge search space spanning data representation, model design, training on domain-specific data, and runtime customization. This workflow is fragmented and difficult to scale across diverse edge applications. We present EdgeCraft, an LLM-driven system that turns high-level intent into deployable edge ML artifacts. Building such a system raises two challenges: (1) How can an LLM be guided to find high-quality solutions that meet dynamic SLOs for task quality, latency, and energy? (2) How can trustworthy target-device verification be obtained at low cost? EdgeCraft addresses these challenges with two designs. (1) A constraint-aware synthesis tree explores alternative candidates and uses measured SLO gaps to guide each improvement. (2) A multi-fidelity verifier progressively combines low-cost checks with full target-device verification to reduce verification cost while preserving reliable verification results. It also records verified failures for reuse, avoiding repeated device work. To support concurrency, EdgeCraft provides a multi-tenant runtime that runs cloud training and target-device verification in parallel while isolating requests. Across 50 public tasks, EdgeCraft exceeds the task-specific Reference in best-observed quality on 40 tasks and finds an SLO-feasible artifact on 45, with the two outcomes overlapping on 38 tasks. Moreover, EdgeCraft achieves competitive performance on our self-collected SEN dataset, suggesting its generalizability to real-world IoT sensing tasks.
|
| 1534 |
QAM: Quadratic-Accurate Checkpoint Merging via Sequential Consistency
2609.35168
|
cs.LG
|
Shihao Wang, Rui Kong, Xinran Chen, Hui Wu, Qipeng Qian |
Saved checkpoints record states along a training trajectory, but generally do not determine the updates at states that would be visited under a different schedule. We study how accurately these checkpoints can reconstruct the endpoint of a sequential reference...Saved checkpoints record states along a training trajectory, but generally do not determine the updates at states that would be visited under a different schedule. We study how accurately these checkpoints can reconstruct the endpoint of a sequential reference with prescribed update strengths. Under a common local transition model, two checkpoint-index moment conditions characterize all convex merges that agree with this reference through second order. We then prove an information limit that for nondegenerate profiles, no algorithm using only a fixed-length gradient-descent (GD) history with step size $h$ can achieve $o(h^3)$ endpoint error uniformly over a fixed class of smooth, strongly convex losses. The lower bound follows from two losses with identical GD checkpoint histories but sequential reference endpoints separated by $\Omega(h^3)$. \textbf{Quadratic-Accurate Merging} (QAM) achieves a matching uniform $O(h^3)$ endpoint error bound. Its explicit coefficients also define the unique profile-dependent merge that exactly matches the sequential GD reference across all fixed quadratic objectives. Across two public Adam checkpoint trajectories (SmolLM3-3B and OpenEuroLLM-Prelude-9B), three windows and three profiles per model, and 15 tasks, QAM shows mixed results for short windows and broader advantages over \textbf{Warmup-Stable and Merge} (WSM) for longer windows. Matched-moment GSM8K diagnostics further show that local consistency alone does not fully determine downstream scores. These results characterize the reconstruction limits of saved histories, provide a coefficient rule that attains the optimal rate, and assess its practical utility.
|
| 1535 |
ProtoSeam: Lifting Classifier Training with Latent Gaussian Mixture Models
2609.35174
|
cs.LG
|
Robert Lampel, Timon Klein, Sebastian Sager |
We propose a lifted reformulation of supervised classification that improves the final accuracy of standard classifiers without changing the architecture at inference time. A network $N=N_2\circ N_1$ is split at a single semantic interface and one learnable pr...We propose a lifted reformulation of supervised classification that improves the final accuracy of standard classifiers without changing the architecture at inference time. A network $N=N_2\circ N_1$ is split at a single semantic interface and one learnable prototype per class is inserted there. Training combines a quadratic consensus penalty that pulls $N_1(x)$ toward the prototype of its class with a classification loss of $N_2$ evaluated on samples drawn around the prototypes, whereat no gradient crosses the interface. At inference the prototypes are discarded and the unmodified network $N_2\circ N_1$ is used. Across CIFAR-10, CIFAR-100, and TinyImageNet with ResNet and vision transformer backbones, lifted training improves test accuracy by up to five percentage points over variants without lifting under a shared tuning protocol. Moreover, we provide theoretical justification of those results.
|
| 1536 |
Subgroup Rank-1 Lattice for Practical High-dimensional Black-box Integral Approximation
2609.35177
|
cs.LG
|
Yueming Lyu |
Estimating integrals of black-box, high-dimensional functions, from expectations and kernel mean embeddings to the softmax kernel in self-attention, is a basic subroutine in machine learning. Rank-1 lattice rules suit this setting: they query the integrand onl...Estimating integrals of black-box, high-dimensional functions, from expectations and kernel mean embeddings to the softmax kernel in self-attention, is a basic subroutine in machine learning. Rank-1 lattice rules suit this setting: they query the integrand only at a fixed point set and need no gradients. When the $n$ points serve as a design matrix $X\in\mathbb{R}^{n\times d}$ for a feature map, however, computing $\Psi(X)^\top v$ or $\Psi(X)w$ for an elementwise nonlinearity $\Psi$ costs $O(nd)$ time and memory for any standard quasi-Monte Carlo point set. We study subgroup rank-1 lattices, whose Korobov generator $(1,t,\dots,t^{d-1})$ uses a scalar $t$ of fixed multiplicative order $m$. Splitting $\mathbb{F}_n^\times$ into cosets of $\langle t\rangle$ reduces both maps to short cyclic correlations evaluated by FFT, giving exact results for arbitrary $\Psi$ in $O(n\log m)$ time and $O(n)$ memory, without forming $X$. Since fixing $m$ falls outside classical component-by-component theory, we prove convergence directly: via resultants with the cyclotomic polynomial $\Phi_m$, the squared worst-case error in the Korobov space decays as $O(n^{-(\alpha-1)/(m-1)})$ for prime $m\ge d+1$, and this threshold is exact. Using the splitting of $n$ in $\mathbb{Q}(\zeta_m)$, averaging over the $m-1$ admissible generators improves the constant by a factor $\Theta(m-1)$. Empirically, the subgroup lattice beats Gaussian and orthogonal random features and scrambled Sobol' and Halton points in 49 of 54 synthetic kernel-estimation settings and all 45 softmax-attention settings on nine real datasets, and builds a sample set with $d=2048$, $n\approx4.1\times10^7$ in 2.3 ms.
|
| 1537 |
ConRAG: Lightweight inference of multi-hop relations
2609.35193
|
cs.LG
|
Kilian B\"anziger, Sonia Laguna, Markus Kreft, Robert Jakob, Kevin O'Sullivan |
Understanding how two entities are connected often requires tracing multi-hop relations across documents to identify intermediate entities and supporting evidence that explain a connection. This is a task that appears frequently in scientific research and othe...Understanding how two entities are connected often requires tracing multi-hop relations across documents to identify intermediate entities and supporting evidence that explain a connection. This is a task that appears frequently in scientific research and other knowledge-intensive analyses. We formalise this setting as multi-hop relation inference: given two known endpoint entities, we aim to recover the bridge entities and evidence-grounded reasoning chains that connect them across a document corpus, and to generate an explanation grounded in the retrieved evidence. Existing multi-hop RAG systems typically seek an unknown answer entity rather than explicitly recovering the connection between two known endpoints and graph-based approaches often rely on costly LLM-extracted knowledge graphs that limit scalability to large document collections. We introduce ConRAG, which builds a lightweight entity-document graph from entity co-occurrence and LLM-based entity filtering. Its connective retrieval infers and semantically ranks paths between two endpoints. On MuSiQue and 2WikiMultiHopQA, ConRAG consistently improves bridge entity and reasoning chain recovery over strong RAG baselines, while reducing graph-indexing token cost by up to roughly 1.5 orders of magnitude. Our results show that endpoint-constrained path retrieval provides an effective and index-efficient approach to evidence-grounded relation discovery.
|
| 1538 |
Adversarial Consistency-Guided Representation Learning for Multi-view Clustering
2609.35212
|
cs.LG
|
Yuchen Lin, Kunpeng Xu, Ying Fang, Lifei Chen |
Multi-view clustering aims to capture cross-view consistency while exploiting view-specific information. However, shared representations learned to capture cross-view consistency may still retain view-identifying information, potentially compromising the consi...Multi-view clustering aims to capture cross-view consistency while exploiting view-specific information. However, shared representations learned to capture cross-view consistency may still retain view-identifying information, potentially compromising the consistency of cross-view clustering structures. To address this issue, we propose ACGRL, an adversarial consistency-guided representation learning framework for multi-view clustering. ACGRL employs a gradient-reversal view discriminator to reduce view identifiability and obtain invariant reference representations. These representations are then frozen to provide fixed references for disentangling view-specific information from cross-view common information in the subsequent learning stage. The fixed reference representations are concatenated with the learned view-specific representations for reconstruction and clustering, with cross-view cluster alignment encouraging consistent clustering assignments. Experiments on four benchmark datasets demonstrate the superior clustering performance of ACGRL compared with representative multi-view clustering methods.
|
| 1539 |
Temporal Heterogeneous Graph Pretraining for Relational Deep Learning
2609.35219
|
cs.LG
|
Yixin Peng, Er Jin, Diego Collarana, Stefan Decker |
Relational deep learning models database rows and foreign-key links as a heterogeneous graph for prediction from record attributes and relational context. These graphs contain two distinct temporal signals: record age changes with the prediction cutoff, while ...Relational deep learning models database rows and foreign-key links as a heterogeneous graph for prediction from record attributes and relational context. These graphs contain two distinct temporal signals: record age changes with the prediction cutoff, while intervals between observed records remain fixed. Prior work often treats time as a single signal or studies temporal representation and pretraining separately. We investigate how explicitly encoding both signals affects temporal pretraining for downstream tasks. Our framework combines Multi-scale Time Encoding, which captures record age using learnable time scales and type-specific projections, with Rotary Time Encoding, which represents signed inter-record intervals through rotary transformations during graph propagation. We pair these encodings with three self-supervised objectives: historical relation recovery, horizon-aware future relation activity prediction, and temporal subgraph contrast. All inputs respect their observation cutoffs. Pretraining proceeds in two stages: subgraph contrast first learns neighborhood representations, followed by refinement through either relation recovery or future activity prediction. We evaluate on five RelBench datasets across 11 classification and regression tasks using heterogeneous GNN and graph Transformer backbones. With both encodings, the best evaluated staged schedules improve over supervised training with the same encodings by 3.02% and 1.06% on the two backbones, respectively, and over controls without pretraining or either encoding by 3.24% and 2.37%.
|
| 1540 |
Long-Horizon Scaling: How Model Capabilities Shape the Returns to Computation
2609.35236
|
cs.LG
|
Haoyu Zheng, Zhengyu Chen, Huaisheng Zhu, Ruishan Fang, Teng Xiao |
Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To...Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab and EdgeBench, two long-horizon benchmarks. We find that starting performance and subsequent growth are associated with different capabilities: within a task category, similar early scores can precede different later gains. To formalize this finding, we model capability-time scaling with category-specific logistic power laws shared across models. Fitted to early trajectories, these curves extrapolate the observed models' category-average scores to later computation. However, rising average scores mask narrowing improvement opportunities: later gains concentrate among fewer improving models. High final scores and continued improvement also have distinct capability profiles. Predicted mean gains estimate each model's fraction of improving tasks; averaging these estimates forecasts the average share of improving models. These uneven returns motivate deciding whether a specific run should continue. We therefore derive a continuation policy to save time and compute with limited score loss. The policy conditions growth predictions on the run's observed progress and weighs immediate and delayed gains against computation costs. In replay with training and price calibration based on other models' histories, the policy saves roughly one-third of full-run time, with relative score losses of 2.4% on AutoLab individual runs and 3.3% on EdgeBench published mean curves. Our repository is available at https://github.com/Chihaya-Anon-chan/long-horizon-scaling.
|
| 1541 |
Disentangling Lung-Cancer CT/LDCT AI: A Systematic Evidence Map of Clinical Tasks, Evidence Chains, and Translational Gaps
2609.35240
|
cs.LG
|
Surajit Das |
Artificial-intelligence studies using computed tomography (CT) for lung cancer are often broadly labelled "prediction" despite addressing clinically distinct tasks. We systematically mapped CT/low-dose CT (LDCT)-centered lung-cancer AI using five-database retr...Artificial-intelligence studies using computed tomography (CT) for lung cancer are often broadly labelled "prediction" despite addressing clinically distinct tasks. We systematically mapped CT/low-dose CT (LDCT)-centered lung-cancer AI using five-database retrieval, full-text eligibility assessment, role-aware modality/omics extraction, clinical-task classification, and a Multi-Tier Evidence Graph (MTEG). The final corpus comprised 293 studies (2016-2026): 230 Detection, 8 future Risk-prediction, and 55 Other studies. Clinical variables (96.2%), 3D CT/LDCT (73.0%), and radiomics (63.5%) predominated, whereas external validation (29.0%), calibration (20.5%), decision-curve analysis (13.0%), longitudinal CT (17.7%), and saliency/attribution XAI (21.5%) were less frequent. The MTEG comprised 377 nodes and 3,444 edges; only 31 studies (10.6%) completed the six-tier substantive evidence chain, with greatest attrition at reasoning/explanation. Overall, the literature is detection-dominated, genuine future risk prediction remains uncommon, and complete translational evidence chains are rare.
|
| 1542 |
AIM-ZO: Activation-Informed Subspace Maintenance for Zeroth-Order LLM Fine-Tuning
2609.35257
|
cs.LG
|
Yue Xie, Zhi Zheng, Yunpeng Ba, Xuyang Wu, Xialiang Tong |
Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic p...Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly informative directions. To make these evaluations more informative, existing ZO methods restrict perturbations to low-dimensional subspaces. Yet the quality of these subspaces is critical: overly compressed or poorly maintained spaces can miss useful update directions. To obtain a high-quality subspace for ZO updates, this paper proposes AIM-ZO, a ZO fine-tuning method based on Activation-Informed Subspace Maintenance. AIM-ZO uses forward activations as local directional information and continuously integrates them into a broad, evolving subspace over training. To access broader gradient-relevant structure while keeping individual perturbations low-dimensional, AIM-ZO activates only a smaller set of shared and sampled directions, decoupling the maintained width from the active width. We evaluate AIM-ZO across 5 LLMs and 11 downstream tasks under matched forward-evaluation budgets; its six-task average exceeds the strongest fully evaluated ZO baseline by 1.26 percentage points on OPT-2.7B and MeZO by 2.85 percentage points on OPT-30B. Our code is available at https://github.com/EkkoXy/AIM-ZO
|
| 1543 |
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
2609.35259
|
cs.LG
|
Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar |
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, makin...On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.
|
| 1544 |
Latency and accuracy tradeoffs in Spiking Neural Networks
2609.35260
|
cs.LG
|
Zhanglu Yan, Zixuan Zhu, Kaiwen Tang, Yuyang Cai, Qianhui Liu |
Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural netw...Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level, SNNs may complete execution in less time than comparable bit-serial QNNs. However, this overlap relies on spikes firing on incomplete inputs, and a spike once generated cannot be withdrawn, so its error persists and reduces accuracy. Waiting for more input before firing would seem to improve accuracy at the cost of reduced overlap. Yet we find and prove that this intuition fails at some layers, where even a small increase in waiting can change spike timing and downstream computation, making the network both slower and less accurate. We therefore propose a Pipeline Delay Search method which selects each layer's delay by balancing task-level accuracy gains against added network latency. We then adapt the selected configurations through spike-based quantization-aware training and bounded tuning of firing thresholds and initial membrane potentials. Together, these steps form Falcon, a framework for Fine-grained Analysis of Latency and Controlled firing which systematically analyzes and optimizes SNN latency under a spatial analog compute-in-memory mapping with shared digital engines. We evaluate Falcon on GSCV2 and SSC, achieving competitive accuracies of 96.31 and 83.02 at modeled network-core latencies of 119.64 and 124.00us, respectively. Together, our analysis and results show that SNNs can compute more yet finish faster, and wait longer yet predict worse, highlighting why Falcon matters for both latency and accuracy.
|
| 1545 |
SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards
2609.35268
|
cs.LG
|
Yingchao Yu, Pengfei Sun, Wenxuan Pan, Wei Chen, Yitian Hong |
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and...Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we formalize as Temporal Credit Carriers (TCCs) and that spiking neural networks (SNNs) naturally fulfill through graded membrane traces and event-driven spikes. Based on this hypothesis, we propose SpikeCredit, an SNN-based framework for RL with sparse rewards that first performs task-adaptive TCC selection and then closes the loop between a fast TCC-reading pathway, where self-motion feedback constraint uses local behavior-grounded cues to constrain transition-level credit recovery, and a slow TCC-writing pathway, where credit-targeted trace alignment feeds recovered credit back into the actor to make future TCC dynamics more credit-readable. Across sparse-reward MuJoCo tasks, SpikeCredit improves Last10 return over sparse SNN baselines by +1169% on Ant, +953% on Hopper, +723% on Swimmer, and +1781% on Walker2d, and exceeds the dense-reward baseline on Swimmer by +113%. Mechanistic analyses further show substantially stronger alignment with dense rewards than the sparse SNN baseline. These results position spiking dynamics as credit-preserving substrates for sparse-reward RL.
|
| 1546 |
eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models
2609.35269
|
cs.LG
|
Mansi, Nikhil Raghavan, Zixia Huang, Kai Sheng Ong, Ji Shen Lim |
The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We...The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We present eval-unlearn, an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models. eval-unlearn integrates twelve published unlearning techniques spanning fine-tuning, closed-form model editing, and inference-time intervention, alongside nine complementary evaluation metrics covering erasure efficacy, adversarial robustness, generative quality, and concept retention. Its plugin architecture lets third-party techniques and metrics self-register without modifying the core framework, and its streaming, batched pipeline supports efficient evaluation of both standard NSFW concepts and arbitrary general concepts. As a further contribution, we release a public leaderboard on HuggingFace along with an interactive tool for real-time evaluation of unlearning techniques. The leaderboard compares nudity concept erasure case study across all twelve techniques, exposing significant accuracy-quality trade-offs that are obscured by heterogeneous evaluation. eval-unlearn is released under the MIT license; the package, code, leaderboard, and documentation are all available at https://eval-unlearn.readthedocs.io.
|
| 1547 |
Multi-Attractor GNNs: Set-Valued Expressivity Beyond Unique Equilibria
2609.35274
|
cs.LG
|
Jialin Liu |
Recurrent and equilibrium graph neural networks (GNNs) often enforce a unique fixed point or use one training target per graph. Yet many combinatorial and scientific problems admit multiple valid solutions, with no preferred one. A designated target can then i...Recurrent and equilibrium graph neural networks (GNNs) often enforce a unique fixed point or use one training target per graph. Yet many combinatorial and scientific problems admit multiple valid solutions, with no preferred one. A designated target can then impose an arbitrary selection rule. For tasks invariant to node relabeling, a symmetric graph may have a symmetric solution set but no symmetric solution. We show that multiple equilibria enable one weight-tied message-passing GNN to represent set-valued equivariant maps: different initializations approach different valid solutions. Under stated regularity assumptions, we first construct globally Lipschitz, permutation-equivariant dynamics that converge almost surely to valid solutions and reach every solution branch with positive probability. We then establish approximate realization by recurrent message passing with continuous component maps, with arbitrarily small update and limiting errors and arbitrarily high probability. This goes beyond standard universality arguments: although message passing alone cannot distinguish symmetric nodes, the evolving state keeps nodes distinguishable at every finite step without auxiliary node identifiers. Such dynamics can be learned without solution labels using problem-specific energies. On Ising ground states, structural module detection in protein graphs, and chemical reaction steady states, the learned updates produce multiple high-quality predictions with high numerical convergence rates. They achieve better average solution quality than the tested unique-equilibrium, single-target, and feedforward baselines, while remaining competitive with much larger diffusion-based solvers.
|
| 1548 |
Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
2609.35291
|
cs.LG
|
Shunchang Liu, Lukas Fluri, Xin Chen, Francesco Croce |
Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However,...Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However, EM has been studied almost entirely in text-only tasks, leaving its manifestation in multimodal models unclear. In this paper, we define and analyze EM in the context of vision-language models. We first induce EM via fine-tuning on narrow multimodal tasks targeting vulnerable code, careless household-object use, and conspiratorial interpretations of ordinary scenes. Across fifteen commercial and open-source models with different scales, we find that narrow multimodal fine-tuning can induce coherent and broadly misaligned behavior that transfers to unrelated tasks, including misaligned opinions, visual factual dishonesty, unsafe image generation, vulnerability to visual jailbreaks, and risky agentic actions. We further find that multimodal EM does not depend on the apparent harmfulness of training data but is sensitive to training-evaluation modality alignment. EM can arise under both supervised fine-tuning and preference optimization and can propagate through intermediate reasoning. Finally, we explore several mitigation strategies, including prompt inoculation, benign continued training, and activation-level steering, which can partially reduce EM. Overall, our findings suggest that multimodal EM reflects a behavioral shift rather than a general loss of capability, extending beyond text to the visual modality.
|
| 1549 |
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
2609.35297
|
cs.LG
|
Arman Bolatov, Artem Riabinin, Nikita Kornilov, Andrey Veprikov, Samuel Horv\'ath |
Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in ...Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every $P$ iterations and Lion steps in between, with a single dual-EMA momentum buffer shared by both. Muon's compute and communication are paid once per $P$ steps, and the optimizer state is half of AdamW's. A single-EMA variant, SignMuon, already improves on Muon. We prove complexity bounds under heavy-tailed noise in which the period sets an interpolation between Muon's and Lion's smoothness and noise constants, and which say when LionMuon is faster than both. On 124M and 355M models trained on FineWeb, LionMuon with $P=2$ and $P=5$ reaches a lower loss than Muon, AdamW, Lion and Signum at the same number of tokens. Under 4-GPU data-parallel training it reaches Muon's final loss with a third less wall-clock on PCIe, and it beats the communication-efficient Muon variants Dion and MuonBP on loss at no more exposed communication, while keeping the exact gradient. Code: https://github.com/brain-lab-research/lion-muon
|
| 1550 |
When Should a Satellite Estimate Be Changed? Stress-Testing Neural Corrections for Evapotranspiration
2609.35314
|
cs.LG
|
Marco Trotta |
Neural residuals can improve satellite evapotranspiration (ET) estimates, but selectors must predict when a correction helps and reject unsupported inputs. We evaluate ten-member models on 16,366 flux-tower observations from 151 stations paired with OpenET, ac...Neural residuals can improve satellite evapotranspiration (ET) estimates, but selectors must predict when a correction helps and reject unsupported inputs. We evaluate ten-member models on 16,366 flux-tower observations from 151 stations paired with OpenET, across nine rolling years and five spatial folds. At one held-out station, Gain accepted corrections on all 32 physically invalid records: it predicted a mean benefit of 0.83 mm/day, but the corrections increased mean absolute error by 21.6 mm/day versus OpenET. On spatially held-out unit errors, SupportGain reduced station-macro MAE versus Gain by 0.148 mm/day under wind x3.6 (simultaneous 95% interval, 0.070 to 0.226), with 9.3% acceptance versus Gain's 51.8%; on clean inputs, its 0.006 mm/day advantage had an interval that includes zero. These fault analyses are exploratory; none of 40 preplanned temporal comparisons passed Holm correction, while a separate predeclared cropland contrast found 0.041 mm/day lower station-macro MAE with crop-only training (95% interval, 0.009 to 0.079).
|
| 1551 |
Collaborative Principle Evolution via Evidence Transfer for Scientific Discovery
2609.35315
|
cs.LG
|
Yingming Pu, Hongyu Chen, Tao Lin |
Large Language Model (LLM)-based agents promise to automate scientific discovery, yet exploring the vast hypothesis space remains costly. Existing principle-evolution methods accelerate this loop, but operate sequentially, which caps exploration breadth and wa...Large Language Model (LLM)-based agents promise to automate scientific discovery, yet exploring the vast hypothesis space remains costly. Existing principle-evolution methods accelerate this loop, but operate sequentially, which caps exploration breadth and wastes wall-clock time on challenging problems. To address this, we formulate collaborative scientific discovery as evidence transfer between parallel principle-evolution branches. We present COEVOLVE, which realizes this transfer through a coordination core over parallel branches. By integrating value-of-information-gated routing and context-discounted likelihood injection, COEVOLVE enables branches to collaborate through shared measurements while keeping their principle posteriors separate. Across six scientific-discovery tasks under a matched evaluation budget, COEVOLVE attains a mean solution quality of 66.5% versus 57.0% for single-branch principle evolution, with a 1.80x mean wall-clock speedup on the GPT-5.6-Terra backbone; on five auto-research tasks delegated to an autonomous research harness, it is the only arm whose mean stays above the published SOTA anchor on every task. These results establish when evidence sharing accelerates parallel discovery and when transfer safeguards are necessary to limit negative or inert transfers
|
| 1552 |
Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents
2609.35319
|
cs.LG
|
Tong Zhang, Zhou Liu, Yihao Liu, Jiahua Bao, Xuchen Li |
On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linki...On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our empirical analysis reveals a supervision-benefit mismatch: large gaps can be benign, while small gaps can be outcome-critical. Teacher-student gaps capture differences at the current turn, whereas the benefit of teacher guidance depends on how the current student interacts with the environment afterward. The student may still succeed despite choosing an action that differs from the teacher's, while a teacher-preferred action may lead to a state from which the student cannot complete the task. Local gaps alone are therefore not enough to determine whether teacher guidance benefits the current student. Effective supervision should instead emphasize guidance that the current student can translate into better final task outcomes. Accordingly, we propose Outcome-Guided On-Policy Distillation (OG-OPD), which applies trajectory-relative weighting to teacher supervision and calibrates these weights using final task outcomes from paired student continuations. This calibration selectively strengthens supervision on the student's original trajectories at turns where teacher guidance benefits the current student. Across ALFWorld, ScienceWorld, and WebShop, OG-OPD consistently outperforms baselines under diverse settings. It improves task success rates by 3.6-17.7 percentage points over vanilla OPD and by up to 7.0 percentage points over the strongest baseline.
|
| 1553 |
Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data
2609.35322
|
cs.LG
|
J\'er\'emie Klinger, Rapha\"el Urfin, Giulio Biroli, Marylou Gabri\'e |
Score-based generative models generate new samples by integrating a time-dependent drift that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time $t$ with a weightin...Score-based generative models generate new samples by integrating a time-dependent drift that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time $t$ with a weighting schedule $w(t)$. Along the backward dynamics, and for multi-modal distributions, trajectories commit to modes of the target within a narrow time window, the \textit{speciation time}. In this work, focusing on high-dimensional data, we decompose the integrated loss into its single-time contributions and analyze each at fixed signal-to-noise ratio $\Lambda(t)$: we show that $\Lambda(t)$ sets the rate at which each feature of a multimodal target - the mode directions and their relative weights - is acquired during training. Crucially, at high $\Lambda(t)$ all mode directions are acquired together, on a single timescale insensitive to their amplitudes, while the relative weights are not learned at all. Only near the speciation time, where $\Lambda(t)$ becomes of order one, do all features become learnable, each on its own timescale: the weights are acquired jointly with the directions, and the directions at rates set by their relative amplitudes. For models trained on time-integrated objectives, the learning dynamics is then governed by how much of the weighting effectively sits near the speciation time, which provides insights on $w(t)$ design choices. These results follow from an exact high-dimensional analysis of the training dynamics of unbalanced and hierarchical Gaussian mixtures. Numerical experiments on image and human genome haplotype generation recover the predicted hierarchy of learning timescales in more complex settings.
|
| 1554 |
Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation
2609.35333
|
cs.LG
|
Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao |
Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur...Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a dual-component architecture: a Compression Transformer that distills extended interaction histories into compact latent tokens, and an AD Transformer that auto-regressively generates actions using a hybrid context of these compressed memories and recent transitions. By maintaining a fixed-size latent buffer, RAD decouples the effective history length from computational complexity, functionally providing the model with a long-horizon memory. Empirical evaluations across diverse environments demonstrate that RAD matches the asymptotic performance of standard AD with significantly reduced context window sizes, offering a scalable solution for efficient in-context decision-making.
|
| 1555 |
Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation
2609.35335
|
cs.LG
|
Petros Tsialis, Steffen Limmer, Tobias Rodemann, Martin Heckmann |
Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction tas...Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625 public tabular and time series datasets. We evaluate the proposed approach in three settings: (i) tabular datasets in comparison with established AutoML heuristics, (ii) cross-domain evaluation across tabular and time series datasets, and (iii) a practical deployment scenario using smaller local models. The results show consistent advantages for LLM-based task type identification, with increasing difficulty in heterogeneous and resource-constrained settings. LLM-based approaches outperform AutoGluon in the tabular setting, reaching 0.98 F1 macro compared to 0.93. In the cross-domain setting, the best model achieves 0.90 F1 macro, while smaller locally deployable models reach 0.75, indicating a trade-off between deployment feasibility and accuracy.
|
| 1556 |
NeuronDiscover: Agent-in-Twin for Mechanistic Discovery in Neuronal Microenvironments with World Action Models
2609.35338
|
cs.LG
|
Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie |
Mechanistic discovery in neuronal microenvironments requires interventions and measurements that separate competing explanations of solute transport and neuronal response. Predictive accuracy cannot settle the question: a real mechanistic change and an error i...Mechanistic discovery in neuronal microenvironments requires interventions and measurements that separate competing explanations of solute transport and neuronal response. Predictive accuracy cannot settle the question: a real mechanistic change and an error in the computational twin leave the same signature in sparse observations. We formalize this twin confounding and reason over a joint mechanism--discrepancy belief, designing experiments that separate the two. NeuronDiscover is an Agent-in-Twin framework whose shared, mechanism-grounded World Action Model (WAM) couples prediction, intervention proposals, and observation design; independently adjudicated outcomes revise a scoped Mechanism--Intervention--Observation--Outcome (MIOY) graph, whose supported relations compile into executable programs carrying discrepancy-adjusted acceptance bounds. We evaluate on simulated brain-fluid tracer-transport worlds adjudicated by an independently frozen finer-mesh reference solver, and on donor-disjoint public current-clamp recordings of cortical neurons. Counting only relations that reach a certified terminal status, and scoring abstentions as unresolved for every method, at a matched budget of 16 experiments over 32 source units NeuronDiscover resolves 4.0 relations per assigned world against 3.4 for the strongest baseline and 3.2 without graph revision, at 5% false support and 82% scope accuracy. Joint mechanism--discrepancy acquisition resolves 3.8 relations versus 2.9 for plug-in expected information gain; discrepancy-adjusted verification lowers accepted-program failure from 15% to 9% at 60% acceptance coverage; and transfer to the recordings yields 1.94 versus 1.53 relations per assigned world. Correctness is adjudicated within declared model worlds and archival recordings.
|
| 1557 |
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
2609.35347
|
cs.LG
|
Xin Li, Hao Jiang, Xin Gao, Annan Wang, Yuchen Xie |
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) ...Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
|
| 1558 |
From Data to Program: Fast & Direct Generative Program Inference from Empirical Data
2609.35348
|
cs.LG
|
Simon Kl\"uttermann, Xueying Ding, Leman Akoglu |
Estimating probability densities from a finite set of samples typically requires dataset-specific model fitting. We introduce PRODiGI, a pretrained data-to-program model that infers an explicit, executable generative program in a single forward pass. Pretraine...Estimating probability densities from a finite set of samples typically requires dataset-specific model fitting. We introduce PRODiGI, a pretrained data-to-program model that infers an explicit, executable generative program in a single forward pass. Pretrained on synthetic datasets paired with their ground-truth programs, PRODiGI accommodates diverse generative families and data dimensionalities through template prediction and non-autoregressive program parameter decoding. Its inferred programs support direct sampling, density and score evaluation, and inspection independently of the pretrained model. We further introduce program-space fine-tuning, which refines differentiable program parameters by matching generated and empirical samples while keeping model parameters intact. Experiments show that PRODiGI achieves lower average density and score MAE than existing pretrained models, while offering multi-fold speedups over its closest competitors. Program-space fine-tuning further reduces generation MMD by 84%. By turning empirical data into explicit, reusable programs, PRODiGI introduces a new direction for fast, interpretable tabular generative modeling.
|
| 1559 |
Quasi Linear Kernel Attention with Infinite Capacity
2609.35349
|
cs.LG
|
Nicolaj Rux, Johannes Hertrich, Sebastian Neumayer |
The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivi...The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivity of attention while enabling quasi linear computation. To quantify expressivity, we introduce a capacity for each kernel, measuring the maximum sequence length for which the attention matrix can approximate the identity. A higher capacity thus indicates greater expressivity. We show that expressive kernels like softmax, Gauss, and Laplace have infinite capacity. In contrast, common quasi linear kernels, such as those derived from finite dimensional feature maps, exhibit finite capacity. As a solution, we propose additive kernels constructed from univariate spline and polynomial exponential kernels. We prove that these maintain infinite capacity while allowing quasi linear computation via sorting. Finally, we implement additive sorting kernels efficiently and benchmark them against modern softmax backends, demonstrating advantages for long sequences.
|
| 1560 |
Interference Beyond Geometry in Concept Extraction
2609.35351
|
cs.LG
|
Val\'erie Costa, Bahareh Tolooshams |
Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions, distinguishing constructive from destructive interference...Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions, distinguishing constructive from destructive interference and frequent weak interactions from rare strong ones. Under local fixed-support assumptions, we characterize how architectural constraints shape interference through four mechanisms: feature orthogonalization, bias compensation, gain adaptation, and encoder-decoder separation. Experiments with sparse autoencoders show that constrained architectures selectively reduce overlap among co-active features, while bias, gain, and encoder freedom allow constructive cross-contributions to remain. Together, these results show that interference in learned representations depends not only on feature geometry, but also on how features are used and on the architecture that produces their codes.
|
| 1561 |
Fiona: Accelerating FHE Inference with Packing-Aware Ternary Weights
2609.35352
|
cs.LG
|
Yiteng Peng, Zhibo Liu, Dongwei Xiao, Shuai Wang |
Fully homomorphic encryption (FHE) enables neural network inference directly on encrypted inputs, but it remains orders of magnitude slower than plaintext in- ference. Applying the server's plaintext weights to encrypted activations involves plaintext-cipherte...Fully homomorphic encryption (FHE) enables neural network inference directly on encrypted inputs, but it remains orders of magnitude slower than plaintext in- ference. Applying the server's plaintext weights to encrypted activations involves plaintext-ciphertext multiplications (PMult) and accounts for more than half of inference time in recent systems. Ternary quantization can replace these multipli- cations with additions and subtractions, but the savings rarely materialize under packed execution. A single PMult applies a weight group fixed by the packing layout and can be avoided only when all its weights share the same ternary value. Ternarizing all groups, however, largely degrades accuracy. We present FIONA, an offline optimizer that selectively ternarizes weights within a given packing layout based on the estimated effect of ternary conversion on the model's performance. FIONA encourages a shared ternary value within each weight group and retains full-precision weights for sensitive groups, so ternar- ized and full-precision paths coexist within a layer. It then compiles these hybrid operators exactly, applying common scaling factors once to accumulated inputs and reusing sums across outputs. Weight ternarization can also narrow the input ranges of downstream polynomials. FIONA fits lower-degree replacements under a cumulative accuracy budget, reducing multiplicative depth and bootstrapping. On VGG11, ViT, and BERT, FIONA reduces PMult operations by 53.4-79.5% and accelerates end-to-end encrypted inference by 2.38x, 1.68x, and 1.84x, re- spectively, with less than 1% accuracy loss across all three models.
|
| 1562 |
d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models
2609.35362
|
cs.LG
|
Ruitao Liu, Qinghao Hu, Song Han |
Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a p...Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. Rather than training such models from scratch, recent work adapts strong pretrained AR models into block dLLMs through distillation. On-policy distillation (OPD) has been widely used for LLM training because it supervises the student on states generated by its current policy, rather than only on fixed offline trajectories. By training on the states the student actually visits, it reduces the mismatch between training and generation and can provide more relevant supervision as the student evolves. Recent work has extended this idea to AR-to-block-diffusion conversion. However, this setting introduces a fundamental mismatch in supervision: the block-diffusion student and the causal AR teacher condition on different information at the same training state. The student predicts from the entire partially denoised block, including visible future context, whereas the standard AR teacher target is defined only from the causal prefix. As a result, the teacher distribution used for distillation is not fully aligned with the information available to the student. We therefore introduce d-OPD, a future-aware on-policy distillation method that corrects the AR teacher distribution to better align with the student-visible state by incorporating visible future information within each block, providing supervision that better matches the information used by the student. Across Qwen3 models from 0.6B to 8B, d-OPD improves the six-benchmark average by up to $4.0$ points over OPDLM and reduces training time by $1.35$-$1.58\times$. The code is available at https://github.com/mit-han-lab/d-OPD.
|
| 1563 |
Do Temporal Link Predictors Need Learned Memory? A Smoothed-Count Baseline with a Handful of Parameters
2609.35364
|
cs.LG
|
Lisi Qarkaxhija, Ingo Scholtes |
Many temporal link predictors summarize past interactions through learned node representations. We examine whether simple counts of recurring interaction patterns can provide competitive predictions without learning these representations. We propose a temporal...Many temporal link predictors summarize past interactions through learned node representations. We examine whether simple counts of recurring interaction patterns can provide competitive predictions without learning these representations. We propose a temporal link predictor based on statistical language modelling. It pools transition and co-occurrence counts across sources to predict links that a source has never formed. We smooth sparse estimates using destination frequencies or Kneser-Ney continuation counts. A shared log-linear rule combines these estimates with popularity, source history, and recency, without node embeddings. In our main evaluation, the model achieves the highest MRR among the compared methods on 7 out of 16 datasets from TGB and TGB-Seq. It also outperforms EdgeBank and Base3 on all 16 datasets and the heuristic family on 14. These gains extend to datasets designed to limit repeated edges. With only 9--13 learned parameters, our model provides a simple and competitive baseline for evaluating future neural temporal link predictors.
|
| 1564 |
First Learn, Then Memorize: The Spectral Bias of Diffusion Models
2609.35377
|
cs.LG
|
Rapha\"el Urfin, Tony Bonnaire, Giulio Biroli, Marc M\'ezard |
Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dy...Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample ($m$ noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for $m=1$. The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size $n$. We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear ($n \asymp d$) and polynomial ($n \asymp d^k$) sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank $r$ tunes the generalization--memorization transition, and an $L_2$ penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.
|
| 1565 |
Identifying Neural Source Dynamics from Unknown Local Interventions
2609.35379
|
cs.LG
|
Ayana Mussabayeva, Jiaqi Sun, Anuar Aimoldin, Olivier Oullier, Kun Zhang |
Electroencephalography (EEG) records mixtures of brain-source activity. Even with a known anatomical forward model, experiments that excite only part of the source-state space leave the dynamics unidentified, and repetition cannot resolve the ambiguity. We sho...Electroencephalography (EEG) records mixtures of brain-source activity. Even with a known anatomical forward model, experiments that excite only part of the source-state space leave the dynamics unidentified, and repetition cannot resolve the ambiguity. We show that unknown local mechanism changes can supply the missing information. We consider linear dynamics among fixed anatomical sources with known source-state initialization patterns. Changing one source's update rule for one transition leaves a rank-one, source-specific signature in subsequent EEG: subtracting matched baseline responses isolates it, and the forward model identifies the source and calibrates its response history. Combining these histories with initialization responses recovers source interactions without baseline reachability and without first identifying the intervention coefficients. We establish sufficient recovery conditions, a direct estimator, and a noise-sensitivity bound conditional on correct source labels. Simulated EEG on anatomy derived from magnetic resonance imaging confirms the information gain: with baseline excitation confined to four of twelve source coordinates, eight unknown changes recover all dynamics in 32/32 systems, whereas baseline realization, baseline regression through an invertible forward model, and changes that leave the tested states unexposed all fail, and explicitly constructed alternative dynamics reproduce every baseline mean. Where baseline information suffices, direct reconstruction is also more reliable than a matched-information spectral estimator. Nonlocal changes and forward-model error limit accuracy even when source labels are correct.
|
| 1566 |
Inductive Feedback for Mixed-Policy Distillation
2609.35390
|
cs.LG
|
Amir Moeini, Huaijiang Zhu, Daniel Havir, Shangtong Zhang |
Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition th...Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the student to match the teacher's predictions on student-generated rollouts. However, this approach can transfer teacher preferences that the feedback did not motivate, while leaving much of the feedback's guidance unused. We find that both problems come from the standard on-policy distillation objective, specifically the divergence it minimizes and the distribution it uses as its target. Our proposed method addresses both limitations. First, to isolate the information conveyed by the feedback from the teacher's inherent preferences, we treat verbal feedback as evidence for or against the hypothesis that a particular token comes next at a given prefix. We then adopt a probabilistic confirmation framework which uniquely determines an ordering over the vocabulary based on the teacher's predictions before and after it receives feedback. Using a confirmation score consistent with this ordering, we construct a target distribution within a trust region of the student. Second, to learn from guidance that student rollouts can leave unused, we derive a simple shared-rollout estimator of a symmetric divergence between the student and target distributions over rollouts, reusing student and feedback-conditioned teacher rollouts in both directions through importance weighting. Empirical evaluations show that our method outperforms the common on-policy distillation recipe and a recent contrastive variant on knowledge-based and agentic benchmarks.
|
| 1567 |
The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
2609.35392
|
cs.LG
|
Yihe Zhou, Tongtian Zhu, Yingxiao Huo, Satya Prakash Dash, Can Wang |
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$\beta...Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$\beta$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-$\beta$ Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.
|
| 1568 |
Persistent Partners Raise Prices Among Learning Agents
2609.35402
|
cs.LG
|
Paul-Peter Arslan, Yubin Kim, Xiao Xiao |
When pricing agents meet repeatedly on a platform, the platform decides who faces whom. We ask whether that choice moves the prices the agents learn, and whether a rise comes with learned punishment. In a pre-registered randomised experiment in the Bertrand du...When pricing agents meet repeatedly on a platform, the platform decides who faces whom. We ask whether that choice moves the prices the agents learn, and whether a rise comes with learned punishment. In a pre-registered randomised experiment in the Bertrand duopoly of Calvano et al., each agent's price is set by a tabular Q-learning module, not by the small language model attached to it, and we randomise whether each agent keeps its partner, sees its rival's prices and can send messages. Keeping the same partner raises the level of profits, averaged over training, by 0.27 of the gap between competitive and monopoly profit (95% CI 0.20 to 0.35, all twenty paired runs positive), our registered primary result, and the resting price by 0.17 of the Nash-to-monopoly range (post hoc). A plain tabular learner reproduces the effect in all 25 further blocks, and there one permanent partner raises the level more than about three do (+0.23 against +0.05, exploratory). Where rival prices are hidden, the price-setting module cannot see a cut, so cannot punish it, yet the resting price rises as much and the rise lasts to the end of training, while with visible rivals it shrinks with longer training (post hoc). Where the rival is visible, a static best responder accounts for a third to a half of what a forced-deviation probe reads as punishment, on the starts where the rival can see the cut, and net of it the registered test of learned punishment is inconclusive. A test that looks only for punishment would thus miss the rise where the rival is hidden, while a check for profitable deviations flags most of those prices (post hoc). In an exploratory extension, untrained Qwen2.5 7B and 14B models under one prompt show the effect when the rival's price is left out of the prompt and inconsistently when it is shown, the 7B result replicating on fresh blocks, while two other model families show none.
|
| 1569 |
ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
2609.35433
|
cs.LG
|
Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh |
Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clip...Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an $\alpha$-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.
|
| 1570 |
SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning
2609.35440
|
cs.LG
|
Bojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li |
Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update lock...Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict the target through its own readout, but has not scaled to billion-parameter pretraining. We identify these private readouts as a key weakness, since they leave each module without information from deeper modules. We propose Shared-Output LOcal learning (SOLO), which replaces them with a shared, read-only copy of the final module's readout, the only one trained on the output of the whole network. Taken from the previous step, the copy transmits information from the final module without passing gradients between modules or reintroducing update locking. SOLO approaches backpropagation on Transformers of 340M to 2B parameters pretrained on 15B tokens, staying within one point in average zero-shot accuracy with a perplexity gap that narrows with scale. Readout ablations attribute SOLO's improvement over private readouts to sharing. Without update locking, each of p pipeline stages holds activations for O(1) micro-batches instead of O(p). The freed memory permits larger micro-batches, which reach up to 1.44x the best measured throughput of pipeline backpropagation on the same partition. To our knowledge, SOLO is the first local learning method to show such memory and throughput gains in billion-parameter language-model pretraining. Local learning thus becomes a practical alternative to backpropagation for large-scale pretraining.
|
| 1571 |
Riccati State Space Models: Non-iterative Parallelization for Nonlinear Sequence Modeling
2609.35441
|
cs.LG
|
M\'onika Farsang, Ramin Hasani, Daniela Rus, Radu Grosu |
State space models (SSMs) achieve efficient sequence processing because their affine state updates are closed under composition and can therefore be evaluated with an associative parallel scan. Nonlinear recurrent models can provide richer, state-dependent dyn...State space models (SSMs) achieve efficient sequence processing because their affine state updates are closed under composition and can therefore be evaluated with an associative parallel scan. Nonlinear recurrent models can provide richer, state-dependent dynamics, but generally lose this compositional structure: parallel evaluation then requires iterative methods that repeatedly linearize and scan the recurrence. We ask, what state-dependent nonlinear dynamics can be designed to remain exactly composable? We answer by introducing RiccatiSSM, a nonlinear SSM, in which each state dimension follows an input-conditioned Riccati differential equation. Its quadratic state dependence makes the local Jacobian explicitly state-dependent, while its exact per-step flow under piecewise-constant inputs is a M\"obius transformation. Since M\"obius maps are closed under composition and compose through $2\times 2$ matrix multiplication, the complete nonlinear state trajectory can be evaluated exactly with a single associative parallel scan, without iterative linearization. We further derive a constrained parameterization that ensures bounded, contractive dynamics, and avoids poles in the fractional-linear state update. Across long-sequence classification, regression, and forecasting tasks, RiccatiSSM achieves competitive predictive performance while reducing runtime by $22{-}33\%$ compared to the nonlinear LrcSSM under matched architectures. These results demonstrate that state-dependent nonlinear dynamics can retain exact composability and be evaluated efficiently within a single parallel scan.
|
| 1572 |
NeuronSifter: Intervention Planning in CNS Microenvironments
2609.35445
|
cs.LG
|
Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie |
Prioritizing central nervous system (CNS) interventions requires predicting how a dose, route, and schedule act on a partially observed microenvironment, then choosing the measurement that would change the decision. Action-conditioned predictors reduce a regim...Prioritizing central nervous system (CNS) interventions requires predicting how a dose, route, and schedule act on a partially observed microenvironment, then choosing the measurement that would change the decision. Action-conditioned predictors reduce a regimen to an identity token or a scalar exposure, discarding where and when the target is engaged; handing a point estimate to a separate planner then discards the joint uncertainty that makes a measurement worth running. We therefore treat decision quality as a property of the intervention interface, not of controller placement. NeuronSifter compiles regimens into state-conditional target-occupancy fields with support masks, propagates them through microenvironment dynamics with an occupancy-conditioned diffusion operator, and selects measurements by their expected reduction in intervention loss, assimilating typed outcomes into the same posterior. In a declared synthetic Alzheimer's disease (AD) evaluation over 64 paired scenario blocks, occupancy conditioning lowers trajectory continuous ranked probability score from 0.165 to 0.110 and raises intervention ordering accuracy from 0.760 to 0.880, and every paired benchmark contrast remains separated after Holm correction. Decision-directed acquisition attains terminal risk 0.160 against 0.166 for a matched numerical Bayesian experimental design planner, and reaches the target risk at 0.796 $[0.732,0.873]$ of an earlier design control's cost, while the corresponding ratio against the matched planner, 0.963 $[0.907,1.025]$, is not separated from equality; point-state and dependence-ablated interfaces instead raise risk to 0.220 and 0.199, and a full-posterior external controller ties exactly. Published AD trials supply a separate retrospective endpoint bridge.
|
| 1573 |
Manifold-Stable Flow Matching
2609.35454
|
cs.LG
|
Amirhossein Nazerian, Ali Pezeshki, Jianguo Zhao |
Flow matching (FM) learns generative dynamics through velocity regression. Geometric FM variants commonly assume a prior supported on the data manifold, requiring geometric knowledge that is often unavailable. Without such knowledge, low regression error alone...Flow matching (FM) learns generative dynamics through velocity regression. Geometric FM variants commonly assume a prior supported on the data manifold, requiring geometric knowledge that is often unavailable. Without such knowledge, low regression error alone does not guarantee manifold adherence. Adherence keeps generated samples within valid configurations and is empirically associated with better task performance. We introduce manifold-stable flow matching (MSFM), which can start from an arbitrary ambient prior, not necessarily supported on the manifold. Using tools from nonlinear dynamics, namely contraction theory, MSFM combines learned tangential transport with prescribed normal contraction. The construction uses analytical projectors for known manifolds and local affine proxies estimated by principal component analysis for unknown data geometry. By implementing contraction theory in both cases of known and unknown manifolds, we guarantee manifold invariance and transverse convergence to the manifold within a desired time window (e.g., one second). We derive a family of compatible probability paths and decompose the training loss into a learnable tangential term and a normal residual. An ellipse experiment attains a mean terminal off-manifold error of order $10^{-6}$. In Push-T robotic experiments, MSFM raises success from $74\%$ to $82\%$. In the Robomimic Square task, success increases from $60\%$ to $72\%$, while rotation-manifold deviation decreases from order $10^{-2}$ to $10^{-7}$. The MSFM terminal geometric errors are controlled by the chosen numerical tolerance. These results demonstrate stronger geometric adherence and higher observed task performance, supporting prescribed normal contraction as a complement to learned generative transport.
|
| 1574 |
Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter
2609.35465
|
cs.LG
|
Pier-Jean Malandrino |
Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of co...Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of code. We present Tetra, a new codebook on the same lattice. A 24-weight block still takes 48 bits, most of which index a 64-state trellis of the Golay code and one shared 16 KiB table. The kernel decodes a block with six table loads and two small lookups inside the matrix-vector product, and reads 2.148 bits per weight. For full models, we retrain one scale per matrix row, store the matrices that lose the most as 4-bit integers, and pay for them with 4-bit embedding tables. Our Qwen3-4B, 8B and 14B files hold 2.73, 2.70 and 2.73 bits per parameter over the whole model. They score 63.37, 69.58 and 75.66 on the full MMLU test set, 4.76, 4.21 and 2.46 points below 4-bit AWQ at 5.3 to 6.0 bits per parameter. They generate 113.8, 95.0 and 57.2 tokens per second in our engine. On GSM8K, through the served kernel, they lose 9.63, 4.62 and 3.26 points to FP16. At 4B our file scores 23.6 points above llama.cpp's IQ2_XXS (2.48 bits per parameter). Every number we measured for a table or figure comes from one NVIDIA L40S GPU. We preregistered the main experiments.
|
| 1575 |
An analysis of Mirror-Descent Soft Actor-Critic
2609.35466
|
cs.LG
|
Denis Zorba, Michal Valko |
Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the...Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the target policy arises from policy mirror descent and compare it with the classical Gibbs target. We derive sufficient conditions for the strong convexity and smoothness of the actor objective, characterised by the curvature of the $Q$-function estimate through the Legendre differential operator, and establish an $\mathcal{O}\!\left(N^{-\frac{1}{5}}\right)$ best-iterate finite-time convergence rate up to actor and critic approximation errors. Moreover, the mirror-descent step size $\lambda$ directly controls the target drift and hence actor tracking error, whereas the analogous Gibbs bound contains a non-vanishing tracking term.
|
| 1576 |
Universal Approximation of Measure-to-Measure Operators by Pushforwards
2609.35483
|
cs.LG
|
Takashi Furuya, Nicholas H. Nelsen, Frank Cole |
Many learning tasks map an input distribution to an output distribution. A natural way to model such an operator is to transform each input sample using a continuous function that may depend on the entire input distribution, and then take the distribution of t...Many learning tasks map an input distribution to an output distribution. A natural way to model such an operator is to transform each input sample using a continuous function that may depend on the entire input distribution, and then take the distribution of the transformed samples. This defines a measure-dependent pushforward model and includes measure-theoretic formulations of transformers. We ask when such models can approximate arbitrary continuous operators between spaces of probability measures. We first show that universal approximation fails when atomic inputs are allowed: some continuous measure-to-measure operators that split or redistribute atomic mass cannot be approximated arbitrarily well by deterministic pushforward models. We then introduce the uniform level set condition, which requires a continuous measure-dependent scalarization whose shrinking level set neighborhoods carry uniformly vanishing mass over the input family. This condition is satisfied, in particular, by compact families of absolutely continuous measures. On every compact family satisfying this condition, we prove that any continuous measure-to-measure operator with outputs of finite $p$-th moment can be uniformly approximated, in the $p$-Wasserstein distance, by continuous measure-dependent pushforwards. Combining our theorem with existing approximation results for measure-dependent in-context maps yields universal approximation by measure-theoretic transformers. We also extend the framework to continuously-varying source measures, yielding a corresponding universality result for a class of pushforward models that are closely aligned with cross-attention architectures.
|
| 1577 |
Physics-Guided Conditional Diffusion Model for Rare Event Synthesis and Diagnosis for the Water-Gas Shift Reaction
2609.35499
|
cs.LG
|
Md Abrar Rafid Siddique, Bibek Aryal, Qiugang Lu |
As the world moves towards sustainable energy sources, hydrogen (H2) can be treated as an eco-friendly alternative to fossil fuels due to its high energy density and zero carbon emissions. The water-gas shift (WGS) reaction is a widely used industrial process ...As the world moves towards sustainable energy sources, hydrogen (H2) can be treated as an eco-friendly alternative to fossil fuels due to its high energy density and zero carbon emissions. The water-gas shift (WGS) reaction is a widely used industrial process for hydrogen production by converting carbon monoxide and steam into hydrogen and carbon dioxide. However, occurrences like severe fouling, catalyst deterioration, and thermal runaway can hamper the reaction kinetics/process safety and decrease the yield of H2. These incidents are rare, and gathering process data under such abnormal conditions is challenging. In this work, we propose a physics-guided conditional diffusion model to generate realistic rare-event trajectories for the WGS reaction. The proposed model integrates a conditional denoising diffusion probabilistic model (CDDPM) with governing laws of the reaction to generate physically consistent process trajectories. The conditioning features allow the model to produce high-quality synthetic profiles for rare-event domains that are typically beyond the training regimes. The generated rare-event trajectories then augment the raw dataset for a balanced distribution between normal and abnormal conditions. We further propose a hazard score to assess the risk severity of the operating condition based on the operating trajectory. Deep learning models are trained with the augmented dataset to diagnose the health status of the reaction. Simulation results show that the proposed physics-guided diffusion model outperforms data-driven models in terms of the quality of synthetic data and diagnosis performance for rare events.
|
| 1578 |
Structured Latent Modeling for Supervised Multimodal Information Decomposition
2609.35502
|
cs.LG
|
Wanting Huang, Sanvesh Srivastava, Weiran Wang |
Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer ...Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or masked objectives at intermediate layers, coupled with source-wise invertible normalizing flows and a supervised, low-rank latent variable model. This architecture explicitly factorizes the joint distribution into shared task-relevant variation, modality-specific predictive variation, and task-irrelevant dependence. Drawing connections to prior multimodal learning assumptions, our approach evaluates how modalities independently and jointly contribute to the target. Ultimately, this framework unites intermediate representation learning with structured likelihood-based guidance, offering a practical latent-variable lens for characterizing continuous multimodal interactions. Empirically, we demonstrate the effectiveness of our approach across diverse multimodal benchmarks, showing robust improvements in predictive performance.
|
| 1579 |
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
2609.35505
|
cs.LG
|
Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang |
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillati...We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp $\tilde{\mathcal O}(\log K)$ regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.
|
| 1580 |
Improving Generative Model Self-Training with Geometrically Modified Outputs
2609.35512
|
cs.LG
|
Patrick Batsell, Thomas Walker, Richard Baraniuk |
Self-training generative models - the continued improvement of a model using its own outputs - is becoming increasingly important as high-quality training data becomes scarce. However, naively finetuning on model-generated samples leads to degradation through ...Self-training generative models - the continued improvement of a model using its own outputs - is becoming increasingly important as high-quality training data becomes scarce. However, naively finetuning on model-generated samples leads to degradation through model collapse and the model autophagy disorder. Negative-guidance self-training methods turn this degradation into a useful signal, using a model finetuned on its own outputs to guide the original model toward improved generation. Existing methods, however, take the negative signal in standard model outputs as given. We instead ask whether this signal can be explicitly strengthened. We introduce Geometrically Modified Outputs (GMOs), which reweight the singular values of the generator's input-output Jacobian to increase the influence of its leading singular directions. This geometric modification amplifies the mode-seeking behavior and distortions of standard outputs, providing a stronger and more targeted negative signal for self-training. Across a range of one-step generative models, GMOs consistently improve the performance of negative-guidance methods, including Neon and SIMS, compared with using standard model outputs.
|
| 1581 |
One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices
2609.35514
|
cs.LG
|
Ruishuo Chen, Weijia Li, Xun Wang, Yu Chen, Leheng Cai |
In ecology, psychometrics, and the analysis of social and financial networks, binary matrices are often analyzed conditional on their observed row and column sums, which restricts the problem to a finite sample space of matrices with the same margins. Two fund...In ecology, psychometrics, and the analysis of social and financial networks, binary matrices are often analyzed conditional on their observed row and column sums, which restricts the problem to a finite sample space of matrices with the same margins. Two fundamental problems are to count this space and to sample uniformly from it. Sequential importance sampling (SIS) addresses both with independent weighted samples and an unbiased count estimator, but its efficiency depends critically on the proposal distribution. Existing proposals are analytically designed, and their accuracy can vary substantially with the margins. We show that the ideal SIS proposal, under which every weight equals the count and the variance vanishes, is exactly the policy of a generative flow network (GFlowNet) with unit reward on every matrix that has the given margins. We therefore propose MarginFlow, a framework that turns the design of the proposal into a learning problem and amortizes it across margins by exploiting their self-similarity. Every partial matrix is itself an instance with reduced margins, so one set transformer that reads the remaining margins serves every margin. We train MarginFlow on a pool of 1904 margins and evaluate it zero-shot on 1190 held-out margins, synthetic and real, from $3\times3$ to $870\times6$. On 1187 of the 1190 margins it matches or beats the best of 31 analytically designed configurations, chosen post hoc for each margin, and its median effective sample fraction is 99.8%. On the 56 margins where that best loses more than one nat of effective sample size, MarginFlow wins every one and raises the median effective sample fraction from 10.3% to 94.1%.
|
| 1582 |
Reward-Aligned Reweighting for On-Policy Distillation
2609.35517
|
cs.LG
|
Haofeng Xu, Junwei Su, Lansong Diao, Wenchao Zhou, Chuan Wu |
On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a pro...On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R$^{2}$-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R$^{2}$-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.
|
| 1583 |
Deep Epistemic Value Functions for Optimistic Exploration
2609.35525
|
cs.LG
|
Leander Diaz-Bone, Marco Bagatella, Jonas H\"ubotter, Andreas Krause |
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and p...Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.
|
| 1584 |
GeoGAE: Scalable Graph-Level Autoencoding via Hyperball Cloud Representations
2609.35527
|
cs.LG
|
Rados{\l}aw Nowak, Anna Bielawska, Bogusz Stefa\'nczyk, Maciej Sanocki, Pawe{\l} Wawrzy\'nski |
Embedding structured objects into Euclidean spaces has enabled a wide range of successful machine learning applications. Such objects include words, documents, image patches, time series, and graph nodes. In contrast, embedding entire graphs remains a challeng...Embedding structured objects into Euclidean spaces has enabled a wide range of successful machine learning applications. Such objects include words, documents, image patches, time series, and graph nodes. In contrast, embedding entire graphs remains a challenging problem. Existing methods either sustain the original order of the graph nodes or match the output nodes to the input ones, both of which create scalability issues. In this work, we propose a graph representation as a cloud of hyperballs, which allows us to define a specific, typically unique, node ordering. Based on this representation, we propose GeoGAE, an autoencoder, in which the Transformer encoder translates a hyperball cloud into a graph-level embedding, and the Transformer decoder translates the graph-level embedding back into the graph. This formulation enables the model to capture both the global graph structure and local relational patterns. We evaluate our method on multiple graph datasets, spanning various domains. The results demonstrate effectiveness of our method in encoding and reconstructing graphs from their embeddings.
|
| 1585 |
Let the Neurons Die: Exploiting ReLU-Induced Model Degradation
2609.35528
|
cs.LG
|
Kexin Li, Wenjun Qiu, Joshua Abraham, Aditi Maheshwari, David Lie |
Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availab...Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availability attacks based on data ordering and poisoning. We begin with the basic dynamic data-ordering attack (DOA), which greedily constructs a training prefix by selecting the next example that minimizes the target layer's post-update weight sum, aiming to push ReLU units toward negative pre-activations without modifying training samples or labels. We then develop two poisoning attacks, IG-DOA and IG-SKA, which use gradient inversion to synthesize class-conditioned samples by matching reference gradients in adverse model states constructed through data ordering or soft knockout, respectively. Soft knockout rearranges weights across adjacent layers to concentrate negative contributions. On a fully connected ReLU network trained on MNIST, ordering 100 of 60,000 training examples reduces test accuracy from 96% to 95% after only five epochs. Adding 200 poisoned samples from a single class reduces test accuracy to approximately 86-88% after five epochs in most evaluated conditions, compared with approximately 96% under clean training. These results demonstrate that ReLU-targeted data ordering and poisoning can impair learning without directly modifying the victim model's parameters.
|
| 1586 |
Optimal Networks for Agentic Information Aggregation
2609.35537
|
cs.LG
|
MohammadHossein Bateni, Zahra Hadizadeh, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Shayan Taherijam |
We study information aggregation in the networked learning model introduced by Kearns, Roth, and Ryu (SODA 2026). There is a fixed distribution over $d$ features and a common label. Agents learn in topological order on a directed acyclic graph. Each observes a...We study information aggregation in the networked learning model introduced by Kearns, Roth, and Ryu (SODA 2026). There is a fixed distribution over $d$ features and a common label. Agents learn in topological order on a directed acyclic graph. Each observes a subset of the features and its parents' predictions, fits a linear predictor to minimize mean squared error, and passes only its prediction forward. The global predictor is the best linear predictor using all features. Kearns, Roth, and Ryu show that the output agent's error approaches the global predictor's error along sufficiently deep paths with suitable feature coverage, while insufficient depth can prevent aggregation even in large networks. In contrast to their main focus on a given graph and feature allocation, we consider the limits of the model under two settings. In the adaptive designer setting, a designer chooses the graph, feature allocation, and output agent knowing the distribution. In the oblivious designer setting, the designer fixes all three before an adversary chooses the distribution. Each agent observes one feature and receives predictions from a limited number of parents. We call the aggregation exact when the output agent matches the global predictor exactly. For $d\ge3$, we show that no finite depth guarantees exact aggregation for every distribution with one parent per agent, even when the designer knows the distribution. In contrast, two parents per agent suffice for exact aggregation even in the oblivious designer setting. A fixed graph, feature allocation, and output agent achieve this for every distribution at depth $O(d\log d)$. Knowing the distribution reduces the depth to $O(d)$. Both constructions use $O(d^2)$ agents, with a very large constant for two parents. We show the bounds on the depth and number of agents are all optimal up to constant factors.
|
| 1587 |
Learning the Robustness Mechanism with Bilevel Optimization
2609.35541
|
cs.LG
|
Yiyang Shen, Qihang Lin, Weiran Wang |
We propose a distributionally robust learning framework where parameters defining the robustness mechanism are learned from held-out data instead of extensively tuned. Using bilevel optimization with both upper and lower level minimax problems, we create two i...We propose a distributionally robust learning framework where parameters defining the robustness mechanism are learned from held-out data instead of extensively tuned. Using bilevel optimization with both upper and lower level minimax problems, we create two instances of our framework to tackle setups with and without group labels in the training set. Theoretically, we provide sample complexity analysis for our robustness mechanism learning paradigm, showing that it achieves generalization guarantees comparable to exhaustive grid search while being more computationally efficient. Empirically, we evaluate our framework under a challenging setup when both intra-group and inter-group test distribution shifts occur at the same time, thereby demonstrating the efficacy and scalability of our method.
|
| 1588 |
Graph World Models for Constrained Epidemic Policy Planning
2609.35545
|
cs.LG
|
Yiqi Su, Rashed Shelim, Lingyi Wang, Walid Saad, Naren Ramakrishnan |
Epidemic policy planning often requires coordination between geographical regions, taking into account mobility-driven spillovers and how to make use of limited resources. Existing methods either lack action-conditioned models of coupled dynamics or cannot gua...Epidemic policy planning often requires coordination between geographical regions, taking into account mobility-driven spillovers and how to make use of limited resources. Existing methods either lack action-conditioned models of coupled dynamics or cannot guarantee per-period feasibility. We present EpiMind, a graph world model framework for constrained epidemic policy planning across regions. A graph-factored recurrent state-space model generates joint policy-conditioned rollouts from regional latent beliefs, while graph-temporal ADMM optimizes regional interventions, enforces shared-resource feasibility through projection, and evaluates temporal specifications under the learned model. EpiMind reduces admission RMSE by 29% relative to graph-free dynamics modeling, plans within 1-5% of the best feasible constant policy with guaranteed shared-budget feasibility, and outperforms all deployable baselines across three resource budgets in real-context evaluation. These results demonstrate that graph-structured policy imagination with explicit constrained coordination supports effective epidemic interventions from learned dynamics.
|
| 1589 |
Simplex Diffusion Models
2609.35553
|
cs.LG
|
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton |
Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps thro...Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information collapse). We propose Simplex Diffusion Models (SDMs), a framework that lifts the diffusion process to the probability simplex to represent beliefs over categories. SDMs admit probability paths with closed-form reverse transitions and can be trained with a simple cross-entropy loss. Contrary to earlier proposals such as Dirichlet Flow Matching which requires integrating an ordinary differential equation, we introduce a DDIM-like sampler with a tunable level of stochasticity. Because SDMs operate on samples on the simplex, they can carry uncertainty across denoising steps, which mitigates information collapse. On OpenWebText, SDMs are competitive with strong Discrete Diffusion baselines, achieving $17.0$ GenPPL at $5.46$ unigram entropy in 64 sampling steps, close to real validation data. Even without Self-Conditioning (SC), SDMs outperform masked and uniform diffusion (with SC or predictor-corrector sampling) on code generation (TinyGSM, $T=0.1$; $49.0\%$ vs. $45.8\%$). Distilled down to 8 steps, SDMs solve $32.1\%$ of GSM8K problems, more than distilled Discrete Diffusion models with 128 steps ($21.4\%$).
|
| 1590 |
From Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel Synthesis
2609.35568
|
cs.LG
|
Longxiao Fan, Tao Zhang, Han Yan, Jiajun Li, Mingcong Song |
High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DS...High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models and memory hierarchies differ substantially from those of GPUs. To address this transfer gap, post-training methods adapt LLMs to NPU programming but depend on scarce expert data and substantial training compute. Memory-learning agents instead adapt through external memory, but their uniform credit assignment gives adopted and unused experiences the same reward target, potentially biasing subsequent retrieval rankings. Moreover, when learned values guide only retrieval, high-value experiences that generalize across operators must be retrieved repeatedly rather than retained in context, thereby increasing retrieval overhead and weakening cross-task guidance. We therefore present SAGE, a persistent self-improving agent for NPU kernel synthesis. Adoption-Traced Utility estimation (ATU) combines explicit adoption records with kernel evaluation outcomes for adoption-aware credit assignment. Utility-Gated Consolidation (UGC) uses positive utility and repeated adoption across operators to select and abstract reusable rules into a bounded resident context. On NPUKernelBench, SAGE achieves a 95.5% execution rate versus 84.1% for the strongest controlled baseline, with 86.9% of solved operators outperforming torch_npu. With GLM-5.3, SAGE achieves a 43.99x speedup over the torch_npu reference on sparse flash attention. These results show that adoption-aware credit assignment and selective consolidation enable agents to accumulate and reuse hardware-specific knowledge across tasks.
|
| 1591 |
Output-aware Residual Stream Pruning for Large Language Models
2609.35579
|
cs.LG
|
Chayne Thrash, Kevin Chen, Soheil Kolouri |
Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directio...Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities, but is difficult to optimize directly. We derive a tractable spectral upper bound that reduces subspace selection to an eigendecomposition of a sensitivity-weighted covariance matrix, retaining the efficiency and structural simplicity of rotation-based pruning methods. Across several instruction-tuned language model families, our method consistently reduces calibration KL divergence relative to activation-only pruning and improves perplexity and downstream task performance over a range of compression levels. Our results show that preserving activation energy alone is insufficient for residual-stream pruning, and that explicitly accounting for how perturbations propagate to the model output provides a more effective criterion for selecting dimensions to remove.
|
| 1592 |
Hardware-Aware Features for CUTLASS Kernel Selection
2609.35587
|
cs.LG
|
Shriram Chandran, Dominic Rinderer, Yakup Budanaz, Alexandru Calotoiu, Marcin Copik |
GPU libraries such as CUTLASS expose tens of thousands of semantically equivalent kernels for a single operation, making exhaustive autotuning expensive and execution-free selection difficult. Existing analytical selectors require hand-designed performance rul...GPU libraries such as CUTLASS expose tens of thousands of semantically equivalent kernels for a single operation, making exhaustive autotuning expensive and execution-free selection difficult. Existing analytical selectors require hand-designed performance rules, while learned selectors operate on raw configuration parameters and must infer hardware consequences from data. We introduce a hardware-aware representation for CUTLASS kernel selection that augments candidate configurations with statically computable estimates of induced hardware behavior. We construct a dataset of 4.9 million CUTLASS kernels and train gradient-boosted and neural learning-to-rank models to rank candidates within each problem. On held-out exhaustive evaluation problems, hardware-aware representations reduce selection regret by up to 40\% relative to structural baselines and 64.2\% relative to NVIDIA's matrix-multiply heuristics. We further evaluate data-efficient cross-precision and epilogue-fusion transfer within CUTLASS GEMM, showing that explicitly representing candidate-induced hardware behavior provides a useful inductive bias for learned kernel selection.
|
| 1593 |
Control-Geometry Straightening for Sampling-Based Latent Planning
2609.35603
|
cs.LG
|
Ziang Fu, Ning Ning |
Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary l...Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary loss that learns planner-friendly representations by directly straightening control geometry for sampling-efficient planning. CGS matches pairwise cosine similarities among actions to those among corresponding latent differences only using local transitions from pixel-action pairs. The loss can be applied across world-model architectures using end-to-end learned or pretrained representations. Under linear-dynamics, our theoretical analysis connects this objective to temporal straightening and more balanced terminal-cost curvature across the full planning horizon, yielding finite-budget guarantees for MPPI, local contraction results for CEM, and convergence bounds for gradient descent. Across four control environments and multiple planners, CGS improves planning with fewer sampled candidates and refinement steps, achieving success-rate gains up to 20 and 12.6 percentage points over LeWorldModel (LeWM) and its temporal-straightening variant (LeWM+TS), respectively, with sampling-based planners using 128 candidates per update. Probes, comparisons with DINO-WM architecture, and planner-side ablations clarify how latent motion organization, state dependence, and dynamical context shape planning behavior. Straightening control geometry thus makes good action sequences easier to find under limited planning budgets.
|
| 1594 |
EvE: An Alternate Optimizer to Adam
2609.35614
|
cs.LG
|
Shashank Raj, Kalyanmoy Deb |
Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pr...Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one candidate via DE, running a short burst of gradient descent only if the DE step fails to improve on the incumbent. Selection is greedy, so on a deterministic objective the best-so-far value is provably monotone non-increasing, and since gradients are used only as a targeted rescue, per-iteration cost stays within a constant factor of a single Adam step regardless of dimension. Under a fixed, evaluation-cost-matched budget, EvE wins or ties Adam on 76% of 70 (problem, dimension) cells across seven scalable benchmarks up to one million variables. On three real neural-network tasks (an MLP on MNIST, and LoRA fine-tuning of a 1.5B-parameter language model on two datasets) EvE finishes the same charged budget 1.7-3.9x faster, at a modest cost in final quality (about one accuracy point on MNIST, 9-11% higher relative test loss on the two fine-tuning tasks; on GSM8K, Adam is about 5 accuracy points more accurate, and fine-tuning lowers accuracy below the base model for both). Inside successive halving on UCI Adult, EvE completes hyperparameter and architecture searches 3.1-3.5x faster, ranking configurations about as consistently with Adam as Adam does with itself across seeds (Kendall's tau 0.66-0.69). EvE is not a total replacement for Adam as a final-stage trainer, but a fast, gradient-aware proxy for the search-heavy, budget-constrained regime one level up.
|
| 1595 |
Behavioral Foundation Models for Quality Diversity
2609.35615
|
cs.LG
|
Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud |
Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, an...Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search space to discover large repertoires of behaviorally diverse and high-performing policies through Quality-Diversity (QD) methods. While QD methods generally search directly in high-dimensional policy parameter space, in this paper, we present BFM-QD, a framework that performs QD search in the compact latent space of a BFM. We further show that the BFM-QD framework provides a closed-form, gradient-free policy improvement operator that approximates a policy gradient update, but requires no critic training and no backpropagation. Across continuous-control benchmarks spanning dense locomotion, sparse navigation, and contact-rich manipulation, BFM-QD consistently outperforms parameter-space baselines, with particularly stark gains in sparse and deceptive settings, where all tested parameter-space QD methods collapse to near-zero performance. These results show the effectiveness of the BFM-QD framework, benefiting from the synergy between dimensionality reduction of the search space and offline pretraining from diverse behavioral data. This positions BFMs as a general-purpose backbone for QD optimization, extending their utility beyond zero-shot task solving to the discovery of diverse behavioral repertoires.
|
| 1596 |
Attention Graphons: A Graph Limit Perspective on Graph Transformers
2609.35620
|
cs.LG
|
Caio F. Deberaldini Netto, Moshe Eliasof, Luana Ruiz |
Graph Transformers produce, for each attention head, a dense $n\times n$ matrix of learned pairwise interactions. We ask a fundamental question: do these attention-induced graphs converge to a stable limit object as $n$ grows, or does the learned interaction p...Graph Transformers produce, for each attention head, a dense $n\times n$ matrix of learned pairwise interactions. We ask a fundamental question: do these attention-induced graphs converge to a stable limit object as $n$ grows, or does the learned interaction pattern remain unstructured and size-dependent? We answer this using dense graph limit theory, treating each attention matrix as a finite sample from an underlying kernel---an \emph{attention graphon}---and studying concentration around this limit under the cut-distance. We derive a worst-case variance bound requiring no assumptions on the graphon, and a sharper regularity-aware bound based on nonparametric estimation theory. To operationalize the theory, we propose a canonicalize-then-block-average pipeline for estimating dataset-level attention graphons, and a variance-based diagnostic for testing whether attention admits a stable continuum description. Experiments across multiple graph benchmarks show that learned attention stabilizes to dataset-specific graphon structure on several datasets; that empirical cut-distance and cut-norm variance decreases with $n$ consistent with our bounds; and that attention graphons transfer to larger graph sizes with error decreasing in $n$.
|
| 1597 |
Cartridges++: KV Cache Compression without Off-Context Derailment
2609.35621
|
cs.LG
|
Sonia Laguna, Joao Monteiro, Marco Cuturi, Pierre Ablin, Eleonora Gualdoni |
Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and a...Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference time. Methods to obtain CKVs range from drop mechanisms that reduce their number of columns, to learned approaches. Among the latter, Cartridges have emerged as a leading compression method, learning compact KV representations through distillation on relevant Q/A pairs. While existing evaluations focus primarily on whether Cartridges and other CKVs yield approximately similar responses to document-related, on-context queries, we investigate the crucial deployment question of whether they can handle off-context queries, something the native KV representation is particularly good at, thanks to the mechanics of attention. We observe a fundamental trade-off: while Cartridges perform better for on-context queries, heuristic-variants preserve better the original LLM's ability to operate off-context. We measure this through their capability to avoid context contamination in their response, retain general knowledge, and follow instructions. We propose Cartridges++, simple modifications to cartridges that retain off-context abilities at small or negligible cost. The router variant decides at inference time whether the query should use the learned long-context memory, while the data-mixing variant allocates a small fraction of training Q/As to queries outside the reference long document. Our study shows that assessing CKVs on document utility alone can mask substantial degradation in broader model capabilities, yet those issues can be fixed with benign changes to CKV inference or training.
|
| 1598 |
Arbitrary-Accuracy Neural Approximation with Optimal Neuron Count and Near-Optimal Bit Complexity
2609.35628
|
cs.LG
|
Zilan Cheng, Li-Lian Wang, Zhongjian Wang |
We study the minimum number of hidden neurons required for arbitrary-accuracy approximation of multivariate H\"older-continuous functions on $[0,1]^d$ and the associated encoding complexity. For $d\geq 2$, we construct a fixed, explicitly defined activation fu...We study the minimum number of hidden neurons required for arbitrary-accuracy approximation of multivariate H\"older-continuous functions on $[0,1]^d$ and the associated encoding complexity. For $d\geq 2$, we construct a fixed, explicitly defined activation function for which a closed-form network with two hidden layers of widths $d$ and $1$ achieves arbitrary accuracy in the uniform norm. We prove that $d+1$ is the exact minimum total number of hidden neurons among standard feedforward networks with locally integrable activations and affine outputs. We further give a simpler construction using a single elementary activation that combines the floor and exponential functions. This construction requires three hidden layers of widths $d$, $1$, and $2$, only two neurons above the minimum. If a skip connection is allowed, widths $d$, $1$, and $1$ suffice. These constructions use explicit grid addressing and integer encoding of quantized function values. For a bounded $\alpha$-H\"older class, they require $O(\varepsilon^{-d/\alpha}\log(1/\varepsilon))$ bits, matching the metric-entropy lower bound up to a logarithmic factor.
|
| 1599 |
DR-net-Mamba: Selective State-Space Modeling for Long-Range ECG Time-Series Denoising
2609.35634
|
cs.LG
|
Basile Morel, Samuel Ruiperez-Campillo, Andreas P. Streich, Julia E. Vogt, Thomas Hofmann |
Electrocardiogram (ECG) recordings are corrupted by non-stationary noise sources that degrade diagnostic reliability, particularly in ambulatory and long-duration recordings. Deep learning denoisers exist, but convolutional architectures are limited by their r...Electrocardiogram (ECG) recordings are corrupted by non-stationary noise sources that degrade diagnostic reliability, particularly in ambulatory and long-duration recordings. Deep learning denoisers exist, but convolutional architectures are limited by their receptive field, transformer-based models scale quadratically with sequence length, and diffusion-based approaches incur prohibitive inference cost. We propose a Mamba-augmented model that inserts selective state-space blocks at the convolutional bottleneck, combining local feature extraction with long-range temporal modeling at linear complexity. We comprehensively evaluate the proposed model with respect to reconstruction fidelity, noise robustness, recording-length scaling, and downstream diagnostic classification across over 40 pathology classes. On synthetic and real datasets, our model achieves the highest SNR and lowest RMSE, with the Mamba advantage increasing with sequence length and in low-SNR regimes. On classification with two independent classifiers, the proposed Mamba-based models achieve the best macro AUROC among all denoisers and improve over their convolutional base models. Calibration is more nuanced and classifier-dependent: denoising improves Binary Cross-Entropy and Brier score on Inception1D but often fails to beat the noisy input on ResNet1D-Wang, and the lead-specific Mamba variant is the only denoiser to improve both calibration metrics over the noisy baseline on both classifiers. Per-class analysis reveals a morphology-dependent benefit: Mamba substantially improves ST/T-change diagnoses, which depend on broad, context-sensitive waveforms.
|
| 1600 |
Bounding Retraining Equivalence and the Deletion Floor in Materials Machine Unlearning
2609.35635
|
cs.LG
|
Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban |
In materials machine learning, closely related retained structures can sustain accurate property predictions even after removing a specific record, rendering post-deletion prediction error an ambiguous metric for machine unlearning. To resolve this ambiguity, ...In materials machine learning, closely related retained structures can sustain accurate property predictions even after removing a specific record, rendering post-deletion prediction error an ambiguous metric for machine unlearning. To resolve this ambiguity, we define the deletion floor as the expected target loss under a specified retraining procedure at the deleted request. Standard indistinguishability constraints yield a sharp interval bounding an update's target loss around this baseline reference. Theoretically, a conditional neighbor bound links a low deletion floor directly to retained fit, prediction regularity, and local label agreement, while an exact ridge identity isolates residual fit from the prediction change induced by record deletion. Empirically, controlled redundancy sweeps show an $\approx 8\times$ drop in median normalized retraining loss when one retained relative remains after deletion. Across two distinct fitting regimes in a paired Materials Project study, the lower-floor regime also exhibits a larger prediction change on more than 50% of the shared requests. Systematic comparisons against approximate updates and the original model decouple deliberate target suppression from preserved overall model utility. Consequently, request-level unlearning evaluations should report reference loss, prediction change, and retained utility together, interpreting post-deletion accuracy against what retraining itself leaves behind.
|
| 1601 |
Transferable Mass Spectrum Prediction via Reference-Guided Test-time Specialization
2609.35649
|
cs.LG
|
Yunhua Zhong, Runting Li, Yifan Li, Pan Liu, Zhiwen Yang |
Tandem mass spectrum prediction supports compound identification across metabolomics, natural-product discovery, and environmental analysis. However, pretrained predictors often degrade under shifts in chemical space and acquisition conditions, while retrainin...Tandem mass spectrum prediction supports compound identification across metabolomics, natural-product discovery, and environmental analysis. However, pretrained predictors often degrade under shifts in chemical space and acquisition conditions, while retraining domain-specific models from scratch is costly. We introduce SPARC, a retrieval-guided test-time specialization framework that adapts a pretrained predictor using a spectral reference library without accessing test-query spectra. For each target query, SPARC retrieves chemically related reference spectra to recalibrate fragment intensities within the learned fragmentation space. During Transfer, SPARC combines reference-guided spectral adaptation with reliability-aware consistency, using reconstruction behavior on retrieved spectra to selectively preserve trustworthy predictions during continual specialization. Across MassSpecGym, NPLIB1 and application-specific GNPS libraries, SPARC improves spectral prediction under multiple transfer settings. These results establish retrieval-guided test-time specialization as a practical strategy for extending pretrained MS/MS predictors to specific chemical and acquisition domains, with continual test-time training providing further refinement during deployment.
|
| 1602 |
Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
2609.35686
|
cs.LG
|
Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang, Luke Zhang |
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that c...Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.
|
| 1603 |
Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
2609.35698
|
cs.LG
|
Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang |
We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based ...We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert trajectories, we prove a $\widetilde{O}\left(\frac{1}{K}+\frac{1}{N}\right)$ bound on the regularized imitation gap for fixed regularization parameters. Our analysis combines an online mirror descent construction for general convex reward classes to control estimation error from finite expert data and stochastic learner feedback, with a sharp analysis of optimistic KL-regularized policy learning. To the best of our knowledge, Dually Regularized AIL is the first algorithm to simultaneously achieve $\widetilde{O}\left(\frac{1}{\epsilon}\right)$ sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts. These results provide a rigorous characterization of the complementary statistical benefits of reward and policy regularization in AIL.
|
| 1604 |
Distillation Defenses Easily Break After Reinforcement Learning
2609.35699
|
cs.LG
|
Shidan Javaheri, Alexander Panfilov, Oliver Britton, Yarin Gal, Yonatan Gideoni |
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then ...Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.
|
| 1605 |
MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
2609.35701
|
cs.LG
|
Chang-Wei Shi, Xu Wang, Wu-Jun Li |
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance u...The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.
|
| 1606 |
A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion
2609.35703
|
cs.LG
|
Fred Xu, Thomas Markovich, Florence Regol, Yizhou Sun |
Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as ra...Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as random graph signals: graph Fourier filters capture structural variation, and a scalar orthogonal-polynomial chaos coordinate captures latent stochastic variation. The resulting doubly-spectral stochastic (DSS) expansion supplies task-matched readouts from one representation: the mean coefficient encodes class evidence for the energy-based OOD score, the higher-order coefficients encode structured logit variation, and quadrature averaging over the chaos coordinate defines the single predictive distribution used for prediction and calibration. A capacity theorem shows that, under a full-rank feature assumption, a restricted subfamily matches the chaos coefficients of any Gaussian-latent random graph signal, with exponentially decaying truncation error under a growth condition; the task-level claims are established empirically. DSS-GNN has two deployment modes: standalone, or as a residual branch beside a deterministic encoder (DSS-Hybrid). Standalone DSS-GNN achieves the lowest Brier score among the compared uncertainty-aware baselines on all 14 node classification benchmarks without post-hoc correction; DSS-Hybrid achieves the best AUROC on most node-OOD settings, competitive cross-graph OOD detection, and the strongest shifted accuracy on all 7 GOOD concept-shift benchmarks under standard empirical risk minimization (ERM). Cross-evaluating both modes on all three tasks shows that each remains effective on the other's tasks, with documented exceptions, and yields explicit deployment guidance.
|
| 1607 |
ScAn-Bench: Evaluating Scaling Analysis Methodology
2609.35707
|
cs.LG
|
Artin Sermaxhaj, Nastaran Alipour, Donat Sinani, Johannes Hog, Neeratyoy Mallik |
Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that...Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future research, we introduce the surrogate benchmarks ScAn-Bench-LLM and ScAn-Bench-VLM based on 4524 and 8024 checkpoints of language and vision-language model pipelines. On our benchmarks, we perform the first systematic evaluation of both data acquisition and extrapolation methodology for scaling analysis across different data modalities.
|
| 1608 |
X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets
2609.35715
|
cs.LG
|
Prithwish Dan, Chenyang Ma, Wei Zhan |
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting dive...Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.
|
| 1609 |
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
2609.35750
|
cs.LGcs.AI
|
Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman |
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most ...Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.
|
| 1610 |
Neural Harmonic Measure Operator
2609.35752
|
cs.LG
|
Jinjin He, Sinan Wang, Yuchen Sun, Bo Zhu |
We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichl...We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on the geometry, not on the boundary data. NHMO parameterizes the density of this measure as a transformer-based boundary kernel supervised by Walk-on-Spheres exit samples, so one trained kernel handles different boundary values on a shape with no retraining. We extend it to Poisson via a classical decomposition, with an auxiliary network amortizing the source-induced correction and avoiding the singular volume quadrature that breaks direct evaluation. At inference, new boundary values and new sources both yield PDE solutions by re-integration against the fitted kernel and lift, with no retraining. NHMO improves over four prior baselines on the MCB-B 3D variable-shape Poisson benchmark across all five categories, and is competitive with major neural-operator baselines on a controlled 2D testbed.
|
| 1611 |
TokenCast: Forecasting Token Consumption During LLM Agent Execution
2609.35760
|
cs.LGcs.AI
|
Chaoqian Ouyang, Ling Yue, Libin Zheng, Hanghui Guo, Shengxiang Xu |
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates ...When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.
|
| 1612 |
Unifying Distributional Training for One-Step Visual Generation
2609.35763
|
cs.LG
|
Chi Zhang, Haoyang Shi, Yueyi Liu, Ruichuan An, Junkang Zhou |
\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from...\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates \textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with \textbf{1.45} $\mathrm{FDr}^6$ on pMF-H and \textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/
|
| 1613 |
Kolmogorov-Arnold networks in nuclear binding energy prediction
2407.20737
|
cs.LG
|
Hao Liu, Jin Lei, Zhongzhou Ren |
This study explores the application of Kolmogorov-Arnold networks (KANs) in predicting nuclear binding energies, leveraging their ability to decompose complex multiparameter systems into simpler univariate functions. By utilizing data from the Atomic Mass Eval...This study explores the application of Kolmogorov-Arnold networks (KANs) in predicting nuclear binding energies, leveraging their ability to decompose complex multiparameter systems into simpler univariate functions. By utilizing data from the Atomic Mass Evaluation (AME2020) and incorporating features such as atomic number, neutron number, and shell effects, KANs achieved a significant lower root mean square error (0.26 MeV), surpassing traditional models. The symbolic regression analysis yielded simplified analytical expressions for binding energies, aligning with classical models like the liquid drop model and the Bethe-Weizsaecker formula. These results highlight KANs' potential in enhancing the interpretability and understanding of nuclear phenomena, paving the way for future applications in nuclear physics and beyond.
|
| 1614 |
Bidirectional Neural Networks for Global Nucleon-Nucleus Optical Model Calculations
2512.22500
|
cs.LG
|
Jin Lei |
Modern nuclear data evaluation increasingly requires not only accurate scattering calculations, but also efficient methods for uncertainty quantification and parameter optimization, tasks that benefit from differentiable solvers amenable to gradient-based algo...Modern nuclear data evaluation increasingly requires not only accurate scattering calculations, but also efficient methods for uncertainty quantification and parameter optimization, tasks that benefit from differentiable solvers amenable to gradient-based algorithms. I present a neural network emulator based on Bidirectional Liquid Neural Networks (BiLNN) that provides a fully differentiable mapping from optical potential parameters to scattering wave functions. The key innovation enabling generalization across the parameter space is the use of phase-space coordinates $\rho = kr$ that normalize the oscillation wavelength regardless of projectile energy, allowing a single network to span 1 to 200~MeV. Trained on Numerov solutions for twelve target nuclei (\nuc{12}{C} to \nuc{208}{Pb}), both protons and neutrons, and partial waves up to $l=30$, the network achieves an overall relative error of 1.2\%. The predicted wave functions yield accurate $S$-matrix elements and elastic scattering cross sections, reproducing diffraction patterns spanning four orders of magnitude. Importantly, the model extrapolates successfully to nuclei not included in training (\nuc{24}{Mg}, \nuc{63}{Cu}, \nuc{184}{W}) with comparable accuracy, demonstrating that it has learned the physics of the optical model rather than memorizing specific targets. The differentiable nature of the trained model opens the door to gradient-based optimization of optical model parameters and efficient uncertainty quantification.
|
| 1615 |
Exterior complex scaling enables physics-informed neural networks for quantum scattering
2602.04553
|
cs.LG
|
Jin Lei |
Physics-informed neural networks (PINNs) have emerged as a powerful tool for solving differential equations, yet their application to nuclear scattering has been hindered by the oscillatory, non-decaying nature of scattering wave functions. In this work, I dem...Physics-informed neural networks (PINNs) have emerged as a powerful tool for solving differential equations, yet their application to nuclear scattering has been hindered by the oscillatory, non-decaying nature of scattering wave functions. In this work, I demonstrate that exterior complex scaling (ECS) transforms scattering boundary conditions into exponentially decaying waves suitable for neural network solutions, enabling PINNs to solve nuclear reaction problems for the first time. I develop a driven-equation formulation where the source term is confined to the real axis, avoiding the need to analytically continue nuclear potentials into the complex plane. The method is validated on nucleon-nucleus scattering (n+$^{40}$Ca at $E_{\text{lab}}=20$~MeV) with 21 partial waves, achieving phase shift accuracy of $\Delta\delta \lesssim 0.1^\circ$ for the strongly absorbed channels ($\ell \leq 4$) and $\Delta\delta \leq 0.60^\circ$ for all channels up to $\ell = 10$, when compared to conventional solvers. I further demonstrate the approach on heavy-ion scattering ($^6$Li+$^{208}$Pb at 40~MeV) with 41 partial waves and strong Coulomb effects, where an auto-adaptive anchor warm-down for weak-source channels yields a mean S-matrix accuracy of $|\Delta|S_\ell|| \approx 3 \times 10^{-3}$ across the full angular momentum range, including the absorption-to-transparency transition region. This work establishes the foundation for extending PINNs to inverse problems where end-to-end differentiability enables direct fitting of optical potential parameters, coupled-channel reactions, and few-body scattering where traditional grid methods face exponential scaling.
|
| 1616 |
FUSION: a skill-based research agent for publicly obtainable nuclear-physics codes
2609.04742
|
cs.LG
|
Jin Lei |
Running an unfamiliar nuclear-physics code is rarely difficult because of the physics alone. One must find and build the program, learn its input conventions, and decide whether a plausible output is actually correct. A general-purpose coding agent helps with ...Running an unfamiliar nuclear-physics code is rarely difficult because of the physics alone. One must find and build the program, learn its input conventions, and decide whether a plausible output is actually correct. A general-purpose coding agent helps with the first two tasks but may make the last one harder: it can write an input file that runs with the wrong physical convention. FUSION addresses this problem with code-specific skills. A skill obtains the code from its public source, starts from a verified input, runs and parses the calculation, records known failure modes, and must reproduce a stated benchmark to a stated tolerance before reporting a result. The current release covers twenty codes, spanning optical models and reactions, nuclear structure, fission and statistical models, astrophysics and R-matrix analysis, and heavy-ion transport. It also includes an offline, searchable collection of 61 167 pages derived from the nucl-th literature. User notes and credentials remain outside the public repository. FUSION is available under the MIT license at https://github.com/jinleiphys/FUSION; documentation is at https://vibeinscience.com. Here I describe the design, the checks behind the current release, and one complete calculation from input to comparison with measured data.
|
| 1617 |
Accurate Sampling from Diffusion Models
2609.29902
|
cs.LG
|
D\'enes Sexty |
A new proposal called DM-SMC (Diffusion Model - Sequential Monte Carlo) is investigated, which samples ensembles defined in terms of an action, using diffusion models trained on samples from the ensemble. The SMC setup allows for accurate sampling in spite of ...A new proposal called DM-SMC (Diffusion Model - Sequential Monte Carlo) is investigated, which samples ensembles defined in terms of an action, using diffusion models trained on samples from the ensemble. The SMC setup allows for accurate sampling in spite of an approximate diffusion model and the finite stepsize used in the numerical solution of the stochastic process. Improved update strategies are also investigated. Results are presented for a $Z_2$ symmetric scalar field theory in 2 dimensions near its 2nd order phase transition.
|
| 1618 |
Measurement-Error-Aware Causal Distributed-Lag Quantile Modeling of Indoor Air Pollution and Short-Term Lung-Function Deterioration
2609.31646
|
cs.LG
|
Shayma Alkobaisi, Anas Ali |
Low-cost indoor air-quality sensors could support personalized asthma prevention, but their nonlinear measurement error, delayed exposure effects, time-varying confounding, and heterogeneous lower-tail responses limit risk estimation. We present CAUSALQUANT-AS...Low-cost indoor air-quality sensors could support personalized asthma prevention, but their nonlinear measurement error, delayed exposure effects, time-varying confounding, and heterogeneous lower-tail responses limit risk estimation. We present CAUSALQUANT-ASTHMA, a measurement-error-aware causal quantile distributed-lag framework for short-horizon peak expiratory flow analysis. Sparse reference measurements train a nonlinear calibration model; stabilized sequential generalized-propensity weights address measured exposure assignment; and a susceptibility-modulated, smooth, noncrossing quantile model estimates lag-specific and sustained-exposure contrasts. Because no authorized cohort simultaneously provided dense indoor sensing, reference co-location, and outcome-compatible longitudinal data, evaluation used five semi-synthetic panels with known counterfactual truth, 150 patients and 12,600 patient-days per realization. Across eight methods, CAUSALQUANT achieved a dose-response integrated absolute error of 0.304 plus or minus 0.094, improving 24.2 percent over the strongest measurement-error and propensity-weighted baseline. It also obtained the lowest overall pinball loss, 1.065, while maintaining zero quantile crossings and 78.1 percent coverage for the nominal 80 percent interval. Sensor calibration reduced held-out exposure RMSE by 33.7 percent. Stress tests quantified degradation under sensor noise, missing personal measurements, and hidden confounding. These findings establish methodological feasibility and reproducibility, not clinical effectiveness; prospective, governance-approved external validation is required before patient-level interpretation or deployment.
|
| 1619 |
Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling
2609.31657
|
cs.LG
|
Ch Muhammad Awais, Marco Reggiannini, Davide Moroni |
Remote-sensing foundation models offer strong representations for SAR imagery, but their behavior under severe long-tail class imbalance is still not well characterized. We benchmark DOFA and SAR-JEPA on the imbalanced OpenSARShip dataset and compare them with...Remote-sensing foundation models offer strong representations for SAR imagery, but their behavior under severe long-tail class imbalance is still not well characterized. We benchmark DOFA and SAR-JEPA on the imbalanced OpenSARShip dataset and compare them with ImageNet-pretrained baselines under a fixed, training-efficient protocol that keeps the backbone frozen. To mitigate imbalance without fine-tuning, we apply four oversampling methods in embedding space exclusively to minority classes and train a lightweight classifier head on the augmented embeddings. Across both foundation models, oversampling improves Macro-F1 and test accuracy relative to their respective baselines, with the largest Macro-F1 gains observed for DOFA using ADASYN (34.39 to 38.56) and for SAR-JEPA using SVM-SMOTE (25.89 to 32.30). We also report class-wise behavior, showing that aggregate improvements can coexist with persistent failures on specific rare classes. Code for embedding extraction and reproducible multi-seed evaluation is provided to support rapid experimentation on free-tier hardware.
|
| 1620 |
Cross-Dataset Transfer and Unknown-Class Detection in Imbalanced SAR Ship Classification
2609.31658
|
cs.LG
|
Ch Muhammad Awais, Marco Reggiannini, Davide Moroni, Giulio Del Corso |
Ship classification from Synthetic Aperture Radar (SAR) imagery is a critical computer vision task, yet the robustness of models under deployment shifts remains unclear. While models are often trained on one dataset and deployed on another, we lack a comprehen...Ship classification from Synthetic Aperture Radar (SAR) imagery is a critical computer vision task, yet the robustness of models under deployment shifts remains unclear. While models are often trained on one dataset and deployed on another, we lack a comprehensive understanding of their cross-dataset generalization. To address this, we evaluate six pretrained models on two SAR ship datasets in three settings: in-domain classification, cross-dataset transfer, and unknown-class detection. For unknown detection, we hold out all classes one at a time. In-domain, SARDet100K gives the best balanced accuracy on both datasets (73.4\% on FUSARShip and 53.2\% on OpenSARShip). In cross-dataset transfer, we observe strong failures: some models show moderate accuracy but near-chance balanced accuracy (for example, 64.5\% accuracy but 33.3\% balanced accuracy for OpenSARShip to FUSARShip). In unknown detection, performance depends on the held-out class and dataset, while MC-dropout variance is often close to random. These findings show that cross-dataset generalization in SAR remains limited and that task-specific uncertainty scores are often more informative than MC-dropout variance for held-out-class detection, although their relative ranking depends on the dataset and held-out class.
|
| 1621 |
Learning Steadily: Accumulating Relative Point Margin Scores for Face Image Quality Assessment
2609.31662
|
cs.LG
|
Guray Ozgur, Tahar Chettaoui, Eduarda Caldeira, Marco Huber, Jan Niklas Kolf |
Face Image Quality Assessment determines the suitability of captured face images for automated face recognition (FR), a critical capability for reliable biometric systems. Existing state-of-the-art FR-integrated FIQA methods suffer from temporal instability: a...Face Image Quality Assessment determines the suitability of captured face images for automated face recognition (FR), a critical capability for reliable biometric systems. Existing state-of-the-art FR-integrated FIQA methods suffer from temporal instability: as the feature space evolves during training, single-epoch quality estimates fluctuate, creating a moving target that undermines reliable quality prediction. We introduce CARPM-FIQA, a stabilization strategy for FR-integrated FIQA that accumulates relative point margin measurements, the ratio between intra-class compactness and inter-class separation, across the entire training trajectory rather than relying on single-epoch estimates. This cumulative averaging approach provides theoretically grounded advantages: reduced variance in quality estimates, improved mean squared error, and enhanced ranking stability with convergence guarantees as training progresses. Through controlled experiments on the SynFIQA dataset with labeled quality groups, we demonstrate that cumulative averaging achieves superior discriminative ability, and ablation studies across different training configurations confirm consistent improvements. Evaluated against twelve FIQA methods on eight challenging benchmarks with four FR models at two FMR thresholds, CARPM-FIQA places 4th (CARPM-FIQA(L)) and 6th (CARPM-FIQA(S)) of 17 compared methods by pAUC-EDC and AUC-EDC averaged across FR models and, after per-benchmark normalization, across benchmarks, staying within a few percent of the best method's normalized average for every FR model, providing a principled solution to training instability while maintaining the performance benefits of FR integration. More broadly, our work demonstrates that temporal aggregation strategies can stabilize training objectives in deep learning systems where target values inherently fluctuate due to evolving feature representations.
|
| 1622 |
Data Processing for Offline Evaluation in Recommender Systems: a Survey
2609.31696
|
cs.LG
|
Alberto Carlo Maria Mancino, Angela Di Fazio, Danilo Danese, Matteo Attimonelli, Daniele Malitesta |
Offline evaluation is the dominant experimental paradigm in recommender systems research, enabling reproducible and cost-effective comparisons on historical interaction data. Yet, while considerable attention has been devoted to recommendation models and evalu...Offline evaluation is the dominant experimental paradigm in recommender systems research, enabling reproducible and cost-effective comparisons on historical interaction data. Yet, while considerable attention has been devoted to recommendation models and evaluation methodologies, the data processing decisions that precede model training have received less scrutiny. These decisions determine the information available to recommendation algorithms and can affect the comparability and reproducibility of experimental results. This survey provides a systematic, cross-domain characterisation of data processing practices for the offline evaluation of recommender systems. We examine the data-centric pipeline, from dataset selection and interaction representation to data preparation, multimodal feature extraction, and train-validation-test splitting. Our analysis spans recommendation paradigms, including collaborative, sequential, session-based, graph-based, knowledge-aware, context-aware, multimodal, federated, cross-domain, contrastive-learning, and LLM-based recommendation. Beyond reviewing existing practices, we introduce a unified framework and taxonomy for describing data transformations and feature-extraction strategies, distinguishing data preparation from the extraction of representations from multimodal side information. Our empirical analysis reveals a landscape dominated by a narrow set of dataset-level transformations, particularly support-driven filtering, while representation-dependent transformations remain less common. We further identify substantial heterogeneity in how auxiliary information is prepared and represented, as well as inconsistencies in the specification of data splitting protocols, where similar labels may conceal different experimental conditions.
|
| 1623 |
Be Careful Who You Trust: Coordination Dynamics under Corrupted Communication in LLM Multi-Agent Games
2609.31704
|
cs.LG
|
Xuanyi Liu, Niall Dalton, Hairi Amin, Xiyuan Yin, Lydia Lim |
Large language models are increasingly used as interacting agents, but it remains unclear how robust their coordination is when public communication is unreliable. We study this question in iterated $N$-player Stag Hunt games played by homogeneous LLM groups u...Large language models are increasingly used as interacting agents, but it remains unclear how robust their coordination is when public communication is unreliable. We study this question in iterated $N$-player Stag Hunt games played by homogeneous LLM groups under controlled programmatic action inversion, which changes both the public transcript and the actions used for execution. Across an experimental grid spanning group sizes, coordination thresholds, corruption levels, and seven LLMs, we observe three main patterns. First, honest agents' pre-flip Stag choices decline as corruption increases, but the sharp fall in public success is primarily mechanical. In the focal $N=5,M=3$ setting, pre-flip success remains 78% at 80% corruption, while public success falls to 12%. Second, honest choices are associated with the public history available at decision time, particularly under high corruption. Third, three threshold-style public-report benchmarks yield similar action-match rates to the LLM agents, showing substantial descriptive agreement between LLM decisions and these benchmarks. Overall, our results show that original choices, public actions, and executed outcomes must be separated when evaluating multi-agent robustness, as corrupted communication can severely and predictably degrade mutually beneficial cooperation.
|
| 1624 |
Statistical Testing for Multiple Instance Learning via Selective Inference with Applications to Computational Pathology
2609.31712
|
cs.LG
|
Noriaki Hashimoto, Shuichi Nishino, Teruyuki Katsuoka, Tomohiro Shiraishi, Daiki Miwa |
Multiple instance learning (MIL) is widely used in computational pathology because it enables weakly supervised analysis of whole-slide images (WSIs) without requiring patch-level annotations. In attention-based MIL, instances with high attention scores are of...Multiple instance learning (MIL) is widely used in computational pathology because it enables weakly supervised analysis of whole-slide images (WSIs) without requiring patch-level annotations. In attention-based MIL, instances with high attention scores are often interpreted as diagnostically important regions and used as visual explanations. However, attention scores alone cannot determine whether selected high-attention instances are significantly different from normal instances, limiting the reliability of attention-based explanations. In this paper, we formulate the evaluation of high-attention instances as a statistical hypothesis testing problem. Specifically, we assess whether a selected high-attention instance significantly deviates from a representative normal reference instance selected based on feature similarity. A major challenge is that both the target instance and the reference instance are selected through data-dependent procedures, rendering standard hypothesis testing invalid. To address this issue, we introduce a selective inference (SI) framework that explicitly accounts for the selection events induced by attention-based instance selection and adaptive reference selection, thereby enabling the computation of valid selective $p$-values conditional on these events. Experiments demonstrate Type-I error control on synthetic and MNIST-based data and practical applicability to pathological WSIs, with higher statistical power than the conventional over-conditioning approach.
|
| 1625 |
Frequency-Domain AI-Generated Image Detection: Exploring Decoder and Channel Attention for Feature Refinement
2609.31723
|
cs.LG
|
Uday Shankar Roy, Mahbuba Jahan Minu |
With the rapid progress of AI, the number of AI-generated images has increased significantly in recent years. However, the increasing variety of image generation models makes detection more difficult. In this work, we use Fast Fourier Transform (FFT) represent...With the rapid progress of AI, the number of AI-generated images has increased significantly in recent years. However, the increasing variety of image generation models makes detection more difficult. In this work, we use Fast Fourier Transform (FFT) representation with EfficientNet-B0 for AI-generated image detection. EfficientNet-B0 provides a lightweight architecture that can be useful for resource-limited applications. Most frequency-domain detectors use a standard encoder to extract features from the FFT spectrum and directly pass them to a classifier. We explored a different approach by investigating ECA, U-Net, and Attention U-Net as alternatives to this direct encoder-to-classifier approach. ECA applies channel attention, while U-Net and Attention U-Net use decoder-based architectures to recover and refine spatial information in the extracted frequency features. We used a balanced subset of the MS COCOAI dataset that includes AI-generated images from five different models. Three runs were carried out for each experiment, and the average values were recorded. Experimental results indicate that EfficientNet-B0 obtained an accuracy of 84.64%, which is 4.50 percentage points higher than the ResNet-50 baseline reported in the dataset paper. EfficientNet-B0 with U-Net provided a small improvement, achieving an accuracy of 84.85%, while ECA did not increase the overall performance. EfficientNet-B0 with Attention U-Net achieved the best overall performance, with an accuracy of 85.51% and an ROC-AUC of 92.99%. This represents an improvement of 0.87 percentage points in accuracy compared to the EfficientNet-B0 baseline and 5.37 percentage points over the ResNet-50 baseline reported in the dataset paper.
|
| 1626 |
GERIS: A Game-Theoretic Framework for Filtering Instance-Dependent Label Noise in License Plate Data Augmentation
2609.31731
|
cs.LG
|
Seyedeh Sara Jalili Shani (Department of Computer Science, University of Alberta, Alberta, Canada), Rouhollah Ahmadian (Department of Mathematics and Computer Science |
In this paper, we propose GERIS, a game-theoretic framework for instance selection in the data augmentation phase of license plate recognition systems. During augmentation, synthetic license plate images are generated and transformed using stochastic noise to ...In this paper, we propose GERIS, a game-theoretic framework for instance selection in the data augmentation phase of license plate recognition systems. During augmentation, synthetic license plate images are generated and transformed using stochastic noise to simulate real-world conditions. However, certain noise configurations lead to highly distorted, unreadable images that degrade model performance by introducing instance-dependent label noise. GERIS formulates a non-cooperative game in which each noise vector competes for inclusion in the training set based on its similarity to labeled data and its contribution to model reliability. By identifying and pruning low-quality instances, GERIS improves the overall quality of the augmented dataset. Unlike traditional black-box learning methods, GERIS offers a transparent, theoretically grounded mechanism for data filtering. Experimental results demonstrate that GERIS outperforms existing instance selection methods in terms of classification accuracy and robustness.
|
| 1627 |
Seeing the Heat: Synthesizing High-Resolution Wood Thermal Responses from Optical Imagery
2609.31737
|
cs.LG
|
Jingren Xie |
The thermal behavior of wood is a critical factor in advanced material assembly. However, pixel-level thermal analysis remains fundamentally constrained by the low resolution and noise inherent to infrared thermography. To address this, we introduce an end-to-...The thermal behavior of wood is a critical factor in advanced material assembly. However, pixel-level thermal analysis remains fundamentally constrained by the low resolution and noise inherent to infrared thermography. To address this, we introduce an end-to-end computational framework that synthesizes high-resolution thermal responses directly from wood RGB images. We first establish a core physical linkage: because spatial color variation in natural wood is driven by cellular anatomy, optical intensity serves as a reliable geometric proxy for the localized solid volume fraction. By leveraging this theoretical insight, we develop an automated finite-element-method data engine that maps pixel-level optical intensity to a 3D thermodynamic voxel grid, generating high-fidelity synthetic thermal responses. We find that 1) when the thermal conductivity along the thickness direction is uniform or linear, wood RGB images and their corresponding thermal responses exhibit extreme morphological similarities, and the lateral thermal diffusion acts as a low-pass filter that smooths out high-frequency details; 2) when the thermal conductivity along the thickness direction is random, such morphological similarities are destroyed, and wood's 3D structure dominantly governs its thermal response. We further utilize these synthetic thermal responses to supervise a neural surrogate model built upon the DINOv3 foundation model. Our results demonstrate that the neural surrogate model successfully internalizes the governing thermodynamic laws, thereby bypassing computationally expensive simulations and enabling high-resolution thermal inference. This methodology effectively bridges the semantic and thermodynamic domains, unlocking systematic, pixel-level analysis of fine-grained wood thermal responses. Project: https://zekifayes.github.io/seeheat
|
| 1628 |
Calibration-Free Surface Normals Estimation in Vision-Based Tactile Sensing using Universal Photometric Stereo
2609.31754
|
cs.LG
|
Zdravko Dugonjic, Stefanie Speidel, Roberto Calandra |
Vision-based tactile sensors are a popular solution for capturing rich contact surface geometry. However, to obtain high-detail contact surface normals and depth, it is necessary to calibrate the sensor by physically pressing a probe with known geometry agains...Vision-based tactile sensors are a popular solution for capturing rich contact surface geometry. However, to obtain high-detail contact surface normals and depth, it is necessary to calibrate the sensor by physically pressing a probe with known geometry against the sensor elastomer and mapping tactile images onto the ground truth probe's shape. This approach does not scale across different tactile sensors, and the calibration effort can be complex depending on the sensor shape and optical system. Instead, we propose a calibration-free procedure for the estimation of contact surface normals using Universal Photometric Stereo neural networks. In a series of real-world experiments, we evaluate our approach on 3 sensors with different optical systems, demonstrating that universal methods are a suitable approach for estimating surface normals at the contact patch from tactile images, thereby alleviating the need for tactile sensor calibration. Controlled experiments with a metal ball show that universal methods match the calibrated method, with a mean angular error of $6.56^{\Large\circ}$. We show that the proposed framework recovers high-frequency surface details of objects with natural textures, achieving an overall mean angular error of $10.66^{\Large\circ}$. Universal method robustly recovers the contact surface normals captured with dome-shaped Digit 360, achieving a low angular discrepancy of $10.18^{\Large\circ}$ relative to the calibrated baseline. This experiment demonstrates that with sufficient illumination settings surface normals could be estimated using a model trained solely on synthetic data. By providing a unified representation of contact surfaces across different vision-based tactile sensor designs, Universal Photometric Stereo neural networks lay the foundation for transferable tactile perception across sensors.
|
| 1629 |
Acoustic domain shift in spoken language identification from systematic domain generalization evaluation to real-world application
2609.31759
|
cs.LG
|
Francois Derrida (X), Rapha\"el Duroselle (X), Thomas Courtat (X), Jean-Fran\c{c}ois Bonastre (X) |
Domain Generalization (DG) aims to develop models that remain robust to conditions unseen during training. While DG has been systematically studied in computer vision through controlled benchmarks and diverse distribution shifts, its evaluation in spoken langu...Domain Generalization (DG) aims to develop models that remain robust to conditions unseen during training. While DG has been systematically studied in computer vision through controlled benchmarks and diverse distribution shifts, its evaluation in spoken language recognition remains less structured. Existing speech datasets provide valuable benchmarks for robustness to real-world acoustic conditions, but are primarily designed around specific scenarios and large scale rather than as general-purpose tools for systematically constructing and evaluating domain shifts. In this work, we introduce a smallscale speech dataset and evaluation protocol for controlled studies of acoustic domain shifts. It enables the evaluation of spoken language recognition models under a variety of acousitc domain shifts. We introduce the speech modality into the DomainBed Domain Generalization framework and evaluate three domain generalization algorithms. We show that in-domain performance is not a reliable predictor of cross-domain robustness and verify that explicit domain invariant algorithms such as MMD or DANN algorithms do not outperform ERM. We further validate the generality of these findings on MMS-LID-126, a state-of-the-art spoken language identification system. We release the code.
|
| 1630 |
What Stops Recursive Self-Improvement in Robotics? Lessons from 123 Rounds of Agentic Skill Discovery
2609.31760
|
cs.LG
|
Jiaming Wang (National University of Singapore) |
Can a robot improve itself the way coding agents now improve software? We built an agentic system to find out. It watches a robot fail, works out which capability is missing, writes new skills or finds and installs external models, tests every change in simula...Can a robot improve itself the way coding agents now improve software? We built an agentic system to find out. It watches a robot fail, works out which capability is missing, writes new skills or finds and installs external models, tests every change in simulation, and repeats, with no human writing robot code. We ran it for 123 improvement rounds on household manipulation tasks. This report describes what we learned. The good news is that the agent can discover capabilities on its own: noticing that its targets were out of view, it asked for an active-viewing model, debugged it, and deployed a working search skill. The bad news is that its improvements did not add up. Changes kept passing their tests, yet the target task, putting condiments on the top shelf of a fridge, never succeeded. We found that the agent was rarely the bottleneck. Three things around it were. First, chained perception modules do not understand relations. Segmenters such as SAM 3 find shelves but not "the top shelf", so the agent filled the gap with ever more geometric rules that never converged, when what it needed was a different kind of model. Second, skill chains lock learning onto the first step. Long tasks mostly fail early, so evidence and fixes pile up there, and later skills are rarely reached, tested, or improved. Third, what the agent learns is decided by the harness. The agent optimized exactly what the evaluator measured, including where it was wrong, and weak tests and misleading memory turned activity into a standstill. We distill these lessons into concrete recommendations for building robot systems that improve themselves, each paired with an experiment that could prove it wrong.
|
| 1631 |
Suitable Measures for the Potential Operational Utility of AI NWP Rainfall Forecasts Over Africa
2609.31775
|
cs.LG
|
Shruti Nath, Docko Sow, Koomi Toussaint Amoussouvi, Fenwick Cooper, Josiah Kiarie Kimani |
Artificial intelligence (AI)-based weather prediction is approaching the skill of physical numerical weather prediction (NWP) systems at a fraction of the computational cost. This is particularly promising for Africa, where rainfall extremes are intensifying a...Artificial intelligence (AI)-based weather prediction is approaching the skill of physical numerical weather prediction (NWP) systems at a fraction of the computational cost. This is particularly promising for Africa, where rainfall extremes are intensifying and many forecasting centres lack the infrastructure to run physical models at extended lead times. We present a calibrated comparison of GraphCast, GenCast and the Functional Generative Network (FGN) against the physical NWP model IFS for rainfall prediction across Africa. Deterministic and probabilistic forecasts are postprocessed using Isotonic Distributional Regression and evaluated with the Continuous Ranked Probability Score against IMERG, RFEv2 and CHIRPS across seasons, wet and dry regimes, elevation zones and lead times. All models retain skill beyond climatology across most seasons and at extended lead times. AI models generally outperform IFS in wet regions, whereas IFS performs better in dry, high-elevation areas, where its finer resolution better represents orographic controls on rainfall. Across observational datasets and seasons, AI models achieve a median improvement of approximately 5% over IFS. GraphCast achieves calibrated skill comparable to the ensemble-based FGN, although FGN provides greater significant skill at longer lead times. These results highlight the potential of calibrated AI weather prediction to provide accessible and computationally efficient rainfall forecasts across Africa, while demonstrating the continuing importance of spatial resolution, ensemble design and regional characteristics.
|
| 1632 |
Panoptic Scene Program Diffusion Transformer
2609.31780
|
cs.LG
|
Chika Maduabuchi |
Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion ...Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture that treats a panoptic scene program as a first-class latent variable rather than an external control signal or post-hoc parse. PSP-DiT jointly denoises image latents and scene-program latents through coupled transformer streams, while panoptic grounding and cycle-consistency objectives tie object instances, attributes, relations, and counts to visual support in the generated image. Under matched training and inference settings, PSP-DiT improves over a strong flat-text baseline across GenEval 2, SANEval-Simple, PSG-Score, and DetailMaster, with the largest gains on counting, attribute binding, role-sensitive relations, and long structured prompts. The method preserves image quality, adds modest inference overhead, and remains robust to imperfect scene programs.
|
| 1633 |
Cross-Modal Knowledge Distillation for Acoustic Pedestrian Detection
2609.31785
|
cs.LG
|
Yonghyun Kim, Chaeyeon Han, Sancho Gatungay, Subhrajit Guhathakurta, Alexander Lerch |
Audio-only pedestrian detection is attractive for urban sensing but limited by weak acoustic cues. An appealing strategy is cross-modal knowledge distillation, in which a video teacher supervises the audio student during training so that the deployed model run...Audio-only pedestrian detection is attractive for urban sensing but limited by weak acoustic cues. An appealing strategy is cross-modal knowledge distillation, in which a video teacher supervises the audio student during training so that the deployed model runs on audio alone. Under this task's severe class imbalance and wide video-audio modality gap, however, what such distillation contributes is unclear. We introduce Trust-Filtered Distillation (TFD), which selectively suppresses teacher supervision on pedestrian samples, and interpret its logit formulation as conditional label smoothing under a shared temperature. Across ten distillation configurations under five-fold cross-session validation on ASPED, most methods yield modest gains in macro accuracy accompanied by small changes in PR-AUC. The main effect is higher no-pedestrian accuracy at the cost of lower pedestrian recall. Adding TFD to logit distillation strengthens this trade-off without improving mean PR-AUC. These findings clarify the benefits and limitations of selective cross-modal supervision by distinguishing operating-point shifts from discrimination gains in imbalanced acoustic detection.
|
| 1634 |
Convergence-Aware Pareto Selection of Covariate Scaling Transformations for Markov Deterioration Hazard Models: Evidence from Bridge Inspection Data
2609.31786
|
cs.LG
|
Takato Yasuno, Keita Kobayashi, Ryuta Sakaguchi |
Maximum likelihood estimation of Markov deterioration hazard models for infrastructure asset management is sensitive to the numerical conditioning of explanatory covariates, yet production pipelines often adopt a single scaling convention without systematic ju...Maximum likelihood estimation of Markov deterioration hazard models for infrastructure asset management is sensitive to the numerical conditioning of explanatory covariates, yet production pipelines often adopt a single scaling convention without systematic justification. We study four covariate scaling transformations---baseline max scaling, min-max scaling, z-score scaling, and Box-Cox transformation with a training-derived positivity shift---applied to an Exponential Hazard Markov (EHM) model estimated via L-BFGS-B on bridge inspection transition data.
|
| 1635 |
Prompting Particle Physics: Tokenized Multi-modal Foundation Models for Combinatorially Many Tasks
2609.31862
|
cs.LG
|
Nilotpal Kakati, Daniel Murnane, Baran Hashemi, Samuel Klein, Jeffrey Krupa |
Reconstruction and simulation at a collider experiment are long chains of specialised algorithms, each tuned to a single step. We explore how one model can serve many of those steps at once, while still producing the intermediate objects (tracks, calorimeter c...Reconstruction and simulation at a collider experiment are long chains of specialised algorithms, each tuned to a single step. We explore how one model can serve many of those steps at once, while still producing the intermediate objects (tracks, calorimeter cells, clusters, particles and jets) that make the chain interpretable. To do so, we represent every object in a jet in a single shared token vocabulary and train one model to map any subset of these modalities to any other. A task is then only a choice of which modalities to provide and which to request: particle flow, detector simulation and charged energy subtraction are all directions through the same set of weights. We train over all modality combinations, with a decoder emitting tokens either autoregressively or in parallel. With tokenisation, both architectures train stably with little tuning. In evaluations, both models produce realistic reconstruction and simulation objects, with the autoregressive model particularly faithful to output from Geant4. The autoregressive model also outperforms a state-of-the-art particle flow algorithm on many typical jet reconstruction metrics.
|
| 1636 |
A Large-Scale Benchmark and Risk Assessment of Traffic Analysis Attacks on Cloud LLM Services
2609.31877
|
cs.LG
|
Shahrooz Pouryousef, Jesus Lopez, Saeefa Rubaiyat Nowmi, Md Mahmuduzzaman Kamol, Moinul Hossain |
Cloud-based Large language model (LLM) services create a network-level traffic side channel that can expose model, prompt, and task behavior despite encryption. From packet sizes, directions, timing, and burst structure alone, a passive local observer can infe...Cloud-based Large language model (LLM) services create a network-level traffic side channel that can expose model, prompt, and task behavior despite encryption. From packet sizes, directions, timing, and burst structure alone, a passive local observer can infer the serving model, the user's prompt category, and the task executed by a collaborative multi-agent system. Yet current evidence is fragmented across separate datasets and settings, limiting reproducibility and comparison. We present, to our knowledge, the first unified measurement study and public benchmark of encrypted LLM traffic across both user--LLM and multi-agent executions. The large-scale benchmark contains 60,000 user--LLM interactions across 10 models and 6 prompt categories, plus 2,838 multi-agent executions covering 10 task categories and two coordination topologies. Using only encrypted packet metadata, we assess the risk of traffic analysis attack by characterizing traffic signatures, identifying the features most associated with leakage, and testing robustness under prompt reformulation, decoding-temperature changes, larger candidate model sets, and partial traffic observation. Model fingerprinting achieves 97.7\% balanced accuracy, prompt-category fingerprinting reaches 76.7\% mean accuracy, and multi-agent task fingerprinting achieves up to 90.7\% accuracy. Prompt reformulation weakens but does not remove model-specific leakage, and task fingerprints remain detectable even from a single agent's traffic.
|
| 1637 |
Knowledge-Driven XRD Phase Identification via Multi-View Retrieval and Explanation
2609.31888
|
cs.LG
|
Doaa Mohamed, Markus Stricker |
X-ray diffraction (XRD) is a experimental technique for determining the phase composition and structure of crystalline materials. However, interpreting XRD patterns is challenging, particularly in high-throughput materials discovery, where many novel materials...X-ray diffraction (XRD) is a experimental technique for determining the phase composition and structure of crystalline materials. However, interpreting XRD patterns is challenging, particularly in high-throughput materials discovery, where many novel materials may need to be characterized and no reference patterns are available. Consequently, machine learning is increasingly used to accelerate and automate the analysis while reducing errors associated with human interpretation. We propose a multi-decision framework for XRD phase analysis that integrates representation learning, similarity-based retrieval, and explainable decision support within a unified reference database. A convolutional autoencoder learns compact latent representations of XRD patterns that preserve structural similarity while remaining robust to variations arising from experimental noise and measurement conditions. By integrating multiple decision pathways within a shared latent space, the framework moves beyond single-label prediction toward ranked and interpretable phase analysis that mirrors expert practice. During inference, complementary decision mechanisms are applied, including latent-space classification and retrieval, explanation-guided similarity using Integrated Gradients, and composition-based similarity search. These mechanisms generate ranked candidate phase lists that are aggregated into a final prediction with an associated confidence score. Experiments on synthetic datasets demonstrate strong predictive performance, achieving 98.85\,\% accuracy for crystal system classification and 95.82\,\% accuracy for space group prediction on the test set, while maintaining robustness under realistic perturbations. The framework supports reliable, analyst-friendly identification of crystal phases and structures in high-throughput and exploratory materials discovery settings.
|
| 1638 |
MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus
2609.31898
|
cs.LG
|
K M Naimul Hassan, Ali Alavi, Donald S. Williamson |
Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Spe...Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise across multiple signal-to-noise ratio (SNR) conditions, enabling attention decoding under realistic listening scenarios. Through a four-speaker attention decoding benchmark, we show that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding. These findings open the door to new applications, analyses, and methodological advances in multimodal AAD. The complete dataset is publicly available at https://huggingface.co/datasets/aspire-osu/maestro-eeg-dataset . The official code repository is available at https://github.com/ASPIRE-OSU/MAESTRO .
|
| 1639 |
GT-VLA: Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation
2609.31904
|
cs.LG
|
Ninghan Zhong, Jing-Chen Peng, Sriram Vishwanath |
Vision-Language-Action (VLA) models have shown strong performance on robotic manipulation, but they often struggle to generalize to unseen tasks, configurations, and long-horizon settings. A key challenge is that VLAs overfit to training scenes and fail to fol...Vision-Language-Action (VLA) models have shown strong performance on robotic manipulation, but they often struggle to generalize to unseen tasks, configurations, and long-horizon settings. A key challenge is that VLAs overfit to training scenes and fail to follow novel language instructions. Off-the-shelf vision-language models (VLMs) often provide stronger generalization, but cannot directly control robot actions. To combine the common sense of VLMs with VLA control, we propose Guided Trace VLA (GT-VLA), a steerable framework that accepts guidance from an external generalist VLM through trace-conditioned action generation. GT-VLA uses a generalist model to identify semantic guidance for the current skill, converts this guidance into a 2D visual trace, and conditions its action policy on the resulting trace-rendered observation. This design separates semantic target acquisition, trace generation, and low-level action execution, allowing high-level guidance to propagate to robot actions. GT-VLA uses a Mixture-of-Experts architecture with skill-specific trace and action modules for robust execution. We evaluate GT-VLA on LIBERO and a physical robot platform, showing improved generalization over recent VLA baselines in both settings. The code and additional supplemental materials are available on our project website at https://ivaniz.github.io/gt-vla/.
|
| 1640 |
Artificial Neural Network Assisted Modelling of Tangent Galvanometer Measurements for the Determination of Horizontal Component of Earth's Magnetic Field
2609.31930
|
cs.LG
|
Saralasrita Mohanty, Sudakshina Prusty, Anshuman Pal, Pradipta Kumar Mishra |
The Tangent Galvanometer (TG) is a standard undergraduate laboratory experiment for estimating the horizontal component of Earth's magnetic field (BH) by measuring the angle of deflection of a magnetic needle corresponding to the current flowing through a circ...The Tangent Galvanometer (TG) is a standard undergraduate laboratory experiment for estimating the horizontal component of Earth's magnetic field (BH) by measuring the angle of deflection of a magnetic needle corresponding to the current flowing through a circular coil. In this study, an artificial neural network (ANN) is used as a complementary data-driven model to predict the value of BH. A dataset comprising 225 observations obtained using 50-turn and 500-turn coils was used for developing the ANN model. After quality control, 223 observations were retained and divided into training (70%), validation (15%), and testing (15%) subsets. The model was optimized using a feed-forward neural network with Tanh activation. Three different models (Models A, B, and C) with different input variables were compared for optimum performance. Model C used five input variables: current, deflection angle, tan(theta), magnetic field produced by the coil, and number of turns. The addition of tan(theta) produced a substantial improvement in prediction performance, which was further improved by including the magnetic field produced by the coil. Model C gave the best test performance, with R2 = 0.99053, RMSE = 0.54076 microT, and MAE = 0.32940 microT. The experimental and ANN-predicted values of BH were also compared with an adopted local geomagnetic reference value of 39.0 microT. The mean experimental and ANN-predicted values were 37.38898 microT and 37.32161 microT, respectively. The results demonstrate the usefulness of ANN as a complementary tool for analyzing experimental variability and nonlinear relationships in an undergraduate physics laboratory experiment.
|
| 1641 |
CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
2609.31957
|
cs.LG
|
Zhenhao Zhang, Zhaoyu Fan, Haohan Ying, Jingwen Hu, Hancen Fan |
Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained ...Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving. It contains 50K puzzles across 20 CAPTCHA types and 5 interaction modes, with every solution verified through execution. CaptchaArena provides 50K screenshot-action trajectories, including 46K with step-by-step reasoning annotations. It also includes fine-grained pixel-mask annotations for irregular targets. Using CaptchaArena, we train CaptchaAgent, a single 9B policy for all 20 CAPTCHA types, with supervised fine-tuning followed by reinforcement learning. The environment verifier directly provides the RL reward. Supervised fine-tuning reaches 70.5 Pass@1, and reinforcement learning further improves it to 71.7, while also improving performance on two external benchmarks. These results demonstrate the value of large-scale, fine-grained computer-use supervision for training interactive CAPTCHA agents. We release CaptchaArena and CaptchaAgent at https://github.com/X0X0X00/CaptchaArena.
|
| 1642 |
SNIP++: Fine-Grained Symbolic-Numerical Alignment for Symbolic Regression
2609.31965
|
cs.LG
|
Benjamin L\'eger, Shubham Gupta, Samy Mammeri, Kazem Meidani, Cem Subakan |
Mathematical expressions and the numerical behavior they produce are two views of the same underlying function, and connecting them is central to scientific discovery. Symbolic Regression (SR) relies on this connection directly: it searches for an expression t...Mathematical expressions and the numerical behavior they produce are two views of the same underlying function, and connecting them is central to scientific discovery. Symbolic Regression (SR) relies on this connection directly: it searches for an expression that reproduces a given behavior. Recent multi-modal models learn this connection by embedding symbolic expressions and their numerical behavior in a shared representation space. We show that this embedding space is only globally aligned: complete expressions correspond to complete behaviors, but the contribution of individual parts of an expression is not represented. This granularity gap leaves the model unable to tell how a local edit to an expression changes its behavior, the central operation in SR. We introduce a compositional alignment method that closes this gap: a structural positional encoding exposes the substructure of an expression to the encoder, and a multi-granularity contrastive objective grounds each subexpression in the behavior it produces before propagating this grounding to the full expression. The resulting representations close much of the modality gap between symbolic and numerical embeddings, reliably distinguish the effects of local edits that the original alignment cannot, and transfer to external SR corpora.
|
| 1643 |
Bridging Stochastic Flow Maps and Boltzmann Generators with Normalizing Flows
2609.31978
|
cs.LG
|
Louis Grenioux, RuiKang OuYang, Luhuan Wu |
Generating independent, equilibrium samples of molecular systems at scale remains a central obstacle in computational statistical mechanics. Boltzmann Generators address this by pairing a generative model with importance sampling to obtain consistent samples f...Generating independent, equilibrium samples of molecular systems at scale remains a central obstacle in computational statistical mechanics. Boltzmann Generators address this by pairing a generative model with importance sampling to obtain consistent samples from the target distribution. We introduce Normalizing Flow Flow Maps (NF$^2$M), which combines the strengths of recent stochastic flow maps with the tractability of classic normalizing flows to build a Boltzmann Generator. Unlike most methods, which correct the generative model only at the end, NF$^2$M reweighs each denoising transition as generation proceeds, avoiding wasted compute on trajectories that are ultimately discarded. At each denoising step, a conditional normalizing flow proposes clean configurations given the current noisy state (a simpler task than sampling directly from the target) and its exact likelihood enables correcting each proposal toward the true denoising transition of the target Boltzmann distribution. This is in contrast to most existing methods, whose likelihoods are approximate or expensive to evaluate, undermining the statistical reliability of the correction. We establish consistency of the corrected transitions and bound how approximation errors propagate through the sampling chain. We evaluate NF$^2$M on peptide systems, demonstrating improved sampling efficiency and sample quality.
|
| 1644 |
Evasion Attacks on Cost-Utility-Based Adversarial Training for Online AutoML in IoT Networks
2609.31981
|
cs.LG
|
Chukwunonso Henry Nwokoye, Wajiha Zaheer, Khalil El-Khatib, Li Yang |
As Internet of Things (IoT) networks increasingly depend on machine learning for anomaly, malware, intrusion detection, and network monitoring, such systems have become attractive targets for evasion attacks. Evasion attacks pose a major security risk because ...As Internet of Things (IoT) networks increasingly depend on machine learning for anomaly, malware, intrusion detection, and network monitoring, such systems have become attractive targets for evasion attacks. Evasion attacks pose a major security risk because an adversary intentionally modifies input data to mislead a trained model into producing incorrect predictions while evading detection. This study evaluates the impact of black-box evasion attacks on a cost-utility-based adversarial training defense strategy in an Online AutoML context for IoT networks. Specifically, evasion attacks were applied to online learners, including Hoeffding Tree (HT), Leveraging Bagging (LB), Streaming Random Patches (SRP), Hoeffding Adaptive Tree (HAT), and Adaptive Random Forest (ARF). By developing naive and adversarially trained (AT) versions of these online learners, we generated clean and adversarial accuracies for each model. The results show that the AT versions of LB and SRP performed best, achieving the highest adversarial accuracy (0.985) and high clean accuracy (0.993) at the highest cost budget of 1.00, with a maximum accuracy reduction of only 0.8%. Finally, drift detection was conducted using the Early Drift Detection Method (EDDM).
|
| 1645 |
Identifiability Limits of Gravitational Wave Phase Deviations: Multiclass Classification with a Multihead Neural Network
2609.31993
|
cs.LG
|
Lavinia Heisenberg, Shayan Hemmatyar |
We study how well gravitational-wave phase deviations can be identified using a neural network with classification and regression heads. The classification head distinguishes general relativity (GR) from six parametrized post-Einsteinian phase families with ex...We study how well gravitational-wave phase deviations can be identified using a neural network with classification and regression heads. The classification head distinguishes general relativity (GR) from six parametrized post-Einsteinian phase families with exponents $b\in\{-7,-5,-3,-1,+1,+2\}$; the regression head predicts the logarithm of the coupling magnitude, $\log_{10}|\beta|$. The network input is a response function quantifying the sensitivity of waveform mismatch to waveform deformations. We use stationary Gaussian noise with the Advanced LIGO design spectrum for 262 sources and a measured Hanford spectrum from the third observing run (O3) for 132 sources, and also test injections into recorded Hanford strain. On the two Gaussian datasets, the network identifies the exponent of deviations clearly above the noise with accuracies of $0.427$ and $0.329$. To interpret these results, we compare the network with an approximate classifier built from predicted waveform residuals for the six phase families and same sources. The two classifiers often confuse the same families. Examining residuals left by different phase corrections, we link these errors to similarities between the underlying deviations. On selected recorded Hanford strain, networks trained on simulated or recorded noise achieve overall classification accuracies close to those on simulated Hanford noise. Of thirteen neural-network variants tested on the Advanced LIGO design dataset, the recurrent network has the highest mean classification accuracy for loud deviations, although all variants remain below the approximate classifier. The network nearly matches that classifier's accuracy, indicating that identification is limited by similarities between neighboring phase families rather than by the network, for known intrinsic source parameters.
|
| 1646 |
What Does the Rank Buy? A Spectral and Distributional Analysis of Low-Rank Adaptation
2609.32002
|
cs.LG
|
Babak Barazandeh |
The rank $r$ in LoRA is widely treated as a capacity control: a smaller rank is assumed to yield a simpler model that generalizes better. We show that, under hard per-factor norm budgets---the idealization of the weight decay and norm control used in practice-...The rank $r$ in LoRA is widely treated as a capacity control: a smaller rank is assumed to yield a simpler model that generalizes better. We show that, under hard per-factor norm budgets---the idealization of the weight decay and norm control used in practice---this intuition breaks down. The reason is structural: under such budgets, the updates LoRA can reach are exactly the matrices of rank at most $r$ inside a nuclear-norm ball, and every complexity and displacement functional we analyze is maximized over this set by a rank-one update---so the rank cap never binds. The consequences follow directly. The linear-readout model class we study is identical for every $r \ge 1$, its Rademacher complexity carries no dependence on $r$, and the distance the adaptation can move the source distribution obeys a rank-independent upper bound that we show is sharp. If rank does not control capacity, where does it act? We identify two places. Statistically, replacing the per-factor budgets with a joint budget on the product restores a data-dependent, rank-sensitive complexity bound---though the gain appears only for well-spread feature distributions, and the worst case remains rank-free. Spectrally, rank sets the price of adaptation: canceling the leading singular directions of the pretrained weight requires both sufficient rank and sufficient budget. We bound the smallest rank achieving a desired source--target alignment, with upper and lower bounds that match under two-sided spectral decay. Together, these results recast rank as governing which updates are reachable and what cancellation costs---not how much capacity the model has.
|
| 1647 |
Continuous-Time Trajectory Generation from Discrete Observations with Stochasticity
2609.32026
|
cs.LG
|
Ruifeng Shang, Shu Liu, Yuhua Zhu |
Physical systems evolve continuously in time, yet their states are typically observed only at discrete times. Generating trajectories consistent with their probability densities from such observations therefore requires capturing the continuous-time evolution ...Physical systems evolve continuously in time, yet their states are typically observed only at discrete times. Generating trajectories consistent with their probability densities from such observations therefore requires capturing the continuous-time evolution rather than only learning transition mappings between consecutive observations. We propose PhiBE-Flow, a framework that directly estimates the probability velocity field induced by the stochastic differential equation (SDE) which governs this continuous-time distributional evolution. PhiBE-Flow learns from discrete observations using a model-free approach requiring neither known SDE coefficients nor score estimation. We establish convergence guarantees for the method, accounting for both time-discretization and finite-sample errors. We evaluate PhiBE-Flow on systems of increasing complexity, from controlled stochastic numerical systems to Navier--Stokes dynamics and real-world videos. Our results show that PhiBE-Flow accurately recovers probability flows of stochastic dynamics, preserves multiscale physical statistics, and improves video generation performance over representative baselines. The code is available at https://github.com/R1fe/PhiBE-Flow.
|
| 1648 |
Correcting the Dropout-LayerNorm Expectation Gap Improves Protein Structure Models
2609.32062
|
cs.LG
|
Isaac Ellmen, David Errington, Matthew I. J. Raybould, Charlotte M. Deane |
Although $\mathbb{E}[\mathrm{dropout}(x)] = x$, here we show that $\mathbb{E}[\mathrm{LayerNorm}(\mathrm{dropout}(x))]$ is not equal to $\mathrm{LayerNorm}(x)$. Accordingly, the pattern of a Dropout layer followed by a LayerNorm, which is common to many AlphaF...Although $\mathbb{E}[\mathrm{dropout}(x)] = x$, here we show that $\mathbb{E}[\mathrm{LayerNorm}(\mathrm{dropout}(x))]$ is not equal to $\mathrm{LayerNorm}(x)$. Accordingly, the pattern of a Dropout layer followed by a LayerNorm, which is common to many AlphaFold2-based protein structure predictors, produces a systematic bias at evaluation time that can hamper performance. To address this, we derive a closed-form, first-order correction for this gap, which we call a Dropout-LayerNorm Correction (DLC). DLC empirically matches the performance boost of large Monte Carlo dropout ensembles. We evaluate its effect across nine protein structure models (ESMFold, OpenFold, ABB3, FlashABB, Ibex, NbForge, Genie1, Genie2, Genie3) on both paired and single-chain antibody structures as well as one protein-ligand docking model (QuickBind). The correction is computationally negligible and improves accuracy in all ten models tested ($\sim0.3\%-13\%$), with a modest but consistent improvement in ESMFold and OpenFold and a substantial improvement in antibody-specific models. This work identifies the mathematical consequence of chaining together Dropout and LayerNorm and provides a free, principled adjustment to improve the evaluation performance of many pretrained models. Code to reproduce the experiments is available at https://github.com/oxpig/DLC.
|
| 1649 |
Period Segmentation in Transition Network Analysis: A Topological Data Analysis Approach
2609.32094
|
cs.LG
|
Hitoshi Inoue, Koichi Yasutake |
Temporal dynamics in learning behavior can be revealed through period segmentation in Transition Network Analysis (TNA). Cristea et al. demonstrated that segmenting courses into halves and quarters reveals how learning strategies evolve and relate to academic ...Temporal dynamics in learning behavior can be revealed through period segmentation in Transition Network Analysis (TNA). Cristea et al. demonstrated that segmenting courses into halves and quarters reveals how learning strategies evolve and relate to academic performance. Building on this approach, we investigate whether Topological Data Analysis (TDA), specifically connected components ($\beta_0$) change points from Zigzag Persistent Homology, can provide data-driven period boundaries that identify intervention windows. Analyzing 22 courses from the Open University Learning Analytics Dataset, we find that $\beta_0$-based segmentation captures greater between-period variation than time-based segmentation (median variance ratio (VR) = 4.10$\times$; 21/22 courses show VR $>$ 1.1). Permutation tests confirmed significance ($p < 0.05$) in 18% of individual courses, with 73% showing positive improvement over random breakpoints. Only one course showed better performance with time-based segmentation. However, analysis of academic outcomes reveals a key insight: final-period behavior shows lower correlation with outcomes in $\beta_0$-based segmentation than in time-based segmentation, reflecting a marked behavioral collapse where engaged learners rapidly disengage. We interpret this as evidence that final-period behavior reflects rather than causes outcomes: students who will pass maintain engagement, while those who will fail disengage. This reframes the value of $\beta_0$: rather than improving prediction, change points identify intervention windows. These are periods where behavioral structure shifts and targeted support may be most effective.
|
| 1650 |
A Unified Optimism-Agnostic Framework for Linear Bandits over Spherical Action Sets
2609.32149
|
cs.LG
|
Arda G\"u\c{c}l\"u, Subhonmesh Bose, John R. Birge |
Linear bandits model sequential decision-making problems with noisy rewards that are linear in the decision variable, where an agent must simultaneously learn about an unknown parameter that governs the mean rewards, while maximizing (expected) rewards over ti...Linear bandits model sequential decision-making problems with noisy rewards that are linear in the decision variable, where an agent must simultaneously learn about an unknown parameter that governs the mean rewards, while maximizing (expected) rewards over time. Two prominent algorithmic families--upper confidence bound (UCB) and Thompson sampling (TS)--achieve a balance of exploration (to estimate said parameter) and exploitation (utilization of knowledge about it) across time. The quality of estimation of that parameter depends on the eigenvalues of a design matrix. In this paper, we begin by showing that if the inference quality obtained from exploration, encoded in the minimum eigenvalue of the design matrix, grows $\gtrsim \sqrt{t}$ with time $t$, while actions remain sufficiently concentrated for exploitation, then an algorithm produces optimal high-probability $\mathcal{O}(\sqrt{T}\log T)$-regret rate over a time-horizon $T$ for spherical action sets. This analysis is algorithm-agnostic and follows an alternative route to the classical optimism-based elliptical-potential argument for regret analysis. Then, we illustrate that variants of UCB and TS satisfy the inference and concentration properties and in turn, enjoy optimal regret rate. In effect, our results provide a modular framework that can be used to analyze linear bandit algorithms and explicitly connect quality of parameter estimation to optimal regret accumulation.
|
| 1651 |
FOCUS: Fixed-Confidence Online Causal Learning Using Sequential Adaptive Interventions
2609.32165
|
cs.LG
|
Haijie Xu, Chen Zhang |
We study fully online fixed-confidence causal discovery without any historical observational data. Starting from zero samples, the learner sequentially selects interventions to recover both the causal DAG and its edge weights under a linear-Gaussian structural...We study fully online fixed-confidence causal discovery without any historical observational data. Starting from zero samples, the learner sequentially selects interventions to recover both the causal DAG and its edge weights under a linear-Gaussian structural equation model. We establish an instance-dependent lower bound for any $(\epsilon,\delta)$-correct algorithm and propose \textsc{FOCUS}, which adaptively allocates interventions through an online max--min game. A key contribution is a computable concentration inequality for the accumulated KL divergence involving causal parameters shared across interventions. We prove that \textsc{FOCUS} is $(\epsilon,\delta)$-correct and that its expected stopping time matches the lower bound in its $\Theta(\log(1/\delta))$ dependence up to an instance-dependent constant. Experiments demonstrate improved structure and edge-weight recovery and confirm the predicted stopping-time trend. Our codes are available on https://anonymous.4open.science/r/FOCUS_code-76E5
|
| 1652 |
Uncertainty Quantification of Next Generation Reservoir Computing with Applications to Memory-Driven Dynamical Systems
2609.32169
|
cs.LG
|
Livia Popa, Sumanta Basu, Martin T. Wells |
Nonlinear dynamical systems with memory arise across science and engineering, yet uncertainty quantification for efficient forecasting methods such as Next Generation Reservoir Computing (NGRC) remains underdeveloped. We study Bayesian ridge and conformal pred...Nonlinear dynamical systems with memory arise across science and engineering, yet uncertainty quantification for efficient forecasting methods such as Next Generation Reservoir Computing (NGRC) remains underdeveloped. We study Bayesian ridge and conformal prediction intervals for NGRC and characterize when their uncertainty estimates agree or differ. In low dimensions, their asymptotic widths are governed by different summaries of the residual distribution, so agreement depends on residual shape rather than dimensionality alone. In high dimensions, regularization introduces a further tradeoff between estimation variance, shrinkage bias, and posterior uncertainty, leading to an explicit transition between regimes where Bayesian intervals are wider or narrower than conformal intervals. We extend these results to quadratic NGRC feature maps and give sufficient conditions for transferring the analysis to temporally dependent forecast windows. Simulations and real-data experiments support the theoretical predictions and illustrate how residual distribution, regularization, dimensionality, and distribution shift affect interval calibration and efficiency. These results provide a principled framework for choosing and interpreting uncertainty quantification methods in reservoir-based forecasting.
|
| 1653 |
Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models
2609.32185
|
cs.LG
|
Wenhao Zhang, Zhongliang Zhou, Shiyuan Zhang, Yiqing Yang, Pinqiao Wang |
Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis predi...Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis prediction, we found that most public pathology VLMs showed minimal differences when changing from feeding the models with whole-slide images, an annotated lesion, or no image at all. To better understand the specific features leveraged by these models, this paper presents two contributions aimed at disentangling these factors. First, we present CleanSlide, a TCGA-based VQA benchmark designed to eliminate image- and question-side contamination. It contains 149K audited multiple-choice questions over 9,985 slides, with patient- and tissue-source-disjoint splits. Every question is audited for option shortcuts, stem leakage, cross-split duplication, and blind solvability. Second, we propose Pair-DPO, a preference loss over counterfactual slide pairs from the same question and source. By controlling for shared confounding factors, Pair-DPO cancels out the question-attributable signal and leaves image evidence as the source of preference. Specifically, each pair consists of two real slides with opposite, verified findings, introducing neither editing artifacts nor unverified labels for diffuse or graded features such as invasion, necrosis, and tumor grade. Experiments show that our method gains 15.29% from image evidence on the CleanSlide, compared with 2.81% for the best published model. On the external CPTAC and BCNB cohorts, our method achieves accuracies of 57.6% and 59.0%, outperforming all other evaluated models by 9.7% and 3.4%, respectively. We will release the benchmark and code.
|
| 1654 |
High-Probability Guarantees for SGD under $\beta$-Heavy-Tailed Gradient Noise
2609.32195
|
cs.LG
|
Qijun Tong, Masahiro Ikeda, Ryota Kawasumi |
Stochastic gradient descent (SGD) is widely used to train machine learning models, but subsampling the training data introduces noise into its updates. The strength and applicability of high-probability guarantees therefore depend critically on how the tails o...Stochastic gradient descent (SGD) is widely used to train machine learning models, but subsampling the training data introduces noise into its updates. The strength and applicability of high-probability guarantees therefore depend critically on how the tails of gradient noise are modeled. Reports of heavy-tailed gradient noise in deep learning motivate relaxing the bounded-noise and sub-Gaussian assumptions commonly used in high-probability analyses of SGD. We use Young functions from Orlicz space theory to describe noise tails in a common framework. We model SGD gradient noise by adopting a Young function that preserves the finiteness of all polynomial moments while allowing tails heavier than sub-Weibull, including lognormal distributions. The resulting class is called $\beta$-heavy-tailed, with $\beta$ controlling the tail heaviness. We establish concentration inequalities for $\beta$-heavy-tailed noise and combine them with a uniform bound on the difference between empirical and population gradients along the SGD trajectory to obtain high-probability bounds on optimization and population-risk stationarity for smooth nonconvex losses under trajectory assumptions. The bounds are not restricted to a particular learning-rate decay rule and make explicit the effects of noise tails and learning-rate schedules. Under the Polyak-{\L}ojasiewicz condition, we bound the risk at the last iterate. We also analyze SGD with gradient clipping under the $\beta$-heavy-tailed noise model.
|
| 1655 |
SparSP: Exploiting Communication Sparsity for Sequence-Parallel Video DiTs
2609.32197
|
cs.LG
|
Desen Sun, Xinrui Zhong, Yuke Wang, Sihang Liu |
Diffusion Transformers have become the dominant architecture for video generation. Their substantial computational cost motivates scaling inference across multi-GPU servers, yet efficient scaling remains challenging on commodity GPUs connected via PCIe, whose ...Diffusion Transformers have become the dominant architecture for video generation. Their substantial computational cost motivates scaling inference across multi-GPU servers, yet efficient scaling remains challenging on commodity GPUs connected via PCIe, whose bandwidth is limited. Although sparse attention substantially reduces computation, its implications for communication remain underexplored. This paper argues that attention sparsity should be treated as a communication primitive. We present SparSP, an efficient sparse sequence parallel communication system that co-designs token distribution, communication routing, and asynchronous execution for sparse video diffusion models. First, Dependency-Aware Placement distributes sequence blocks according to diffusion models' sparse attention patterns. Second, Demand-Directed KV Routing transfers KV blocks directly to requesting GPUs without intermediate relays. Third, a Decoupled Transfer Runtime separates communication from GPU computation to reduce resource contention and maximize effective bandwidth. Our evaluation shows that SparSP improves attention performance by 1.38- 1.5$\times$, achieves an average 1.17$\times$ (up to 1.69$\times$) end-to-end speedup across three representative servers and three video diffusion models, and reduces communication volume by 12.54-23.05%. Moreover, we achieve an average 1.53-1.76$\times$ bandwidth improvement over NCCL.
|
| 1656 |
Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation
2609.32201
|
cs.LG
|
Hantao Yu, Sandy Han, Udaya Ghai, Ferhat Erata, Joe Lilien |
On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leib...On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instructions that target common student mistakes observed on the training samples, and show that such simple instructions can outperform gold as the OPCD privilege. In autoformalization tasks, using a matched formatting instruction as the privilege could outperform gold in OOD accuracy by a large margin. In 7 out of 8 experiments using ProverQA, ProofWriter, and ProntoQA as datasets, and Qwen3-Thinking and Olmo3-Thinking families as models, matched instruction privileges outperform gold in OOD by 4 to 17 points, while remaining on par with gold in-domain. Each instruction is only a few sentences (and thus contains much less information compared to all instance-specific gold) and is applied uniformly to every training sample. These results indicate that a general instruction, which applies equally to source and target domain examples, can be substantially more transferable than instance-specific gold in OPCD while maintaining in-domain performance.
|
| 1657 |
Kernel-Based Steering of CLIP with Vision-Language Model Preferences
2609.32203
|
cs.LG
|
Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Sina Mansouri, Mahnoosh Alizadeh, Farzan Farnia |
Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilit...Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that learns from elicited pairwise judgments without accessing teacher embeddings or collecting new human similarity annotations. ASK constructs positive semidefinite target kernels within small image groups and combines visual kernel matching with an image--text distributional anchor. Low-rank adapters jointly update the visual and text encoders while regularizing predictions toward frozen CLIP. After adaptation, retrieval uses CLIP image embeddings and cosine similarity, with no VLM calls. Experiments across five image domains, four CLIP backbones, and six judges evaluate teacher agreement, retrieval, and recognition retention. For ViT-B/16, mean retrieval mAP on classes excluded from adaptation increases from 53.8 to 75.0, compared with 71.7 for DINOv2 targets with KL anchoring. Mean zero-shot accuracy with jointly adapted encoders increases from 61.8\% to 62.4\%, averaged over 12 benchmarks and the five adaptation domains. Prompting provides an additional capability: selecting which visual distinctions the student learns. Human-annotated evaluations across four datasets support this criterion-specific control.
|
| 1658 |
DegreeSpar: Structured Degree Sparsity for Efficient Secure Transformer Inference
2609.32204
|
cs.LG
|
Yifei Cai, Zhuoran Li, Xiaozuo Shen, Hongyi Wu, Chunsheng Xin |
Secure Transformer inference protects sensitive inputs but incurs substantial cryptographic overhead, with nonlinear operations such as Softmax and GeLU becoming major bottlenecks. Existing compression methods reduce nonlinear complexity, sequence-dependent co...Secure Transformer inference protects sensitive inputs but incurs substantial cryptographic overhead, with nonlinear operations such as Softmax and GeLU becoming major bottlenecks. Existing compression methods reduce nonlinear complexity, sequence-dependent computation, or model structure through separately defined compression variables. Under aggressive compression, however, these independently optimized perturbations can accumulate: at a matched compression level, stacking representative approximation, token-pruning, and model-pruning methods reduces ViT-S accuracy from 80.20% to 76.41%. We introduce DegreeSpar, which formulates secure Transformer compression as structured sparsification over nonlinear polynomial degrees. Polynomial degree directly controls the cost of secure nonlinear evaluation, while computation-aligned zero-degree structures expose token-level and model-dimension computation as removable within the same optimization space. DegreeSpar further incorporates approximation-aware training for low-degree Softmax and GeLU, enabling aggressive degree reduction and creating the optimization headroom required for structured computation removal. Across vision and language Transformers, DegreeSpar consistently improves the accuracy-latency trade-off across model scales, tasks, and sequence lengths, achieving speedups from 2.29x to 6.63x over the corresponding baselines. Under the same network setting, DegreeSpar achieves 92.68% accuracy on BERT/SST-2 in 110.55 s, compared with 92.66% in 167.26 s for CipherPrune, the closest prior hybrid secure-inference approach. These results establish structured polynomial degree as an effective shared optimization space for secure Transformer compression.
|
| 1659 |
Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation
2609.32239
|
cs.LG
|
Biprodip Pal, Kaushik Roy, Yanming Zhu, Brendan Tidd, Alan Wee-Chung Liew |
Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutual...Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter aggregation destructive. We present FedDRMan, a federated subspace-guided distillation framework for heterogeneous robot manipulation. At each communication round, the server model provides a frozen teacher for local behavior cloning, while low-rank multimodal subspace and action-distribution distillation preserve globally useful representation geometry and policy behavior. To address heterogeneous aggregation, FedDRMan groups clients by update compatibility and maintains a persistent model for each cluster. The server then spectrally rebalances each compatible aggregate to mitigate attenuation of weaker task-relevant robot-policy update directions. Extensive experiments on LIBERO across diverse non-IID settings, heterogeneity levels, client participation variation, together with ablations and aggregation analyses, show that FedDRMan substantially improves knowledge transfer and consistently outperforms strong federated baselines achieving a peak mean success rate of 80.7%, 11.6 percentage points above the strongest evaluated federated baseline.
|
| 1660 |
Certifying Interventional Agreement Among Observationally Equivalent Causal Models
2609.32247
|
cs.LG
|
Sourena Khanzadeh, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama |
Observationally equivalent causal models can still disagree about what happens under intervention, because interventions create inputs that never occur in observational data. We introduce Interventional Separation Selection (ISS), which repeatedly queries the ...Observationally equivalent causal models can still disagree about what happens under intervention, because interventions create inputs that never occur in observational data. We introduce Interventional Separation Selection (ISS), which repeatedly queries the true system with an admissible intervention on which the surviving candidate models disagree, discards the candidates the outcome contradicts, and stops once no intervention within a cost bound separates the survivors. If the true system is among the candidates, this stopping condition certifies that every survivor agrees with it on every admissible intervention within the bound, a guarantee that no observational learner can give, however much data it sees. The stopping condition depends only on the survivors, so it can be checked without knowing the truth. For continuous variables the candidates form an infinite version space, and mixed-integer linear programs decide the stopping condition exactly over all of it, with agreement holding up to a tolerance. On a three-digit colored MNIST causal abstraction task in which ink hue tracks digit size, plain convolutional networks trained on examples reach zero held-out error, yet disagree with shape-based labels on 26% of single-digit edits, as often as hue-based labels do. Auditing the causal abstractions of networks observed only on such images, ISS certifies what each network perceives with 13.6 interventions per image on average, and each certificate, checked against every admissible intervention, holds whenever the network's true abstraction is among the candidates. When a network bypasses a unit that every candidate abstraction relies on, certificates covering interventions on that unit can be silently void, and twenty random validation interventions refute 69% of them.
|
| 1661 |
Active Data Acquisition with Side Information via Discrete Diffusion Priors
2609.32252
|
cs.LG
|
An Vuong, Thinh Nguyen |
Acquiring data is costly: higher measurement fidelity costs power and storage and risks collecting irrelevant content, while aggressive cost reduction can discard information that later analysis needs. We address this trade-off with an information-theoretic fr...Acquiring data is costly: higher measurement fidelity costs power and storage and risks collecting irrelevant content, while aggressive cost reduction can discard information that later analysis needs. We address this trade-off with an information-theoretic framework that acquires data relevant to a broad set of tasks rather than to one model. A mask policy, conditioned on side information, chooses which pixels to measure so as to maximize the mutual information between a discrete image and its partial observation under a budget; since the image entropy does not depend on the mask, this is equivalent to minimizing the conditional entropy. A frozen discrete denoising diffusion model (D3PM) supplies the posterior, and we use it in two ways: as an entropy surrogate for training a one-shot mask generator, and as the criterion for sequential greedy acquisition. The one-shot generator outperforms random masks only with care, including an unbiased gradient estimator for binary masks. With sequential acquisition, on MNIST the prior makes $8\times$ fewer errors than random at a $10\%$ budget, and on CIFAR-10 it gains $0.9$--$3.4$~dB. On fastMRI, our proposed technique using a static mask outperforms the well-known methods such as variable density and LOUPE.
|
| 1662 |
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
2609.32259
|
cs.LG
|
Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Murali Annavaram |
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache ...Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$--$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.
|
| 1663 |
Overfitting of Spectral Gradient Descent: How Matrix Geometry shapes Generalization and Implicit Bias
2609.32270
|
cs.LG
|
Guillaume Braun, Ichiro Hashimoto, Masaaki Imaizumi |
We study the generalization of spectral gradient descent (SpecGD) in overparameterized matrix classification with corrupted labels. Each input combines a shared low-rank signal with a rank-one sample-specific perturbation, referred to as a shortcut, that enabl...We study the generalization of spectral gradient descent (SpecGD) in overparameterized matrix classification with corrupted labels. Each input combines a shared low-rank signal with a rank-one sample-specific perturbation, referred to as a shortcut, that enables memorization but does not generalize. We contrast collapsed shortcuts, which share a singular direction, with dispersed shortcuts, which occupy distinct singular directions. Changing only this geometry can reverse the relative generalization of GD and SpecGD: collapsed shortcuts can favor SpecGD, while dispersed shortcuts can favor GD. In the dispersed regime, exact shortcut orthogonality eliminates the signal from the late-stage SpecGD direction, while vanishing random correlations collectively generate a small but generalization-relevant signal through a second-order effect. To identify the direction selected by SpecGD, which the spectral max-margin problem alone does not determine, we combine a refined analysis of its dual with the exponentiated-gradient dynamics of normalized loss weights. Finally, we show that a single SpecGD step can already interpolate and generalize well, while continued training converges to a direction with substantially worse generalization.
|
| 1664 |
Neural ODEs Meet Concurrent Learning: Stable Online Learning with Lyapunov Guarantees
2609.32289
|
cs.LG
|
Omkar Sudhir Patil |
Neural ODEs learn dynamics from trajectory losses, but their adjoint gradients lack the regressor-times-parameter-error structure on which Lyapunov analyses of online adaptation rest, so training on streaming data comes without stability guarantees. We show th...Neural ODEs learn dynamics from trajectory losses, but their adjoint gradients lack the regressor-times-parameter-error structure on which Lyapunov analyses of online adaptation rest, so training on streaming data comes without stability guarantees. We show that this structure is in fact present: the adjoint gradient decomposes exactly into a positive semi-definite trajectory operator acting on the parameter error plus a nonlinear perturbation with explicit, horizon-dependent bounds. A quadratic Lyapunov function then certifies online Neural ODE training over sliding windows under computable gain and horizon conditions, and the same certificate extends to stored data: its drift branch recovers concurrent learning, and its trajectory branch yields NODE-CL, a stored-segment Gauss-Newton method built on batched forward sensitivities that needs no state-derivative estimates. On four DeepMind Control Suite domains, NODE-CL attains the lowest median prediction error on three under velocity measurement noise, where observer-based concurrent learning degrades by up to 8x; with clean measurements it is best on the pendulum and within a factor of 1.6 of the best stored-data baseline on the cartpole and reacher.
|
| 1665 |
Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging
2609.32303
|
cs.LG
|
Jingyuan Huang, Zuming Huang, Yucheng Shi, Zhongzhi Li, Xiaoming Zhai |
Domain experts trained from a shared checkpoint can be merged into one model through on-policy distillation (OPD), where they act as teachers supervising a student on its own trajectories. One upstream choice is rarely examined: whether to build each expert wi...Domain experts trained from a shared checkpoint can be merged into one model through on-policy distillation (OPD), where they act as teachers supervising a student on its own trajectories. One upstream choice is rarely examined: whether to build each expert with supervised fine-tuning (SFT) or reinforcement learning (RL). Yet equally strong teachers need not be equally good teachers. We probe this choice through controlled single-teacher OPD, a building block of multi-teacher OPD: in Agentic, Reasoning, and Perception, comparably performing SFT and RL teachers are trained from Qwen3.5-9B, each guiding a student initialized from it. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points in Agentic, Reasoning, and Perception, respectively, and recover more of their teachers' performance gains over the base model. The contrast is clearest in Agentic, where the best SFT-guided student recovers only 44.44% of its teacher's gain, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our analysis points to an explanation: RL teachers stay much closer to the shared initialization in parameter space than SFT teachers and are therefore easier for their students to follow.
|
| 1666 |
All On-Board: Fully On-Chip Neuromorphic Q-Learning with Embedded CartPole Simulation
2609.32317
|
cs.LG
|
Steven C. Nesbit, Giovanni T. Michel, Gerd J. Kunde, Edward Kim, Andrew T. Sornborger |
As AI models grow in size and usage, their energy demands increase dramatically, raising sustainability and economic concerns. Neuromorphic hardware, inspired by the energy efficiency of the brain, seeks to address this challenge by offering low-power, fast-pr...As AI models grow in size and usage, their energy demands increase dramatically, raising sustainability and economic concerns. Neuromorphic hardware, inspired by the energy efficiency of the brain, seeks to address this challenge by offering low-power, fast-processing alternatives to conventional computing. Such hardware is particularly well-suited to control systems deployed in resource-constrained environments, which are best trained via reinforcement learning (RL). This contribution presents the design and implementation of a fully on-chip, closed-loop Loihi 2 RL agent. Our neuromorphic circuit consists of a fully embedded Q-learning algorithm and an on-chip simulation of the CartPole-v0 environment on Loihi 2. Our Q-learning algorithm trained the same number of successful agents as the CPU implementation in only half the execution time and with two orders of magnitude less dynamic power. These findings demonstrate the viability of RL on neuromorphic hardware and highlight its promise for building energy-efficient, real-time, embedded AI systems.
|
| 1667 |
SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation
2609.32391
|
cs.LG
|
Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang |
Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. ...Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training frameworks schedule only the benchmark's own events, leaving each benchmark and agent pair to build a custom scheduling loop. We present SCLATE, an execution substrate where benchmarks and unmodified agents each add their events to one open event scheduler through an adapter. A hybrid simulated clock runs these events on a shared timeline, flowing in real time while the agent works and skipping idle gaps, which compresses a month-long scenario into hours. SCLATE also serves as a rollout engine that runs any agent's harness and memory unmodified, recording the tokens and log probabilities of every model call through an in-container proxy. We port seven benchmarks to SCLATE and compare ten unmodified harness and memory configurations head to head on ten models. The comparison shows that an added memory system does not reliably beat the harness's native memory and that models differ widely in how they use the same harness and memory. We then post-train Qwen3.5-4B through unmodified harnesses and memory systems. The model learns to use both, reading 6.8x fewer file lines with a 16.7-point higher SWE-bench Verified pass rate, and writing richer memory records, while its held-out MetaClaw accuracy rises by up to 11.8 points.
|
| 1668 |
Convergent Plug-and-Play Image Restoration with Annealed Noise Levels
2609.32393
|
cs.LG
|
Samuel Hurault |
Plug-and-Play (PnP) methods solve imaging inverse problems by incorporating deep denoisers into iterative optimization algorithms. Although practical implementations often decrease the denoiser noise level $\sigma$ along iterations, most existing convergence a...Plug-and-Play (PnP) methods solve imaging inverse problems by incorporating deep denoisers into iterative optimization algorithms. Although practical implementations often decrease the denoiser noise level $\sigma$ along iterations, most existing convergence analyses assume a fixed denoiser. In this work, we establish convergence guarantees for a broad family of Plug-and-Play algorithms with annealed noise level, spanning deterministic methods (RED--GD and PnP--PGD) and stochastic methods (SNORE, equivariant RED, and a variant of PnP--Flow). For each method, we identify an explicit, nonconvex objective associated with the terminal denoising level and prove asymptotic stationarity of the iterates with respect to this objective. Our analysis does not prescribe any decay rate for the noise schedule, and our assumptions cover both learned gradient-step denoisers and exact MMSE denoisers. Overall, our theoretical results bridge the gap between existing PnP convergence theory and the decreasing-denoising practices used by state-of-the-art image restoration methods. We empirically demonstrate the benefits of such schedules and illustrate the predicted convergence behavior on several imaging inverse problems, including inpainting, super-resolution, demosaicing and tomography.
|
| 1669 |
Memory as a cache: Exact context reuse and deletion by construction
2609.32395
|
cs.LG
|
Shengyao Wang, Jiang Liu |
The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes...The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose context representation is a cache by construction. A block-local encoder maps each block to memory rows independently of other blocks, and a reader conditions generation on their union through cross-attention. For every parameter setting, memory composes exactly at fixed block indices, deleting a block is an exact $O(b)$ update for $b$-token blocks, and the memory state is independent of the edit path. At $4\times$ the training context, under the shared recipe, SMem retrieves planted needles beyond any trained-length window (exact match 0.14-0.28 at distances of 31 and 63 blocks), where learned-position, RoPE, and Block-Attention-style transformers all score at most 0.02. A fully cached context is served by computing one block alone at a near-constant 3.1-6.2 ms, whereas cold prefill grows with context; batched decode stores 34-38% fewer KV rows and runs 1.4-1.7$\times$ faster when bandwidth-bound; and deletion beats suffix recomputation by 8.5$\times$ at 512 blocks and 452$\times$ at 4096 blocks (32-256$\times$ the trained length, probing the cost model rather than a served regime). The cost is a perplexity gap of -4.7% to +2.8% (negative favors SMem) against a parameter-matched transformer with the same positional scheme, at 160M-1.5B on FineWeb-Edu across two recipes and a learning-rate search. SMem also composes with RoPE: at 160M and 410M the composite matches or leads the matched transformer and closes 29-59% of SMem's gap to a RoPE transformer. Dropping prefix entanglement thus keeps perplexity comparable while making the cache exactly composable and editable.
|
| 1670 |
AECSF: Adaptive Ensemble Conditional Score Filtering for High-Dimensional Nonlinear Data Assimilation
2609.32411
|
cs.LG
|
Yangwen Zhang, Shiwei Ni, Xiaoping Zhang, Xiaofei Guan, Lili Ju |
Bayesian state estimation for high-dimensional nonlinear dynamical systems entails a fundamental tension between statistical fidelity and computational tractability, as particle weights can collapse, while Gaussian ensemble updates can miss non-Gaussian poster...Bayesian state estimation for high-dimensional nonlinear dynamical systems entails a fundamental tension between statistical fidelity and computational tractability, as particle weights can collapse, while Gaussian ensemble updates can miss non-Gaussian posterior structure. Score-based diffusion filters offer a sampling-based alternative, but existing training-free score filters often rely on heuristic likelihood corrections, which can compromise posterior accuracy by neglecting uncertainty about the system state associated with each noisy reverse particle. To address these issues, we propose AECSF, a training-free adaptive ensemble conditional score filter. AECSF constructs an analytically tractable score estimator from the conditional Tweedie identity, which recasts noisy posterior score estimation as estimating the conditional mean of the system state given a noisy reverse particle and the observation. To estimate these conditional means efficiently, AECSF employs a shared adaptive weighted proposal ensemble, while particle-specific conditional weights yield an estimate for each noisy reverse particle without separate proposal sampling. The proposal ensemble is updated using reverse-particle information within the same reverse-diffusion run to improve conditional-mean estimation. Theoretically, we characterize when a fixed weighted proposal measure yields the exact noisy posterior score. Under stated assumptions, we establish a bound relating conditional-mean estimation errors to reverse-sampling endpoint error. Numerical experiments demonstrate that AECSF improves the accuracy of posterior sampling and nonlinear filtering in high-dimensional problems with limited forecast ensembles.
|
| 1671 |
RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation
2609.32416
|
cs.LG
|
Jiawei Zhang, Xiangrong Zhang, Rui Song, Huanbin Zhou, Chengye Song |
Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the t...Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thus misassign training credit on states where the teacher fails. We propose RE-0, a recursively verified policy improvement framework: rather than assuming that the teacher globally outperforms the student, RE-0 requests local corrections from the teacher on the student's own failure histories and checks in the environment whether each correction is genuinely beneficial; verified corrections yield immediate improvement. Building on this, we propose RE-OPD, which turns verified interventions into supervision for on-policy distillation. Only counterfactually verified teacher interventions provide distribution-level supervision, weighted by their measured local benefit, and the improvement they induce is projected back into the standalone student, so both where supervision is applied and how much credit the teacher receives co-evolve with the student policy. We further prove that the student's per-round gain is lower-bounded by its verified intervention gain up to verification and projection error terms. Experiments across multiple Code-as-Policy embodied tasks show that RE-0 improves both teacher-assisted execution and the standalone student, and generalizes to novel robots and scenes.
|
| 1672 |
PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers
2609.32429
|
cs.LG
|
Yanlong Chen, Yining Chen, Song Zhang, Amirhossein Habibian, Yawei Li |
Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group su...Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The affine offsets represent the energy in this subspace without widening the range within the group. We formulate rotation design as a Ky Fan trace maximization and derive a closed-form solution that is provably optimal for this alignment objective. Compact Householder transformations and their compact-WY representation enable gradient-free construction and efficient application at both foldable and online sites. A predictive range law further connects unaligned activation energy and group size to quantization-relevant variation. Experiments on Llama, Qwen, and Mistral span dense models up to 70B parameters and a 30B mixture-of-experts model. Under W4A4KV4, PrismQuant sets the state of the art on Llama-3.2-3B among the compared methods in both perplexity and accuracy. On Llama-3.1-70B, it attains 3.85 perplexity and 72.46% average zero-shot accuracy, only 0.22 percentage points below full precision. In the deployment study on Llama-3.1-8B, our optimized implementation achieves 1.51x prefill and 1.22x CUDA Graph decode speedups over matched FP16 baselines, with 56.34% lower decode peak memory and only 2.35% additional Graph decode latency over Hadamard. Code is available at https://github.com/ForeverBlue816/PrismQuant.
|
| 1673 |
Controllable GNN Explanations via Multi-Metric Preference Selection
2609.32436
|
cs.LG
|
Rachit Verma, Yashraj J. Deshmukh, Anirban Dasgupta |
Mechanisms for generating GNN explanations are crucial for building trust and mitigating biases in Graph Neural Networks (GNNs), especially in high-stakes scenarios. Most current methods optimize only for fidelity under the sparsity constraint. However, this d...Mechanisms for generating GNN explanations are crucial for building trust and mitigating biases in Graph Neural Networks (GNNs), especially in high-stakes scenarios. Most current methods optimize only for fidelity under the sparsity constraint. However, this discounts the need for interpretable explanations (those that consist of familiar motif patterns) and stable explanations (those that remain unchanged under structural perturbations). We propose a novel approach that optimizes GNN explanations across these metrics, exposing their relative weighing as a control. Experiments on various real-world datasets, including MUTAG, BA-2Motif, BAMultiShapes, and PROTEINS, suggest that our method produces higher-fidelity explanations than a state-of-the-art baseline on MUTAG and PROTEINS across all evaluated budgets, and on BA-2Motif at larger budgets, while being faster in the regime of small explanation budgets. We also explore how, given an input motif library containing standard motifs for the corresponding domain, the method can be used to determine the relative importance of those motifs in generating the explanations, and how this information can be used to further improve the quality of the output explanations. We also examine the relationship between different metrics through their induced tradeoff surface, and explore its dependence on the nature of the motif library.
|
| 1674 |
De-biasing Skeleton-based Action Recognition with Convex Hull Adaptive Shift
2609.32454
|
cs.LG
|
Mengyuan Liu, Yuhang Wen, Yi Zhang, Songtao Wu, Hong Liu |
Skeleton sequences can represent both individual actions and multi-entity interactions, encompassing human bodies, hands, objects, and robots. Existing approaches to recognize skeleton-based actions and interactions usually adopt a late fusion strategy, which ...Skeleton sequences can represent both individual actions and multi-entity interactions, encompassing human bodies, hands, objects, and robots. Existing approaches to recognize skeleton-based actions and interactions usually adopt a late fusion strategy, which expects individuals are independent and identically distributed to train a robust weight-shared entity encoder. However, observed entity bias in various skeletal data violates this assumption, leading to suboptimal optimization of backbone models that might produce wrong recognition results. This bias arises from the world coordinate system's initial configuration, where the choice of origin often creates bias in representation. To this end, we propose a Convex Hull Adaptive Shift based normalization method to reduce Entity bias (CHASE), improving performance across a variety of skeleton-based action and interaction recognition tasks. To adaptively apply plausible shifts to the input skeletons, we formulate a plug-and-play parameterized network that ensures the relocated world origin lies within the skeleton convex hull, which avoids non-convergence by limiting the search space. To further minimize entity bias, we incorporate an auxiliary objective that leverages pair-wise distribution distances to guide network optimization. To support both single- and multi-entity actions, we propose a sub-entity strategy that offers a consistent formulation for both scenarios. Moreover, CHASE demonstrates compatibility with various intra-skeleton modalities, such as bones and velocities, highlighting its adaptability. Essentially, our method works as a normalization approach to reduce entity bias, enabling subsequent classifiers to achieve improved recognition performance across diverse settings. Extensive experiments on 7 datasets verify our approach by seamlessly integrating with various backbones and significantly boosting their performance.
|
| 1675 |
Agnostic Smoothed Online Regression with Adversarial Responses
2609.32478
|
cs.LG
|
Xuanyu Chen, Yue Yu |
We study smoothed online prediction with bounded adversarial responses. This widely studied framework bridges i.i.d. sampling and adversarial covariate selection through a smoothness parameter $\textsf{C}_{\textsf{cov}}$, which bounds conditional covariate den...We study smoothed online prediction with bounded adversarial responses. This widely studied framework bridges i.i.d. sampling and adversarial covariate selection through a smoothness parameter $\textsf{C}_{\textsf{cov}}$, which bounds conditional covariate densities relative to a fixed, unknown base measure. We propose \textsc{Hedge-Cover}, an information-theoretic algorithm that achieves sublinear regret $\widetilde{O}(\sqrt{\text{Pdim}(\mathcal{F}) \textsf{C}_{\textsf{cov}} T})$ for function classes with bounded pseudo-dimension. The algorithm aggregates a carefully constructed family of experts using \textsc{Hedge}, with a prior that links regret to the number of disagreements between a consistent selector and a target function. We bound this number by exploiting covariate smoothness. This answers an open problem posed in \cite{blanchard2025agnostic} on the minimax optimal adaptive regret of the smoothed online regression problem. We establish a matching lower bound for the class of linear predictors. The main intricacy of the lower bound lies in explicitly constructing a challenging sequential covariate distribution supported on mutually orthogonal hyperplanes. This construction may be of independent technical interest. Finally, we revisit the well-specified setting and quantify the effect of response noise. For conditionally $\nu^2$-subGaussian responses, we extend the existing lower bound under realizable responses by showing that the minimax expected regret is $\Omega((1\vee \nu)\sqrt{(\textsf{C}_{\textsf{cov}}-1)dT})$ for a function class of VC dimension $d$. A corresponding upper bound for ERM matches this dependence on $\nu$.
|
| 1676 |
JEPA Learns What the Mask Leaves Unrecoverable
2609.32481
|
cs.LG
|
Peng Xie, Amr Alanwar |
Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly support...Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose support lies inside the hidden region falls in the measurement's null space and leaves no trace in the data. The JEPA loss asks only that the encoded context suffice for the target, so a target that a low-level prior can recover admits a shortcut, one the moving-average target encoder can make self-consistent. What removes the shortcut is the coarse-scale content the mask leaves unrecoverable, provided enough context stays within reach of each target. We score that content before training and test the account's distinctive predictions in 151 pre-training runs. On ImageNet-100, strip masks match blocks in area and contiguity yet are recoverable, and they land at 40.3% linear top-1, beside random masks at 40.8%, against 64.3% for blocks; within one geometry family, the placements that leave the least unrecoverable content lose 6.5 points to those that leave the most, over five seed pairs; pixel targets span 7 points where latent targets span 25; and against a frozen target the gap between random and block masks, 19 points on the same kind of GPU, closes to 1.5, so the geometry acts through the target the encoder produces for itself. On UCF101 the masking ratio decides which condition, content or reach, binds; removing whole frames, unrecoverable in space but recoverable from neighbouring frames, is worst at both ratios; and on V-JEPA's own masks, batching them intact instead of truncated changes little (36.0% against 35.1%), whereas making 100 target tokens inside the blocks visible lifts them to 48.7% and hiding 100 context tokens outside the blocks does not (33.7%).
|
| 1677 |
When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections
2609.32488
|
cs.LG
|
Maojun Sun, Yancheng Yuan, Jian Huang, Ruijian Han |
Dense retrieval powers retrieval-augmented generation, semantic search, and question answering, yet the theoretical basis for choosing between shared and dual query-document projections remains unclear. We introduce a bias-variance theory for low-rank bilinear...Dense retrieval powers retrieval-augmented generation, semantic search, and question answering, yet the theoretical basis for choosing between shared and dual query-document projections remains unclear. We introduce a bias-variance theory for low-rank bilinear scoring. Shared projections induce positive-semidefinite operators, whereas dual projections realize arbitrary low-rank operators. We derive their exact approximation gap and prove a local Gaussian boundary: dual has lower risk exactly when squared directional signal exceeds the estimation cost of its additional degrees of freedom. This boundary motivates the Cross-fitted Asymmetry Risk Selector (CARS), which estimates reproducible directional signal from training pairs; its Gaussian counterpart admits exact selection-power and regret formulas. Guided by the theory, we run retrieval experiments across multiple datasets and embedding models. The mean Dual-minus-Shared NDCG@10 advantage more than doubles as query rotation increases from 0 degrees to 90 degrees. In the rank-sample-size grids, Shared wins 13 of 16 cells at n=32, whereas Dual wins all 32 cells at n=1024 and n=2048. Consistent with this shift, all 168 comparable operator-risk curves move toward Dual as training data grow. Compared to the two fixed-geometry baselines, CARS reduces held-out regret by 49-96% and achieves 90.1% mean geometry-selection accuracy.
|
| 1678 |
TreeRef-BFN: Equivariance-Free De Novo Molecule Generation based on 2D Topology and Internal 3D Geometry
2609.32502
|
cs.LG
|
Ruiqing Sun, Sen Yang, Dawei Feng, Bo Ding, Yijie Wang |
De novo 3D molecular generation jointly models molecular size, topology, and geometry. Most methods pre-sample molecular size and generate Cartesian coordinates, limiting variable-size conditional tasks such as fragment completion and scaffold decoration while...De novo 3D molecular generation jointly models molecular size, topology, and geometry. Most methods pre-sample molecular size and generate Cartesian coordinates, limiting variable-size conditional tasks such as fragment completion and scaffold decoration while often relying on equivariant architectures. Internal-coordinate methods avoid rigid-body redundancy but typically require a known molecular graph or autoregressive construction, which may accumulate errors. We propose TreeRef, a tree-based molecular representation that assigns molecular topology and topology-dependent local 3D geometry to a naturally variable-size tree. RingRef nodes encode ring closures while preserving the tree structure, while Null nodes allow molecular size to emerge directly from node occupancy. Based on TreeRef, we develop TreeRef-BFN, a Bayesian Flow Network with a standard Transformer backbone that globally couples these locally defined variables and jointly generates discrete molecular variables and continuous local geometry. A single pretrained TreeRef-BFN supports unconditional generation and variable-size structure-conditioned 3D generation through masking alone, without retraining. Empirical studies demonstrate strong chemical validity, molecular stability, and diversity, accurate local geometric distributions, fast sampling, and competitive property-conditioned generation, establishing TreeRef-BFN as an efficient and flexible framework for 3D molecular generation.
|
| 1679 |
Fast Differentiable SVD on GPU via Polar Decomposition
2609.32505
|
cs.LG
|
Uliana Parkina, Askar Tsyganov, Sergei Kudriashov, Sergey Samsonov, Maxim Rakhuba |
We present a fully GPU-oriented SVD pipeline based on polar decomposition, motivated by iterative methods that rely solely on matrix multiplications, such as the Newton-Schulz iteration. We show that this approach enables up to a $2\times$ speedup compared to ...We present a fully GPU-oriented SVD pipeline based on polar decomposition, motivated by iterative methods that rely solely on matrix multiplications, such as the Newton-Schulz iteration. We show that this approach enables up to a $2\times$ speedup compared to standard implementations. Furthermore, we derive a numerically stable backward pass for the polar decomposition and leverage it to obtain a fully differentiable SVD. Our methods are released as open-source implementations in both PyTorch and JAX: https://github.com/fallnlove/cans_svd.
|
| 1680 |
REFINE: A Resilient Evolution Framework for Intelligent Enterprise Alert Triage in Security Operations Centers
2609.32516
|
cs.LG
|
Huimin Chen, Quan Long, Yanhao Wang |
Security Operations Centers (SOCs) process large volumes of alerts daily. Alert triage prioritizes high-risk threats while reducing manual review of benign alerts. LLM agents can reason over logs and threat intelligence, but struggle to keep aligned with organ...Security Operations Centers (SOCs) process large volumes of alerts daily. Alert triage prioritizes high-risk threats while reducing manual review of benign alerts. LLM agents can reason over logs and threat intelligence, but struggle to keep aligned with organization-specific, rapidly evolving SOC operational standards. We introduce REFINE, an LLM-agent framework for enterprise alert triage. REFINE encodes analyst expertise as structured skills and continuously adapts using analyst disposition feedback. It enforces recall = 1.0 as a hard constraint during evolution to maximize auto-closure of false positives, and identifies judgment blind spots by combining alert distributions with model error boundaries. Evaluated on four real industrial SOC scenarios across four MITRE ATT&CK phases with temporal split: REFINE achieves recall=1.0 on all evolution sets. On future test windows, it retains recall=1.0 in three scenarios; the degraded case reaches 0.807 recall, still outperforming self-evolution baselines (0.49-0.58).
|
| 1681 |
LocalProp: Neuro-Localized Memory-Efficient Backpropagation
2609.32517
|
cs.LG
|
Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu |
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since ...The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly localized. Therefore, we propose LocalProp, a training procedure that locally updates the weights of a model. Our neuro-localized weight updates follow the "pre-training then fine-tuning" paradigm, where the pre-training is based on I-JEPA. After locally updating the weights, a pruning operation is performed, followed by a short final fine-tuning phase. Pruning helps by sending the learning signal from higher blocks to lower blocks. We perform experiments on several datasets, including large-scale benchmarks such as ImageNet, and empirically show that LocalProp reaches good performance at a fraction of GPU peak memory. By varying the number of jointly optimized blocks, we identify gradient-propagation span as a practical control over the accuracy-memory trade-off.
|
| 1682 |
STR: Supervised Transcoder Replacement for Reducing Steering Side Effects
2609.32519
|
cs.LG
|
Haonan Yu, Junhao Liu, Zhenyu Yan, Haoran Lin, Xin Zhang |
Model steering can strengthen a target behavior while degrading other useful behaviors. We introduce Supervised Transcoder Replacement (STR) to reduce these side effects for existing steering methods, including those fitted without a protection objective. STR ...Model steering can strengthen a target behavior while degrading other useful behaviors. We introduce Supervised Transcoder Replacement (STR) to reduce these side effects for existing steering methods, including those fitted without a protection objective. STR learns a replacement for the multilayer perceptron (MLP) computation at the steering layer through supervision for target control, non-target preservation, and fidelity without steering. Selected steering methods then fit directions on the frozen replacement while retaining their own fitting objectives. We evaluate three steering methods across Gemma and Llama models using Corrigibility preferences and four harmful-request safety datasets. SALAD-Bench supplies protection training data and a separate in-distribution evaluation split; HarmBench, AdvBench, and StrongREJECT are reserved for out-of-distribution testing. STR substantially reduces steering side effects on the in-distribution evaluation and extends this protection to the unseen safety datasets while retaining effective target control. For target-only supervised steering vectors, pooled out-of-distribution attack success rate falls from 42.46% to 14.42% on Gemma-3-4B and from 34.97% to 12.91% on Gemma-3-12B. These results show that replacement training can benefit steering methods fitted without protection objectives.
|
| 1683 |
AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses
2609.32527
|
cs.LG
|
Sikai Huang, Zhiwen Yang, Kai Yu, Jiayuan Chen, Stan Z. Li |
Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what...Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot separate target-specific predictions from a shared background response. Second, common metrics remain high under gene shuffling, so gene-level accuracy is never verified. Third, a score at one training size says nothing about coverage, which depends on representation-space proximity and response-constraining power. We propose AmbiModBench, a specificity-aware, gene-resolved and coverage-aware benchmark. It pairs every score with a training-mean reference fitted on the same split, screens each readout by gene-coordinate permutation, and links embedding distance to response variation. Across K562, RPE1 and Norman, strong absolute scores largely reflect shared background rather than target-specific learning. Widely used readouts track response magnitude distributions rather than the affected genes. Detectable gain follows representation-space coverage rather than training-set size. Nonetheless, on RPE1 the protocol yields a reproducible target-specific gain across five additional splits and three gene selections, which absolute scores alone cannot distinguish from shared background.
|
| 1684 |
DepthBench: Measuring How Residual Connections Enable More Computational Depth
2609.32534
|
cs.LG
|
Keyu Wang, Yangyi Huang, Jiale Kang, David Gonz\'alez-Mart\'inez, Weiyang Liu |
Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (\text{e.g.}, LayerNorm Scaling) or residual connections (\text{e.g...Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (\text{e.g.}, LayerNorm Scaling) or residual connections (\text{e.g.}, mHC, AttnRes) to enable better information flow and depth utilization. However, it remains unclear whether they truly translate increased architectural depth into effective computational depth, and whether their reported gains stem from better access to information across depth, or unaccounted-for confounding factors. In this paper, we introduce \textbf{DepthBench}, a controlled benchmark for studying computational depth across various architectures. We systematically vary the width--depth aspect ratio ($d_{\text{model}}/n_{\text{layer}}$) from shallow--wide to deep--narrow shapes, while keeping the model size and pre-training recipe fixed. Across 10 representative architectures, we find that the benefit of allocating more capacity to depth is strongly architecture-dependent. Standard Pre-LN and most of its norm- and scaling-based variants provide little benefit and can even degrade performance as models become deeper and narrower, whereas HC and Full AttnRes improve consistently even at extreme deep shapes. These gains extend beyond pre-training loss and consistently translate into improved domain-specific performance and effective computation. Controlled layer-level analyses further show that the gains of HC and Full AttnRes are associated with more effective utilization of additional layers, revealing distinct mechanisms of computational depth across architectures. Overall, our results identify residual connection design as a key determinant of whether depth can serve as a meaningful scaling axis by enabling additional architectural depth to translate into effective computation.
|
| 1685 |
Interpretable Physics Informed WiFi Indoor Localization: Learning an Effective Access Point Geometry and Using It to Prune
2609.32539
|
cs.LG
|
Arshia Eftekhari zadeh, Rezvan Nasiri, Hadi Moradi |
Deep learning models can achieve high accuracy for indoor localization, but their black-box nature limits interpretability and the reuse of learned information. We propose a hierarchical deep learning framework for WiFi fingerprint-based indoor localization th...Deep learning models can achieve high accuracy for indoor localization, but their black-box nature limits interpretability and the reuse of learned information. We propose a hierarchical deep learning framework for WiFi fingerprint-based indoor localization that jointly predicts user location and learns an effective geometry of the surrounding access points (APs). Physics-informed decoders infer this geometry directly from RSSI measurements and labelled user positions, without requiring the true AP coordinates during training. The learned geometry is then used to rank and prune APs. On the UJIIndoorLoc dataset, the proposed chained model achieves a mean 3D localization error of 7.07 m, reducing error by 26% to 36% compared with baseline models. Previously published methods evaluated on the same official split report errors 10.6% to 31.0% higher. Pruning 35% or 50% of the APs causes only a small loss in localization accuracy. The inferred geometry also enables Fisher-information-based AP ranking even when fingerprint databases do not contain surveyed AP coordinates. Experiments on the Tampere/TUT and UTSIndoorLoc datasets show that geometry-guided AP selection performs comparably to selectors built directly from labelled data. These results show that physics-informed interpretability can improve indoor localization while also supporting effective feature selection.
|
| 1686 |
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion
2609.32540
|
cs.LG
|
Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos |
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconst...Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value (KV) cache by additional forwards to build the cache without advancing an output latent. However, every denoising forward itself already computes the in-flight KV of the current chunk. We introduce FlashForward, which directly reuses this cache to avoid the heavy cache-update-only model forwards. After the current chunk completes one denoising stage, its stage-specific cache is already available for the next chunk. Assigning one GPU to each stage therefore lets different chunks occupy different stages concurrently. This early availability has a quality cost: the resulting stage-matched history is noisy, causing appearance and motion drift among chunks. To complement it, FlashForward produces sparse auxiliary clean anchor latents before the corresponding region is generated so the generation trajectories can be stabilized by this two-sided conditioning. The two memories operate at different temporal scales: sparse clean anchor KV supplies coarse, long-range two-sided structural guidance, while dense stage-matched history preserves fine, recent evolution. With up to four GPUs, FlashForward runs $1.16$--$1.69\times$ faster than HiAR and $1.42$--$2.92\times$ faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbone scales at 480p and 720p. On VBench, for the 1.3B model at 480p, it achieves higher scores and remains stable at longer durations, demonstrating that FlashForward generates high-quality and temporally consistent videos across durations of 20s, 35s and 65s at a much faster generation speed.
|
| 1687 |
Are Vision-Language-Action Models Robust to One-Step Observation Perturbations?
2609.32550
|
cs.LG
|
Shojiro Yamabe, Jun Sakuma |
Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an e...Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an episode. However, momentary observation corruption, in which observations are severely perturbed only briefly within an episode, remains an underexplored safety threat. To address this gap, this work investigates robustness to one-step perturbations applied at a single time step per episode. Our experiments reveal that these perturbations substantially degrade VLA performance and that their impact depends on the action chunk execution length. Based on them, we propose CARE, which dynamically selects the execution length based on consistency with the previously predicted action chunk. CARE improves robustness with low computational overhead while preserving clean performance by selecting shorter execution lengths only under perturbations.
|
| 1688 |
Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC
2609.32591
|
cs.LG
|
Yi Xian Goh, Sze Jue Yang, Hao Luan |
Data-driven model predictive control (MPC) combines learned world models with online trajectory optimization, achieving strong performance in continuous control. However, the per-step cost of sampling and evaluating hundreds of candidate trajectories restricts...Data-driven model predictive control (MPC) combines learned world models with online trajectory optimization, achieving strong performance in continuous control. However, the per-step cost of sampling and evaluating hundreds of candidate trajectories restricts deployment to control frequencies well below what real-time robotics demands. Motivated by the dual-process theory of human cognition, which distinguishes between fast, intuitive processing (System 1) and slower, deliberative reasoning (System 2), we ask whether every decision requires the same degree of computational deliberation. We propose Fast-TD-MPC, a lightweight framework that adaptively routes between fast policy execution and test-time planning, reserving costly deliberation for states where it is most needed. Fast-TD-MPC delivers competitive task performance across 103 continuous control tasks while achieving up to ~4x faster inference. Under external disturbances, Fast-TD-MPC selectively falls back to planning, maintaining robustness comparable to the original planner.
|
| 1689 |
Learning What to Evaluate: Correlation-Aware Decoupling for Multiobjective Bayesian Optimization
2609.32632
|
cs.LG
|
Ashwin Renganathan, Peter Bachman |
Multiobjective Bayesian optimization (MOBO) with Gaussian process (GP) surrogates is a sample efficient approach to solving multiobjective optimization problems. In MOBO, a Bayesian decision theoretic acquisition function guides the adaptive selection of new c...Multiobjective Bayesian optimization (MOBO) with Gaussian process (GP) surrogates is a sample efficient approach to solving multiobjective optimization problems. In MOBO, a Bayesian decision theoretic acquisition function guides the adaptive selection of new candidate inputs, on which objectives and constraints are evaluated to update the surrogate model sequentially. Existing approaches maintain independent GP models for the objectives and constraints, with new observations evaluating all objectives and constraints in a coupled fashion. However, the objectives and constraints often contain inherent correlations which, if exploited, can enable ${decoupled}$ evaluations where only a subset of them are evaluated at each round. We present a new approach that leverages a multitask GP model to jointly learn all objectives and constraints, and propose a total correlation metric that enables identifying an ${optimal}$ subset of objectives and constraints to be evaluated at every round, even under uniform evaluation costs. Theoretically, we show that our acquisition policy is asymptotically consistent despite decoupling and that our proposed decoupled subset selection rule maximizes the expected posterior entropy reduction about unevaluated tasks under mild conditions. Empirically, we show that our approach outperforms coupled and decoupled baselines in the state of the art.
|
| 1690 |
Schur-Neural KF: Learned Schur-Consistent Corrections to the Extended Kalman Filter
2609.32640
|
cs.LG
|
Min Kim, Lianghao Cao, Soon-Jo Chung, Andrew M. Stuart |
We present Schur-Neural KF (SN-KF), a learning-based correction to the extended Kalman filter (EKF) that preserves the probabilistic conditioning interpretation of the EKF. The method perturbs the predictive state-measurement cross-covariance and the Cholesky ...We present Schur-Neural KF (SN-KF), a learning-based correction to the extended Kalman filter (EKF) that preserves the probabilistic conditioning interpretation of the EKF. The method perturbs the predictive state-measurement cross-covariance and the Cholesky factor of the measurement noise covariance so that the resulting joint predictive covariance is always positive semidefinite. The positive semidefiniteness is ensured by a Schur complement-based parametrization. We instantiate the parametrization with a recurrent neural architecture whose matrix outputs are modulated by amplitude gates. We prove that incorporating a measurement does not increase the filter's state uncertainty, and show that no measurement can induce an arbitrarily large state correction relative to its statistical surprise. We also present a perturbative analysis suggesting SN-KF's structural strength in the data-scarce regime. We provide two numerical experiments to illustrate the practical benefits of SN-KF. In a two-radar experiment, enforcing Schur-consistency provides a much broader failure-free hyperparameter region and reduces RMSE for small training subsets, consistent with our theoretical analysis in the data-scarce regime. In the unicycle experiment, SN-KF achieves the best precision, recall, false alarm rate, and gated RMSE under innovation-based sensor-fault rejection.
|
| 1691 |
Prediction Limits and Koopman Closure of Geometry-Induced Soft State Abstractions
2609.32652
|
cs.LG
|
Mohit Kumar, Somayeh Kargaran |
We study when geometry-induced soft state abstractions admit accurate finite-dimensional linear dynamics. Each state is represented by simplex-valued coordinates obtained from class-specific Kernel Affine Hull Machine (KAHM) reconstruction scores, and a matrix...We study when geometry-induced soft state abstractions admit accurate finite-dimensional linear dynamics. Each state is represented by simplex-valued coordinates obtained from class-specific Kernel Affine Hull Machine (KAHM) reconstruction scores, and a matrix is used to predict the next-state coordinates. Our main result is a computable lower confidence bound on the minimum root-mean-square prediction error over all matrices satisfying a prescribed spectral-norm limit. The bound combines within-class variation of successor coordinates with the deviation of soft coordinates from their one-hot reference labels, and can be evaluated from independent state-successor pairs without fitting a prediction matrix. For fixed coordinates and evaluation distribution, the certificate converges almost surely to a population lower bound as the sample size grows; any tolerance below this limit is eventually certified unattainable. Reconstruction-score margins further control the soft-to-hard assignment error. Under deterministic dynamics and exact coordinate closure, eigenvectors of the closure matrix and its reduced transpose induce Koopman and adjoint Koopman eigenfunctions, respectively. A four-state KAHM construction shows that identical soft coordinates can permit exact closure under one dynamics map yet force positive prediction error under another. Experiments on Duffing, Van der Pol, CartPole, MountainCar, and Acrobot compare direct soft-coordinate prediction with state-space DMD/EDMD baselines and report prediction, representation-variation, and spectral diagnostics. The benchmarks assess fitted models but do not numerically evaluate the exclusion certificate.
|
| 1692 |
MixBench-TS: A Multivariate Time Series Forecasting Benchmark Where Channel Mixing Pays Off
2609.32656
|
cs.LG
|
Ibram Abdelmalak, Mischa Putzke, Jungmin Choi, Tom Hanika, Vijaya Krishna Yalavarthi |
Multivariate Time Series Forecasting (MTSF) models that mix information across channels assume that the past of one channel carries information about the future of another. Yet they are evaluated on a small fixed set of standard datasets whose cross-channel st...Multivariate Time Series Forecasting (MTSF) models that mix information across channels assume that the past of one channel carries information about the future of another. Yet they are evaluated on a small fixed set of standard datasets whose cross-channel structure is rarely examined. We ask two questions: "How can we reliably measure lagged, non-linear, and joint coupling in MTSF datasets?" and "Do the standard datasets actually have such coupling?" To answer the first, we test four candidate measures on synthetic datasets with planted ground-truth coupling: Granger Causality (GC), Transfer Entropy (TE), lagged Mutual Information (MI), and the CD gain, a model-based measure we introduce that compares a channel-dependent (CD) model to its channel-independent (CI) variant. Only lagged MI and the CD gain recover every planted coupling. For the second question, the answer is a definite no, as the standard datasets have a median of only 23% lagged-coupled channel pairs and a median CD gain of -4.9%, compared to 78% and +4.6% on chaotic ODE systems. We therefore propose MixBench-TS, a benchmark of 10 real-world datasets with a median of 55.5% lagged-coupled pairs and a median CD gain of +1.7%. Across six state-of-the-art models tuned under one protocol, CI models win on 10/10 (MSE) and 8/10 (MAE) standard datasets, but on only 3/10 and 2/10 MixBench-TS datasets. We recommend using our benchmark for evaluating new CD models. Moreover, we propose profiling new datasets with lagged MI and the CD gain before using them to evaluate multivariate models. Code and data are available at https://anonymous.4open.science/r/mixbench-ts-B027.
|
| 1693 |
What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation
2609.32670
|
cs.LG
|
Arshia Eftekhari zadeh |
When a language model explains an answer it has already given, does it reuse the computation that produced the answer or reconstruct a story from the answer alone? Attribution, transportability and recoverability are each compatible with causal use without est...When a language model explains an answer it has already given, does it reuse the computation that produced the answer or reconstruct a story from the answer alone? Attribution, transportability and recoverability are each compatible with causal use without establishing it. We propose an evidence standard: pair each positive statistic with a variable specific null that removes the tested variable's identity while matching relevant nuisance dimensions as far as possible, and audit unmatched dimensions. We apply this standard to a known cause. A cue naming a wrong option raises the rate of choosing that option by 64 to 68 percentage points across three models. Explanations mention the cue in 1.8 percent of items or fewer in three of four models tested. Three estimator classes yield favorable statistics, but none establishes causal sensitivity to the cue contrast under its own control in the three-model analysis. In the strongest case, a recovered cue direction reaches $R^2$ of 0.95 and exceeds a geometry matched random direction in all three seeds, while a direction fitted by the same pipeline with cue labels scrambled reproduces 61 to 76 percent of its effect at comparable realized edit magnitude. A fourth model passes one interchange endpoint, but unequal edit magnitudes and a contrast that changes both cue identity and cue-answer agreement limit its interpretation. These experiments leave causal access unresolved. They establish an evidentiary requirement: favorable mechanistic statistics must survive controls for variable identity and nuisance structure. Reusable controls separate generic from identity specific transport effects, fit null directions with scrambled labels, and audit realized intervention magnitudes.
|
| 1694 |
When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds
2609.32677
|
cs.LG
|
Ke Wang, Zijie Zhao, Zhiyi Yuan, Changlun Li |
Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the...Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks policies well overall. We formalize this gap as Improvement Fidelity, which asks whether proxy improvements preserve the sign and ordering of deployment improvements over the updates an improvement process actually proposes. We show that global policy accuracy need not guarantee update fidelity: operator shift and deployment response can create update-level errors, while candidate margins determine whether those errors change the replacement decision. We introduce PIVOT-KG, a paired, decision-aware validator that allocates scarce high-fidelity evaluation according to the expected reduction in selection regret per unit cost. Across 90 held-out roots in Leduc, Kuhn, and Melting Pot, proxy and deployment optimal sets are disjoint in 51 cases. In an eight-candidate HighwayEnv stress test, PIVOT-KG reduces mean improvement-selection regret from 0.0435 under the exact Uniform validation rule to 0.0055 at the primary budget. Together, these results show why reliable self-improvement should evaluate proposed improvements in the worlds they induce, while providing a practical rule for allocating scarce deployment evidence when it can affect the replacement decision.
|
| 1695 |
PINNMorph: Evolving Online Adaptation Policies for Physics-Informed Neural Networks
2609.32685
|
cs.LG
|
Xu Yang, Mingyang Yu, Jun Zhang, Keqian Li, Jing Xu |
Physics-informed neural networks (PINNs) provide a learning-based framework for solving partial differential equations (PDEs), yet their training behavior can change substantially throughout optimization. Residual distributions, gradient interactions, regional...Physics-informed neural networks (PINNs) provide a learning-based framework for solving partial differential equations (PDEs), yet their training behavior can change substantially throughout optimization. Residual distributions, gradient interactions, regional learning difficulty, and model-capacity requirements may evolve over time, while the network architecture and major training mechanisms are typically determined before training. We propose PINNMorph, an online PINN adaptation framework based on large language model (LLM)-guided policy evolution. PINNMorph maintains a population of state-conditioned adaptation policies that map execution diagnostics to controlled interventions over topology modification, additive representation augmentation, objective balancing, gradient handling, adaptive sampling, and optimizer-phase control. At each intervention opportunity, candidate programs are instantiated from the current policy population, selected according to the observed training state, and applied directly to the PINN under training. The resulting model inherits its existing parameters and training state and continues optimization along the same trajectory. Execution outcomes are subsequently used to evaluate interventions and evolve the policy population. Unlike pre-training architecture search or fixed adaptation rules, PINNMorph jointly adapts the current PINN and the policies governing its interventions using feedback from actual training. Experiments on 13 PDE benchmarks show that PINNMorph achieves lower solution errors than SA-PINN, ConFIG, RoPINN, HARMONIC, and PINNsAgent across all evaluated problems. Ablation studies further examine the effects of online adaptation, state-conditioned intervention selection, and execution-feedback-driven policy evolution.
|
| 1696 |
Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection
2609.32691
|
cs.LG
|
Animesh Shaw |
LLM agents that invoke privileged tools are vulnerable to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack the agent's actions. A growing body of work evaluates defenses against IPI, but the validity of that ...LLM agents that invoke privileged tools are vulnerable to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack the agent's actions. A growing body of work evaluates defenses against IPI, but the validity of that evaluation is rarely examined. We audit an IPI benchmark and its harness and identify four defect classes -- silent payload non-delivery, attack success scored by tool identity rather than arguments, false-rejection rate conflated with model incapacity, and the absence of an audit trail -- each of which yields a plausible, publishable, and incorrect number. We quantify the distortion by re-scoring identical execution traces under the defective and corrected definitions: on real agent behaviour, the tool-identity scorer reports a 21.7% attack-success rate where the true argument-level rate is 1.2%. In the sharpest case, an open model previously reported at 62.8% registers 0% under the corrected harness -- the prior figure largely an artifact of undelivered payloads and identity-level scoring. We release a harness whose construction makes each defect unrepresentable -- machine-checkable payload placement, argument-level attacker predicates, per-scenario environments, and mandatory trace persistence -- and use it to report three quantities the field does not: whether a compromised agent discloses the attack, the full security/utility operating curve of an LLM-judge defense, and tool-calling capability disentangled from defensive over-blocking. A corrected harness further overturns a reported "capability barrier": a model deemed incapable of tool use is in fact fully capable, its earlier result an artifact of environment mismatch. We argue that evaluation validity is a prerequisite for, not a footnote to, defense claims in agentic security, and provide an instrument that enforces it.
|
| 1697 |
Affine Geometry of Gaussian ReLU Networks via Conditional Kac-Rice Formulas
2609.32695
|
cs.LG
|
Recep \"Ozkan, Christian Hirsch |
We study how the affine geometry of finite ReLU networks is created at random initialization and reorganized by supervised training. We call a sign-changing zero of a hidden preactivation an activation switch and a point where the scalar network output is nond...We study how the affine geometry of finite ReLU networks is created at random initialization and reorganized by supervised training. We call a sign-changing zero of a hidden preactivation an activation switch and a point where the scalar network output is nondifferentiable a scalar kink. For one-dimensional input, conditioning on the preceding layers makes each preactivation Gaussian and affine on the cells of a random finite partition. This yields an exact finite-width conditional Kac-Rice formula for the expected number of activation switches along an input interval. For fixed depth and proportionally growing widths, the resulting switch intensities converge to explicit deterministic limits. A visibility estimate shows that the expected number of switches that do not produce scalar kinks is negligible, yielding an explicit leading formula for the expected number of scalar kinks and hence affine regions. In higher input dimensions d >= 2, the analogous conditional surface formula yields the leading expected (d-1)-dimensional Hausdorff measure of the scalar kink set. On the Breast Cancer Wisconsin data, the initialization formula accurately predicts switch counts along held-out segments. After training, switch counts decrease along within-class segments and increase along between-class segments. Thus, training redistributes rather than merely contracts affine complexity.
|
| 1698 |
Flat-Consensus Diffusion for Robust Data Reshaping under Noisy Evaluator
2609.32696
|
cs.LG
|
Hongyu Cao, Kunpeng Liu, Fei Xie, Sandip Ray |
Data shape determines how features are structured, how patterns are separated, and how distributions cover the underlying domain. Poor data shape can make models learn noise rather than generalizable structure. This paper studies robust feature-centric data re...Data shape determines how features are structured, how patterns are separated, and how distributions cover the underlying domain. Poor data shape can make models learn noise rather than generalizable structure. This paper studies robust feature-centric data reshaping: generating feature transformations that remain useful, stable, and reproducible under noisy evaluation and imperfect data conditions. We view reshaping operation sequence search as reward-guided diffusion generation, and robust reshaping as searching for regions in the latent reward landscape rather than isolated high-reward transformations. The key challenge is dual instability: noisy evaluators distort local reward guidance, while stochastic generative trajectories can converge to inconsistent solutions. We propose FCDiff, a flat-consensus diffusion framework that addresses both failures through a micro-macro decomposition. The micro layer replaces point-estimate reward guidance with Gaussian-smoothed, Monte Carlo averaged gradients, steering generation toward locally flat reward regions. The macro layer aggregates independently guided trajectories with a weighted Frechet-mean barycenter, selecting consensus-supported basins and filtering stochastic outliers. Across an 8-dataset headline cohort under heavy-tailed evaluator noise, FCDiff attains the best aggregate rank on lower-tail reliability and robustness against both search-based AutoFE and robustness-oriented generative baselines, with statistically significant accuracy gains over every generative baseline. Our results show that robust data reshaping requires searching for flat, consensus-supported regions rather than sharp single-trajectory optima.
|
| 1699 |
One-Step Generative Modeling via Unbalanced Optimal Transport
2609.32708
|
cs.LG
|
Yirong Shen, Mengfei Xia, Junpeng Jing, Lu Gan, Cong Ling |
Drifting models enable one-step generation by amortizing distribution transport into training, but this efficiency places greater demands on the transport field estimated at each update. In large-scale training, the field is computed from finite mini-batches o...Drifting models enable one-step generation by amortizing distribution transport into training, but this efficiency places greater demands on the transport field estimated at each update. In large-scale training, the field is computed from finite mini-batches of generated and real samples, which provide only imperfect approximations to the underlying distributions. Balanced optimal transport enforces exact mass matching within every mini-batch, making the estimated field sensitive to the particular composition of the real-data batch. We find that generated and real samples should be treated asymmetrically: letting the mass assigned to real samples adapt while keeping every generated sample fully transported improves generation across six feature-space metrics in controlled ablations, and is more robust to the relaxation strength than relaxing both marginals simultaneously, which falls below balanced transport under stronger relaxation. Motivated by this observation, we propose Unbalanced Optimal Transport Gradient Flow (UOT-GF), which keeps the generated-sample marginal fixed and relaxes only the real-data marginal. Under identical settings at DiT-B/2 on ImageNet-256, UOT-GF improves Fr\'echet Inception Distance (FID) from 1.53 to 1.46 over the balanced W-Flow baseline; scaling the same recipe yields 1.34 and 1.22 FID at L/2 and XL/2, the best FID among the one-step models we compare. We further derive the induced UOT transport force, establish a kinetic Vlasov--Fokker--Planck formulation whose overdamped zero-temperature limit recovers the drifting dynamics, and characterize non-target stationary states together with sufficient conditions for convergence.
|
| 1700 |
Revisiting AdaGrad in Stochastic Convex Optimization: Last Iterates, High Probability, and Lower Bounds
2609.32729
|
cs.LG
|
Weiming Ou, Xiao Wang |
AdaGrad and AdaGrad-Norm are widely used adaptive methods, but their precise behavior in stochastic convex optimization remains less understood. We first show that AdaGrad-Norm and AdaGrad do not admit any universal \textbf{last-iterate} rate, even under sub-G...AdaGrad and AdaGrad-Norm are widely used adaptive methods, but their precise behavior in stochastic convex optimization remains less understood. We first show that AdaGrad-Norm and AdaGrad do not admit any universal \textbf{last-iterate} rate, even under sub-Gaussian noise and bounded iterates. We then prove that bounded variance alone is too weak: even with bounded iterates, it cannot yield \textbf{high-probability average-iterate} rates, and without bounded iterates it may not even guarantee convergence in expectation. Besides, we construct tight $\Omega(\log T/\sqrt T)$ lower bounds for both AdaGrad-Norm and AdaGrad under sub-Gaussian noise, showing that $\log T$ in existing average-iterate upper bounds is unavoidable. Finally, we show that this $\log T$ loss disappears once bounded-iterate condition is imposed: under a general ABC condition, both methods achieve rates ${O}(1/\sqrt{T})$.
|
| 1701 |
Domain Adaptation with Target Information via Doubly-Anchored Distributionally Robust Optimization
2609.32730
|
cs.LG
|
David Kepplinger, Anand N. Vidyashankar |
Domain Adaptation (DA) often lacks worst-case guarantees, while Distributionally Robust Optimization (DRO) based only on source data centers its ambiguity set at the source law and ignores available target structure. To bridge this gap, we introduce a doubly-a...Domain Adaptation (DA) often lacks worst-case guarantees, while Distributionally Robust Optimization (DRO) based only on source data centers its ambiguity set at the source law and ignores available target structure. To bridge this gap, we introduce a doubly-anchored DRO framework whose ambiguity set is the intersection of $\phi$-divergence balls centered at the source law and a source-completed target reference law, the latter pairing the target covariate law with the source conditional law. We derive dual-induced adversarial bridge geometries for symmetric and asymmetric divergence pairings, notably introducing a Kullback--Leibler/squared-Hellinger (KL/HD) bridge. This asymmetric formulation yields a Lambert-$W$ geometry in which source-side exponential risk tilting and target-side Hellinger stabilization enter through distinct terms, attenuating, but not bounding, the effect of large likelihood ratios. Furthermore, without imposing covariate shift, we establish finite-sample generalization bounds for the minimizer of a structural, loss-agnostic density-bridge risk under bounded-overlap conditions; these bounds do not apply directly to the loss-aware DRO min--max estimator. We translate our framework into a bridge-weighted Nadaraya--Watson estimator, proving uniform consistency for the source regression function and pointwise asymptotic normality, with target recovery when the source and target regression functions coincide, as under covariate shift. Finally, an empirical evaluation on a domain-shifted Fashion-MNIST dataset illustrates the finite-sample stability of the asymmetric KL/HD bridge under severe synthetic target-covariate corruption.
|
| 1702 |
Plan-to-Synthesis: Cross-City Human Mobility Generation via Semantic Latent Flow Matching
2609.32732
|
cs.LG
|
Zhoufu Wang, Baoshen Guo, Zhiqing Hong, Junyi Li, Kailai Sun |
Human mobility generation aims to synthesize realistic point-of-interest (POI) visitation trajectories and has become an important tool for travel behavior modeling, transportation management, and urban planning. Existing diffusion-based methods achieve high f...Human mobility generation aims to synthesize realistic point-of-interest (POI) visitation trajectories and has become an important tool for travel behavior modeling, transportation management, and urban planning. Existing diffusion-based methods achieve high fidelity but require per-city generation, given the inherent heterogeneity of geospatial locations and POI categories, while large language model-based methods generalize across cities but remain too costly at scale, especially for long-horizon trajectory generation. To address this, we propose SeMoFlow, a Semantic human Mobility generation framework based on latent Flow matching. We first encode heterogeneous POIs from different cities into a shared cross-city representation space via hierarchical Semantic IDs, where shared prefixes capture transferable semantics, and successive codes progressively refine the representation toward individual POIs. Building on the semantic IDs, SeMoFlow follows a plan-to-synthesis hierarchical generation paradigm, in which an autoregressive planner generates coarse-grained semantic and recurrence patterns, and a flow matching realizer synthesizes fine-grained suffix latents. The generated latents are subsequently decoded and grounded to concrete POIs. Extensive experiments on large-scale multi-city datasets show that SeMoFlow achieves higher trajectory fidelity than existing baselines, preserves city-specific mobility motifs, and supports both joint multi-city generation and effective cross-city transfer.
|
| 1703 |
Self-Evolving Multi-Agent Symbolic Discovery for Financial Fundamental Analysis
2609.32746
|
cs.LG
|
Kelvin J. L. Koa, Filip Orestav, Shengqiong Wu, Michael J. Wooldridge, Ke-Wei Huang |
While symbolic regression (SR) has been successfully used in science to discover new equations, its use in financial valuation is hindered by several limitations. Whereas the natural sciences provide objectively correct relationships, financial valuation const...While symbolic regression (SR) has been successfully used in science to discover new equations, its use in financial valuation is hindered by several limitations. Whereas the natural sciences provide objectively correct relationships, financial valuation constitutes a distinct class of symbolic discovery problems, as it admits multiple valid perspectives, operates under non-stationary market conditions, and involves noisy, continuous performance signals. In this work, we propose Multi-Agent Fundamental Analysis with Symbolic Adaptive learning (MUFASA), a hierarchical multi-agent framework for symbolic discovery in finance. MUFASA introduces (1) disentangled equation discovery via specialized agents representing distinct valuation perspectives, (2) a meta-coordinator that performs hierarchical-level reasoning over market context information, and (3) a memory mechanism that reasons over statistical performance summaries (e.g., accuracy, stability, and tail risk) to guide learning under noisy feedback. Experiments across datasets from multiple countries show that MUFASA achieves state-of-the-art performance on the valuation task compared to classical finance methods, financial large language models, and SR approaches, while simultaneously producing interpretable equations, which we share with the community. We also make publicly available the distilled learnings across evolution iterations and context-dependent strategy weights, which might offer useful insights for future research on financial fundamental analysis.
|
| 1704 |
Action Shaping: Policies Absorb What They Can Express
2609.32752
|
cs.LG
|
Yanjun Chen, Jinghan Wang, Xiaoyu Shen, Wenjie Li, Wei Zhang |
Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the co...Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its minimal instance is a zero-initialized linear head behind a learnable gate, added to an actor that trains through a learned action-value function, with no penalty or schedule. The gate rises and then falls on its own, for deterministic and stochastic actors alike, and on 20 tasks removing the head costs almost nothing. The condition is exact reproduction, not capacity: a nonlinear head with more parameters is not absorbed, and in a paired control, one linear path added to a nonlinear base head restores absorption. Exact reproduction gives the loss a flat direction that gradient noise drifts along, and the offset's amplitude indicates, before removal, what dropping the head will cost. Action shaping thus gains the counterpart of the shaping theorem, a condition for absorption, together with the mechanism behind it and a diagnostic that reads it. Policies absorb what they can express, and only that.
|
| 1705 |
SAGE: Semantic Audio Generative Encoder
2609.32755
|
cs.LG
|
Francesco Brigante, Luca Cerovaz, Davide Marincione, Giorgio Strano, Luca Zhou |
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space...Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.
|
| 1706 |
Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models
2609.32757
|
cs.LG
|
Drandreb Earl Juanico |
VLM bounding-box localization is both language generation and spatial commitment. Parseable fields such as bbox_2d make localization easy to score, but dimensions that emit coordinate tokens need not repair localization after visual evidence is damaged. We stu...VLM bounding-box localization is both language generation and spatial commitment. Parseable fields such as bbox_2d make localization easy to score, but dimensions that emit coordinate tokens need not repair localization after visual evidence is damaged. We study this readout/recovery separation in Qwen3-VL-4B-Instruct on single-object COCO grounding. We compare clean coordinate-token readout rankings with corruption-derived repair rankings, using object-mask endpoint replacement for recovery and clean-input flooring for depth localization. In Qwen3-VL, coordinate-token rankings are inert through layer 24, load-bearing from layers 32-35, and peak at layer 34; corruption-derived rankings harm layers 16-24 but become beneficial near layer 35/final. A Kimi-VL-A3B diagnostic shows a matching output-proximal transition despite a different box format. Object-mask recovery separates rank budgets: $k=250$ shows necessity, $k=500$ shows Top-$k$ restoration above random, and $k=d/2$ is largely capacity-driven. Partial-occlusion sweeps reveal that high-overlap coordinate-token sets can hurt at $k=1000$ and help mainly at half-width, while population corruption-derived sets provide no reliable fixed repair set. Edge-attribution patching shows coordinate-token paths are high precision but low recall for detection recovery, and RMSNorm quasi-layer controls do not close the endpoint-repair gap. Endpoint coordinate triage is therefore a useful circuit prior, but occlusion recovery requires a separate benchmark.
|
| 1707 |
An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning
2609.32762
|
cs.LG
|
Mino Nakura, Sriram Krishna, Yufei Wang, Shubham Tulsiani, Zackory Erickson |
Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. W...Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choices allow visuomotor policies to generalize across viewpoints. We find that viewpoint generalization improves when dense visual tokens are retained and the action head participates in geometric reasoning. On a suite of simulated tasks that span a wide range of camera poses, we show that these design choices yield a policy that remains performant across viewpoints. As a practical consequence, a policy trained with these design choices also transfers zero-shot from simulation to the real world under random camera configurations.
|
| 1708 |
Forecasting Intraday USD/CAD Exchange Rate with News-Derived Monetary-Policy Signals
2609.32773
|
cs.LG
|
Maya Kodeih, Aliaa Alnaggar, Mucahit Cevik |
Monetary-policy announcements and central-bank communications play a central role in foreign exchange markets, yet their qualitative, unstructured form makes their forecasting value difficult to quantify. While prior research has largely focused on sentiment e...Monetary-policy announcements and central-bank communications play a central role in foreign exchange markets, yet their qualitative, unstructured form makes their forecasting value difficult to quantify. While prior research has largely focused on sentiment extracted from financial news, comparatively little is known about the relative contribution of different dimensions of monetary-policy communication. Existing studies primarily evaluate whether textual information improves overall forecasting performance but provide limited insight into which communication channels drive such improvements. To address this gap, this paper introduces a statistical attribution methodology that decomposes monetary-policy communication into interpretable channels and quantifies their incremental forecasting contribution under false-discovery-rate control. Monetary-policy news is transformed into structured communication signals using large language models (LLMs) and temporal feature engineering. These signals are evaluated using rolling-window experiments with tree-based machine-learning models. The results show that monetary-policy communication contains measurable predictive information. Attribution analysis shows that predictive value is concentrated in a small subset of signals, with communication timing providing the strongest individual feature-level contribution, targeted communication-activity measures also contributing positively, and LLM-derived sentiment providing complementary information at the group level. The findings indicate that communication-based forecasting value extends beyond sentiment alone and that attribution, rather than aggregate accuracy alone, is central to evaluating news-derived signals.
|
| 1709 |
Copper-Policy: Focus on the Representation for Robust Robot Manipulation
2609.32779
|
cs.LG
|
Zexin Feng, Yixu Feng, Lingyu Xiao, Shang Su, Kexin Zheng |
World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a ques...World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve control? We introduce Copper-Policy, which learns a compact World representation with the policy rather than relying on a predefined target space. Through temporal joint-embedding prediction, it predicts future observation embeddings conditioned on task intention without reconstructing pixels. This prediction and action decoding shape the representation jointly, while the policy retains access to current-frame spatial detail for execution. Representation analyses show that the learned features better separate task-driven change from perturbations and provide complementary information for control. Compact prediction targets reduce training tokens per sample, enabling a 2B-parameter model trained in 9.67 hours on 8$\times$ RTX 5090 GPUs and 6$\times$ faster than Fast-WAM on matched A100 GPUs. Copper-Policy outperforms every compared method without embodied pretraining on RoboTwin and several embodied-pretrained VLAs on LIBERO-Plus (80.85%). On three challenging real-robot tasks, it performs comparably to $\pi_{0.5}$ and attains a higher average score. Together, these results show that Copper-Policy combines strong control performance with efficient training.
|
| 1710 |
OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing
2609.32780
|
cs.LG
|
Long Qian, Bingke Zhu, Jiaqi Wei, Yingying Chen, Jinqiao Wang |
Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved...Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved: which intermediate visual representations should be exposed to language computation for a given question? We introduce OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector. For each image-prompt pair, the projector selects a sparse set of intermediate visual depths and reuses the resulting global preference to guide both local patch fusion and dynamic visual injection into the language model. This design enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks. Across eight image-based benchmarks, OmniMoE-VL achieves an average score of 85.9 with 28B total and 9B activated parameters. Controlled comparisons show that the routed visual interface provides the dominant architectural gain, while matched route and component controls, same-image route analysis, and route interventions further support the value of coupling and question-conditioned visual access.
|
| 1711 |
CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models
2609.32781
|
cs.LG
|
Long Qian, Bingke Zhu, Jiaqi Wei, Yu Li, Yingying Chen |
Diffusion vision-language models generate answers by gradually resolving masked tokens, making accurate conditional prediction in partially resolved states central to post-training. Masking completed answers yields coherent contexts and targets, but prescribed...Diffusion vision-language models generate answers by gradually resolving masked tokens, making accurate conditional prediction in partially resolved states central to post-training. Masking completed answers yields coherent contexts and targets, but prescribed masks do not reflect the model's reveal decisions. Its trajectories capture these decisions, yet their provisional visible tokens can conflict with the target response. Outcome-based reinforcement learning follows these trajectories but provides only response-level feedback, which loses contrast when sampled rewards tie. To align coherent token-level supervision with the model's reveal decisions, we introduce Counterfactual Trace On-Policy Distillation (CT-OPD), which combines completed teacher responses with trajectory masks from the current student. CT-OPD retokenizes each teacher response in the student's vocabulary and extracts unresolved-position masks at successive stages of the student's reverse process. For each mask, it discards provisional rollout values and reconstructs the partial state from the teacher endpoint, so the supervised positions follow the current trajectory while the visible context and targets remain consistent with the same response. The student is trained on these reconstructed states with its native categorical loss, and trajectories are refreshed as the model evolves. Across dense and sparse diffusion architectures, CT-OPD consistently enhances multimodal understanding and reasoning capabilities, with gains of up to 9.80 points on the nine-benchmark average. On the unified understanding-and-generation architecture, it also improves both visual understanding and image generation, showing that the same principle transfers across architectures and modalities. Ablations further attribute these gains to coherent reconstruction and current-model trajectory masks.
|
| 1712 |
CollisionGAT: Controller-Agnostic One-Step Collision Screening for Multi-Agent Motion
2609.32783
|
cs.LG
|
Alan Debbas, Edwin Meriaux, Gregory Dudek |
Before a team of robots moves, each proposed step must be checked for collisions with other robots and with obstacles. We present CollisionGAT, a graph-attention network that reads the current and proposed states of moving agents together with locally relevant...Before a team of robots moves, each proposed step must be checked for collisions with other robots and with obstacles. We present CollisionGAT, a graph-attention network that reads the current and proposed states of moving agents together with locally relevant stationary obstacles and returns one collision-risk score per moving agent. Any controller can use these scores to accept, repair, replan, or postpone a proposed step. We mount CollisionGAT on a continuous path-following controller and on GATeD, an obstacle-blind D* Lite planner that uses typed vetoes to update its planning graphs. Exact geometric checks supply the training labels and independently audit every executed step.
|
| 1713 |
Rank Collapse Is Recoverable, Growing $|Q|$ Is Not: Out-of-Sample Early Warning for Value Divergence in High-UTD Soft Actor-Critic
2609.32819
|
cs.LG
|
Tianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao |
Raising the update-to-data (UTD) ratio breaks off-policy critics in two ways grouped as "plasticity loss": collapsing representations and growing value magnitude $|Q|$. We separate them in Soft Actor-Critic (SAC) with scaled critics (width 2048, no normalizati...Raising the update-to-data (UTD) ratio breaks off-policy critics in two ways grouped as "plasticity loss": collapsing representations and growing value magnitude $|Q|$. We separate them in Soft Actor-Critic (SAC) with scaled critics (width 2048, no normalization). Collapse is survivable: at UTD ratio 16, HalfCheetah critics with most units dormant keep learning, and the training guard, which stops runs whose loss or $|Q|$ explodes, never flags them. Within one high-UTD SAC configuration, runs start close together, and how far a critic's $\log_{10}|Q|$ has climbed by step 15k, its early growth, ranks the runs by how soon the guard flags them. At 15k, a flagged run's $|Q|$ sits a median of over a hundredfold below its flag level, yet the climb's rate already orders the flags (Harrell's C and out-of-sample AUC 0.78 on Walker2d, 0.98 on Ant, at UTD ratio 4). Dormancy does not. Aborting on this rate saves about a tenth of held-out Walker2d compute and stays net-positive live. A LayerNorm critic lowers the rate, removes the flag on Walker2d at UTD ratio 4 and lowers the return.
|
| 1714 |
Learning Shuffle Ideals with Membership Queries and Contrastive Examples
2609.32820
|
cs.LG
|
S. Mahmoud Mousawi, Pierluigi San Pietro, Sandra Zilles |
This paper studies learning of shuffle ideals with membership queries as well as with contrastive queries---a form of membership query that reveals not only whether a selected word $w$ is in the target language or not, but also provides a most similar word $w'...This paper studies learning of shuffle ideals with membership queries as well as with contrastive queries---a form of membership query that reveals not only whether a selected word $w$ is in the target language or not, but also provides a most similar word $w'$ that belongs to the target language if %and only if $w$ does not. For both settings, it is shown that even some very simple classes of shuffle ideals cannot be learned efficiently. By contrast, we obtain positive learnability results for classes of shuffle ideals that meet certain structural conditions. In the case of membership queries, these structural conditions are related to the previously studied notion of universal words, and raise new questions in word combinatorics. In the case of contrastive queries, the structural conditions relate to partitioning sets of words that are listed in shortlex order.
|
| 1715 |
Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs
2609.32821
|
cs.LG
|
Yuanyi Wang, Yanggan Gu, Su Lu, Guanghao Zhu, Pengkai Wang |
Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fu...Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fundamental question that remains unclear: \emph{does routing drift after MoE merging actually indicate routing failure, and what evidence should justify repair?} We investigate these questions across DeepSeekMoE, OLMoE, and Qwen3-MoE proposing a routing analysis toolkit for controlled counterfactual interventions and token-level analysis. By crossing source and merged router inputs and parameters, we attribute most expert reassignments to input shifts rather than parameter changes at the same layer. However, source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and different expert selections can produce directionally similar mixture outputs. We therefore operationalize routing failure as \textit{task loss recoverable under a specified routing intervention, with non-routing parameters fixed.} These tests detect recoverable loss under deliberate router corruption, whereas source-route restoration does not establish reliable task benefits in the evaluated merged models. Motivated by these, we propose \emph{Selective Router Repair (SRR)} as a case study, and find that source-specialist token-likelihood advantages do not reliably identify beneficial local corrections. Together, these findings show that \textbf{routing drift alone is insufficient evidence of routing failure}: source-informed corrections must be judged by their task-level intervention effects. The analysis toolkit and SRR code are released.
|
| 1716 |
Radiomap Blind Prediction under Incomplete Observation: Error Characterization and Correctable Propagation-Prior Learning
2609.32836
|
cs.LG
|
Xiaojie Li, Yu Han, Han Fang, Shangqing Liu, Guangxu Zhu |
Radiomap blind prediction aims to infer radiomaps from observable representations of the propagation environment and base station configuration without field measurements. In practice, the observable representations are inherently incomplete. Thus, the target ...Radiomap blind prediction aims to infer radiomaps from observable representations of the propagation environment and base station configuration without field measurements. In practice, the observable representations are inherently incomplete. Thus, the target radiomap is not fully determined by the inputs when generalizing to unseen configurations or environments. Under incomplete observation, we establish a population-level theory of deterministic radiomap blind prediction that identifies the conditional mean as its optimal target and separates prediction error into reducible predictor approximation and irreducible uncertainty caused by missing physical information. The framework further characterizes the train-test risk gap and the uncertainty reduction enabled by observation enrichment. Building on it, we reveal the dual role of propagation priors: they provide physically grounded guidance, yet their implementable forms may bias the attainable predictor. This motivates RadioDecomp, which treats a prior-guided predictor as a correctable base and learns its remaining predictable discrepancy through residual refinement. To evaluate RadioDecomp across distinct propagation-prior designs, we instantiate it with a feature-guided monolithic base and a LoS-Shadow structured base, yielding RadioFR and RadioLSR, respectively. Across random, cross-configuration, and cross-environment settings, experiments confirm the benefit of propagation-related representations and show that both instantiations improve upon their respective bases. Further controlled studies on base capacity, training-support coverage, and observation coarsening corroborate the proposed analysis.
|
| 1717 |
Mend the Measurement Gap: Latent User Preference Modeling for Short-Form Video Recommendation
2609.32839
|
cs.LG
|
Shuo Chang, Yueqi Wang, Zihuan Diao, Ali Montazer, Jiangguo Zhang |
Recommender systems rely heavily on heterogeneous behavioral feedback to infer user preference. Although abundant, these signals are imperfect measurements: the same observed behavior can arise from different underlying states, such as genuine enjoyment, passi...Recommender systems rely heavily on heterogeneous behavioral feedback to infer user preference. Although abundant, these signals are imperfect measurements: the same observed behavior can arise from different underlying states, such as genuine enjoyment, passive consumption, or inattention. The challenge is especially acute in short-form video, where watch-based signals are strongly affected by measurement confounders such as video duration - the same watch time can imply different levels of preference for videos of different lengths, while ratio-based metrics can systematically favor short videos. As a result, optimizing raw engagement can amplify measurement artifacts rather than improving user value. We propose a Factorized Latent Value Model (FLVM) for measuring user preference from heterogeneous behavioral feedback. The model treats observed behaviors as noisy measurements of a low-dimensional, factorized latent value state and uses structured output heads to model heterogeneous feedback signals. A restricted baseline path captures predictable variation from measurement-confounding features such as video duration, user propensity, and session context, while a routed latent path estimates preference-relevant value advantage. The resulting latent value score can be integrated into an existing recommender system as a ranking feature or ranking score. On YouTube Shorts, a major short-form video platform, this model improves offline metrics and lifts a primary viewer enjoyment metric by 2.67% in online A/B tests.
|
| 1718 |
VCRE-Fib: View-Conditioned Regional Evidence for Fine-Grained Ultrasound Grading of Schistosoma japonicum-Associated Liver Fibrosis
2609.32840
|
cs.LG
|
Ziyang Xu, Shuli An, Hao Zhou, Haitian Zhong, Tingting Wu |
Accurate assessment of Schistosoma japonicum-associated liver fibrosis is essential for disease management and long-term follow-up in endemic regions. Ultrasound provides non-invasive imaging, but complex local echogenic patterns and anatomical structures make...Accurate assessment of Schistosoma japonicum-associated liver fibrosis is essential for disease management and long-term follow-up in endemic regions. Ultrasound provides non-invasive imaging, but complex local echogenic patterns and anatomical structures make fine-grained grading challenging. Existing deep learning methods can predict fibrosis scores, yet directly incorporating acquisition views and regional cues into grading while retaining spatial information for inspection remains an open problem. Here we present VCRE-Fib, a view-conditioned regional evidence framework that integrates anatomical context, local information, and global image assessment for fine-grained ultrasound grading. The framework forms view-conditioned local grading evidence before spatial pooling, uses weak localization to guide its aggregation, and combines it with global predictions. Image-only inference jointly returns a fibrosis score, acquisition view, and candidate abnormal-region map. We developed and evaluated the method on a re-curated cohort of 108,709 ultrasound images from 6,373 patients across 35 centers. On a patient-disjoint test set of 4,107 images from 240 patients across four centers, VCRE-Fib reduced the prespecified composite grading risk by 7.115% relative to SFibAI trained and evaluated on the same data split. Image-level mean absolute error decreased from 0.391 to 0.378, alongside lower patient-max, patient-median, and center-balanced risks. The full model also achieved lower composite grading risk than variants that separately removed view conditioning or weak localization. These results support incorporating anatomical context and regional evidence into ultrasound grading while exposing spatial predictions for inspection alongside severity estimates.
|
| 1719 |
SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals
2609.32846
|
cs.LG
|
Yavuz Yarici, Ghassan AlRegib |
Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that...Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes into three components: redundancy shared between modalities, uniqueness specific to each modality, and synergy available only from their joint observation. Recent frameworks extend contrastive learning to capture all three components, yet synergy remains undertrained in practice. We propose SynCo (Synergy Contrastive Learning), a method that directly addresses synergy undertraining through dedicated supervision on an interaction residual. SynCo fits a linear projector to predict the fused representation from independently computed unimodal features, and the resulting interaction residual, which removes the linearly unimodal-predictable component, receives dedicated contrastive supervision at negligible computational cost. On the controlled Trifeature benchmark, SynCo achieves state-of-the-art synergy capture with a $+5.98\%$ gain over the baseline, and on real-world benchmarks from MultiBench, DARai, and MM-IMDb, SynCo consistently outperforms or matches prior methods across diverse modality combinations and task types. The method operates as a plug-in to existing contrastive multimodal frameworks without modifying the underlying fusion architecture and can further improve synergy capture when combined with other methods.
|
| 1720 |
Sliced Orlicz-Wasserstein
2609.32847
|
cs.LG
|
Binh Thuan Tran, Khai Nguyen |
We propose sliced Orlicz-Wasserstein (SOW) distance which is a generalization of sliced Wasserstein (SW) distance. SOW replaces the $L^p$ norm in SW with a Luxemburg norm cost induced by an Orlicz function $\phi$. First, we prove that SOW distance is a metric ...We propose sliced Orlicz-Wasserstein (SOW) distance which is a generalization of sliced Wasserstein (SW) distance. SOW replaces the $L^p$ norm in SW with a Luxemburg norm cost induced by an Orlicz function $\phi$. First, we prove that SOW distance is a metric on the space of measures with finite Orlicz norm, and show that it recovers the SW distance when the Orlicz function is $\phi(x)=x^p$. Next, we derive the topological properties of the SOW distance. In particular, we show that convergence under SOW implies weak convergence, and the converse is true under the compact support condition. We then present the theoretical results for estimating the SOW distance. We derive sample complexity for both the distance itself and the powered functional of the distance, and prove their minimax optimality. In addition, we discuss the computational algorithm for approximating the SOW distance by Monte-Carlo estimation and bisection search, as well as the associated approximation error and computational complexity analysis. Our experimental results reveal the superior computational efficiency of SOW compared with Orlicz-Wasserstein (OW) distance. Also, in the experiments, we demonstrate the favorable flexibility of SOW distance over SW in detecting differences between distributions by comparing their performance in two-sample tests and evaluating generative models on image datasets.
|
| 1721 |
Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?
2609.32857
|
cs.LG
|
Duc T. Nguyen, Thanh Ha Do, Phuong M. Cao, Hieu Pham |
Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the sa...Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in several studies, but what they already solve and where additional complexity is needed remain unclear. We study this behavior through the structure of prediction error under the mean-squared error (MSE) objective. Differences in average expression across genes can account for a substantial part of aggregate prediction performance, while a key unresolved error lies in recovering variation within each slide. Decomposing MSE into slide-level and within-slide components, we find that the within-slide component has lower residual-normalized parameter sensitivity in controlled neural experiments. This motivates Component-Guided Loss (CGL), which increases supervision of the within-slide component. CGL-Linear is a closed-form affine instantiation that achieves overall state-of-the-art performance across HEST-1k cohorts and gene-panel sizes. The same within-slide supervision improves existing neural models. These results suggest that substantial gains can come from aligning the training objective with prediction-error structure rather than increasing model complexity.
|
| 1722 |
Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning
2609.32870
|
cs.LG
|
Xing Han, Yuxin Wang, Chen Chen, Wei Dai, Gautham Krishna Gudur |
Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generati...Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action--outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.
|
| 1723 |
Machine learning for the LHC physics program: a 2025-2026 stocktake
2609.32874
|
cs.LG
|
Jesse Thaler |
The first sentence of this abstract--and the introduction to these proceedings--was authored by a human, but the bulk of this document was generated by an agentic AI system. In this talk, I take stock of machine learning (ML) for the LHC physics program over t...The first sentence of this abstract--and the introduction to these proceedings--was authored by a human, but the bulk of this document was generated by an agentic AI system. In this talk, I take stock of machine learning (ML) for the LHC physics program over the twelve months from May 2025 to May 2026. The corpus is the HEPML Living Review, split at May 2025 into 1,756 earlier papers and 569 later ones. An AI pipeline surveyed the 569 abstracts, ranked them by citations, recency, theme, and collaboration involvement, and read 103 papers in full (95 from after the split, plus 8 earlier baseline papers), producing a structured note for each. The notes were then synthesized into six claims about the state of the field, checked by independent reviewer agents, and re-verified against the source papers. The headline claim is that (1) ML for high-energy physics (HEPML) stopped being a research area that builds tools and became infrastructure that the LHC physics program depends on: ATLAS and CMS now publish physics results that depend on neural networks, and the archived ALEPH data have re-entered production. The other five claims are: (2) simulation-based inference and foundation models are two revolutions starting to merge; (3) AI agents are the genuinely new front, with 47 papers in twelve months and no adopted measurement yet; (4) "do we trust it?" is the fastest-growing agenda, with one recent paper in five about uncertainty, calibration, or interpretability; (5) what is slowing down is informative, since equivariance was absorbed into a tool and model-specific phenomenology ceded ground to model-agnostic searches; and (6) theory ML crossed a capability threshold in multiple research areas. I close with what is settled, what is incoming, and what is open, and briefly discuss the concerns raised by this way of working with AI.
|
| 1724 |
Neural Network-Assisted Refinement of Traditional Schemes for One-Dimensional Scalar Conservation Laws
2609.32887
|
cs.LG
|
Imre Fekete, Ferenc Izs\'ak, Vendel P. Kup\'as |
A clear link is established between conventional numerical methods and neural network approximations for solving one-dimensional scalar conservation laws. The focus is on the construction of an appropriate flux term in the case of convex flux functions for imp...A clear link is established between conventional numerical methods and neural network approximations for solving one-dimensional scalar conservation laws. The focus is on the construction of an appropriate flux term in the case of convex flux functions for improving the classical schemes. The first neural network developed here is able to rediscover Godunov's method, while the second one emulates the behavior of a second-order slope-limiter function. In this way, by merging them, second-order reconstruction-based schemes can be developed. The networks presented here employ a minimal number of parameters, significantly reducing the complexity compared to previous approaches. These networks can also be linked consecutively to get a deep one corresponding to multiple time steps. Training them with an appropriate loss leads to stable schemes, improving even the classical methods without increasing their complexity.
|
| 1725 |
EEG-Based Motor Imagery BCI Algorithms and Technologies: A Review
2609.32930
|
cs.LG
|
Mohammad Hossein Koohi Ghamsari, Seyede Fatemeh Ghamkhari, Siavash Bayat, Ahmed Hemani |
Brain-computer interfaces (BCIs) have emerged as transformative technologies that enable direct communication between the brain and external devices. Among various BCI paradigms, EEG-based motor imagery (MI) has gained prominence due to its simplicity, non-inv...Brain-computer interfaces (BCIs) have emerged as transformative technologies that enable direct communication between the brain and external devices. Among various BCI paradigms, EEG-based motor imagery (MI) has gained prominence due to its simplicity, non-invasiveness, and potential to restore motor function and facilitate rehabilitation for patients with motor impairments. This paper presents a comprehensive review of the most practical processing algorithms developed over the past decade for decoding brain sensorimotor cortex signals. Specifically, this paper discusses the integration of artificial intelligence (AI)-based algorithms, particularly machine learning and deep learning techniques, and their contributions to improving the performance and efficiency of MI-BCI systems in detail. Furthermore, the paper reviews state-of-the-art hardware platforms and emerging converging technologies, including system-on-chip (SoC) architectures, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), wearable devices, the Internet of Things (IoT), and augmented/virtual reality (AR/VR), and discusses their integration with advanced signal processing algorithms to enable next-generation MI-BCI systems. By highlighting current achievements of EEG-based MI-BCI technology and predicting future research directions that could further enhance real-time capabilities, this paper aims to provide valuable insights for researchers and practitioners, fostering innovation in high-performance EEG-based MI-BCI systems.
|
| 1726 |
Distributed Hydrological Modeling in the Feature Space
2609.32971
|
cs.LG
|
Mohamad Hakam Shams Eddin, Maria Luisa Taccari, Yikui Zhang, Shijie Jiang, Juergen Gall |
Accurate forecasting of river discharge and floods is very challenging. River dynamics are affected by storage, meteorological forcing, and flow propagation at different spatial and temporal scales. Forecasting thus requires a framework that considers the upst...Accurate forecasting of river discharge and floods is very challenging. River dynamics are affected by storage, meteorological forcing, and flow propagation at different spatial and temporal scales. Forecasting thus requires a framework that considers the upstream-to-downstream flow through river networks across grid cells and catchments. This modeling is known in hydrology as distributed modeling and routing. Existing deep learning approaches either ignore this topology, operate on lumped catchments, or route predicted physical quantities through a separate graph or physical routing model. We instead introduce feature-space routing: a topology-aware state-space operator embedded directly in the forecasting dynamics. At every forecast step, the operator gathers latent states from upstream grid cells and causally updates the downstream state according to the known river network. This preserves the physical connectivity of the river system while allowing the propagated state itself to be learned end-to-end and allows the model to predict river discharge considering both local dynamics and neighboring upstream contributions. To address uncertainty and provide probabilistic forecasts, we minimize the fair continuous ranked probability score (fCRPS) as a training objective. Our experiments on the European Flood Awareness System (EFAS) and observational data for river discharge forecasting demonstrate that encoding the physical structure of river networks explicitly in the feature space substantially improves the forecasting skill, particularly in an ungauged setting. Our approach achieves state-of-the-art results on both reanalysis and observational data and is able to forecast maps of river discharge at 1 arcminute and 6-hourly resolution up to 10 days lead time.
|
| 1727 |
CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning
2609.33007
|
cs.LG
|
Shivam Aarya, Zhang Xi-Jia, Chengyue Huang, Junhyun Kim, Huishu Xue |
Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: d...Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general-purpose multimodal foundation models into deployable robot policies by using the foundation model itself as an autonomous demonstrator. While sufficiently capable models can generate successful zero-shot manipulation trajectories, repeatedly invoking them during physical execution is slow and expensive, limiting their utility as scalable data generators. As a solution, we introduce CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan. We evaluate across RoboCasa tasks and on physical Franka and bimanual YAM-arm platforms, measuring task success, model calls, token usage, collection time, and cost. We further train Diffusion Policy and ACT on matched sets of human-teleoperated and foundation-model-generated demonstrations to evaluate the downstream learning value of autonomously collected data. We find that CAPEX increases the number of successful demonstrations by 4.3x while reducing the cost per successful demonstration by 80%. Policies trained on CAPEX-generated data approach the performance of those trained on matched human demonstrations; with longer training, this gap largely closes for policies trained from scratch. These results suggest that foundation models can serve as scalable sources of reusable robot experience. Project page: https://capex-paper.github.io/
|
| 1728 |
Decentralized Optimization with Cross-Coupled Mixed Affine Constraints
2609.33021
|
cs.LG
|
Ewsey Obzherin, Ilya Khomchenko, Natalia Shelegeda, Nhat Trung Nguyen, Demyan Yarmoshik |
We study decentralized optimization with cross-coupled mixed affine constraints, where local and shared variables interact through two affine channels. We show that the intrinsic difficulty of combining separately well-conditioned channels is governed by their...We study decentralized optimization with cross-coupled mixed affine constraints, where local and shared variables interact through two affine channels. We show that the intrinsic difficulty of combining separately well-conditioned channels is governed by their Friedrichs angle. This geometry induces a cross-coupling factor that cannot be removed by channelwise preconditioning and governs the additional affine-oracle and communication complexity. We develop an accelerated decentralized method with matching minimax guarantees in the smooth strongly convex regime and extend the framework to smooth and nonsmooth convex objectives. Experiments confirm the predicted dependence on cross-channel geometry and network conditioning.
|
| 1729 |
Byzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal Prediction
2609.33037
|
cs.LG
|
Prasanjit Dubey, Aritra Guha, Xiaoming Huo |
Retrieval-augmented generation (RAG) lets language models answer questions more accurately by consulting relevant documents. Many valuable collections, such as medical records, cannot be pooled because of privacy rules. Federated RAG leaves each collection wit...Retrieval-augmented generation (RAG) lets language models answer questions more accurately by consulting relevant documents. Many valuable collections, such as medical records, cannot be pooled because of privacy rules. Federated RAG leaves each collection with its owner, or node, which scores candidate answers from its own documents; a central hub combines the scores. Some nodes, called Byzantine, may be compromised, faulty, or misled by instructions hidden in documents, and report arbitrary scores. Conformal prediction returns a set containing the correct answer with a chosen probability, using a cutoff set in a calibration step on questions with known answers. An unknown group of nodes, no larger than a declared bound, may misreport both in this step and at query time. Existing methods assume every node is honest or protect only the calibration step. We observe that the honest nodes are the same in both steps. The hub therefore has all nodes score the same calibration questions, and keeps a candidate only if some plausible group of honest nodes, using its own scores in both steps, would keep it. We prove that the resulting sets contain the correct answer with the chosen probability in finite samples, whatever the Byzantine nodes report. No method using the same information can return smaller sets without risking the loss of an answer the honest nodes support. If nodes fail at random, the guarantee weakens only by the probability that more nodes fail than declared. In simulations, on real question-answering tasks including medical exams, and with language models as nodes, some hijacked, our sets reached the target whenever no more nodes misbehaved than declared, while plain averaging could miss it. They were also clearly smaller than those of simpler methods with the same protection, most of all when the declared bound was generous, so a cautious bound costs little.
|
| 1730 |
BudgetVerify: Budget-Tiered Verification for Financial QA
2609.33052
|
cs.LG
|
Janet Jenq, Hongda Shen |
Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute. We propose BudgetVerify, a budget-tiered generator-verifier fr...Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute. We propose BudgetVerify, a budget-tiered generator-verifier framework that routes each generated answer to one of three verification tiers: no verification, lightweight check-and-revise, or higher-cost solve-first-then-compare verification. The router is trained from offline correctness and token-cost outcomes and, at test time, selects a verification tier using information available before verification, including the question, context statistics, the generated answer, and associated generator metadata. The selected tier either returns the generated answer directly or invokes the corresponding verifier. Across six commercial and open-weight base models, BudgetVerify consistently produces more efficient accuracy-cost Pareto frontiers than fixed verification policies by selectively allocating stronger verification only when it is useful. Although absolute performance varies across models, these efficiency gains and the resulting qualitative frontier shape are consistent across generator models.
|
| 1731 |
Parameter-Efficient 3D Segmentation of Liver and Liver tumors: Depthwise factorization Scales Better Than Dense Convolution with Spatial Dimensionality
2609.33077
|
cs.LG
|
Adham M. Alkhadrawi, Mohammed A. B. Mahmoud |
Three-dimensional dense convolutional networks are the strongest performers on volumetric medical image segmentation, but their parameter counts scale poorly: moving a dense k x k convolution to k x k x k multiplies its weights by k. We observe that depthwise ...Three-dimensional dense convolutional networks are the strongest performers on volumetric medical image segmentation, but their parameter counts scale poorly: moving a dense k x k convolution to k x k x k multiplies its weights by k. We observe that depthwise separable factorization does not share this penalty. Because the cubic kernel term applies only to the depthwise stage while the pointwise projection, which dominates the parameter count, is unchanged, the same architecture grows by 5 % from 2D to 3D where a dense convolutional U-Net grows by 200 %. We exploit this asymmetry to build a 3D U-Net with 536,990 parameters, 24x fewer than an identical dense 3D U-Net. On MSD Task03 Liver (the Medical Segmentation Decathlon liver task, derived from LiTS), evaluated per case under five-fold cross-validation over all 131 public volumes, the model reaches a tumor Dice of 0.577 (95% CI [0.518, 0.633]) and a liver Dice of 0.947 (95% CI [0.941, 0.952]). Its liver Dice exceeds previously reported performance. Its tumor Dice exceeds their low-resolution configuration (0.4701) by 0.107 and their 2D configuration (0.5394), at approximately one twenty-fourth of the parameters and roughly half the in-plane resolution. Trained under identical conditions on a common held-out split, it exceeds a dense 3D U-Net on liver by +0.031 Dice (paired p = 0.006) and on tumor by +0.041 (95 % CI [+0.005, +0.081], paired p = 0.056), suggesting the factorization also acts as a regularizer in the small-data regime characteristic of medical imaging. We further show, on both LiTS and a 2D endoscopy benchmark, that a large fraction of the network's learnable spatial filters can be replaced by fixed shifts at no cost in accuracy, but that replacing all of them is measurably worse, the placement of spatial capacity matters more than its total amount.
|
| 1732 |
Evolving Dexterous Robots from Scratch
2609.33101
|
cs.LG
|
Zihan Guo, Shuzhe Zhang, Muhan Li, Peiyang Li, Sam Kriegman |
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and ma...Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
|
| 1733 |
Can Tabular Foundation Models Amortize Statistical Inference?
2609.33114
|
cs.LG
|
Kai Ye, Shijin Gong, Hongyi Zhou, Valentina Zangirolami, Chengchun Shi |
For decades, statistical inference has largely been developed one problem at a time. Given a scientific target, such as a treatment effect or a regression function, statisticians design a problem-specific estimator together with a procedure for quantifying its...For decades, statistical inference has largely been developed one problem at a time. Given a scientific target, such as a treatment effect or a regression function, statisticians design a problem-specific estimator together with a procedure for quantifying its uncertainty. This paper proposes a different paradigm. We focus on a classical problem in statistical inference, confidence interval construction, and develop TabCon, an amortized inference system built on a tabular foundation model that produces confidence intervals for new datasets through a simple forward pass. The key methodological ingredients of TabCon are a sparse mixture-of-experts architecture and reinforcement-learning-based post-training that calibrate the resulting confidence intervals to a desired coverage level. Across a wide range of benchmark datasets, TabCon attains near-nominal coverage while producing short confidence intervals. At inference time, it also offers considerably greater computational efficiency, running 50 times faster than the classical bootstrap procedure, even when the latter uses only 50 bootstrap samples.
|
| 1734 |
Sharp Critical Minimax Laws and No-Learning Thresholds in Continuous-Time Adaptive Control
2609.33154
|
cs.LG
|
Chen Jia |
We study episodic continuous-time control with an unknown vector control gain, scalar state, quadratic action cost, and smooth convex terminal cost. In the scalar Gaussian experiment, let $\Delta(H)$ denote the minimax improvement over zero control and set $\d...We study episodic continuous-time control with an unknown vector control gain, scalar state, quadratic action cost, and smooth convex terminal cost. In the scalar Gaussian experiment, let $\Delta(H)$ denote the minimax improvement over zero control and set $\delta=\sqrt2H^2-1$. We prove the critical law $$ \Delta(H)\asymp \delta^4\sqrt{\log(1/\delta)} \qquad (\delta\downarrow0), $$ with a matching lower and upper bound. The lower bound follows from a uniform deficit--energy inequality valid for fully adaptive controls with unbounded amplitudes, while a moving soft-threshold feedback attains the rate. For local parameters $\theta=N^{-1/4}h$, $|h|\le H$, we show that the normalized minimax regret over $N$ episodes is within $O(N^{-1/2})$ of a fixed-horizon Gaussian sequential control problem, without an additional dimension factor. The terminal task enters the limit only through $c_g=\operatorname{Var}(g(Z))$. The Gaussian problem exhibits an exact no-learning phase transition: $$ C_d^T(H)=TH^2/2 \iff H^4T\le d^2/2, $$ yielding the asymptotic task boundary $c_gH^4=d^2/2$.
|
| 1735 |
Notes on Generative Modeling for Feedback Control and Planning
2609.33164
|
cs.LG
|
Karthik Elamvazhuthi |
In these notes, we view control as a dynamically constrained sampling problem on the state-space of a control system. With this viewpoint, we extend methods from generative modeling, such as flow matching, normalizing flows and denoising diffusions to control ...In these notes, we view control as a dynamically constrained sampling problem on the state-space of a control system. With this viewpoint, we extend methods from generative modeling, such as flow matching, normalizing flows and denoising diffusions to control problems. Concepts such as controllability, optimal control and trajectory planning play an important role in guiding the extension and understanding well-posedeness of the corresponding algorithms, with application to steering systems to target states or distributions and sampling from reachable sets. The notes are intended as an accessible introduction for readers with a background in control theory and robotics.
|
| 1736 |
Dynamic Manipulation with World-Action Models via Counterfactual Planning
2609.33172
|
cs.LG
|
Sunwoo Park, Wonbin Lee, Seonghyun Jin, Youngmin Kim, Jangho Park |
World-Action models (WAMs) trained on static demonstrations often fail to manipulate moving targets even when they possess the required manipulation skills. We attribute this failure to target-response collapse: as execution advances, the policy becomes increa...World-Action models (WAMs) trained on static demonstrations often fail to manipulate moving targets even when they possess the required manipulation skills. We attribute this failure to target-response collapse: as execution advances, the policy becomes increasingly biased toward the learned continuation of its ongoing behavior and less responsive to target relocation. To bridge the gap between what the model has learned and what it can generate from the current context, we formulate dynamic manipulation as counterfactual planning by decoupling the context used for plan generation from the physical state used for execution. Our framework, Dynamic Predictive Planning (DPP), first uses the WAM's predictive rollout to estimate when an interaction is expected to occur, and combines this timing estimate with observed target motion to predict the target's future interaction position. DPP then constructs a counterfactual observation that places this predicted target position in a familiar robot context, allowing the model to invoke an existing manipulation skill rather than generate a recovery behavior from an unfamiliar robot-target configuration. The resulting plan is connected to the robot's actual state during execution. DPP enables real-time dynamic manipulation on a single consumer GPU without additional training on dynamic data. Experiments in simulation and on a real robot demonstrate consistent improvements across diverse target motions, with simulation performance surpassing all evaluated baselines, including methods additionally trained on dynamic data. Project page: https://methoder00.github.io/DPP/
|
| 1737 |
Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks
2609.33180
|
cs.LGcs.AI
|
Xiaojing Sun, Yuhan Zeng, Zihua She, Xiao Wang |
As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is pro...As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is proposed next. However, when these finite resources are repeatedly reused, new candidates are proposed based on feedback from the same evaluation set, so the search trajectory can adaptively overfit and empirical improvement may not reflect genuine population improvement on the underlying task distribution. Some existing methods account for multiple comparisons but assume that candidates are chosen independently of the evaluation set, and therefore do not control this adaptive dependence. To address this, we propose REUSE (Risk-controlled Evaluation Under Sequential Evolution), a certified evaluation and promotion framework that allows a fixed evaluation set to support repeated adaptive decisions while providing statistical guarantees. For a user-specified error level $\alpha$, with probability at least $1-\alpha$, every promoted modification is a genuine population improvement on the underlying task distribution. REUSE achieves this by strictly limiting the evaluation feedback returned to the search process and accounting for possible promotion histories within the error budget. We develop detailed statistical theory for RSI evaluation in this setting, including simultaneous error control, valid lower bounds on cumulative improvement, and a characterization of the fundamental limits of adaptive evaluation reuse. In live self-improvement experiments, REUSE commits substantially fewer false promotions than evaluation frameworks from current RSI systems and error-controlled baselines, reducing the proportion of false promotions from up to 20.7% to 0%, while achieving final true population performance comparable to the best baselines.
|
| 1738 |
Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations
2609.33186
|
cs.LG
|
Amitakshar Biswas, Yuhan Li, Ruoqing Zhu |
Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step an...Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing weight. In ideal settings, this mixed Bellman formulation can provide a natural bias--variance trade-off between approximation error under a restricted value-function class and the increased variance arising from multi-step importance weighting. To solve this mixed residual optimization, we adopt a minimax formulation involving a critic function. Unlike standard approaches that rely on a fixed functional class, we construct a data-dependent critic representation using predicted future feature directions which effectively induces a kernel adapted to the underlying transition dynamics. This allows the critic to focus on directions that are most relevant for the estimated Bellman error. To control overfitting, we use sample splitting to construct the critic and estimate the value function on separate data subsets. Simulation studies and MetaWorld tasks illustrate the effect of the mixing parameter and show that intermediate residual combinations can improve value estimation in challenging settings.
|
| 1739 |
MorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision Transformers
2609.33207
|
cs.LG
|
Rachmad Vidya Wicaksana Putra, Amirhesam Jafari Rad, Muhammad Shafique |
Spiking Vision Transformers (SViTs) are developed as an energy-efficient alternative to conventional ViTs for computer vision tasks at the edge. However, huge parameter counts and complex multi-head self-attention (MHSA) operations make it challenging to achie...Spiking Vision Transformers (SViTs) are developed as an energy-efficient alternative to conventional ViTs for computer vision tasks at the edge. However, huge parameter counts and complex multi-head self-attention (MHSA) operations make it challenging to achieve high energy efficiency in SViT inference, especially in tightly constrained applications. To maximize efficiency gains of SViT processing, we propose MorphAtt, a novel digital accelerator that expedites SViT inference through streamlined processing. Specifically, it processes MHSA operations using cascaded hardware modules: a Spiking Query-Key-Value generator (SpikeQKV), a low-complexity Spiking Multi-Head Self-Attention engine (SpikeAtten), and Reparameterization Convolution (RepConv) modules. To mitigate traffic congestion in on-chip memory accesses and data reuse, specialized inter-module buffers are integrated within the dataflow. Under synthesis using 32nm CMOS technology, MorphAtt achieves 792-1605 GOPS of throughput, while incurring ~39-55 mW of power consumption and 1.5 mm^2 of area, which lead to 20.3-29.1 TOPS/W of energy efficiency. These results also demonstrate that our MorphAtt offers better performance and efficiency trade-offs than state-of-the-art, thereby enabling highly energy-efficient vision-based AI systems at the edge.
|
| 1740 |
Schr\"odinger--F\"ollmer Actor--Critic: Diffusion Policy Improvement with Finite-Sample Analysis
2609.33239
|
cs.LG
|
Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Jincheng Ying |
Diffusion policies represent multimodal action distributions, but an advantage-weighted update does not specify how to sample from the resulting target distribution. We propose Schr\"odinger--F\"ollmer Actor--Critic (SFAC), an offline-to-online reinforcement l...Diffusion policies represent multimodal action distributions, but an advantage-weighted update does not specify how to sample from the resulting target distribution. We propose Schr\"odinger--F\"ollmer Actor--Critic (SFAC), an offline-to-online reinforcement learning (RL) method for Kullback--Leibler (KL)-regularized policy improvement through conditional diffusion. A minimax Bellman critic estimates the advantage function, which defines an exponentially tilted target policy. A Doob $h$-transform expresses this update as a correction to the reference diffusion drift. We derive a posterior-mean representation of the correction and estimate it using paired self-normalized importance sampling (SNIS). Supervised regression on these drift targets updates the neural actor without critic action gradients. In the small-update regime, the KL-regularized update follows the natural policy-gradient direction, and the Doob correction represents the same local change in the space of diffusion drifts. Under suitable conditions, we derive finite-sample bounds that separate the effects of critic estimation, neural drift regression, finite-sample SNIS, diffusion discretization, and inherited actor error on expected average policy suboptimality. Synthetic experiments assess the accuracy of approximation to prescribed advantage-tilted targets and sensitivity to sampling budgets. On six offline-to-online continuous-control tasks, a reference-anchored implementation achieves higher final-window returns than those of its corresponding offline initialization.
|
| 1741 |
CORTEX: A Verified Experience Layer for Generalist Agents
2609.33260
|
cs.LG
|
Garapati Keerthana, Manik Gupta |
An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is st...An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and when it should be discarded. We introduce CORTEX (Contextual Orchestration and Reuse of Task EXperience), a general AI systems framework that connects specialized agents through an external layer of verified experience. Each episode records its task conditions, source and tool state, decisive predicates, proof trace, verifier, and outcome. A meta-controller chooses exact replay, checked adaptation, fresh synthesis, or escalation. Accepted episodes can become task patterns and procedural strategies through a challenge-driven development loop. This gives the system an implicit competence layer that can grow without changing model weights. We formalize system contracts for exact replay and source-version separation, and derive when reuse saves computation. A controlled two-domain implementation tests the exact-replay core on 1,000 synthetic cases. Complete-family holdouts test procedural transfer on 1,000 new-family cases across eight clinical and policy splits, with complete fresh-evidence grounding and perfect invariance to irrelevant-field and insertion-order perturbations. The transfer trace exposes the work required for verified strategy execution. These results establish an initial path toward general intelligence through reusable procedures, typed experience, and developmental transfer.
|
| 1742 |
Modelling non-linear aeroelastic loads in long-span bridges with extreme learning machines
2609.33274
|
cs.LG
|
Gledson Rodrigo Tondo, Samir Chawdhury, Sergio Andres Castro Giraldo, Guido Morgenthal |
Accurate modelling of aerodynamic loads is essential for predicting instabilities and ensuring the safety of long-span bridges. A methodology is introduced for modelling aerodynamic self-excited forces in bridge-deck cross-sections using extreme learning machi...Accurate modelling of aerodynamic loads is essential for predicting instabilities and ensuring the safety of long-span bridges. A methodology is introduced for modelling aerodynamic self-excited forces in bridge-deck cross-sections using extreme learning machines (ELMs). ELMs, as single-layer feedforward neural networks, offer efficient training and accurate predictions. Forced-oscillation datasets from computational fluid dynamics (CFD) or wind-tunnel experiments are used for training, enabling systematic data selection to capture non-linear aerodynamic behaviour often missed by semi-analytical approaches. Once trained, the model predicts self-excited loads for any arbitrary motion composed by frequencies and amplitudes within the training domain. Comparisons with analytical, semi-analytical, and CFD results show superior accuracy in capturing non-linear force components and close agreement for aerodynamic loads and flutter wind speeds. Training required about 1.1\% of the time of a conventional neural network, and coupled flutter analysis runs in seconds, providing orders-of-magnitude speed-ups over CFD. These results indicate that ELM-based frameworks are accurate, practical, and efficient alternatives for modelling self-excited loads, particularly when preliminary CFD or wind-tunnel data are available. The presented approach offers a reliable data-driven technique for aeroelastic load modelling in long-span bridges.
|
| 1743 |
The Statistical Benefits of Multiple Responses for Learning from Demonstrations
2609.33291
|
cs.LG
|
Chandramauli Chakraborty, Cong Ma |
Many generative systems return multiple candidate responses and are evaluated according to the best one. Recent work shows that, when demonstrations are optimal, pass@$k$ can reduce the sample complexity of learning from demonstrations by a logarithmic factor ...Many generative systems return multiple candidate responses and are evaluated according to the best one. Recent work shows that, when demonstrations are optimal, pass@$k$ can reduce the sample complexity of learning from demonstrations by a logarithmic factor in $k$. We ask what happens when the demonstrator is not assumed to be optimal. We find that multiple responses provide a qualitatively stronger benefit in this setting. In a finite reward-class model with no reward feedback, moving from pass@$1$ to any pass@$k$ with $k\ge2$ changes the worst-case dependence on target accuracy from $1/\varepsilon^2$ to $1/\varepsilon$, uniformly over demonstrator quality. Under standard evaluation, where an unknown reward is fixed before training, increasing $k$ provides an additional and distinct benefit: the optimal dependence on a reward class of size $N$ improves from $\log N$ to $\log N/\log k$. We further show that these two effects can be separated. Under robust evaluation, where one learned policy must compete with the demonstrator simultaneously for every reward in the class, the fast $1/\varepsilon$ dependence persists, while the $1/\log k$ improvement can disappear. We establish matching upper and lower bounds in the corresponding regimes and give a greedy multiplicative-weights learner achieving the upper bounds without any assumption on demonstrator quality.
|
| 1744 |
When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions
2609.33312
|
cs.LG
|
Dipankar Sarkar |
Moving ML-mediated decision making onto privacy-preserving clients decentralises the economic decision along with the inference. Shared budget constraints then depend on information that cannot be globally current, creating an information-structure failure tha...Moving ML-mediated decision making onto privacy-preserving clients decentralises the economic decision along with the inference. Shared budget constraints then depend on information that cannot be globally current, creating an information-structure failure that conventional pacing is not designed to solve. We study this information misalignment in an auction-logic-faithful on-device simulation with 36 campaigns and 50 devices. Accounting is in dimensionless integer score units; no currency semantics are claimed. Across 30 paired demand paths, proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks under the original 20-times budget pressure. The effect does not depend on that severe a budget: at two-times pressure, 50-tick overspend remains 106.95%. A visible-budget no-sale guard makes zero-lag compliance exact at this score-unit granularity, yet leaves 11.88% overspend at one tick because other devices' debits remain invisible. A declared bursty, heterogeneous-device sweep retains a strictly increasing mean lag curve. We derive a finite-window expected excess-debit bound under conditional charge caps and find positive paired slack in every bounded-value cell. A second, incentive misalignment arises when the ML/pacing score transformation is allowed to change payment units: 98.23% of rival auctions at one tick admit a profitable deviation. An executable implementation-level counterexample isolates the runner-up's multiplier in the winner's price. Critical-base-bid payment is per-auction DSIC conditional on current multipliers, but does not establish dynamic truthfulness and does not repair base-value ranking disagreement.
|
| 1745 |
Agentic Multi-Turn Reasoning: A Fairness Approach
2609.33323
|
cs.LG
|
Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill, Jackson Cothren, Marios Savvides |
Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental ch...Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or $\Phi$-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.
|
| 1746 |
QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG
2609.33351
|
cs.LG
|
Hyojun Ahn, Emily Jimin Roh, Soohyun Park, Walid Saad, Hyung-Chul Lee |
Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum...Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent retrieval adaptation (QuPID) repairs this by making the circuit input-dependent through data re-uploading and by comparing measurement readouts, vectors of local Pauli expectations, rather than states. The result is a small readout for adapting frozen image features to a local archive with limited data: training simulates the circuit classically, and inference runs on a GPU with fixed learned parameters. We characterize the class as a structured factorization of input-modulated quadratic feature maps, bound the frequency support of its re-uploading channel, and give a parameter-count generalization bound that motivates its small budget. Under a shared frozen backbone and a label-free protocol, QuPID's 60 parameters give higher precision-at-5 (P@5) on ChestX-ray14 and MURA than frozen medical encoders, and than adapters and low-rank adaptation (LoRA) with up to 5.25 million trainable parameters. On ChestX-ray14, the P@5 gain over the frozen encoder is +0.116, the lead over retuned adapters is widest at 512 adaptation examples (+0.040), and the full-budget margin over an equally compact classical rotation-plane head is +0.023 with a 95% interval excluding zero. Medical imaging is the primary testbed; the pattern recurs on two non-medical benchmarks, in report generation, and under simulated gate noise and finite-shot readout.
|
| 1747 |
Geometry-Adaptive Mechanisms for Private Synthetic Data
2609.33363
|
cs.LG
|
Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian |
Generating differentially private synthetic data with meaningful Wasserstein utility guarantees is challenging in high dimensions. For datasets of size \(n\) on $[0,1]^d$ with $d\ge2$, existing pure \(\varepsilon\)-differentially private mechanisms achieve exp...Generating differentially private synthetic data with meaningful Wasserstein utility guarantees is challenging in high dimensions. For datasets of size \(n\) on $[0,1]^d$ with $d\ge2$, existing pure \(\varepsilon\)-differentially private mechanisms achieve expected $1$-Wasserstein error of order $(\varepsilon n)^{-1/d}$, reflecting the curse of dimensionality. While this rate is optimal in the worst case, it can be overly pessimistic when the data are supported on a lower-dimensional set. We formalize this through a multiscale packing-growth dimension $k$, which captures the geometric complexity of the support via the growth of packing numbers across scales. We propose \emph{Adaptive Pruned-PMM}, a pure $\varepsilon$-differentially private mechanism that combines private depth selection with our pruned variant of the Private Measure Mechanism (PMM) of He et al.\ (2023). The mechanism supports deeper, geometry-adapted hierarchies with expected running time $O\!\left(d(n+d)\log(\varepsilon n)\right)$, which is near-linear in $n$ for fixed dimension and privacy budget. Under an external multiscale packing-growth condition with dimension $k$, we show that, for fixed positive privacy budgets and fixed geometry, the expected $1$-Wasserstein error is of order $(\varepsilon n)^{-1/k}$ for $k>1$ as $n$ grows. We also prove a lower bound under a corresponding internal packing-growth condition, showing that the exponent $1/k$ is sharp within this framework.
|
| 1748 |
API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary
2609.33371
|
cs.LG
|
Patrick Kenney, Hadi Ahmadi, Denis Lusson, Donald Nguyen, Gurbinder Gill |
Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a dat...Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where it may persist in conversation history, logs, memory stores, generated code, and error payloads. Prompt injection and excessive agency then convert passive disclosure into unauthorized action. This paper formalizes the credential-exposure threat chain for agentic systems; synthesizes evidence from a platform secret-store incident, vendor-reported secret-sprawl measurement, and OWASP and NIST guidance; and describes a vault-mediated execution architecture in which the model selects a connector identifier while a trusted request boundary supplies authentication. We evaluate a production implementation, Corvic Security Vault, in two controlled black-box experiments. Across 16 probes spanning seven control domains, every probe met its expected outcome: an authenticated GitHub API request succeeded while the credential stayed absent from process environment values, caller-visible request headers, tested filesystem locations, three third-party echo services, and two unrelated API origins; both cloud instance-metadata endpoints were unreachable. We also report a negative result, a connector whose stored header mapping did not satisfy its provider's authentication contract, showing that centralized custody does not by itself guarantee correct configuration. Vault mediation removes several disclosure paths but is necessary rather than sufficient: least privilege, deterministic action authorization, human approval, telemetry redaction, and rotation remain independently required. The study is purposive and small, a functional security evaluation rather than a certification.
|
| 1749 |
Identifying the Predictable Drift of a Semimartingale from Marginal Laws
2609.33372
|
cs.LG
|
Jakub Marecek, Enrico Biffis, Abigail Langbridge, Robert Shorten |
A special semimartingale admits a unique decomposition $X=X_0+M+A$ into a local martingale $M$ and a predictable finite-variation part $A$. We consider the identification of $A$ when $X$ is observed only through repeated cross-sections. The estimand is then th...A special semimartingale admits a unique decomposition $X=X_0+M+A$ into a local martingale $M$ and a predictable finite-variation part $A$. We consider the identification of $A$ when $X$ is observed only through repeated cross-sections. The estimand is then the projection of the sampled predictable compensator onto the observable feature filtration, namely the current state together with whatever randomness is shared across the population, so that at a fixed diffusion coefficient the marginal flow identifies the drift only up to a Markovian projection. If the drift is an affine functional of an observed lag window, the joint problem is a convex quadratic programme whose solution is the pseudo-panel regression of econometrics. Our principal concern is the case, which we believe not to have been treated before, in which the drift is the output of a hidden linear dynamical system whose dynamics are themselves to be identified from the marginals. The joint problem is then a bilinear quadratically constrained programme, which we solve to certified global optimality by spatial branch and bound; with unpenalised state disturbances and a drift basis growing with the grid it is NP-hard already in latent dimension one, by reduction from $\ell^1$ rank-one matrix approximation, whereas the complexity of the deterministic system at fixed latent dimension remains open. A block-coordinate decomposition offers a cheaper alternative. For the estimator itself, we obtain rates at a fixed mesh, separated into Monte-Carlo, estimation and grid contributions.
|
| 1750 |
What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection
2609.33375
|
cs.LG
|
Jiajun Xu, Menglu Li, Xiao-Ping Zhang |
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discri...The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.
|
| 1751 |
Recursive Harness Distillation across Agents for Robot Manipulation
2609.33378
|
cs.LG
|
Seungyeon Kim, Junhoo Lee, Minkyu Kim, Baekseung Kim, Nojun Kwak |
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong ...A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventions through interaction with these policies. We propose Recursive Harness Distillation to accumulate this experience as reusable guidance across agents. A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback. The resulting playbook enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters. In real-world manipulation, the harness improves success from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook. The same playbook also benefits the strong agent, which reaches 79.2% success. These results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.
|
| 1752 |
Local LMO is Secretly a Projection Method!
2609.33383
|
cs.LG
|
Peter Richt\'arik, Ammar Mahran |
The local linear minimization oracle (Ferris and Zavriev, 1996; arXiv:2605.08850), or Local LMO, solves constrained convex problems without having to compute a projection: it minimizes a linear model over the intersection of the feasible set with a ball around...The local linear minimization oracle (Ferris and Zavriev, 1996; arXiv:2605.08850), or Local LMO, solves constrained convex problems without having to compute a projection: it minimizes a linear model over the intersection of the feasible set with a ball around the current iterate. We show that, whenever the ball radius does not exceed the Polyak radius, the Local LMO step is the Euclidean projection of the current iterate onto the intersection of the feasible set with a half-space that separates the iterate from the solution set; this projection is nonetheless computable by a linear oracle alone. We demonstrate that Local LMO belongs to a broader family of projection methods which may be indexed by the depth of the localizing half-space. For objectives with $\vartheta$-H\"older continuous gradient, every method from this family whose half-space lies sufficiently deep drives the best of its first $K$ iterates to optimality at the universal rate $\mathcal{O}(K^{-(1+\vartheta)/2})$, matching the non-accelerated universal gradient method of Nesterov (2015). Run at the Polyak radius, Local LMO attains the same rate when the constrained optima are also unconstrained ($\|\nabla f(x_\star)\| = 0$), and the rate $\mathcal{O}(K^{-1/(2-\vartheta)})$ otherwise.
|
| 1753 |
Let CSP Be Your ANCHOR: Adaptive Crystal Search over Frozen Structure Priors
2609.33407
|
cs.LG
|
Emma Lei Hovmand, Jonas Elsborg, Melih Kandemir, Arghya Bhowmik |
De novo crystal generation (DNG) models decide where to search in composition space and how to generate structures with one set of weights. We argue that discovery is better served by separating the two. A crystal structure prediction (CSP) model is a physical...De novo crystal generation (DNG) models decide where to search in composition space and how to generate structures with one set of weights. We argue that discovery is better served by separating the two. A crystal structure prediction (CSP) model is a physical prior that should be improved by likelihood training, while rewards, including novelty measured against the search's own history, should act on a search over compositions. We introduce ANCHOR, a GRPO composition policy trained with multi-objective rewards around a frozen CSP model, and continuous adaptive novelty (CAN), a graded novelty score against known structures and a growing discovery history. Using the frozen CSP model as a fixed ruler under one evaluator, we test where adaptation should act. Replacing DNG compositions with ANCHOR's policy on the same CSP backbone raises MSUN from 11.4% to 47.6% and SUN from 1.1% to 22.1% at 99.9% formula uniqueness. Fine-tuning DNG models directly on the same rewards instead moves their composition marginal without raising their on-hull fraction. We show that KL-regularized fine-tuning of a DNG model can only reweight chemistry the pretrained model already supports by a bounded factor, while unregularized DNG fine-tunes move toward known or less stable chemistry. Even a stability-only reward routed into ANCHOR's CSP backbone roughly halves SUN relative to the frozen backbone, whereas likelihood training on structures found during search can improve a CSP backbone. Under MatterGen's evaluation pipeline, ANCHOR raises state-of-the-art MSUN from 29.2% to 41.3%, transfers without retraining to two further CSP backbones, and reaches 47.1% after distillation into Crystalite-CSP. As with any model optimised against a potential, its on-hull rate depends on that potential.
|
| 1754 |
Recovering Lower-Dimensional Semialgebraic Support of a Measure from its Moments
2609.33420
|
cs.LG
|
Ruben Karapetyan, Shenyuan Ma, Ales Wodecki, Jakub Marecek |
Recovering probability measures from their moments has numerous applications, esp. in connection with the method of moments in statistics and optimization. In the setting where measure need not be finitely atomic, but its support is known to be compact and sem...Recovering probability measures from their moments has numerous applications, esp. in connection with the method of moments in statistics and optimization. In the setting where measure need not be finitely atomic, but its support is known to be compact and semialgebraic with codimension at least one, the problem is still open. We combine moment-matrix kernel information with the Christoffel--Darboux kernel to provide a discrete approximation of the support. To validate the proposed approach, we test our algorithm on analytically computed moments and pseudo-moments arising from polynomial optimization problems without unique global minimizers. This complements well-known recent work on recovery of measures with algebraic support, where the kernel of a moment matrix can reveal polynomials vanishing on the support, and on recovery of sufficiently regular full-dimensional supports, where estimators constructed by thresholding the Christoffel--Darboux kernel are known to converge asymptotically to the support.
|
| 1755 |
A Visual Classification Dataset and Model Evaluation for Historical Manuscript Illustrations
2609.33449
|
cs.LG
|
Yoav Evron, Michal Bar-Asher Siegal, Michael Fire |
Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the mate...Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the material itself remains difficult to explore at scale. Extraction systems can find illustrations on manuscript pages, but without meaningful categories, large collections remain hard to search and explore. We address this gap by introducing a manually labeled dataset of 15,000 illustrations from manuscripts dating back hundreds of years across 22 categories, and evaluating modern vision models for image classification on this task. The problem is challenging due to stylistic diversity, degradation, and semantic ambiguity, with many images that fit more than one category. We compare fine-tuned CNN and Transformer-based classifiers, zero-shot CLIP, embedding-based classifiers, and direct vision-language models. Results show that fine-tuned image classifiers perform best overall, with ConvNeXt reaching 88.9% accuracy and 81.3% macro-F1. Using CLIP embeddings with XGBoost provides a strong alternative. In contrast, zero-shot CLIP and direct vision-language classification perform substantially worse, highlighting the limits of general-purpose models in this domain. Beyond overall performance, the analysis reveals which categories are visually separable and where errors reflect genuine semantic overlap, suggesting that some limitations arise from the taxonomy itself.
|
| 1756 |
Sharp training-conditional coverage for conformal prediction under covariate shift
2609.33456
|
cs.LG
|
Mehrdad Pournaderi |
Weighted split conformal prediction reweights calibration scores by the likelihood ratio between the test and training covariate distributions and guarantees marginal coverage under covariate shift. We study its coverage conditional on the calibration data. An...Weighted split conformal prediction reweights calibration scores by the likelihood ratio between the test and training covariate distributions and guarantees marginal coverage under covariate shift. We study its coverage conditional on the calibration data. An elementary argument, based on a single concentration inequality at a fixed population quantile, gives explicit training-conditional bounds without unspecified constants, and shows that the relevant scale is not the supremum of the likelihood ratio but a variance proxy built from the chi-squared divergence of the shift and from the average of the ratio over the part of the test population, of probability equal to the miscoverage level, where it is largest. A two-point lower bound shows that the root-m rate and the chi-squared contribution are intrinsic to the shift. Run at an explicitly inflated level, the weighted quantile becomes a deterministic PAC prediction set. We compare it with randomized rejection sampling and with importance-weighted learn-then-test and, through a certified choice of a clipping level for the likelihood ratio, map the regime in which each gives the narrower valid set. The analysis extends to estimated likelihood ratios and to tail functionals estimated from an unlabeled source sample, which yields a fully finite-sample certificate.
|
| 1757 |
Protected Cores Are Not Enough: Certifying AI-Proposed Revisions of Temporal Specifications
2609.33461
|
cs.LG
|
Ruggero Lanotte |
Runtime monitoring traditionally evaluates a specification that is fixed before execution or externally modified when requirements change. In learning-enabled and data-intensive systems, however, the temporal relationships represented by a specification may th...Runtime monitoring traditionally evaluates a specification that is fixed before execution or externally modified when requirements change. In learning-enabled and data-intensive systems, however, the temporal relationships represented by a specification may themselves evolve. Allowing an AI component to directly replace a formal specification is unsafe: it may overfit transient behavior, weaken protected requirements, or activate statistically unsupported revisions. We introduce an intersymbolic architecture in which an untrusted AI proposer suggests temporal specification revisions and a symbolic governor controls their activation. Two results organize the framework. First, origin-version semantics makes the outcome of each obligation invariant to later revisions. Second, aggregate certification can conceal systematic failures on protected triggers; simultaneous aggregate and core-conditional post-selection certification controls both targets. A structural invariant preserves designer-protected components, and a proposer-independent lifetime error bound supports repeated activation decisions. The statistical bound concerns the predictable means of completed certification samples; interpreting it as future operational validity requires an additional stability assumption. Controlled synthetic experiments use a frozen supervised AI proposer to illustrate the masked-core failure at one decision and across repeated governed revisions. The proposer is a supervised regressor trained offline on synthetic tasks and frozen before use; it ranks candidates by predicted aggregate margin and never observes the protected-trigger success rate, so the masked-core failure arises from optimising the aggregate rather than from an adversary constructed by hand.
|
| 1758 |
Mask-Induced Displacement in Audio XAI via Logit Trajectory Decomposition
2609.33486
|
cs.LG
|
Nico Garc\'ia-Peguinho (School of Electronic Engineering and Computer Science, Queen Mary University of London), David Kelly (Department of Informatics, King's College London), Fabrizio Smeraldi (School of Electronic Engineering and Computer Science |
Perturbation-based XAI methods for audio classifiers often estimate feature importance by masking spectrogram regions and crediting output changes to the retained signal. Yet they typically assume the fill (the mask replacement) is negligible. We propose logit...Perturbation-based XAI methods for audio classifiers often estimate feature importance by masking spectrogram regions and crediting output changes to the retained signal. Yet they typically assume the fill (the mask replacement) is negligible. We propose logit-space trajectory decomposition to examine this assumption. An on-axis component captures output along a line connecting the fully filled (occluded) spectrogram to the fully retained original; an off-axis component captures perpendicular displacement. We evaluate across three fills, three audio classifiers, and 1,100 AudioSet clips. We demonstrate that no fill is acoustically neutral under full occlusion: Zero fill activates silence and Gaussian Noise fill activates broadband noise. Under partial masking, 41-77% of output displacement is off-axis, with the direction of the residual stable across mask retention fraction and specific to each model-fill combination. Attribution methods are unevenly exposed to off-axis displacement through their sampling and weighting strategies, revealing apparatus-dependence. Where the off-axis residual is stable and low-dimensional, its structure affords mitigation.
|
| 1759 |
Domain-Adapted Diffusion Models for Conditional Independence Testing
2609.33490
|
cs.LG
|
Yanfeng Yang, Junda Zhao, Yijie Gao, Jiaqi Yang, Xinyu Shi |
Conditional independence (CI) is a fundamental concept in statistics and machine learning. Recent advances in conditional generative modeling provide flexible tools for generative-model-based CI tests, which rely on an estimated conditional distribution to gen...Conditional independence (CI) is a fundamental concept in statistics and machine learning. Recent advances in conditional generative modeling provide flexible tools for generative-model-based CI tests, which rely on an estimated conditional distribution to generate randomized samples. However, errors in estimating this distribution accumulate in existing Type I error bounds, and consistency of the generative estimator alone does not guarantee asymptotic Type I error control. To address this limitation, we formulate conditional generative modeling as a domain adaptation problem and leverage auxiliary data from multiple source domains to improve estimation in the target CI testing domain. We propose Domain-Adapted Diffusion (DA-Diff), a multi-source domain adaptation framework for conditional diffusion models based on weighted empirical risk minimization over both target and source domains. We establish the convergence rate of DA-Diff and show how transferable source data can improve target-domain estimation through an increased effective sample size while controlling transfer bias. Building on DA-Diff, we further propose Domain-Adapted Conditional Independence Testing (DA-CIT) and show that its Type I error satisfies $P(p \leq \alpha) \leq \alpha + o(1)$. Experiments demonstrate that DA-Diff improved conditional generation quality compared with transfer-learning diffusion baselines, while DA-CIT provides strong Type I error control and competitive power.
|
| 1760 |
Neural Scaling Laws of Transformer Operator Network
2609.33533
|
cs.LG
|
Haoran Yan, Zhongjie Shi, Yuanzhe Xi, Peng Chen, Wenjing Liao |
Transformers have emerged as powerful architectures for learning solution operators of physical systems. Empirically the prediction error has been observed to decrease when the data size and model size increase, suggesting neural scaling behavior. Yet a theore...Transformers have emerged as powerful architectures for learning solution operators of physical systems. Empirically the prediction error has been observed to decrease when the data size and model size increase, suggesting neural scaling behavior. Yet a theoretical understanding of such scaling laws for transformer-based operator learning remains limited. In this work, we develop a theoretical framework for characterizing the approximation and generalization errors of transformer-based operator learning. Our analysis builds on a local-to-global approximation principle that is naturally aligned with the softmax attention mechanism and yields discretization-invariant output functions. On approximation theory, we derive a universal approximation error of transformer-based operator learning for H\"older-regular operators. On generalization theory, we establish a power scaling law between the prediction error and the training data size. The rate of convergence represented by the scaling exponent explicitly reflects the dimensions of the input and output domains, the regularity of the underlying functions and operators, and crucially, the intrinsic dimension of the input function class. By exploiting this intrinsic low-dimensional structure, our analysis yields a power-law generalization rate for operator learning, in contrast to the logarithmic-type power-law rates appearing in existing analyses of operator learning with feedforward neural networks. Numerical experiments validate the predicted power-law scaling and confirm that the convergence rate varies systematically with the intrinsic dimension of the input function class.
|
| 1761 |
OSCC: Certified Observation-Safe Coupling Optimization for Gradient-Noise Control in Imperfect-Information Learning
2609.33543
|
cs.LG
|
Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang |
Coupled rollouts can reduce the noise of counterfactual action comparisons, but two issues prevent standard common-random-number constructions from serving as a general learning primitive in imperfect-information environments. First, an invalid coupling may ex...Coupled rollouts can reduce the noise of counterfactual action comparisons, but two issues prevent standard common-random-number constructions from serving as a general learning primitive in imperfect-information environments. First, an invalid coupling may expose hidden state, synchronize endogenous policy randomness, or misalign chance events after counterfactual histories diverge. Second, in multi-action policy optimization, lower return-contrast variance is not by itself the relevant objective: the optimizer depends on the return covariance matrix after projection through the local policy-gradient geometry. We introduce observation-safe counterfactual coupling (OSCC), a framework that defines an admissible class through marginal preservation, information-state safety, branch-local policy randomness, semantic event alignment, and trace-before-oracle replay. We derive a gradient-aware coupling criterion showing that, for marginal-preserving couplings, policy-gradient noise changes are determined by policy-Jacobian-weighted off-diagonal return covariance. This motivates OSCC-Select, a calibration-only selector that chooses among independent, root-only, continuation-only, and fully coupled rollouts using separate safety and gain certificates. Its gain target combines projected gradient noise with measured physical sampling cost and falls back to independent sampling whenever a simultaneous lower confidence bound does not certify improvement. On 100,000 fixed-root Leduc comparisons, the fully coupled CP-GRPO instantiation reduces return-contrast variance from 41.1158 to 18.1441, a 55.87% reduction, while preserving the declared branch marginals. With three actions, OSCC-Select chooses continuation coupling and attains gradient-noise trace 0.0783 versus 0.0917 for return-variance selection. Increasing calibration from 64 to 2,048 groups raises certification from 0.327 to 0.995.
|
| 1762 |
Calibrated Derivative-Process Sensitivity for Gaussian-Process Variable Selection
2609.33549
|
cs.LG
|
Jia Cai |
Automatic relevance determination (ARD), the default tool for variable selection in Gaussian-process (GP) regression, ranks inputs by inverse lengthscales -- which measure how fast a function varies, not how much an input contributes to prediction -- and offer...Automatic relevance determination (ARD), the default tool for variable selection in Gaussian-process (GP) regression, ranks inputs by inverse lengthscales -- which measure how fast a function varies, not how much an input contributes to prediction -- and offers no calibrated rule for deciding which inputs to keep. The prediction-centred alternative, the derivative sensitivity $\nu_j = \mathbb{E}[(\partial f/\partial x_j)^2]$, is available in closed form from a fitted GP, but turning it into a selection rule is harder than it looks: at a null input the estimator is a degenerate quadratic form, so Wald and Bernstein-von Mises cutoffs are anti-conservative, and the natural residual bootstrap is mis-scaled. We show that a studentized multiplier bootstrap of the GP derivative process repairs both, prove its validity through an invariance principle for quadratic forms, and obtain asymptotic family-wise and false-discovery-rate control across inputs. Over 100 replications the rule controls FDR wherever inputs are truly null, while uncalibrated derivative rankings breach the target by up to 2x and a Bernstein-von Mises cutoff by 2.2x; at matched FDR it loses no power; it holds under a Mat\'ern kernel and input correlation up to 0.99; on real data with planted and authentic null inputs it admits 5-12x fewer spurious inputs; it costs 5-18% of the GP fit; and a block-averaged variant retains validity at cost linear in $n$.
|
| 1763 |
Sparsity by Default: The Theory and Practice of ARD in Gaussian Process Regression for Variable Selection
2609.33550
|
cs.LG
|
Jia Cai |
Automatic relevance determination (ARD) is the standard device for input selection in Gaussian process (GP) regression. By giving the covariance kernel a separate lengthscale for every input and learning those lengthscales by maximizing the marginal likelihood...Automatic relevance determination (ARD) is the standard device for input selection in Gaussian process (GP) regression. By giving the covariance kernel a separate lengthscale for every input and learning those lengthscales by maximizing the marginal likelihood, ARD lets the data decide which coordinates matter: irrelevant inputs receive very large lengthscales and are effectively switched off. We trace this mechanism to the Bayesian Occam's razor embodied in the marginal likelihood, derive the gradient through which it prunes inputs, and emphasize that ARD delivers effective rather than exact sparsity. We review the algorithms used in practice and the rules that turn lengthscales into selections, and we survey the asymptotic theory, distinguishing the fixed-domain identifiability obstruction on the lengthscales from the high-dimensional selection-consistency guarantees recently established for hierarchical GP priors, and noting what remains open for plain ARD. We compare ARD with spike-and-slab priors, sparse axis-aligned and global-local shrinkage priors including the Bayesian lasso and horseshoe, penalized likelihood kriging, sensitivity and projection criteria, and additive kernels. We argue that ARD endures because of its seamless integration with kernel learning, universal software support, and low cost, and we close with its limitations and remedies.
|
| 1764 |
Terminal-Register Certification for Finite-Measurement Learning of Multiscale Quantum States
2609.33567
|
cs.LG
|
Bhvain Makwana, Kashyap Patel, Manjunath Joshi, Jaideep Mulherkar |
Structured quantum-state learning not only depends on an expressive ansatz but also on an operational certificate that stays meaningful with finite measurements and imperfect implementation. We study pure one dimensional states learning by an inverse binary mu...Structured quantum-state learning not only depends on an expressive ansatz but also on an operational certificate that stays meaningful with finite measurements and imperfect implementation. We study pure one dimensional states learning by an inverse binary multiscale entanglement renormalization ansatz (MERA). In the learning procedure, the qubits removed during coarse graining are controlled coherently and measured together at the terminal register. We confirm that an ideal sequential and terminal measurement schedule delivers the same complete bit string distribution under matched causal operations, while normalized postselection can amplify perturbations inversely with prefix acceptance. A noise aware theorem introduces an individual calibrated total variation implementation budget to the finite shot certificate. The protocol is estimated on an open boundary transverse field Ising ground state. A frozen 8-qubit schedule using $560$ million simulated training measurements per run achieves fidelity above $0.99$ in all $60$ held-out runs, with a mean fidelity of $0.996886$. 1080 circuit-noise cells and 6480 confidence-coverage rows are covered by fixed-circuit robustness validation without a locked soundness violation. We then address architectural fairness at $n=16$ using three new studies. In a 120-run exact-gradient multistart diagnostic, MERA has higher fidelity in 58/60 paired restarts and lower long-range error in 60/60, although no run met the prespecified stationarity criterion. Finally, a causal cone-complete, parameter matched local circuit achieves $2.62\times$ greater aggregate gate exposure yet loses all 30 paired comparisons in fidelity, long-range error, energy, and entropy.
|
| 1765 |
Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning
2609.33595
|
cs.LG
|
Boyuan Zhang, Yingjun Du, Xiantong Zhen, Ling Shao |
Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictio...Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictions. One-step accuracy therefore does not capture how prediction errors propagate under recursive rollout. We decompose multi-step rollout error into the errors introduced at individual steps and their propagation through subsequent transitions. We show that state-affine dynamics are precisely the differentiable transitions with state-independent Jacobians, eliminating the nonlinear propagation residual and making the error propagation operators depend only on the action sequence. Guided by this result, we introduce SALT (State-Affine Latent Transition), an action-conditioned state-affine dynamics model in which the action modulates both the state transformation and the additive update. We train SALT through recursive multi-step rollout supervision, feeding each predicted latent state back into the transition so that training matches how the model is used during planning. Across four visual planning environments, SALT exhibits $1.48$--$2.19\times$ higher one-step prediction error than the matched LeWM baseline, yet improves closed-loop success in every environment by $10.0$ percentage points on average. On OGBench-Cube, the fraction of episodes that fail with a sharp rise in model-predicted cost after execution decreases from $23.3%$ to $2.0%$.
|
| 1766 |
The Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM Traces
2609.33597
|
cs.LG
|
Georgios Feretzakis, Alexandros Papaspyridis |
Side-channel evaluators routinely inspect leakage tests while acquisition is still running, and extend or stop the campaign based on what they see. Fixed-horizon screening such as the Welch $t$-test with threshold $|t|>4.5$ gives no error guarantee for this...Side-channel evaluators routinely inspect leakage tests while acquisition is still running, and extend or stop the campaign based on what they see. Fixed-horizon screening such as the Welch $t$-test with threshold $|t|>4.5$ gives no error guarantee for this monitored decision rule. We study anytime-valid leakage detection based on testing by betting: SKIT-type swap-pair e-processes whose false-alarm probability is controlled uniformly over time under an explicit conditional symmetry null. In matched comparisons that share the frozen witness, rows and payoff, first-crossing detection needed 1.68-2.00$\times$ the traces of a fixed-horizon randomization test at 80% detection on synthetic streams, and 1.68-2.38$\times$ on degraded recordings from an open ML-KEM electromagnetic dataset with the primary Ridge witness at $\alpha=0.05$. With the same primary witness and level, on undegraded reference and pqm4 recordings the monitored procedure stopped early: its median stopping point was 62-72 and 146-316 evaluation traces, i.e. 2-8% of a conservative 4096-trace budget. Under exact designed nulls on the recorded backgrounds, repeated-look $|t|>4.5$ screening over all 13000-20000 samples raised a false alarm in 2.7-12.9% of replicates, against 0.0-4.7% for terminal-only screening and no rejection by a sample-wise e-Bonferroni process, which in a prespecified follow-up detected natural-label associations in 4 of 4 backgrounds after 840-3288 traces. All recordings come from one device, and natural-label results are descriptive; we state the assumptions each claim requires.
|
| 1767 |
lapanda: A Matrix-Free Differentiable Solver for Nonconvex Constrained Optimization Layers
2609.33630
|
cs.LG
|
Yuankun Chen, Zifei Nie, Kangyu Lin, J\'{a}n Drgo\v{n}a, Liang Wu |
Differentiable optimization brings the structural guarantees of mathematical optimization to network pipelines, allowing them to be trained end-to-end. However, its application remains challenging for nonconvex constrained problems, as existing differentiable ...Differentiable optimization brings the structural guarantees of mathematical optimization to network pipelines, allowing them to be trained end-to-end. However, its application remains challenging for nonconvex constrained problems, as existing differentiable solvers often suffer from limited modeling expressiveness due to their reliance on specialized problem structures, while also incurring substantial computation time and memory overhead in both the forward and backward passes. To address these challenges, we propose lapanda, a matrix-free differentiable solver for nonconvex optimization with general constraints. It reformulates the problem to a sequence of augmented Lagrangian subproblems, each handled by a first-order inner solver through a proximal averaged quasi-Newton algorithm with adaptive linesearch, thus enabling efficient forward optimization. We establish local well-posedness of the solution map and convergence of the outer iterations, and further derive a sensitivity alignment between the original problem and the final subproblem in the backward pass, demonstrating that the subproblem sensitivity, which can be computed efficiently in a matrix-free manner, provides a principled approximation to the exact optimizer sensitivity. We evaluate lapanda on nonconvex constrained Rosenbrock benchmarks, imitation learning with several representative constrained optimal control problems, and embedded robotic obstacle-avoidance tasks. Compared with state-of-the-art differentiable solvers, lapanda delivers substantial reductions in computation time and memory footprint while maintaining reliable constraint satisfaction and learning performance.
|
| 1768 |
Scalable and Data-Driven Decision Support in the Maintenance, Repair, and Overhaul Process
2609.33641
|
cs.LG
|
Houkun Zhu, Helena Ebel, Dominik Scheinert, Florian Schmidt, Jens Altenkirch |
Several businesses apply maintenance, repair, and overhaul (MRO) principles to the life-cycle of their existing products. In cases like casted gas turbine component Product Lifecycle Management (PLM), repairing components in frequent intervals can extend the l...Several businesses apply maintenance, repair, and overhaul (MRO) principles to the life-cycle of their existing products. In cases like casted gas turbine component Product Lifecycle Management (PLM), repairing components in frequent intervals can extend the lifetime expectation of the product, provide higher cost efficiency compared to newly produced components, and even improve the part design during the repair cycle. Another aspect of repair concerns sustainability, as products often contain rare materials. The emissions produced by the repair process are usually smaller than mining materials and casting new components. To optimize the repair process further, we propose the Smart Expert System (SES), which assists engineering experts with machine learning-based decision support throughout the repair process. We elaborate on its IT architecture and present machine learning models employed for representative MRO use cases. The SES is evaluated using actual industry data from a leading gas turbine company and demonstrably fulfills formulated requirements concerning the suitability of the overall decision support and the stability of the enclosing IT architecture.
|
| 1769 |
Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies
2609.33653
|
cs.LG
|
Duo Wu, Haifeng Wang, Rongwei Lu, Jinghe Wang, Tianyi Xiong |
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by le...Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA$_0$, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA$_0$ using newly collected rollouts as the policy evolves. Experiments show that eVTA$_0$ provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.
|
| 1770 |
LieDiscover: Adaptive Symbolic Library Construction for Explicit Open-form Symmetry Discovery
2609.33663
|
cs.LG
|
Xinxin Li, Jianming Ma, Xingyu Cui, Da Li, Juan Zhang |
Discovering underlying symmetries from data has emerged as a crucial challenge in scientific discovery. Existing data-driven methods for symmetry discovery fail to determine the exact number and mathematical form of unknown infinitesimal generators. Recent exp...Discovering underlying symmetries from data has emerged as a crucial challenge in scientific discovery. Existing data-driven methods for symmetry discovery fail to determine the exact number and mathematical form of unknown infinitesimal generators. Recent explicit methods represent generators using a predefined function library and identify them through algebraic optimization, but they often struggle to capture complex symmetries involving high-order polynomials or transcendental functions. To address this limitation, we formulate symmetry discovery as a joint optimization problem over the function library and coefficients. We propose a novel framework that leverages an encoder-decoder architecture to dynamically generate symbolic expressions and expand the library. This generation process is optimized via reinforcement learning, which accelerates the exploration of the symbolic search space through step-wise rewards. Experiments demonstrate that LieDiscover can successfully uncover open-form infinitesimal generators involving high-order polynomials or transcendental functions, which remain intractable for existing methods. The discovered symmetries also improve performance in downstream PDE solving and discovery tasks.
|
| 1771 |
RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games
2609.33669
|
cs.LG
|
Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang |
Residual policy adaptation provides a lightweight way to modify a strong reference policy, but a shared scale and a fixed subgroup partition can hide heterogeneous degradation and become fragile when the deployment mixture of information states changes. We int...Residual policy adaptation provides a lightweight way to modify a strong reference policy, but a shared scale and a fixed subgroup partition can hide heterogeneous degradation and become fragile when the deployment mixture of information states changes. We introduce RSD-Poker, a structure-adaptive and shift-robust certification framework that freezes a bank of residual families and scales, learns a policy-visible partition on an independent structure split, and freezes that partition before calibration labels are joined. Each candidate-group pair receives a weighted simultaneous upper certificate for anchor-relative risk and a lower certificate for weak-response utility. A robust group-to-candidate map is then selected over a predeclared uncertainty set of deployment group proportions. Under independent calibration units drawn from each frozen group's law, a candidate bank and partition fixed before calibration, and invariant within-group conditionals, the selected map satisfies its declared mixture-robust risk budget and utility certificate with probability at least $1-\zeta_{risk}-\zeta_{util}$. The information contract supports both a teacher-backed transform and a teacher-free observation-only student. The retained deterministic 24-state audit remains an exact replay diagnostic: empirical-zero selects $\alpha=0.08$, raising the weak-response proxy from 4.2082 to 4.2889 with $0/12$ held-out threshold crossings. On stratified held-out states, the learned-partition dual selector raises weak utility from 4.4074 under global dual certification to 4.4936 and lowers held-out violation from 0.0215 to 0.0078; its mixture-robust variant reaches violation 0.0059. Across five observation-only checkpoints, risk-calibrated residuals attain weak utility $4.3659\pm0.0177$ and violation rate $0.0178\pm0.0057$.
|
| 1772 |
COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails with Heterogeneous CPU-NPU Confidence Cascading under Explicit Latency and False-Positive Constraints
2609.33671
|
cs.LG
|
Hao Chen |
When must a foundation-model safety gateway generate tokens, and when should it directly output a calibrated decision? We study calibrated standalone direct-decision foundation models for real-time pre-ingestion safety guardrails, jointly addressing probabilit...When must a foundation-model safety gateway generate tokens, and when should it directly output a calibrated decision? We study calibrated standalone direct-decision foundation models for real-time pre-ingestion safety guardrails, jointly addressing probability calibration, dual-use false-positive control, and heterogeneous CPU-NPU routing under explicit latency SLOs. Pre-ingestion guardrails must screen prompts prior to target-LLM prefill with low false alarms on benign compliance inquiries; however, shallow classifiers are brittle to phrasing shifts, hidden-state probes require coupling to a target LLM, and generative guards incur high decoding latency and dual-use false positives. We present COGNIT-Guard, coupling a validation-calibrated CPU fast gatekeeper with confidence-gated escalation to an NPU-resident 322M bidirectional direct-decision model (Laya-322M) under an asymmetric false-positive penalty. On the clean unseen DUCS-Bench test split ($N=607$), COGNIT-Guard achieves 98.85% accuracy (McNemar $p = 1.19 \times 10^{-4}$ vs. ML), reduces benign FPR to 0.42% ($1/238$; Fisher's exact $p = 8.23 \times 10^{-4}$ vs. ML), and attains 1.12% ECE and 0.0104 Brier score. On Huawei Ascend 910C NPUs, pure NPU inference runs in 21.77 ms mean latency (45.90 QPS), while the live serial CPU-NPU cascade ($\theta^*_{\mathrm{deploy}}=0.70$) achieves 41.63 ms mean latency (P50: 39.47 ms, 99.23% accuracy, 0.00% FPR). Evaluation on SafetyBench-ZH ($N=2,100$) and comparison against a bi-encoder direct-decision baseline (CLM-8B) disentangle in-domain gains, OOD alignment tax (60.33% $\to$ 56.81% on Laya; 55.10% on domain CLM-8B), and experience replay recovery, restoring OOD accuracy to 64.10%-65.05% and reaching 99.67%-99.84% in-domain accuracy with 0.00%-0.42% FPR.
|
| 1773 |
TopoMamba: A Load-Support Relation-Guided Multi-Directional State-Space Model for Topology Optimization
2609.33688
|
cs.LG
|
Bin Lou, Yuxuan Cheng, Huaizhi Zong, Junhui Zhang, Bing Xu |
Deep learning has emerged as an efficient alternative for predicting high-performance material distributions in topology optimization. Existing methods struggle to accurately capture load-transfer information, limiting out-of-distribution generalization, while...Deep learning has emerged as an efficient alternative for predicting high-performance material distributions in topology optimization. Existing methods struggle to accurately capture load-transfer information, limiting out-of-distribution generalization, while their model architectures often incur high computational costs. To address these challenges, this paper proposes TopoMamba, a topology prediction framework incorporating a load-support relation-guided multi-directional state-space model. Coupling physical fields with load-support relations enables more effective modeling of mechanical dependencies. A load-support relation-guided spatially adaptive fusion mechanism dynamically adjusts multi-directional scan features according to spatial conditions. Mamba is coupled with the solid isotropic material with penalty method to enhance structural mechanical performance while maintaining computational efficiency. Results on two-dimensional topology optimization benchmarks demonstrate that TopoMamba achieves superior topology prediction accuracy, out-of-distribution generalization, and computational efficiency over state-of-the-art models. The proposed load-support physics-guided framework enables efficient optimization of more complex structural systems.
|
| 1774 |
Prompt-Anchored Residual Adaptation for Biomedical Vision-Language Models
2609.33701
|
cs.LG
|
Jingxuan Kang, Qianying Yue, Che Liu, Chen Qin |
Pretrained biomedical vision-language models achieve strong zero-shot performance in biomedical image classification. However, downstream biomedical classification often depends on subtle visual differences between classes that may not be fully captured by pre...Pretrained biomedical vision-language models achieve strong zero-shot performance in biomedical image classification. However, downstream biomedical classification often depends on subtle visual differences between classes that may not be fully captured by pretrained representations. Few-shot adaptation addresses this mismatch by optimizing a task-specific predictor on a small labeled support set. Because the selected examples capture only part of the visual variation within the target classes, the adapted predictions can depend strongly on their composition. We propose Prompt-Anchored Residual Adaptation (PARA), which retains the frozen prompt prediction as a support-invariant semantic anchor and incorporates a visual prediction learned from the support set through an anchor-relative residual. The residual step is computed in a closed form from frozen support embeddings using anchor discrepancy and support agreement. Support-set dependence also limits evaluation: comparisons are fair within a shared draw but remain conditional on its composition. To obtain more reliable comparisons, we introduce a repeated-support protocol that separates support-selection variation from optimization randomness and reports both average and worst-20% performance. PARA achieves state-of-the-art performance in both few-shot classification and base-to-novel generalization.
|
| 1775 |
Non-Adaptive Learning of Sparse Erd\H{o}s--R\'enyi Graphs via Affine Splitting
2609.33704
|
cs.LG
|
Hoang Ta |
Graph learning from edge-detecting queries concerns the reconstruction of an unknown edge set on a known vertex set. Each query reports whether a specified vertex subset contains at least one edge. We study non-adaptive schemes, in which all queries are fixed ...Graph learning from edge-detecting queries concerns the reconstruction of an unknown edge set on a known vertex set. Each query reports whether a specified vertex subset contains at least one edge. We study non-adaptive schemes, in which all queries are fixed before any outcomes are observed, with the goal of achieving exact recovery using few queries and fast decoding. For general graphs on $n$ vertices with at most $k$ edges, non-adaptive recovery requires $\Omega(\min\{k^2\log n,n^2\})$ queries in the worst case, even when a small error probability is allowed. In this paper, we consider Erd\H{o}s--R\'enyi ($\mathrm{ER}$) graphs $G\sim \mathrm{ER}(n,q)$, with expected edge count $\bar{k}=q\binom{n}{2}$. Our scheme uses $O(\bar{k}\log n)$ queries and achieves exact recovery in $O(\bar{k}\log n)$ decoding time with probability tending to one throughout the regime $\bar{k}\to\infty$ and $\bar{k}=o(n^2)$. This improves the previous $O(\bar{k}^{1+\delta}\log n)$ decoding guarantee for any fixed $\delta>0$, while maintaining the same query order. The guarantee also extends beyond the previously studied regime $\bar{k}=\Theta(n^{2\theta})$ with fixed $\theta\in(0,1)$. Our approach builds on the binary splitting method used in prior work, which organizes vertices into a hierarchy of successively smaller groups. We introduce three main changes: (i) we use random affine hash functions over a finite field to process each candidate pair in constant time; (ii) we apply the splitting procedure directly to the full graph, avoiding the need to combine solutions to multiple smaller graph-learning subproblems; and (iii) we bound the total decoding workload directly rather than deriving separate high-probability bounds on candidate counts at each level.
|
| 1776 |
Reliable Replay through Spatial Coherence in Online Continual Learning
2609.33725
|
cs.LG
|
Haixiang Sun, Jiefu Zhang, Yinghao He, Yang Xu, Vaneet Aggarwal |
Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same ...Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same update and can overemphasize isolated responses. We introduce SPatial coHErent risk control for REplay (SPHERE), a general replay-allocation method applicable across a broad range of learning settings. SPHERE uses a representation kernel to aggregate signed prospective loss changes, attenuating unsupported spikes while retaining coherent increases. It then formulates allocation as entropy-regularized transport, redistributing uniform source mass toward supported high-risk regions while penalizing long-distance transfers. We derive replay coefficients from the transport objective's sensitivity to the original loss changes and blend them with uniform replay to maintain baseline rehearsal. Our analysis establishes conditions under which kernel aggregation improves risk estimation and bounds transport-value inflation due to residual noise and smoothing bias. Experiments demonstrate that SPHERE improves accuracy and reduces forgetting across noisy-label vision tasks, continual language-model instruction tuning, and code-generation reinforcement learning with incomplete test rewards.
|
| 1777 |
A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions
2609.33727
|
cs.LG
|
Luyang Fang, Haoran Lu, Jiazhang Cai, Tao Wang, Huimin Cheng |
Knowledge Distillation (KD) has emerged as a vital paradigm for transferring the capabilities of high-capacity models to efficient ``student'' counterparts, addressing critical challenges in computational cost, deployment constraints, and privacy-sensitive set...Knowledge Distillation (KD) has emerged as a vital paradigm for transferring the capabilities of high-capacity models to efficient ``student'' counterparts, addressing critical challenges in computational cost, deployment constraints, and privacy-sensitive settings. Although KD is widely used in practice, it is often viewed primarily as an engineering technique, with a unified statistical perspective remaining less developed. This review bridges that gap by presenting a unified Bayesian formulation of KD that formulates teacher predictions as prior information. This provides a principled interpretation of how teacher information is incorporated into student learning and establishes a rigorous connection to uncertainty quantification. We demonstrate how this foundational lens reconciles classical distillation with modern extensions in generative and foundation-model systems, showing that contemporary developments remain rooted in these same statistical principles. By synthesizing theory with emerging methodologies and diverse applications, this review provides a conceptual roadmap and identifies critical open problems for the future of the field.
|
| 1778 |
Constrained Edit Fields for Training-Free Flow Editing
2609.33735
|
cs.LG
|
Jingxuan Kang, Yinsong Wang, Che Liu, Chen Qin |
Text-guided image editing aims to perform a desired edit while preserving source content unrelated to it. Pretrained rectified-flow models enable training-free editing of real images through modifications to their sampling trajectories. However, responses at l...Text-guided image editing aims to perform a desired edit while preserving source content unrelated to it. Pretrained rectified-flow models enable training-free editing of real images through modifications to their sampling trajectories. However, responses at locations unrelated to the desired edit can still accumulate along the editing trajectory and become visible in the final result. To overcome this, we propose Constrained Edit Fields (CEF), which assigns each spatial location a continuous edit responsibility that quantifies its relevance to the desired edit. CEF estimates edit responsibility directly from the source image when the relevant content is present. For edits whose target content is absent from the source, CEF first generates an unconstrained proposal to reveal its realized spatial support and then estimates responsibility from that proposal. At each editing step, CEF decomposes the base edit field into prompt-induced and trajectory-induced components, enabling edit responsibility to preserve instruction-relevant updates while suppressing unintended trajectory-induced changes. Evaluated on all 700 PIE-Bench examples, CEF achieves state-of-the-art Structure Distance, background LPIPS, and background MSE with both Stable Diffusion 3.5 Medium and FLUX, while retaining competitive instruction alignment. On Stable Diffusion 3.5 Medium, it reduces these metrics over the previous best results by 10.2%, 21.2%, and 48.0%, respectively.
|
| 1779 |
Weighted Spline-Expanded Networks with Distributional Balancing for Continuous Treatment Effects
2609.33740
|
cs.LG
|
Shucheng Liu, Chan Park, Guanhua Chen |
Estimating causal effects with continuous treatments in observational studies is challenging due to confounding, model misspecification, and high-dimensional covariates. We propose the Weighted Spline-Expanded Network (WSENet), an end-to-end neural framework t...Estimating causal effects with continuous treatments in observational studies is challenging due to confounding, model misspecification, and high-dimensional covariates. We propose the Weighted Spline-Expanded Network (WSENet), an end-to-end neural framework that addresses these challenges by combining covariate balancing, structured treatment embedding, and bias-corrected outcome estimation. WSENet first applies Distance Covariate Optimal Weights to induce distributional independence between covariates and treatment without relying on parametric models. It then learns the conditional outcome via a structured network that fuses outcome-relevant representations of covariates with a spline-expanded treatment input, enabling smooth and flexible modeling of the dose-response relationship. To mitigate residual bias, we introduce Weighted Targeted Regularization, a correction technique based on efficient influence functions that yields a doubly robust estimator. Extensive evaluations on semi-synthetic and real-world datasets, including high-dimensional genomic and environmental health data, demonstrate that WSENet consistently outperforms existing baselines in both accuracy and stability.
|
| 1780 |
Multi-Marginal Inverse Optimal Transport for Contrastive Learning Via Explicit Anchor-Positive-Negative Coupling
2609.33741
|
cs.LG
|
Ngoc-Hai Nguyen, Thuan Nguyen, Prakash Ishwar, Shuchin Aeron |
Inverse Optimal Transport (OT) based methods for representation learning learn representations such that the global OT coupling between a pair of data marginals in the representation space, concentrates on the positive pairs. This is in contrast to previous me...Inverse Optimal Transport (OT) based methods for representation learning learn representations such that the global OT coupling between a pair of data marginals in the representation space, concentrates on the positive pairs. This is in contrast to previous methods that primarily focused on pairwise matching. However, these methods $\textit{DO NOT}$ utilize negative pairs and hence are not truly contrastive in their approach. We show that this leads to issues of dimensional collapse and hence degraded downstream performance. To alleviate this, we develop a novel multi-marginal (MM) inverse OT (IOT) contrastive learning (CL) approach called Neg-MMIOT-CL, which learns representations such that the global multi-marginal OT (MMOT) coupling between a triple of data marginals, with respect to a carefully designed ground-cost between triplets of data points in the representation space, concentrates on the anchor-positive-negative $\textit{triplets}$. For a latent class model, we empirically show that Neg-MMIOT-CL alleviates dimensional collapse. Furthermore, for a specific choice of ground cost for all triplets in representation space, we prove that the optimal representation configuration for Neg-MMIOT-CL exhibits equiangular property for within-class and across-class representations, which translates to Neural-Collapse when the representation dimension is larger than the number of classes minus one -- a result that is $\textit{previously established only}$ for pairwise contrastive learning methods. Finally, we propose Neg-IOT-CL-PushPull, that is a computationally efficient alternative to Neg-MMIOT-CL, alleviating the high cost of computing MMOT plans needed during implementation. We apply these methods on both synthetic and real-world datasets and show significant improvements over existing OT-based contrastive learning methods.
|
| 1781 |
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
2609.33757
|
cs.LG
|
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou |
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at fron...Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
|
| 1782 |
EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?
2609.33762
|
cs.LG
|
Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li, Chaofan Tao |
LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory...LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the same coding-agent workload it speeds up one deployment, slows down another, and changes nothing on a third, even where loading a token back is several times cheaper than recomputing it. The reason is that cached state must survive until it is used again. While one agent waits for its tool, the server processes the contexts of all other agents, so an agent's prefix is reused only if the host tier holds the reusable context of the whole agent pool, which we call the reuse working set. A smaller tier keeps writing state that is evicted before anyone reads it. We present EfficientAgent, which sizes and manages the host tier by this working set. A stack-distance model estimates the working set from agent histories to size the host tier; its predictions, made before the experiments, located the capacity at which offloading starts to pay. When the tier is too small, a runtime policy stops writing large refills of evicted context and keeps extending prefixes that are still cached; when the tier is large enough, it writes everything. On SWE-bench Verified coding agents, a host tier sized to the estimated working set cuts recomputed prompt tokens by 93% and end-to-end time by 39%. With a small fixed tier, the policy cuts recomputation by 35%; with a large tier, it avoids the 4.3-fold increase caused by always filtering writes. Across three GPU types and two models, offloading pays off when the GPU has little compute per byte of host bandwidth and the host tier holds the working set. Code is available at https://github.com/KunmingSHAO/efficientagent_release.
|
| 1783 |
Taming the Greeks: Option Portfolios with Inductive Biases
2609.33767
|
cs.LG
|
Wee Ling Tan, Stephen Roberts, Stefan Zohren |
We present an end-to-end deep learning framework for systematic options trading that directly embeds hedging behavior through explicit control of portfolio-level risk exposures. While neural networks trained to optimize risk-adjusted performance have been show...We present an end-to-end deep learning framework for systematic options trading that directly embeds hedging behavior through explicit control of portfolio-level risk exposures. While neural networks trained to optimize risk-adjusted performance have been shown to outperform traditional rules-based strategies, such approaches remain agnostic to the sensitivities of the resulting portfolios with respect to specific underlying risk factors. We propose a general training objective that combines a performance-driven loss with a differentiable risk-sensitivity penalty, enforcing neutrality to selected risk dimensions. Unlike reinforcement learning methods that approximate optimal hedging policies via simulated market dynamics, our framework operates entirely on historical data and jointly optimizes risk-adjusted returns and targeted risk constraints in a single learning problem. We instantiate the framework on static delta-neutral straddle portfolios with the penalty directed at first-order directional exposure, and evaluate two penalty variants -- an exposure-normalized penalty and a Greek-ratio drift penalty. Empirical results on Nasdaq 100 equity options demonstrate that appropriately calibrated regularization simultaneously improves out-of-sample risk-adjusted performance relative to an unregularized baseline while reducing realized directional exposure.
|
| 1784 |
RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation
2609.33796
|
cs.LG
|
Yingming Zhou, Adarsh Vatsa, William Eiers |
Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct Ceda...Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first dataset that supports both training and semantic evaluation for formally verifiable Cedar policy synthesis. It contains 5,800 scenarios across 44 domains and 1,408 representing a single synthetic organization, each with a verified target policy and an executable verification plan. On this data we introduce RAISE, which trains policy synthesizers from formal verification in two stages, verified supervised fine-tuning (SFT) followed by a reinforcement learning (RL) stage that learns from verifier signal. We find that SFT succeeds largely by letting models express authorization logic they already have, since untrained models rarely write valid Cedar but often reason correctly when they do. After SFT, how the verifier's information is used matters more than how much of it is used. Of six RL instantiations that consume progressively richer verifier signal, only RAISE-OC improves meaningfully on SFT; it turns failed checks and symbolic counterexamples into guided exploration and learns from the result with off-context GRPO. With about 5.4K verified scenarios and LoRA fine-tuning, RAISE-OC trains Qwen3.5-9B to surpass zero-shot GPT-6 Astra and Claude Opus 5 by 13.33 and 16.26 percentage points in semantic success on held-out scenarios, and training transfers to the independently constructed CedarBench.
|
| 1785 |
Autonomous phase discovery
2609.33802
|
cs.LG
|
Shiyu Zhou, Yuxuan Zhang, Sebastian Wetzel, Roger Melko, Xiu-Zhe Luo |
Understanding quantum phases of matter has long relied on physicists' intuition and mathematical tools such as symmetry and topology. Remarkably successful as these approaches have been, they provide no universal way to explore a Hamiltonian space whose organi...Understanding quantum phases of matter has long relied on physicists' intuition and mathematical tools such as symmetry and topology. Remarkably successful as these approaches have been, they provide no universal way to explore a Hamiltonian space whose organizing principle is not known in advance. In this work, we introduce a fully autonomous system combining differentiable programming and unsupervised learning for quantum phase discovery. The search evaluates ground-state data along an adaptive trajectory rather than on a predetermined parameter grid. We demonstrate the system with three different solvers and benchmark it against random sampling at equal ground-state-evaluation budgets. On a generalized cluster chain hosting up to $200$ distinct phases, the search finds up to $25$ more phases at the same budget, and matches random sampling given thirty times its budget. On a $50$-parameter Chern insulator, it reaches sectors not obtained by the simple harmonic constructions considered here, in a family whose inverse problem remains open, while recovering all sectors found by sampling. Our results establish autonomous, gradient-driven exploration of Hamiltonian space as a practical route to discovering quantum phases without phase labels or a prescribed target phase.
|
| 1786 |
Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
2609.33812
|
cs.LG
|
Eduardo Ari\~no de la Rubia (Central European University), Szilard Pafka (Epoch) |
Repeated runs of the same coding agent are known to give different benchmark scores. We ask what that variation means for a team running an agent on its own task, by intensive replication on one machine-learning task: an agent improves the training code of an ...Repeated runs of the same coding agent are known to give different benchmark scores. We ask what that variation means for a team running an agent on its own task, by intensive replication on one machine-learning task: an agent improves the training code of an XGBoost classifier for airline delays, and a holdout it never sees scores the result. Across 584 runs, we compare six agents on six open-weight model endpoints, run six agent-model pairings 52 times each under fixed settings, and repeat three of them on a larger model from the same family. Identical runs of one pairing varied more than the pairings differed from one another, so comparisons of a few runs ranked them unreliably; resolving the agent differences we observed would take tens to more than a hundred runs of each. Runs on the larger model scored clearly higher, but by less than one run-to-run standard deviation, and the gap was more than twice as large with one agent as with the others. Fewer than one run in twenty broke the task's data rules, but those runs held the highest scores. Rejecting those runs first and keeping the best compliant result among a few attempts reliably improved the delivered model, even though a few runs could not rank the agents. On flights from a later year, the delivered models kept only a third of their gain over the starting code. At list prices, cost differed more than twentyfold between two agents on the same model, mostly through the prompt cache. Agents and models should be evaluated as pairings, over repeated attempts, with compliance reported beside quality. Data, code and every delivered program: https://github.com/earino/identical-runs-different-results
|
| 1787 |
Annealed Sinkhorn with Momentum: Certified Unregularized Optimal Transport in Linear Memory
2609.33814
|
cs.LG
|
Samuel J. K. Chin, Maximilian Schiffer |
We characterize Bregman Douglas-Rachford splitting (BDRS) for unregularized discrete optimal transport and develop an anytime primal-dual certificate in linear memory. We first establish that BDRS coincides with warm-started Inexact Proximal point method for e...We characterize Bregman Douglas-Rachford splitting (BDRS) for unregularized discrete optimal transport and develop an anytime primal-dual certificate in linear memory. We first establish that BDRS coincides with warm-started Inexact Proximal point method for exact Optimal Transport (IPOT) using a single inner Sinkhorn iteration. By eliminating the primal transport plan from the updates, we derive an equivalent dual formulation that reveals BDRS as annealed Sinkhorn under an implicit inverse-linear temperature schedule, with an additional log-scaling momentum term and a cooler kernel. While this explains the role of the temperature parameter in BDRS as an initial temperature, it also reduces the solver's memory requirement from quadratic to linear. Utilizing this annealing perspective, we introduce overrelaxed BDRS, which combines annealing and overrelaxed scaling within a single recursion. We derive a primal-dual certificate for both methods that can be evaluated in linear memory without transport plan construction, thus providing a computable stopping rule. On the DOTmark benchmark, combining momentum with the cooler kernel produces substantially smaller optimality gaps than annealed Sinkhorn under the same schedule. For pixel-level color transfer between $1024\times1024$ images, BDRS attains a lower repaired transport cost than MDOT-TNT with a 9$\times$ speed up, reaching a relative duality gap of $1.59\%$ in 24 minutes. We further demonstrate a color transfer with $4238\times2365$ images, yielding 10 million pixels per image and approximately one hundred trillion implicit transport entries, reaching a best relative duality gap of $2.41\%$ and $2.80\%$ within 35 hours in each direction on a single NVIDIA L40S GPU.
|
| 1788 |
Achieve What You Imagined: Learning to Align Actions with Visual Plans
2609.33832
|
cs.LG
|
Yuheng Qiao, Ziran Wei, Xiaohan Wang, Daqiang Guo, Yichen Luo |
World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausibl...World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a frozen action-conditioned world model to predict action-conditioned consequences and construct feedback based on consistency between the two future predictions and alignment with the terminal goal. Leveraging this feedback, we employ Flow Policy Optimization (FPO) to optimize the action head of the WAM. This framework avoids online robot interaction and additional training of task-specific reward models. Across four real-world UR5 manipulation tasks, our method increases the mean success rate from 43.4% to 75.1%, compared with 61.4% for $\pi_{0.5}$. These results show that cross-model prediction discrepancy can provide useful feedback for improving robot policies under the evaluated manipulation tasks. Website: https://imagine-to-achieve.github.io/
|
| 1789 |
One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs
2609.33833
|
cs.LG
|
Sen Nie, Jie Zhang, Zhongqi Wang, Shiguang Shan, Xilin Chen |
Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this ...Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. We propose O-Attack, a highly transferable black-box attack framework. This framework builds on our insight that surrogate models contain a broad, high-level, cross-modally aligned semantic space. This space extends beyond final-layer outputs and provides multiple semantically consistent representations that remain underexploited by existing attacks. Within this space, O-Attack anchors aligned representations, progressively broadens semantic conditions, and optimizes perturbations through semantic consensus to promote consistent target alignment. By fully exploiting this space with the same surrogate models as M-Attack, O-Attack raises attack success rates on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Extensive experiments across 24 MLLMs show that O-Attack outperforms six state-of-the-art methods in black-box transferability, with consistent effectiveness across prompts and improved efficiency and imperceptibility. This work exposes the practical safety risks posed by black-box adversarial attacks against frontier MLLMs, underscoring the need for more rigorous robustness evaluation and more effective defenses.
|
| 1790 |
CLIMB-flow: Coupled Linear Inverse posterior sampling via Multiscale-Based flow
2609.33834
|
cs.LG
|
Zeqiu Yu, Ruizhi Yuan, Mathews Jacob |
Diffusion models are now widely used in Bayesian inverse problems in imaging as priors, where latent diffusion models are often used for larger scale problems to keep the computational complexity and model-size manageable. Unfortunately, the auto-encoder based...Diffusion models are now widely used in Bayesian inverse problems in imaging as priors, where latent diffusion models are often used for larger scale problems to keep the computational complexity and model-size manageable. Unfortunately, the auto-encoder based compression results in loss of spatial detail. In addition, the optimization is converted to a non-linear problem. In this paper, we introduce a posterior sampling algorithm customized for the pyramidal/cascaded architecture, which relies on a coarse to fine hierarchical strategy to generate images in the pixel domain. We present CLIMB-Flow which alternates between three steps: an end-point estimation from the current coarse and noisy image, data-consistent update of the clean image, and re-noising it back to the level the network expects. Together these steps sample the posterior at that scale using an approximate Gibbs sampling from two conditional distributions. Experiments on ImageNet, CelebA, AFHQ and fastMRI span inpainting, deblurring, super-resolution and accelerated MRI, with PSNR gains of 1.37-7.66 dB over the strongest competing method on CelebA and pixel-domain reconstruction up to 512x512.
|
| 1791 |
An Active-Bottleneck Mechanism for Weak-to-Strong Generalization
2609.33835
|
cs.LG
|
Mohammad Zeinalpour, Amir Najafi |
Weak-to-strong generalization (W2SG) occurs when a student trained on a teacher's predictions outperforms that teacher. We study when this happens under fully converged, ridgeless two-stage learning, with no early stopping, no explicit regularization, and no a...Weak-to-strong generalization (W2SG) occurs when a student trained on a teacher's predictions outperforms that teacher. We study when this happens under fully converged, ridgeless two-stage learning, with no early stopping, no explicit regularization, and no assumption that the student is more expressive than the teacher. In two-stage linear regression, a teacher is fit from $n$ labeled examples and a student is trained solely on the teacher's predictions on $m$ fresh, unlabeled inputs. Although both stages share the same hypothesis class and the same training rule, we show that the student outperforms the teacher exactly when $m$ lies in an explicit intermediate range: too few pseudo-labels leave the student without enough signal, too many let it inherit the teacher's noise. Under power-law covariance, we derive this range in closed form as a function of the spectral decay and noise level, including regimes where the improving region splits into two disjoint intervals of $m$. We then study a random-feature model in which the student has strictly more features than the teacher, and identify two regimes, again given by explicit thresholds: one where improvement occurs only for $m$ in a bounded interval, and one where it occurs only once the student width $N_S$ exceeds an explicit threshold. Both regimes are governed by a single "active-bottleneck" principle: whichever of $m$ or $N_S$ is scarcer controls how much teacher error is filtered out, while increasing the other resource only reduces estimation noise. Together, these results show that finite data and finite width can themselves regularize a two-stage learner, with no explicit mechanism doing so.
|
| 1792 |
ViBR-WM: Visual Bayesian Regression for World Modeling
2609.33844
|
cs.LG
|
Jifan Li, Ning Ning |
Modeling temporal dependence and uncertainty is central to forecasting with world models. The Visual Bayesian Regression World Model combines visual features, physical histories and known covariates through interpretable regression, within a modular architectu...Modeling temporal dependence and uncertainty is central to forecasting with world models. The Visual Bayesian Regression World Model combines visual features, physical histories and known covariates through interpretable regression, within a modular architecture supporting trend, seasonal and cycle dynamics. Visual compression reduces representation dimension, while Bayesian variable selection reduces active regression dimension. Posterior prediction combines forecasts across predictor subsets using their posterior probabilities as weights and accounts for parameter uncertainty and future disturbances. The model forecasts joint visual--physical states recursively and physical targets directly. Across four forecasting tasks spanning object motion, vegetation greenness and solar power, ViBR-WM achieves lower mean overall physical-target error than Temporal Straightening, ConvLSTM, PredRNN and SimVP on every task. Repeated fitting and resampling support these overall gains.
|
| 1793 |
Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation
2609.33872
|
cs.LG
|
Sichao Liu, Zekun Wang, Lixuan Tang, Yiming Li, Xiaohan Wang |
Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes ...Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions will achieve desired final states, causing execution errors to accumulate during long-horizon manipulation. We present Robot-GST, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment. Our approach constructs a high-fidelity robotic environment from RGB-D observations using 3D Gaussian Splatting and SAM3D, enabling ``simulation and evaluation before acting''. It integrates visual observations and language instructions with spatio-temporal reasoning for long-horizon task planning using large vision-language models. To bridge high-level planning and real-world execution, we introduce Gaussian-aware final-state estimation through geometric sampling and state-based trajectory planning. Before execution, candidate action sequences are simulated and evaluated in the Gaussian-SAM environment to filter infeasible behaviours. We validate our approach on representative manipulation tasks involving rigid, soft, and deformable objects, including cube placing, toy packing, and duck rearrangement, demonstrating that geometry-aware spatio-temporal reasoning and state-aware execution improve manipulation reliability across different object categories. Our results suggest that combining geometry-aware reconstruction with high-quality rendering and simulation provides a scalable approach for evaluating robotic manipulation behaviours. Website: https://robot-gst.github.io
|
| 1794 |
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
2609.33875
|
cs.LGcs.AI
|
Yuanhao Li, Hongbo Wang, Xuhong Chen, Yiming Cao, Xunzhu Tang |
Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable execut...Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.
|
| 1795 |
Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges
2609.33878
|
cs.LG
|
Donghao Huang, Jinling Pei, Zhaoxia Wang |
Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels con...Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels contaminate pseudo-labeled data, while conservative labeling reduces coverage. We investigate whether agreement between two local large language model judges improves pseudo-label reliability. A label is retained only when the judges agree, with separate ordered thresholds for selections and abstentions that guarantee disjoint positive and negative label sets. Retrospective replay on 2,000 expert-annotated queries shows that higher selection thresholds can improve positive-label purity, whereas higher abstention thresholds increase false no-match labels. At thresholds (0.86, 0.80), Muse Glimmer 30B and Gemma 4 31B jointly label 1,633 queries (81.7% coverage) at 96.88% purity; positive and negative purities are 99.47% and 93.38%. This exceeds either constituent model at the same thresholds by more than two percentage points, with lower coverage. A split-half check finds only 0.14 percentage points of threshold-selection optimism. A symmetric threshold of 0.86 adds 40 erroneous no-match labels, while 46 false abstentions persist even with no confidence threshold. Across five matched within-model comparisons, higher reasoning effort yields no clear F0.5 gain and increases median latency by 1.8-5.0 times. These results motivate separate thresholding and auditing for positive and negative pseudo-labels. The study establishes label purity, not student utility; fresh-data curation and student fine-tuning remain necessary to demonstrate downstream value.
|
| 1796 |
DCEmbed: Scalable Optimization over Neural Surrogates
2609.33879
|
cs.LG
|
Akshay Sreekumar, Nicolas Christianson, Priya L. Donti, Ellen Vitercik, Ram Rajagopal |
Neural surrogates can accelerate large-scale optimization by replacing expensive or intractable model components with efficient learned approximations, but solving the resulting embedded problems can remain prohibitively costly. For instance, standard exact en...Neural surrogates can accelerate large-scale optimization by replacing expensive or intractable model components with efficient learned approximations, but solving the resulting embedded problems can remain prohibitively costly. For instance, standard exact encodings of neural networks with ReLU activations allow the problem to be solved by mixed-integer solvers, but add large numbers of binary variables to accommodate the nonlinearity of the activations, which can render the problem computationally prohibitive. To address this, we propose DCEmbed, a heuristic for optimization problems with embedded neural surrogates that leverages the difference-of-convex (DC) representation of the network and avoids adding activation binaries. Exploiting shared structure within the DC representation of a ReLU network, we derive a reduced-size, exact formulation for its convex components that can be embedded in optimization problems using just two linear inequalities and one continuous auxiliary variable per hidden neuron. Using this, the problem is solved via an iterative penalty convex-concave procedure, where only the concave portions of the neural terms are approximated at each stage. The original objective, constraints, and any discrete decisions are retained, allowing standard convex or mixed-integer optimization solvers to optimize the host and surrogate jointly at each iteration. In experiments on quadratic programs, mixed-integer resource allocation, and neural two-stage stochastic programming, our method demonstrates much faster progress toward high-quality feasible solutions than approaches using exact mixed-integer embeddings. In particular, DCEmbed achieves $4\times$ lower normalized primal integral than the best exact baseline on the resource allocation problem, while in two-stage stochastic programming it reaches the global surrogate optimum $\sim 5\times$ faster than Gurobi ML.
|
| 1797 |
Neuron-Level Architecture Growth: A Controlled Evaluation for EEG Time-Series Decoding
2609.33880
|
cs.LG
|
Adam Mounir, Stella Douka, Arnault H. Caillet, Bruno Aristimunha, Sylvain Chevallier |
Convolutional EEG decoders are trained at a fixed width, usually set by their authors on other data. Growing methods add neurons during training where the loss could decrease the most, but whether they improve compared to a reference width is untested on EEG. ...Convolutional EEG decoders are trained at a fixed width, usually set by their authors on other data. Growing methods add neurons during training where the loss could decrease the most, but whether they improve compared to a reference width is untested on EEG. Here, we grow three convolutional backbones on 12 motor-imagery datasets under three protocols and compare each with its reference model per subject. The growing ShallowFBCSPNet scores 2.9 points above its reference model with only half the parameters (0.57x), SCCNet changes by at most 1.2 points. Deep4Net growing models show decreased accuracy, but they require adaptation that prevent to compare faithfully the results. These differences follow the selection step, which keeps a candidate neuron relying on a dynamic threshold from singular values decomposition. Overall, these results suggest that growth helps when its criterion can rank the candidate neurons, and that the rate of skipped neuron addition tells where a decoder can be grown small from scratch.
|
| 1798 |
Prospective Interpretation Risk: Principled Communication Control Between LLMs
2609.33885
|
cs.LG
|
Wanrong Yang, Rehan Deen, Julian Ma, Yuheng Fan, Yaoyu Jin |
Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in hetero...Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in heterogeneous systems, where capable receivers can reconstruct different tasks from the same message. We model this as a sender-receiver problem with a latent receiver type and define prospective interpretation risk (PIR): the probability that a receiver reconstructs a task other than intended. Rather than model an LLM's full input-output behaviour, we use black-box probes relating messages, intended tasks, and receiver-specific reconstructions, yielding scalable supervision while separating interpretation from downstream capability failure. Offline, heterogeneous frozen receivers provide supervision for receiver-conditioned risk and the effects of predefined mutable message features. At deployment, history induces a posterior over receiver types, guiding message revision and selection. We introduce value of interpretation information (VoII), querying for receiver information only when its expected communication benefit exceeds its cost. Our theory characterises when receiver information has decision value and bounds such queries. Empirically, interpretation-failure rates vary by 4-13x across receivers. Receiver information reduces PIR calibration error by 68% relative to a receiver-agnostic predictor, largely by correcting receiver-specific risk levels. PIR-guided revision reduces interpretation failure by 44% relative to the original message and 40% relative to a generic rewrite, mostly through a repair that helps every receiver. VoII outperforms information-gain and random querying at matched cost on the interpretation objective it optimises, lowering interpretation failure from 3.84% to 3.79% while querying 18.2% of episodes.
|
| 1799 |
Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers
2609.33895
|
cs.LG
|
Tongtong Liang, Siqi Kou, Ziqiao Xi, Esha Singh, Kun Zhou |
In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers op...In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement *residual-stream burden* and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet $256^2$. Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.
|
| 1800 |
EEG-Fusion: Failure-Informed Source-Free Expert Routing for Robust Motor Imagery EEG Decoding
2609.33962
|
cs.LG
|
Abdul Basit, Saim Rehman, Muhammad Shafique |
Subject-independent motor-imagery (MI) EEG decoding can exhibit subject-level failures even when average performance appears acceptable: under subject shift, a decoder can become an overconfident near-one-class predictor. This is especially problematic in sour...Subject-independent motor-imagery (MI) EEG decoding can exhibit subject-level failures even when average performance appears acceptable: under subject shift, a decoder can become an overconfident near-one-class predictor. This is especially problematic in source-free deployment, where target-user labels are unavailable during adaptation and expert selection. We present \textit{EEG-Fusion}, a failure-informed decision-level fusion framework that treats source-free MI decoding as label-free reliability estimation over heterogeneous experts. EEG-Fusion applies subject-wise Euclidean alignment and normalization-only test-time adaptation, then routes each target subject to a neural, covariance-based, or physiological-feature expert using a reliability gate trained on source-held-out folds to predict expert performance and collapse risk from label-free stream diagnostics. The gate uses confidence, entropy, prediction diversity, expert agreement, and predicted class balance; collapse is measured as the maximum predicted class fraction. In 9-fold leave-one-subject-out (LOSO) evaluation with three seeds, relative to a no-alignment raw EEGNet source-free anchor, EEG-Fusion improves subject macro-F1 from 0.417 to 0.529 on BCI IV-2a local protocol, from 0.314 to 0.482 on BNCI2014-001, and from 0.607 to 0.708 on BNCI2014-004; corresponding collapse-index reductions are 0.199, 0.227, and 0.169. In a 9-subject Cho2017 external subset, EEG-Fusion improves macro-F1 from 0.516 to 0.630. These results suggest that label-free reliability estimation can reduce subject-level failure modes in source-free MI-EEG deployment.
|
| 1801 |
Greenpixie's AI Token Methodology: Assessing the Energy, Water and CO2-eq Impact of AI Tokens for Open and Closed Weight Models
2609.33965
|
cs.LG
|
Joshua Horswill, Ross Hunter, Matt Clifford, James Hall |
We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) tokens. Graphics processing unit (GPU) energy usage is measured during inference ben...We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) tokens. Graphics processing unit (GPU) energy usage is measured during inference benchmarking with open-weights models on a wide range of text-based tasks. The remaining server energy contribution from non-GPU hardware is estimated from the inference wall time. Bayesian linear regression is used to model the relationship between energy per token and LLM size, request traffic, and hardware deployment configuration. Proprietary frontier LLMs of unknown size and deployment are binned into size buckets based on naming conventions and performance priors, and the space of possible LLM configurations is sampled with Monte-Carlo methods to give a representative average energy per token and uncertainty. We also describe how these energy measurements can be used to estimate the carbon-dioxide equivalent ($\mathrm{CO_2\text{-}eq}$) emissions, both usage and embodied, and water consumed per token of AI inference. This methodology provides actionable data that enables reductions in cost, electricity usage, $\mathrm{CO_2\text{-}eq}$ emitted and water consumed in cloud and Software as a Service (SaaS).
|
| 1802 |
ThinkNet: Compact Architecture Selection and Validation-Gated Ensembles for Subject-Independent MI-EEG Decoding
2609.33967
|
cs.LG
|
Abdul Basit, Saim Rehman, Muhammad Shafique |
Practical assistive and rehabilitative brain--computer interfaces require subject-independent motor-imagery EEG (MI-EEG) decoders that generalize to new users under limited target-user data and constrained compute. However, held-out-subject performance can be ...Practical assistive and rehabilitative brain--computer interfaces require subject-independent motor-imagery EEG (MI-EEG) decoders that generalize to new users under limited target-user data and constrained compute. However, held-out-subject performance can be overstated when test-subject information influences preprocessing, model selection, or ensemble selection. We present \textit{ThinkNet}, a validation-controlled framework that combines train-only normalization, validation-guided evolutionary search, and validation-gated inference to identify compact decoders and inference policies for held-out subjects. We evaluate four-class BCI Competition IV-2a (session T) decoding with nine Leave-One-Subject-Out (LOSO) folds, three seeds, seven fixed decoder entries, and a broader search over ten representative decoder families; the held-out subject is never used for normalization, hyperparameter, architecture, or ensemble-policy selection. In the fixed benchmark, the validation-selected compact decoder achieved 44.35$\pm$15.41\% accuracy with 4.9K parameters, 19 KB FP32 weights, and 0.99 ms batch-1 Orin CUDA inference. Across the broader search, compact models ($\leq$25K parameters) achieved higher mean held-out accuracy than mid-size and large alternatives after selected retraining (40.10\% vs. 35.09\% and 34.78\%). Validation-gated ensembling improved over validation-selected single-model inference, reaching 43.98$\pm$16.25\% in the fixed benchmark and 43.31$\pm$15.88\% for the compact six-family ensemble. A non-deployable oracle analysis revealed a 6.1-point family-selection gap and near-zero validation--test correlation, showing that validation reliability remains a key bottleneck under subject shift. Thus, ThinkNet is a validation-controlled framework for compact MI-EEG model and inference-policy selection, rather than a single-architecture benchmark.
|
| 1803 |
Two-Sample Testing for Inhomogeneous Random Graphs in Non-Integral $L_r$ Norms
2609.33968
|
cs.LG
|
Soham Dan |
Testing whether two populations of networks share the same edge probabilities is a basic problem in network inference. How hard it is depends on the norm used to measure the difference. For the inhomogeneous Erd\H{o}s--R\'enyi (IER) model, the optimal sample c...Testing whether two populations of networks share the same edge probabilities is a basic problem in network inference. How hard it is depends on the norm used to measure the difference. For the inhomogeneous Erd\H{o}s--R\'enyi (IER) model, the optimal sample complexity is known for every integer $L_r$ norm and for $1\le r<2$. For non-integral $r>2$, however, the known upper and lower bounds do not match, and the lower bound was conjectured to be tight. We study this gap for two-sample testing on aligned vertices. We propose a test that runs two published statistics, of orders $2$ and $\lceil r\rceil$, on the same data and rejects if either one rejects. Its thresholds come from H\"older interpolation, so that both statistics have the same sample cost. We prove that this test attains the conjectured rate. Combined with earlier results, this shows that for every fixed $r\ge1$ the minimax sample complexity is of order $n^{\max\{4/r-1,\,2/r\}}/\epsilon^2$, even when the separation changes with $n$. In simulations with $n$ between 32 and 256, the number of graphs needed for 80\% power at level $0.05$ grows with $n$ at a rate consistent with the theory. For $r=2.5$, for example, the fitted exponent is $0.78$, against the theoretical value $0.8$. Interestingly, the two statistics split the work as the interpolation argument suggests: the higher-order statistic is more powerful when only a few edges change, and the $L_2$ statistic when many edges change.
|
| 1804 |
HARMONIA: Interpretable Graph Learning through Mixtures of Neural Bases
2609.33972
|
cs.LG
|
Quan D. Bui, Nguyen Do, An Nguyen Dang, Huyen Nguyen, Nhu Duc Minh Nguyen |
Existing interpretable graph additive models still face limitations in either computational scalability or modeling flexibility. In terms of structural modeling, previous approaches either face quadratic scaling costs or sacrifice explicit source-to-target con...Existing interpretable graph additive models still face limitations in either computational scalability or modeling flexibility. In terms of structural modeling, previous approaches either face quadratic scaling costs or sacrifice explicit source-to-target contribution decomposition. In terms of feature components, they rely either on per-feature neural networks or on single shared bases with limited feature specialization. We address both problems by introducing HARMONIA: Interpretable Graph Learning through Mixtures of Neural Bases, an interpretable-by-design framework. For feature modeling, HARMONIA introduces a Mixture of Neural Bases (MoNB), which routes features to specialized basis experts, enabling parameter sharing without sacrificing feature-specific specialization. For structural modeling, HARMONIA uses Relative Random Walk Probabilities (RRWP) to capture multi-hop and multi-path relationships, and proposes Sparse RRWP Aggregation (SRA) to compute these interactions through sparse graph propagation without quadratic pairwise complexity. HARMONIA retains a simple additive form in which predictions decompose into feature responses modulated by structural influence. Empirically, HARMONIA achieves stronger explanation recovery than existing interpretable graph baselines while maintaining competitive predictive performance and scaling to graphs with millions of nodes. These results show that interpretable graph learning can remain both faithful and scalable without sacrificing predictive effectiveness.
|
| 1805 |
Adapting neural operators for mechanics decisions under changing operating conditions
2609.33978
|
cs.LG
|
Prashant K. Jha, Koffi Enakoutsa, Ian Galloway, Henry Anderson |
Neural operators can accelerate repeated nonlinear mechanics calculations, but their accuracy can deteriorate as operating conditions move beyond the training range. This work studies whether high-fidelity solutions acquired during use can be reused to adapt a...Neural operators can accelerate repeated nonlinear mechanics calculations, but their accuracy can deteriorate as operating conditions move beyond the training range. This work studies whether high-fidelity solutions acquired during use can be reused to adapt a neural operator and improve subsequent mechanics-based command selection. Two hard-magnetic soft-material systems are simulated using high-fidelity finite-element (FE) models, providing reference solutions for evaluating surrogate predictions and selected commands. A neural operator predicts deformation from known material, loading, and magnetic-field inputs, while an empirical error estimator determines which predictions may be used for command selection. Selected FE evaluations supplement these predictions, and their complete loading paths are retained for periodic updates of the neural operator and estimator. In both examples, the fixed operator loses substantial accuracy when stiffness and loading move outside the training range. Updates using 16 acquired paths recover much of the lost accuracy while preserving accuracy in the nominal regime. Under the same FE evaluation budget, the updated operators also improve command selection, although the benefit varies with the operating condition. Error estimation is less consistent, with inaccurate predictions sometimes accepted and accurate predictions rejected. These results demonstrate that reusing high-fidelity loading paths can extend the useful operating range of a neural operator. However, improved forward accuracy alone does not guarantee reliable prediction acceptance, highlighting prediction-specific error assessment as a separate requirement for trustworthy decision making.
|
| 1806 |
The Privacy Fallacy of Crowdsourced Fine-Tuning: Extracting Proprietary Data via Topic-Based Poisoning
2609.33985
|
cs.LG
|
Sae Furukawa, Alina Oprea |
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks. Crowdsourcing user conversations is an established approach to collecting SFT data at scale while reducing the need for costly manual annotation. However, it also a...Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks. Crowdsourcing user conversations is an established approach to collecting SFT data at scale while reducing the need for costly manual annotation. However, it also allows untrusted users to contribute data to the fine-tuning pipeline. We investigate an underexplored privacy risk arising from this setting: can a malicious user poison a small fraction of the crowdsourced data to amplify extraction of previously unseen instructions contributed by other users? We show that this is possible using only black-box, output-only access to the deployed model. Experiments across four models and two datasets demonstrate substantial increases in training-data extraction: with only 50 poisoned examples, near-verbatim extraction reaches $3.71\times$ the rate without poisoning for Qwen2.5-14B on OpenMathInstruct and $3.08\times$ for Llama-3.1-8B on AceReason. Data filtering also proves largely ineffective in detecting poisoned samples: even the best-performing method achieves only 0.378 in F-1 score, leaving the majority of poisoned samples undetected. These findings demonstrate that seemingly benign crowdsourced contributions can amplify leakage of other records while remaining difficult to identify through data filtering.
|
| 1807 |
Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control
2609.34018
|
cs.LG
|
Denis Shcherba, Adrian Abel, Eckart Cobo-Briesewitz, Paul Mattes, Wojciech Samek |
Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillatio...Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is trained to reproduce the actions of the privileged expert. This requires the student to jointly infer the task-relevant state and relearn the expert's action mapping that is already available. An alternative is to reuse the state-based expert and learn only a perceptual interface that reconstructs its missing state inputs. However, minimising the state estimate error alone does not necessarily minimise the downstream control error induced by these estimates. To bridge this gap, we train a visual state estimator using both direct state supervision and an action-consistency loss backpropagated through the frozen, differentiable expert. A scheduled objective first establishes a physically meaningful state estimate and progressively emphasises errors that affect the expert's actions. Across five goal-conditioned manipulation tasks, retaining the expert consistently outperforms direct pixel-to-action imitation from the same expert demonstration corpus. We further demonstrate sim-to-real transfer on a physical Panda robot, achieving 76% success without retraining the underlying expert.
|
| 1808 |
Jev in Medicine: A Benchmark Evaluation. Preliminary Results
2609.34024
|
cs.LG
|
Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho |
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Je...Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
|
| 1809 |
ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence Modelling
2609.34034
|
cs.LG
|
Matei-Ioan Stan, Oliver Rhodes |
A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities that have secured the Transformer's status as the de facto stan...A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities that have secured the Transformer's status as the de facto standard in sequence modelling. Any realistic contender must be data-adaptive, able to capture long-range dependencies, and GPU-parallelisable, but also non-linearly recurrent to enable complex reasoning. Based on evidence suggesting the auditory cortex operates on fixed timescales, this work proposes the ADaptive with Prescriptive Timescales Network (ADPTNet) as a potential solution to achieving all four properties simultaneously. ADPTNet is built around local topological conjugates, obtained by a novel combination of linear attention and Riemannian optimisation, applied to static global dynamics. This enables non-linear yet predictable long-term behaviour. Dynamical systems theory proofs provide theoretical guarantees for the parametric control of ADPTNet's timescales (its Lyapunov spectrum). ADPTNet improves performance on Selective Copying over Hawk, the existing method balancing long-range memory and adaptability, while also improving state tracking over linear SSMs like Mamba. On sequential CIFAR-10, ADPTNet matches linear SSM accuracy and outperforms existing selective models (incl. the Transformer), using fewer parameters. We also introduce a neuromorphic SpikingADPTNet, which achieves a new state-of-the-art accuracy on the Spiking Speech Commands dataset ($83.56\%\pm0.15$). Finally, ADPTNet's constant timescales enable two efficient, Jacobian-free extensions to the DEER parallel simulation algorithm (Conv and Forward DEER) that retain the same average convergence. Conv DEER adds no computational overhead beyond the network's forward pass and enables non-linear RNN parallelisation via iterated convolutions for the first time.
|
| 1810 |
3D Point Tracking with State Space Models
2609.34035
|
cs.LG
|
Masahiro Ogawa, Qi An, Atsushi Yamashita |
Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point trac...Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.
|
| 1811 |
Singularities of Non-negative Matrix Factorization and their application to Bayesian inference
2609.34043
|
cs.LG
|
Naoki Hayashi, Yota Maeda, Yasushi Esaki |
Non-negative matrix factorization (NMF) is a singular statistical model whose Bayesian asymptotics are governed by the real log canonical threshold (RLCT). We study the local geometry of the factorization map and derive an upper bound for the RLCT of NMF. Let ...Non-negative matrix factorization (NMF) is a singular statistical model whose Bayesian asymptotics are governed by the real log canonical threshold (RLCT). We study the local geometry of the factorization map and derive an upper bound for the RLCT of NMF. Let $H$ be the model inner dimension and $H_0$ the non-negative rank of the true $M\times N$ matrix. Assuming that the true matrix admits a strictly positive factorization of inner dimension $H_0$ in the interior of the parameter domain, we prove, for smooth positive priors, that $\lambda\leq \{(H-H_0)\min(M,N)+H_0(M+N-H_0)\}/2$. This bound strictly improves the previous bound when $H_0\geq3$. The proof uses a local analytic normal form that separates independent linear coordinates from a residual matrix product. When $H=H_0$ also equals the ordinary rank of the true matrix, we obtain the exact value $\lambda=H_0(M+N-H_0)/2$. Under the standard assumptions of singular learning theory, these results bound the leading coefficients of the expected Bayesian generalization error and the Bayesian free energy.
|
| 1812 |
Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines
2609.34045
|
cs.LG
|
Murtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya |
Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only t...Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding membership removes what they depend on: a swarm holds each part of the model on several peers and routes around a slow one. A bounded session must use every device it admits. Its pipeline advances at the pace of whichever device received a share it cannot serve quickly, so the division has to be right before serving begins. We propose Kafila, whose protocol assembles a ring from behind NATs, preferring direct paths and relaying where traversal fails, while its planner measures each device's memory bandwidth, capacity and reachability, divides the model exactly for a fixed ring order, and places the head, which holds the embedding and output projection, together with that division rather than beforehand. On machines with different capabilities across three fleets, from a shared LAN to five devices spanning two continents, Kafila shortens the slowest pipeline stage by up to $5.2\times$ against the even split of pipeline parallelism, as in GPipe, and up to $3\times$ against the memory-proportional split of personal-device inference, as in exo, keeps 75 to 87 per cent of the committed hardware doing work where those divisions fall below half, and serves a model no uniform split can place on the fleet at all. What that is worth to a user depends on how much of a token is computation rather than network. Where the members share a network the same division returns $1.56\times$ the throughput of a uniform split and $1.25\times$ of a memory-proportional one, and under four concurrent users that lead compounds to $3.2\times$ rather than fading, each user served at almost the rate of one.
|
| 1813 |
The Statistical Cost of Causal Discovery with Feedback
2609.34050
|
cs.LG
|
Sunmin Oh, Seungsu Han, Gunwoong Park |
What determines the unavoidable sample cost of learning cyclic causal structure? For cyclic linear non-Gaussian models, we study exact condensation recovery from observational data: identifying the strongly connected component (SCC) partition and all edges bet...What determines the unavoidable sample cost of learning cyclic causal structure? For cyclic linear non-Gaussian models, we study exact condensation recovery from observational data: identifying the strongly connected component (SCC) partition and all edges between components. We establish the first information-theoretic lower bounds on sample complexity for this target. For $p$ variables, maximum SCC size $s_{\max}$, and maximum external-parent count $d_B$, any estimator requires order $s_{\max}\log(ep/s_{\max})+d_B\log(ep/d_B)$ samples in the worst case over a regular model class. These bounds distinguish the costs of SCC membership and external-parent selection. Under principal invertibility and without correlation faithfulness, we establish a population block-exogeneity principle that identifies unknown root SCCs through residual independence and inclusion minimality. A sparse-adjustment characterization shows that small adjustment sets suffice to identify SCCs and their direct external parents, without regressing on all previously recovered variables. These characterizations yield BlockExo, which attains a structurally matching sample bound without knowing $s_{\max}$ or $d_B$ under suitable conditions. Simulations support the structural dependence of our sample bound and demonstrate BlockExo's sample-efficient recovery in comparisons with other methods for cyclic causal discovery.
|
| 1814 |
The Double-Edged Sword of Information: Revealed versus Hidden Lotteries in School Choice
2609.34074
|
cs.LG
|
Parinaz Naghizadeh, Jingyan Wang |
In school choice, a lottery number is often used by the matching mechanism to break ties when there are more students who prefer the same school than the number of seats available. There has been growing theoretical and empirical interest in understanding the ...In school choice, a lottery number is often used by the matching mechanism to break ties when there are more students who prefer the same school than the number of seats available. There has been growing theoretical and empirical interest in understanding the impact of revealing the lottery number to students. In practice, in recent years, the NYC Public Schools started revealing the lottery number to students to improve transparency. Theoretical findings from prior literature also suggest that revealing the lottery number strictly improves the number of matches under the deferred acceptance algorithm. However, these theoretical results are based on the over-simplifying assumption that all students share the same preference ranking for schools. In this work, we relax this assumption and allow students to have heterogeneous preference rankings. Under a game-theoretic model where student strategies form a Bayesian Nash equilibrium, we characterize scenarios where revealing the lottery number can either improve or worsen the matching outcome, measured by two metrics of match rate and social welfare. We further consider revealing partial information about the lottery, and demonstrate non-monotonic effects in the amount of information available to students. These results together illustrate complex tradeoffs induced by the lottery revealing policy.
|
| 1815 |
AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
2609.34085
|
cs.LG
|
Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska |
Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world mode...Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by $16\times$ and the embedding dimension by $4\times$, achieving a $100\times$ inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS$^{\dagger}$ without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA
|
| 1816 |
Forecast-Necessary Causal Discovery for Nonlinear Political Panel Data: Feedback, Functional Form, and the Dynamics of Democratization
2609.34115
|
cs.LG
|
Michael Coppedge, Dmitry Zaytsev, Valentina Kuskova |
A non-significant coefficient in a dynamic panel model need not imply the absence of a relationship. It may instead reflect heterogeneous effects averaged toward zero, reciprocal dynamics overlooked by a recursive specification, or relationships masked by the ...A non-significant coefficient in a dynamic panel model need not imply the absence of a relationship. It may instead reflect heterogeneous effects averaged toward zero, reciprocal dynamics overlooked by a recursive specification, or relationships masked by the omission of correlated covariates. Standard linear estimators cannot distinguish among these possibilities. We develop an inferential workflow for political panel data that resolves this ambiguity by combining flexible autoregressive estimation, forecast-necessity testing, functional characterization, and same-data linear benchmarking. The workflow first identifies relationships required for out-of-sample prediction, then characterizes their functional form across political contexts, and finally, distinguishes differences arising from estimator flexibility from those due to model specification. Applied to the causal sequence model of democratization on the V-Dem panel of 113 countries, the workflow reproduces the model's central finding - the protective belt of civil society, the rule of law, and institutionalized parties - while recovering reciprocal relationships from democracy to its institutional supports that a linear model cannot detect. Most importantly, three weak published direct effects, of which two are null, and one is marginally significant, receive three different diagnoses: one dissolves under the full specification, one reflects heterogeneous effects averaged toward zero, and one was masked by the reduced variable set. The workflow corrects the published record in both directions, removing one relationship and recovering two. More broadly, the workflow provides a framework for evaluating dynamic political theories under a model class capable of representing nonlinear and reciprocal mechanisms while preserving relationship-level interpretation and explicit inferential standards.
|
| 1817 |
GT-PSSM: Unified Probabilistic Framework for Stochastic Dynamics Modeling and Dependency Learning in Multivariate Time Series Anomaly Detection
2609.34161
|
cs.LG
|
Wonmo Koo, Jaeyeong Lee, Taeseong Yoon, Heeyoung Kim |
Multivariate time series anomaly detection (MTAD) is crucial for ensuring the safe and reliable operation of complex systems. Many existing methods learn normal patterns by training reconstruction or forecasting models on predominantly normal data. However, a ...Multivariate time series anomaly detection (MTAD) is crucial for ensuring the safe and reliable operation of complex systems. Many existing methods learn normal patterns by training reconstruction or forecasting models on predominantly normal data. However, a large portion of these approaches rely on deterministic models and their associated point-wise output errors for anomaly scoring. Since real-world multivariate time series are inherently stochastic due to measurement noise and intrinsic system randomness, purely error-based scores can be unreliable, as large errors may arise from benign fluctuations rather than true anomalies. Probabilistic approaches address this limitation by quantifying uncertainty in model outputs. In particular, probabilistic state-space models (PSSMs) provide a principled framework by modeling stochastic system dynamics through latent state transitions and measurement noise via emission models. Despite this advantage, existing PSSM-based MTAD methods often struggle to capture long-range temporal dependencies and inter-variable dependencies, as they typically rely on noise-sensitive recurrent architectures and lack explicit cross-variable structure modeling. To address these limitations, we propose Graph-Transformer-Enhanced Probabilistic State-Space Model (GT-PSSM), a novel PSSM-based MTAD method that tightly integrates PSSM-based probabilistic modeling of stochastic dynamics with Graph Transformer-based learning of temporal and inter-variable dependencies. By jointly modeling stochasticity, long-range temporal dependence, and variable interactions within a unified probabilistic framework, GT-PSSM enables more robust anomaly detection.
|
| 1818 |
ReplayLens: Auditing Agents' Use of Outcomes
2609.34177
|
cs.LG
|
Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu |
When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black...When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship in the stored history at a time, holds the remaining interface fixed, and measures the resulting decision. Four interventions target four relationships. Outcome reassignment swaps which scores belong to which actions. Pair transport moves intact action-score pairs to new record slots. Consistent renaming relabels actions in both history and menu. Key-slot reassignment changes both score attachment and position. A constructive separation shows why the audit is needed: two memory writers with identical endpoint accuracy respond differently to the same replay, so conventional evaluation cannot resolve the underlying dependence. On black-box LLM interfaces, swapping scores changes decisions while moving intact pairs does not, separating score attachment from record order. A bounded-memory study exposes ingestion-order sensitivity that endpoint comparison misses. In sequential experiment planning, altered historical scores redirect exploration and reduce final utility despite fresh measurements. A code-debugging agent with sealed hidden tests shows the same pattern outside model selection. ReplayLens provides a relationship-level audit for deciding whether logged experience can be merged, reordered, or reindexed safely.
|
| 1819 |
AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning
2609.34188
|
cs.LG
|
Yingbo Zhao, Zeyu Yang, Zhoufan Zhu |
Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues ...Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.
|
| 1820 |
Functional Autoencoders for Amplitude-Phase Representation Learning
2609.34207
|
cs.LG
|
Peida Wu, Xinyang Xiong, Pengcheng Zeng |
Functional data are intrinsically infinite-dimensional, and often exhibit phase variation, where corresponding events occur at different times across observations. Existing linear dimension reduction methods struggle with nonlinear amplitude variation, while f...Functional data are intrinsically infinite-dimensional, and often exhibit phase variation, where corresponding events occur at different times across observations. Existing linear dimension reduction methods struggle with nonlinear amplitude variation, while functional autoencoders without an explicit warp entangle temporal misalignment with shape. We propose the Amplitude--Phase Functional Autoencoders (AP-FAE), an unsupervised framework for functional data that spans both univariate and multivariate cases, with emphasis on the multivariate setting, and factorizes the latent space into separate amplitude and phase embeddings derived from all channels. A smooth functional decoder reconstructs channel-specific amplitude functions in canonical time, and a shared monotone, endpoint-preserving warp captures phase variation. We prove a bound linking amplitude recovery to registration, reconstruction, and noise errors, and validate it numerically. Across synthetic data and six real-world benchmarks, AP-FAE outperforms state-of-the-art baselines on most clustering and alignment metrics and on all reconstruction metrics. Clustering with amplitude embeddings alone consistently surpasses joint amplitude--phase clustering, confirming the benefit of explicit disentanglement. Code is available at https://anonymous.4open.science/r/APFAE-418C/}{https://anonymous.4open.science/r/APFAE-418C/.
|
| 1821 |
Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures
2609.34215
|
cs.LG
|
Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu |
Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agre...Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purely ordinal measure that records which action wins without reflecting the absolute level of performance. When all actions fail, they tie at zero reward, and independent runs produce the same tied set with high probability, creating an illusion of stability that masks near-zero recovery success. We formalize this limitation through a set-path symmetry result, proving that for equal-cost Bernoulli actions the success probabilities (0.9, 0.8) and (0.2, 0.1) yield identical best-action-set distributions at every sample size. No procedure based solely on which action wins can distinguish these two regimes. We further prove that certifying exact population ties is impossible in finite time, and that the assignment of outcomes to checkpoints carries information beyond marginal outcome distributions. The pooled success probability is the missing scalar that resolves the ordinal ambiguity. Experiments on 864 frozen RecoveryBench episodes and two planning cohorts totaling 3,456 responses confirm the theoretical predictions. Agreement and held-out quality can move in opposite directions, and permuting checkpoint-to-action bindings changes 8 to 13 percent of cell-level conclusions. Based on these findings, we propose reporting four diagnostic quantities (agreement, all-zero fraction, held-out success, and pooled success) that expose this failure mode with no additional data collection.
|
| 1822 |
Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
2609.34242
|
cs.LG
|
Chidera Biringa, Lucas Yannul, Xiaowen Wang, Marco Ayala, Nicholas Yi |
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent...AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.
|
| 1823 |
Certified Multi-Source Integrity for Structured Agent Actions
2609.34245
|
cs.LG
|
Anmol Pandey, Aditya Jain, Liang Chen, Carsten Maple, Christo Panchev |
LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can driv...LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself to extract attacker-chosen values. Current defenses gate on a source's trust label or certify free-text answer quality. None certifies the integrity of a coupled, policy-bound structured action under a corruption budget that accounts for shared upstream sources. We characterize when such an action is safely certifiable and give the maximally live safe certifier. It admits an action only when each field clears the rule its evidence structure supports: a bounded corruption radius over corruption-distinct evidence classes, counted by a minimum hitting set so that re-publishing or laundered copies cannot manufacture a quorum, deterministic reconciliation for complementary fields, and a trusted anchor where the evidence leaves a field single-sourced. We formalize two robustness notions, validate each mechanism by ablation, and measure how often the multi-source precondition holds on sanctions designations (70,966 entities) and software supply-chain provenance (450 packages). Under upper-bound proxies, genuine corroboration is a minority phenomenon in both, and naive attestation counting overstates it, since witnesses that look independent collapse to two corruption-distinct domains once shared origin is counted. Across five current models in a real agent loop, a realistic injection fools every model but one and a naive agent then executes the fraudulent action on most attacks. The certifier admits no unsafe action and recovers the correct value where corroboration permits, while action-gating and provenance baselines are broken in every world of our harness by some attack in its space.
|
| 1824 |
RoboICL: Embodied In-Context Learning with GPT-6 Astra
2609.34261
|
cs.LG
|
Fangcheng Liu, Yeqing Shen, Anda Cheng, Weishi Mi, Chao Tang |
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \e...General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that narrows these gaps without robot-specific parameter updates or a learned VLA. RoboICL separates \emph{demonstration context}, which provides recorded examples when available, from \emph{interaction memory}, which accumulates the model's own actions and observed outcomes. Both use a shared observation--action--receipt--observation grammar. To preserve experience across task stages, RoboICL combines sampled demonstration blocks with bounded anchored memory. Fixed anchors keep earlier rollout interactions available for in-context learning, while the latest interaction supports immediate error correction. Across 30 RoboDojo tasks, using zero shot for Open and one demonstration elsewhere, RoboICL improves on official zero-shot \gptastra{} by 20--27 progress-score points in every category. It leads the leaderboard baselines on Memory and Open, achieves comparable performance to the strongest Precision baseline, and remains competitive on Long-Horizon. Its 30-task Overall score is 50.64, versus 33.68 for the strongest baseline. On a separate ten-task subset, RoboICL scores 60.60, within 2.00 points of the $\pi_{0.5}$ + \gptastra{} hybrid approach. On three real-robot tasks, mean progress rises from 14.45 at zero shot to 63.33 at one shot and 78.89 at three shots. On two development tasks, optional Jev-gated action reuse reduces \gptastra{} calls by 33--48\%. Code is available at \href{https://github.com/Mosi-AI/RoboICL}{https://github.com/Mosi-AI/RoboICL}.
|
| 1825 |
Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes
2609.34262
|
cs.LG
|
Weijun Luo, Kelvin Luu, Xinyi Liu, Guangze Luo, Miguel Romero Calvo |
Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchma...Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchmark surfaces that once seemed harmless can become exploitable, making benchmark validity an ongoing maintenance problem. We introduce a process-verification framework that audits passing trajectories, distinguishes evidenced reward hacking from verifier weakness, and localizes exploitable surfaces for repair. Across 3,810 passing trajectories from 29 model-benchmark cohorts, confirmed violations often increase with model generation but not monotonically. On SWEBench Pro V1.0, confirmed violation rates rise from 24% to 73% between Opus 4.7 and Fable 5 on matched tasks; later cohorts fall to 11% for Fable 5.1 and 0% for GPT-6 Astra. These comparisons are descriptive: configurations were not normalized, and the latest models also pass fewer exploitable tasks. Violations concentrate around a small set of recurring surfaces, especially unintended access to reference solutions through git history. Three repair case studies across two benchmarks show why blocking a recorded exploit is insufficient: the same protected information can remain accessible through another route. Therefore, we combine minimal patches with exploit replay and fresh agent evaluation, auditing new passes under the original standard. No evaluated attempt against the final patches reached the protected channel, and every post-patch pass was judged legitimate. Benchmark integrity requires ongoing maintenance: audit passing behavior, repair the enabling surface, and re-evaluate both exploit access and legitimate solvability.
|
| 1826 |
Query Expansion and Key Specialization in Transformer Attention Geometry
2609.34273
|
cs.LG
|
Vidit Gupta, Siddhesh Nadkarni, Mihik Chaudhari, Vinaya Sawant, Prachi Tawde |
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional...The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that $PR_Q - PR_K$ is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of $QK^\top$ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.
|
| 1827 |
MoSPR: Histology-to-Gene Expression Prediction with Morpho-Spatial Macrostates and Low-Rank Molecular Programs
2609.34280
|
cs.LG
|
Dongmyung Shin, Geongyu Lee, Yesung Cho, Park Jong Bae |
Predicting molecular profiles from histopathology remains challenging because whole-slide images contain spatially organized, heterogeneous tissue patterns, while gene expression comprises thousands of correlated targets. We introduce MoSPR (Morpho-Spatial Pro...Predicting molecular profiles from histopathology remains challenging because whole-slide images contain spatially organized, heterogeneous tissue patterns, while gene expression comprises thousands of correlated targets. We introduce MoSPR (Morpho-Spatial Program Regression), a linear framework that couples an adjacency-informed histology representation with a low-rank molecular basis. MoSPR clusters frozen patch embeddings into morphology microstates, aggregates their spatial adjacencies across the training cohort, and groups microstates with similar adjacency patterns into shared macrostates. Each slide is then represented by global morphology and macrostate-specific deviations, which are linearly mapped to coefficients of a training-derived low-rank gene-expression basis. Across three cancer cohorts from The Cancer Genome Atlas, MoSPR achieves the highest mean gene-expression prediction scores among all evaluated methods. Without pathway-level supervision, pathway scores derived from its predicted expression profiles rank first in eight of nine comparisons across three pathway collections. Ablation studies on the breast cancer cohort show complementary gains from adjacency-derived macrostate representation and low-rank molecular prediction. Moreover, with half of the training data on this cohort, MoSPR exceeds the full-data gene-prediction score of the strongest competing baseline. Finally, its linear formulation enables exact decomposition of each predicted expression profile into global and macrostate-specific molecular contributions, providing an interpretable link between spatially coherent macrostate regions and their associated molecular programs. Our code is available at https://github.com/Radisen-Panthera/MoSPR.
|
| 1828 |
Dexterous Tactile World Model
2609.34286
|
cs.LG
|
Ziyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan, Zhengxiang Yu |
World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterou...World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.
|
| 1829 |
Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale
2609.34292
|
cs.LG
|
Jun-qiang Lu |
We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The comm...We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.
|
| 1830 |
Predictive Semantic Safety: From Visual Physical Reasoning to Safety-Critical Control
2609.34356
|
cs.LG
|
Taekyung Kim, Salem Fradi, Yanning Dai, Mateusz Ostaszewski, J\"urgen Schmidhuber |
Physical interactions can create future hazards that are not apparent from the robot's current geometric surroundings. We present a framework termed Predictive Semantic Safety (PSS), which connects visual physical reasoning to backup-based safety filtering. A ...Physical interactions can create future hazards that are not apparent from the robot's current geometric surroundings. We present a framework termed Predictive Semantic Safety (PSS), which connects visual physical reasoning to backup-based safety filtering. A vision-language model (VLM) predicts physical events and their timing or directly predicts object displacements. An explicit motion model converts event hypotheses into object trajectories. Split conformal prediction calibrates position errors jointly across specified objects, observation times, and future times; geometric shape bounds convert the resulting position regions into predicted object occupancy. PSS evaluates a prescribed backup maneuver against this occupancy and derives input-affine constraints for minimally modifying the nominal input while preserving backup feasibility under the robot dynamics and input limits. MuJoCo experiments with a Unitree Go1 consider falling fixtures, impact-driven support loss, and contact propagation. PSS achieves a safe episode rate of 99.3%, compared with 43.3% for a Backup Control Barrier Function baseline that only uses current obstacle geometry.
|
| 1831 |
P2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response Prediction
2609.34391
|
cs.LG
|
Haojie Yang, Ran Su |
AIVC (AI Virtual Cell) is a learned simulator of cellular behavior across conditions. Predicting how a cell population responds transcriptionally to a genetic perturbation is a core task. Perturb-seq records that response by destructive sequencing, so a contro...AIVC (AI Virtual Cell) is a learned simulator of cellular behavior across conditions. Predicting how a cell population responds transcriptionally to a genetic perturbation is a core task. Perturb-seq records that response by destructive sequencing, so a control cell and a perturbed cell are never observed as a pair, and cells under one condition remain heterogeneous and noisy. Regression on individual cells absorbs sampling variation into the estimated effect, whereas interpretation requires the reproducible population effect. P2P (Perturbation-to-Perturbation) takes a stochastic cell-set view as its supervision unit. Two views drawn from the same condition share a reproducible population effect and differ by view-specific variation. A permutation-invariant set encoder summarizes the control population, a structured encoder represents perturbation tokens, cellular context, dose, and combination interactions, and a gate blends empirical condition-effect memory with a neural residual. A heteroscedastic head predicts the population mean and gene-wise response variance. Under one protocol and five seeds, P2P attains the lowest expression RMSE and the highest Effect Pearson, DEG F1, and DEG average precision on each of Adamson, Norman, Replogle K562, and Replogle RPE1 relative to GenePert, LinearPert, SLIM, Scouter, and scPILOT. On Replogle K562, Effect Pearson rises from 0.643 to 0.702 and DEG F1 rises from 0.067 to 0.178 relative to Scouter, the strongest baseline on both metrics.
|
| 1832 |
Understanding Generalization Requires Universal Induction
2609.34458
|
cs.LG
|
Aram Ebtekar, Marcus Hutter, Danica J. Sutherland |
Classical statistical theory is insufficient to explain the successes of general-purpose AI models, because it depends on handcrafted inductive biases that it cannot justify. No Free Lunch (NFL) theorems force any learner that beats chance on some environments...Classical statistical theory is insufficient to explain the successes of general-purpose AI models, because it depends on handcrafted inductive biases that it cannot justify. No Free Lunch (NFL) theorems force any learner that beats chance on some environments to underperform on others. We might hope that past experience informs which environments to expect, but NFL applies equally to meta-learning. Thus, any method that makes meaningful predictions necessarily begins with an inductive bias external to the data. Choosing to bias toward short programs yields Solomonoff induction (SI), whose performance is competitive against all computable learners - albeit up to "constants" that become large when comparing against specialized methods that exploit background information. We therefore relativize SI to an information vantage point, biasing toward short programs with access to all preexisting information. This reframes the inductive bias: instead of seeking some absolute notion of simplicity, we favor accessibility with respect to our vantage point. An algorithm can only outpredict the relativized SI to the extent that its code contains additional information about the data, and no algorithm can generate such information. While SI is incomputable and hence not a practical algorithm, it provides a formal optimum for inference in the limit of infinite compute, and there is evidence to suggest that frontier AI systems roughly approximate it. Thus, the only known answer to meta-NFL is rooted in algorithmic information theory, which we should expect to play a fundamental role in explaining the generalization behavior of modern (and future) AI systems.
|
| 1833 |
When Does Structured Knowledge Help Neural Theorem Proving?
2609.34460
|
cs.LG
|
Sareh Nabi, Roland Vogl, Marzieh Nabi |
Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer math...Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer mathematicians rely on for discovery (analogies, generalizations, cross-domain bridges) remains implicit. We introduce MathAgent, which builds this layer as a knowledge graph, MathKG, and uses it to augment LLM theorem provers. MathKG connects 364 Mathlib theorems and definitions by 9,434 typed semantic edges inferred via LLM-based relation extraction anchored to verified Mathlib declarations. We run a controlled ablation across four augmentation modes (no context, knowledge-graph context, Mathlib retrieval, both) and five models: Qwen3-8B/32B, their Lean-specialized derivatives Goedel-Prover-V2-8B/32B, and Claude Sonnet 4.6, on miniF2F, plus PutnamBench and MathOlympiadBench for Sonnet. Three findings emerge. (i) Specialization dominates augmentation: Lean fine-tuning adds 33-38 percentage points of solve rate in every mode, and a specialized 8B model beats a $4\times$ larger general one by 29-35 points, while no augmentation mode improves solve rate by more than 3 points. (ii) Augmentation is capability-conditioned: knowledge-graph context helps small models but hurts large ones, with the specialized model gaining more relative to its general base at every scale. (iii) Yet the augmentation modes solve different problems: an oracle selecting the best mode per problem solves 6% to 58% more than the unaugmented prover, a complementarity effect that strengthens on harder problems (32% more on PutnamBench). These results motivate adaptive strategies that select augmentation by model capability and problem. Code, data, and artifacts are available at https://github.com/sarehnabi/mathagent
|
| 1834 |
ARS: Agentic Reward System for Robot Learning
2609.34484
|
cs.LG
|
Sheng Hu, Weiyi Lu, Lingbing Zeng, Gan Weng, Weiwei Zhang |
Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward...Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training. Given an offline trajectory and a task instruction, ARS uses adaptive visual inspection for both event proposal and verification. A subagent proposes a task-relevant event timeline, which a primary agent verifies and revises before estimating per-frame progress. ARS can incorporate optional terminal outcome labels and visual references to inform its judgments. It can also audit progress estimates from external reward models. We evaluate ARS with a 27B VLM on a controlled semantic-mismatch benchmark and downstream policy learning in simulation and on a real robot. The benchmark reveals that several evaluated reward baselines assign spurious progress to wrong-object manipulation even in simple pick-and-place scenes. ARS better suppresses these errors and outperforms these baselines in simulation policy learning. We further demonstrate that ARS supports long-horizon policy learning from mixed-quality offline experience on real-robot multi-screw fastening in a full-scale laboratory replica of an industrial washing-machine assembly line. These results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot reward modeling. Code is at https://github.com/midea-ai/ars
|
| 1835 |
AgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud Continuum
2609.34586
|
cs.LG
|
Michalis Kasioulis, Moysis Symeonides, George Pallis, Marios D. Dikaiakos |
Deploying LLM-enabled agentic applications across the Edge-to-Cloud continuum remains challenging due to hardware heterogeneity, deployment complexity, limited observability, and the lack of systematic evaluation methods. Existing solutions address agent devel...Deploying LLM-enabled agentic applications across the Edge-to-Cloud continuum remains challenging due to hardware heterogeneity, deployment complexity, limited observability, and the lack of systematic evaluation methods. Existing solutions address agent development, observability, or benchmarking separately, offering limited support for the full lifecycle of distributed agentic applications. This paper presents AgentWare, an AgenticOps framework that automates the provisioning, deployment, observability, and evaluation of agentic applications across Edge-to-Cloud infrastructures. AgentWare introduces an end-to-end lifecycle pipeline that automatically prepares heterogeneous execution environments, transforms user-defined agent implementations into distributed applications, deploys agent components across the continuum, and performs unified collection of execution traces, infrastructure telemetry, and evaluation metrics. The framework further supports automated semantic evaluation through LLM-as-a-Judge workflows and generates reproducible reports covering correctness, performance, resource utilization, and energy consumption. We demonstrate the applicability of AgentWare through a distributed book assistant agent deployed across real Edge-to-Cloud infrastructure under multiple deployment and model configurations. The results show that AgentWare enables systematic experimentation and evaluation of distributed agentic applications while significantly reducing the manual effort required for deployment, instrumentation, and analysis.
|
| 1836 |
Probabilistic Geodesic Flow Matching on Location-Scale Families
2609.34613
|
cs.LG
|
Zeyuan Yu, Zhi Chang, Shiwei Lan |
Flow matching (FM) has recently emerged as a promising framework for generative modeling due to its conceptual simplicity and strong empirical performance. In FM, samples are transported along a vector field parameterized by a neural network, inducing a probab...Flow matching (FM) has recently emerged as a promising framework for generative modeling due to its conceptual simplicity and strong empirical performance. In FM, samples are transported along a vector field parameterized by a neural network, inducing a probability path that evolves from a simple noise distribution to the target data distribution, governed by an ordinary differential equation (ODE). However, existing FM approaches predominantly rely on probability paths derived from optimal transport (OT) between Gaussian distributions, which may be suboptimal for capturing complex data with inhomogeneous structures such as heavy tail or sharp contrast. In this work, we generalize FM to the broader class of location-scale families for handling data inhomogeneity and introduce a novel class of probability paths defined as geodesics on the manifold of probability distributions. We name this approach probabilistic geodesic flow matching to distinguish it from prior geodesic (Riemannian) FM methods defined in input space. We argue that Euclidean OT-based paths are not necessarily optimal in probability space and may limit modeling flexibility. Through synthetic benchmarks and scientific datasets at different scales, we demonstrate that the proposed method more effectively captures complex distributions, leading to improved or comparable performance compared with SOTA geometry-motivated generative models.
|
| 1837 |
CLAD: Constrained Abstract Domain for Neural Network Verification
2609.34628
|
cs.LG
|
Hai Duong, Thanh Le, ThanhVu Nguyen |
Neural network verification (NNV) formally verifies that a network satisfies a specified property for all inputs within a defined region. Modern NNV tools employ abstract domains to compute a sound over-approximation of the network's behavior from the given in...Neural network verification (NNV) formally verifies that a network satisfies a specified property for all inputs within a defined region. Modern NNV tools employ abstract domains to compute a sound over-approximation of the network's behavior from the given input region, thus the tightness of these abstractions essentially determines efficiency. A long line of increasingly precise domains has been developed, but they all describe the valid input region in the same restrictive way, e.g., an Lp-norm ball. A practical input region is rarely a simple Lp ball, but rather a combination Lp ball with additional constraints. Verifying a network over such a region with existing abstraction produces a loose over-approximation, which results in either failing to verify a property or spurious counterexamples. We introduce Constrained Lagrangian Abstract Domain (CLAD), a new abstract domain that computes a sound over-approximation of neural networks over input regions defined by a combination of convex constraints. CLAD propagates these constraints and tightens bounds over the true feasible region. However, bounding a neuron over the intersection of these constraints has no closed-form solution, so CLAD relaxes each constraint into the objective with a Lagrange multiplier and solves the resulting max-min problem with a projected primal-dual method, alternating a projected gradient step on the input with a multiplier update. CLAD supports any convex constraint with a subgradient, e.g., from automatic differentiation. We evaluate CLAD on 1,944 instances across four convolutional networks with motion-blur structured perturbations with halfspace or L2-ball constraints. On standard unconstrained Linf property, CLAD verifies as many instances as GCPCROWN at a similar runtime. On constrained properties, CLAD verifies 60\% more instances than GCPCROWN on L2-ball properties, and 22% more in total.
|
| 1838 |
OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming
2609.34653
|
cs.LGcs.AI
|
Zongshang Shen, Wangsong Yin, Daliang Xu, Mengwei Xu, Xuanzhe Liu |
On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV ...On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm level, OmniPick logically retains critical multimodal context based on unit boundaries and modality importance to preserve task accuracy; (2) At the system level, OmniPage physically partitions the cache by retention likelihood and dynamically compacts surviving sparse tokens, minimizing both memory fragmentation and data-movement overhead. Extensive evaluations across three streaming benchmarks and two consumer-device architectures show that OmniTide achieves up to $12.72\times$ kernel speedups and $2.40\times$ lower stream-loop latency. On StreamingBench, it improves accuracy by up to 18.0 percentage points over sliding-window baselines at comparable session cost. OmniPage further reduces the physical KV span by up to 26.7% relative to native logical eviction, unlocking real-time, infinite-context streaming on edge devices.
|
| 1839 |
On the Numerical Reliability of Differentiable Physics-Based Optimization for Robotic Material Manipulation
2609.34666
|
cs.LG
|
Xintong Yang, Minglun Wei, Yu-Kun Lai, Ze Ji |
Differentiable physics is increasingly used in robotic material manipulation for system identification, trajectory or skill optimization, demonstration generation, and robot or end-effector design. These applications depend on gradients propagated through long...Differentiable physics is increasingly used in robotic material manipulation for system identification, trajectory or skill optimization, demonstration generation, and robot or end-effector design. These applications depend on gradients propagated through long, contact-rich simulation rollouts. We study the numerical reliability of those gradients using two Material Point Method (MPM) system-identification benchmarks derived from elastoplastic and granular manipulation. The benchmarks provide controlled cases for three effects that also arise in broader differentiable physics-based optimization. GPU many-to-one sums whose order depends on thread scheduling changed long-horizon gradients and reversed the sign of one parameter gradient relative to a deterministic reference. Finite-difference checks became less reliable for longer rollouts because repeated-run loss variation grew much faster than the loss change produced by the tested parameter perturbations. Observation and loss definitions changed optimization behaviour and the solution preferred by an independent metric. These results motivate reproducible accumulation, finite-difference validation that compares perturbation-induced loss changes with repeated-run variation, and explicit reporting of objective construction when differentiable simulation is used for robotic optimization.
|
| 1840 |
Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks
2609.34667
|
cs.LG
|
Etienne Boursier, Nicolas Flammarion |
Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning f...Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has $m+1$ features, $m$ of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only $\mathcal{O}(d)$ fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.
|
| 1841 |
Hierarchical Clustering and Signal Denoising on Digraphs
2609.34670
|
cs.LG
|
Yi Wang, Sippanon Kitimoon, Hrushikesh N. Mhaskar, Xiaosheng Zhuang |
In this paper, we propose a representation of a digraph (directed graph) as a Hermitian matrix derived from its adjacency matrix. This representation characterizes both the connectivity and the edge orientation of the digraph. Based on the spectral decompositi...In this paper, we propose a representation of a digraph (directed graph) as a Hermitian matrix derived from its adjacency matrix. This representation characterizes both the connectivity and the edge orientation of the digraph. Based on the spectral decomposition of the Hermitian matrix, a digraph clustering algorithm with $k$-means is introduced to produce a partition on the graph. Applying this algorithm (bottom-up) recursively to a digraph with partially labeled vertices yields a spectral hierarchical digraph clustering (\myproj) algorithm that produces consistent nested partitions of the digraph, or equivalently, a tree structure. Furthermore, based on the in-degree and out-degree of each cluster in the digraph clustering, a pair of hierarchical interval partitions (filtrations) can be derived in a top-down manner to produce a pair of nested knot sequences. These knot sequences facilitate the construction of multilevel spline quasi-interpolants, enabling a noisy graph signal to be decomposed into a coarse approximation and inter-level details, followed by adaptive thresholding and reconstruction. Experiments on synthetic and real-world digraphs demonstrate the superiority of our {\myproj} algorithm for digraph clustering across diverse graph structural properties (homophily and heterophily) and supervision settings. Moreover, experiments on digraph signal processing using multilevel spline quasi-interpolants further demonstrate the effectiveness of signal recovery on digraphs in terms of RMSE and SNR.
|
| 1842 |
Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts
2609.34684
|
cs.LG
|
Hyungjoon Kim, Wonbin Son, Mi Young Lee, Jun Young Lee, Seungmin Rho |
Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configur...Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing contextual cues to contribute to prediction. In this paper, we introduce an evaluation framework that separates prediction accuracy, target-state responsiveness, and context stability using physically validated observations that cross target coordinates with robot contexts. We demonstrate that high natural-trajectory accuracy can coexist with weak controlled target-state responsiveness in fixed representation-readout pairs. Comparisons and interventions involving representations, readouts, and training data show that the three properties provide distinct diagnostic information. Furthermore, adding responsiveness and context sensitivity to a failure predictor based on initial state error and physical variables reduces policy-failure prediction error on new initializations relative to the specified baseline while same-observation controlled MAE is also informative. These findings motivate evaluating target-state responsiveness and context stability alongside natural prediction accuracy, and examining their relationship to actual policy behavior and task outcomes.
|
| 1843 |
ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems
2609.34701
|
cs.LG
|
Tarun Chintada, Neelamadhav Gantayat, Ishaan Romil, Renuka Sindhgatta, Soujanya Soni |
Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdow...Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that prevent task completion. While existing observability frameworks provide traces and logs, diagnosis and remediation are largely performed after execution completes, limiting opportunities for recovery during runtime. We present ResonAct, a runtime self-healing framework that enables continuous monitoring, diagnosis, and remediation of multi-agent systems through streaming operational metrics. ResonAct ingests execution traces, agent interactions, and tool invocations into a streaming analytics layer that continuously derives task progress, context health, and tool reliability metrics. These metrics serve as runtime control signals for detecting anomalous execution patterns and localizing root causes using a structured failure model. Based on the diagnosed failure, ResonAct dynamically selects remediation policies and performs actions. The framework operates as an external control plane, enabling intervention without modifying application agents or orchestration logic. We evaluate ResonAct across enterprise workflow scenarios and AppWorld benchmarks. The results show that the streaming metric-based analysis identifies execution degradations and localizes faults. Furthermore, policy-driven remediation improves task completion rates by up to 10.00 percentage points, with detection precision ranging from 70.59% to 82.91%, recall from 63.09% to 100%, recovery rates from 10.48% to 46.67%, and runtime overhead ranging from $-0.25%$ to 14.12% across the evaluated configurations.
|
| 1844 |
Information-Theoretic Analysis of Next-Token Prediction under Markovian Data
2609.34731
|
cs.LG
|
Masoud Kavian, Abdellatif Zaidi, Milad Sefidgaran |
We develop an information-theoretic framework for generalization in next-token prediction under temporally dependent data. We consider independent trajectories generated by finite-memory Markov processes and distinguish algorithmic dependence, quantified by mu...We develop an information-theoretic framework for generalization in next-token prediction under temporally dependent data. We consider independent trajectories generated by finite-memory Markov processes and distinguish algorithmic dependence, quantified by mutual information, from temporal dependence, characterized by mixing. For cross-entropy loss, we derive an expected generalization bound using the Donsker--Varadhan variational representation and a McDiarmid-type concentration inequality for Markov chains. A refinement captures the joint effect of context length and temporal mixing through the mixing properties of the history-state process. We then extend the bound through a rate--distortion formulation, replacing mutual information with the minimum information rate required to represent the learned model within a prescribed distortion in the generalization gap, yielding informative guarantees for deterministic algorithms over continuous hypothesis spaces. For margin-based prediction, we derive explicit bounds for linear and self-attention next-token predictors via noisy low-dimensional compression, revealing the roles of context length, model complexity, sample size, margin, and temporal mixing. Experiments on TinyStories and ETTh2 show that longer contexts can reduce both training and test losses, but typically reduce training loss more, enlarging the generalization gap. A complementary ETTh2 analysis identifies an effective predictive-memory scale near 24 hours, with no statistically supported improvement beyond this scale, offering a plausible explanation for test-performance saturation at larger contexts.
|
| 1845 |
Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?
2609.34742
|
cs.LG
|
Bohao Xing, Xin Liu, Kaishen Yuan, Deng Li, Rong Gao |
Recent studies suggest that large language models encode emotion concepts as structured internal representations, but most existing work focuses on text and a single architecture. Therefore, we ask, do emotion concepts generalize across sources, modalities, an...Recent studies suggest that large language models encode emotion concepts as structured internal representations, but most existing work focuses on text and a single architecture. Therefore, we ask, do emotion concepts generalize across sources, modalities, and architectures in vision--language models (VLMs)? To address this, we construct CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes. For each stimulus source, we extract a separate set of six Ekman emotion vectors from each of three VLMs. We report four main findings as follows: 1) Image-derived emotion vectors form a low-dimensional geometry similar to that of text-derived vectors. Valence is relatively stable across sources, while arousal varies more. 2) Text- and image-derived emotion vectors have modest cosine similarity but still show held-out cross-modal correspondence. Text-derived vectors can also steer image interpretation. 3) Cross-architecture correspondence remains even when native cosine is near zero. Transformations estimated from generic ImageNet activations recover both correspondence and causal transfer without using the six emotion vectors or their labels. 4) After aligning representations across architectures, we construct a shared emotion subspace that preserves affective geometry and selective steering effects. The corresponding consensus emotion vectors also generalize to a held-out fourth architecture at two model sizes. These results suggest that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even when individual vector directions differ.
|
| 1846 |
Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks
2609.34756
|
cs.LG
|
Alexandre Decl\`eves, Etienne Boursier, Nicolas Flammarion |
Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoreticall...Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the lens of sparse linear regression and two-layer diagonal linear networks. In our setting, pretraining provides information through the support (and signs) of the initialization predictor, which may contain coordinates relevant to the downstream task. We show how pretrained information reshapes the implicit bias and training dynamics, and can thereby reduce the sample complexity of recovering the target parameters and support. In particular, for a clean initialization with correctly inherited signs, we show that the required sample size is comparable to that of a weighted Lasso estimator that explicitly exploits the pretrained support through a suitably chosen regularizer. Our results thus show how information encoded in pretrained weights can be implicitly exploited by gradient-based fine-tuning, reducing the amount of data needed to recover a downstream task.
|
| 1847 |
Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models
2609.34765
|
cs.LG
|
Minchan Kang, Kyeonghye Park, Seungyeon Sa, Seoyoung Cho, Daeshik Kim |
Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate se...Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize reconstruction loss with respect to the full-precision model, potentially over-preserving FP behavior and calibration-specific bias. Rather than treating quantization solely as an error to be minimized, we observe that it can also provide beneficial regularization for certain layers and modalities. Motivated by this observation, we propose Balanced Fitting, a quantization effect-based framework that balances precision and regularization beyond reconstruction-based optimization. By measuring layer- and component-wise quantization effects for weights, vision activations, and text activations, Balanced Fitting combines fine-grained fitting for sensitive components with coarser fitting to exploit potential regularization benefits. Experiments on multiple LVLMs show that our method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance. The source code is publicly available at https://github.com/kmc3661/BFQ
|
| 1848 |
CoHuB: A Simulation Benchmark for Multi-Humanoid Collaboration
2609.34782
|
cs.LG
|
Hyunjin Park, Jebeom Chae, Minwoo Park, Sunghyun Park, Hanjun Yoo |
Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under ego...Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduce CoHuB (Collaborative Multi-Humanoid Benchmark), a simulation benchmark for multi-humanoid collaboration under egocentric visual observations. CoHuB provides 10 tasks, eight with two humanoids and two with three humanoids, spanning diverse collaboration patterns. We also provide synchronized demonstrations collected through a multi-operator VR teleoperation pipeline, in which each operator controls one humanoid from its egocentric view. Experiments with representative visuomotor policies reveal substantial challenges across different forms of coordinated perception and control. CoHuB provides a foundation for developing and evaluating multi-humanoid collaboration policies.
|
| 1849 |
BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification
2609.34785
|
cs.LG
|
Yuheng Wu, Berk Gokmen, Sujeeth Jinesh, Lauren McLane, Aarav Wattal |
Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid design...Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE, an agentic framework for multi-turn joint hardware design and verification through functional behavior modeling. We define Behavior IR to express task functionality as executable behavior models without prescribing implementation timing beyond the specification. The agent iteratively develops a register-transfer-level (RTL) design and a behavior model as the design's verification reference. Our evaluator, BEHAVE-Sim, checks both artifacts separately against a hidden golden behavior model using input stimuli generated by random sampling and solver-guided search. BEHAVE thus supports power, performance, and area (PPA) exploration across task-permitted latencies and microarchitectures. During training, the same evaluator provides verifiable reinforcement learning (RL) rewards from specification-behavior pairs without reference RTL. For self-improvement, the agent continually searches for high-level implementations relevant to its capability gaps, constructs and checks specification-behavior pairs, and trains on the expanded task pool. We release BEHAVE-Train and BEHAVE-Eval with 600 human-reviewed specification-behavior pairs for realistic hardware workloads. Starting from 60 seed tasks and acquiring 100 new tasks, self-improvement raises Qwen3.8-27B's RTL pass@1 on BEHAVE-Eval from 55.0% to 75.0%, reaching performance comparable to RL using a 540-task pool.
|
| 1850 |
Finite-Time Concentration and Convergence Rates for Projected Two-Time-Scale Stochastic Approximation with Markov Noise
2609.34791
|
cs.LG
|
Rahul Singh, Vivek S. Borkar, Eric Moulines |
We study finite-time concentration and convergence rates for projected two-time-scale stochastic approximation driven by a controlled Markov chain. The averaged fast map is contractive, while the slow iterate is projected onto a compact convex polyhedron. The ...We study finite-time concentration and convergence rates for projected two-time-scale stochastic approximation driven by a controlled Markov chain. The averaged fast map is contractive, while the slow iterate is projected onto a compact convex polyhedron. The associated projected ordinary differential equation may have a discontinuous vector field at the boundary, preventing a direct application of standard analyses based on Lipschitz vector fields. Using the Skorokhod map, we establish explicit high-probability bounds for tracking the moving fast equilibrium and the projected slow dynamics. These bounds separate martingale fluctuations, Markov-noise residuals, and the bias due to time-scale separation. A Lipschitz Lyapunov function satisfying a uniform decrease condition over fixed time intervals yields almost-sure convergence, with explicit last-iterate rates when the decrease admits a power lower bound. Under uniform Lyapunov contraction, polynomial step sizes yield joint fast-tracking and slow Lyapunov-error exponents arbitrarily close to $1/3$. Under the additional assumption that the reduced slow update map is a Euclidean contraction, logarithmically separated step sizes improve the joint rate to $O(n^{-1/2}\log n)$ almost surely, including for boundary equilibria. The same rate holds under a distinct geometric condition involving a strictly attracting face of a box and a fast equilibrium that is constant on that face. An actor-critic application achieves an almost-sure value-gap rate of $O(n^{-1}\log n)$ relative to the optimum within the constrained policy class. Further applications include projected TD(0) and projected stochastic gradient descent. We also extend the analysis to projection of the fast recursion under Euclidean contractivity.
|
| 1851 |
Physics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water Infrastructure
2609.34804
|
cs.LG
|
Jeff Nijsse, Shu Su, Benjamin Oholeguy, Sreenivas Sremath Tirumala |
Federated learning enables industrial operators to train shared intrusion detection models without disclosing proprietary operational telemetry. However, existing defenses operate strictly in update space, leaving aggregators blind to data poisoning; model upd...Federated learning enables industrial operators to train shared intrusion detection models without disclosing proprietary operational telemetry. However, existing defenses operate strictly in update space, leaving aggregators blind to data poisoning; model updates derived from fabricated telemetry remain indistinguishable from honest contributions. We repurpose cyber-physical process invariants, such as conservation laws and actuator couplings, from runtime detection heuristics into a verifiable admission requirement for federated updates, mined automatically from clean operational data. We evaluate this admission gate across two physical water testbeds (SWaT, WADI) and a distribution benchmark (BATADAL), testing seven aggregation rules against telemetry fabrication, exposure-only replay poisoning, and an invariant-aware adaptive adversary. Across three testbeds the mined invariants reject none of 100 honest shards and all naively fabricated ones, including optimised perturbations that FoolsGold admits in full. On real telemetry, five mined invariants detect 12 of SWaT's 35 attacks, while nine invariants detect 20, with no honest shard rejected. With nine rules, the physics gate recovers 69--100% of the targeted-attack recall lost to replay poisoning, and 54--100% of that lost to fabricated telemetry, across five standard aggregators. To reconcile physical admission control with federated data privacy, we show invariant compliance using zero-knowledge proofs (zk-SNARKs) to allow clients to prove batch adherence without revealing operational telemetry.
|
| 1852 |
On Temporal Binding in Large Audio Language Models
2609.34806
|
cs.LG
|
Paul Primus, Gerhard Widmer |
Reasoning about temporal structure of audio recordings requires Large Audio Language Models (LALMs) to associate sound events with their temporal position. Understanding the underlying mechanisms is a first step toward diagnosing failures and identifying model...Reasoning about temporal structure of audio recordings requires Large Audio Language Models (LALMs) to associate sound events with their temporal position. Understanding the underlying mechanisms is a first step toward diagnosing failures and identifying model components that may need improvement. Using mechanistic interpretability, we investigate how temporal information is represented and bound to sound events in three open-source LALMs. We find that across all three, event-specific location becomes concentrated in event name representations at intermediate modality integration layers. These representations encode coarse event position along a low-dimensional, curved relative time trajectory. Steering event name representations along this trajectory systematically shifts before/after beliefs, providing evidence that these representations contribute to coarse temporal reasoning. In contrast, the same interventions do not reliably shift predicted onset timestamps, suggesting that coarse temporal reasoning and precise metric event localization rely on distinct mechanisms.
|
| 1853 |
From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models
2609.34809
|
cs.LG
|
Rong Yu Xu, Prayag Tiwari, Shaolei Zhang |
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, togethe...Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.
|
| 1854 |
Accelerator Choice Is Not Enough: AlphaFold2 Inference on Cloud TPUs
2609.34818
|
cs.LG
|
Lorenzo Pazienza, Ihab El Bani |
AlphaFold2 is written in JAX, so the same inference code compiles and runs unchanged on CPUs, GPUs and Google Cloud TPUs. That portability makes the accelerator look like the main decision a user has to make. We show that it is not. Running one AlphaFold2 infe...AlphaFold2 is written in JAX, so the same inference code compiles and runs unchanged on CPUs, GPUs and Google Cloud TPUs. That portability makes the accelerator look like the main decision a user has to make. We show that it is not. Running one AlphaFold2 inference workload across a Colab CPU runtime, an NVIDIA T4 GPU and a dedicated eight-chip Cloud TPU v5e slice, we find a large hardware advantage for the TPU, 0.47 s per call in steady state on a single chip against 13.1 s on the T4 in the same measurement campaign, and three ways in which the software layer decides how much of it a user actually gets. The default execution path uses one chip of the eight, and at list prices the idle capacity makes the slice cost about as much per prediction as the GPU. Batching with JAX's vmap never exceeds single-query throughput, while mapping queries across chips with JAX's pmap gives eight chips 6.5-7.9x the throughput of one on a matched grid; automatic sharding leaves the per-chip footprint unchanged, consistent with replication, most plausibly because AlphaFold2 carries no sharding annotations. Our retained trace analysis of a first call at a new input shape reports about three quarters of the traced span in JAX tracing and compilation rather than execution. Reruns five weeks later reproduced neither cloud baseline, the GPU one off by roughly a factor of two, so the hardware ratio above is specific to one campaign.
|
| 1855 |
MW-Nowcast: Six-hour ensemble nowcasting of extreme precipitation
2609.34836
|
cs.LG
|
Ning Wang, Zuliang Fang, Weixin Jin, Zhongjian Lv, Shuang Qin |
Extending reliable nowcasting of extreme precipitation could provide critical additional time for warnings and emergency response during high-impact events such as flash floods. Radar-based generative machine-learning models have enabled skilful hyperlocal pre...Extending reliable nowcasting of extreme precipitation could provide critical additional time for warnings and emergency response during high-impact events such as flash floods. Radar-based generative machine-learning models have enabled skilful hyperlocal precipitation nowcasting, but accurate prediction of intense precipitation remains confined to the first few hours. Because storm-scale structure is predictable for longer than individual cells, a natural strategy is to predict that structure while generatively modelling only the uncertain local growth, decay, reorganisation and initiation of storms. Here we present Microsoft Weather Nowcast (MW-Nowcast), a six-hour ensemble radar nowcasting model that jointly learns a deterministic predictor to capture organised precipitation structure shared across ensemble members, and a generator to produce diverse local residuals around this shared prediction. Across independent test data from the United States, Europe and China, MW-Nowcast achieves higher detection skill than leading methods for heavy and extreme precipitation throughout the 6 h horizon. For the most intense rainfall, MW-Nowcast doubles the available warning time across all three regions, delivering 6 h forecasts with skill previously limited to 3 h for the leading generative baseline. A cost-loss decision analysis shows that MW-Nowcast retains substantial value for a broad range of applications even at 4-6 h, where alternative methods offer little benefit. These additional hours can give forecasters and emergency managers the time to warn and act before extreme rainfall strikes, helping to protect lives and property.
|
| 1856 |
From Soft Targets to Reward Signals: How Assignment and Reward Objectives Interact
2609.34850
|
cs.LG
|
Jiangtao Lin, Bangyang Wei, Siyi Liu, Yihang Ding, Yuhan Dong |
Soft preference targets specify supervision strength, and reward objectives convert that strength into learned reward signals. A central design question remains: how does assigning a fixed set of preference strengths to different response pairs change the rewa...Soft preference targets specify supervision strength, and reward objectives convert that strength into learned reward signals. A central design question remains: how does assigning a fixed set of preference strengths to different response pairs change the rewards produced by different objectives? We introduce assignment geometry to study this interaction. Mean-matched smoothing controls target dispersion, while within-stratum reassignment changes correspondence and preserves the complete target distribution. Across five reward objectives, intact correspondence retains the largest clean preference margins among the compared soft targets within a common accuracy-equivalence budget. Attenuation orderings change with the reward objective, revealing different responses to the same target assignments. Independent reassignments and a related source construction reproduce the retention direction. An attenuation-retention profile compares these combinations through margin magnitude, edit response, and accuracy. Against independently calibrated scaling, APLOT uniform targets deliver additional attenuation on both aggregate and presentation edits. These findings establish a joint design space in which target placement and reward objective shape reward properties beyond preference accuracy.
|
| 1857 |
When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model
2609.34861
|
cs.LG
|
Minchan Kang, Kyeonghye Park, Seoyoung Cho, Daeshik Kim, Yucheol Cho |
Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In p...Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT
|
| 1858 |
Conformal Prediction and Conditional Coverage for Tabular Foundation Models
2609.34887
|
cs.LG
|
Sungwoo Park, Sunghee Park, Won Chang |
Tabular foundation models (TFMs) provide predictive distributions for regression, but their prediction regions can exhibit undercoverage or overcoverage even when point predictions are accurate. We introduce C-USIM (Conditionally-Uniformized Score Integration ...Tabular foundation models (TFMs) provide predictive distributions for regression, but their prediction regions can exhibit undercoverage or overcoverage even when point predictions are accurate. We introduce C-USIM (Conditionally-Uniformized Score Integration Method), a lightweight application of highest predictive density split conformal prediction that accommodates multimodal predictions. Given calibration and test outputs, it requires no additional training or model inference. It provides finite-sample marginal validity under our assumptions. We bound conditional-marginal coverage gaps using distribution-estimation error and score discreteness, and examine coverage heterogeneity through percentile rank-score plots. Experiments with TabPFN and TabICL show improved marginal coverage accuracy and lower average conditional and group coverage errors. Under a fixed data budget, allocating more observations to calibration can reduce marginal coverage error despite less accurate point predictions.
|
| 1859 |
SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living
2609.34907
|
cs.LG
|
Debolina Chowdhury, Suman Samui, Sujoy Saha |
Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environm...Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environment. This work introduces \dataset{}, a seven-class bathroom acoustic-event dataset containing 21{,}387 annotated clips recorded across five environments, and proposes SincDPNet, a compact raw-waveform classifier with a learnable sinc filter bank followed by a depthwise-separable convolutional body. Each sinc filter is controlled by two frequency parameters, allowing the learned passbands to be inspected directly in hertz while keeping the front end small. To reduce room-specific leakage, recording sessions and environments are separated before overlapping windows are assigned to the training, validation, and test partitions. We further use multi-objective Bayesian optimization as a design tool to examine the validation performance--model-size trade-off across 24 configurations. The selected designs span different operating points: the best-performing model achieves 80.2\% accuracy and 0.760 macro-F1 with 14{,}040 parameters, while the compact $N_f=25$ configuration uses only 2{,}848 parameters and achieves 75.7\% accuracy, 0.661 macro-F1, and 0.716 MCC on the held-out environment. Analysis of the learned filters and confusion patterns shows that spectral overlap contributes to confusion among water-related events, while the \textit{Door}/\textit{Walker/Crutch} errors also reflect similarities in their transient temporal structure.
|
| 1860 |
Physics-Informed Neural Networks for Depth-Averaged Avalanche Dynamics
2609.34916
|
cs.LG
|
Pradyumn Singh Sikarwar, Vishal Sharma, Gaurav Bhutani |
Accurate prediction of avalanche motion is essential for hazard assessment in mountainous terrain. This study develops and evaluates a physics-informed neural network (PINN) framework for the Savage-Hutter model of depth-averaged granular flow, progressing fro...Accurate prediction of avalanche motion is essential for hazard assessment in mountainous terrain. This study develops and evaluates a physics-informed neural network (PINN) framework for the Savage-Hutter model of depth-averaged granular flow, progressing from 1D analytical verification to 2D experimental validation. First, three 1D problems of increasing complexity were verified against the analytical solution: height prediction with prescribed velocity, velocity prediction with prescribed height, and coupled prediction of both fields using the conservative formulation. The decoupled tests accurately reconstructed the spatio-temporal evolution of each field when the other was prescribed. The coupled formulation learned both fields without prescribed data, achieving mean height and velocity RMSEs of 0.043 and 0.079 in non-dimensional units. A hyperparameter sensitivity study evaluated the effects of network depth, width, collocation density, learning rate, and epochs. The framework was then extended to 2D and validated against laboratory experiments of a cylindrical granular pile collapsing on an inclined plane, with TITAN2D providing numerical comparisons. Purely physics-based training converged to the trivial zero solution; augmenting the loss with 10 sparse training points from final deposit profiles produced a physics-informed, data-assisted hybrid framework. Peak flow depth, depth-averaged velocity, RMSE, and wetted-area IoU evaluated global and local agreement. Global height RMSE ranged from 2.7 to 6.7 mm across four experimental cases, while mean wetted-area IoU ranged from 69 to 81 %, demonstrating consistent performance across variations in pile mass and slope angle.
|
| 1861 |
JazzSAMBA: A Synchronous and Asynchronous Multi-take Band Audio Dataset of Jazz Standards for Live Music Models
2609.34931
|
cs.LG
|
Phillip Long, Jacob Nguyen, Jace Hosto, Gage Hosto, Jett Takazawa |
Machine learning has made strong progress on music tasks, both as assistive tools and as creative partners. However, most systems train on multitrack corpora that emphasize pop and rock. Jazz, with improvisation at the core of its practice, still lacks a well-...Machine learning has made strong progress on music tasks, both as assistive tools and as creative partners. However, most systems train on multitrack corpora that emphasize pop and rock. Jazz, with improvisation at the core of its practice, still lacks a well-annotated corpus of clean per-stem combo recordings on standards. We introduce JazzSAMBA (Jazz Synchronous and Asynchronous Multi-take Band Audio) to fill this gap: the first originally recorded jazz-combo multitrack dataset of standards with asynchronous (overdubbed) and synchronous (live ensemble) protocols, preferred and alternate takes chosen by the musicians, and timed annotations for bars, chords, sections, and soloists. JazzSAMBA covers 76 standards by eight musicians on drums, bass, piano, trumpet, and saxophone, with per-stem audio, mixtures, and MIDI. It can support chart-conditioned accompaniment, combo source separation, and form-aware music information retrieval. We demonstrate the dataset on two tasks: a jazz combo source-separation baseline and a chart-conditioned accompaniment ablation. The dataset, code, and samples are linked from the project demo page.
|
| 1862 |
From One-Shot Generation to Incremental Music Composition: Adapting a General-Purpose Instruction LLM for Persistent Symbolic Editing
2609.34994
|
cs.LG
|
Andr\'e Ricardo Ducca Fernandes, Jean-Pierre Briot, Simone Diniz Junqueira Barbosa1, H\'elio C\^ortes Vieira Lopes |
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instr...Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: not as a one-shot music generator, but as a reusable operator over an evolving symbolic score. We formulate incremental composition as a sequence of operation-aware state transitions over persistent ABC notation, with explicit requirements on what each operation may change and what it must preserve. The interaction includes two artifact-initialization variants and three editing operations -- chord addition, inpainting, and transposition. We instantiate the formulation by adapting Llama 3.1 8B Instruct with Low-Rank Adaptation (LoRA) on 496,038 operation-aware dialogue records derived from Irish traditional music. The comparison with the unadapted model is used to test the feasibility of learning this interaction contract, not to claim novelty for fine-tuning itself. Across 500 dialogues per model (1,750 attempted output states), checker admission rises from 29.37% to 99.37%, while compliance conditional on admission rises from 0.7205 to 0.9798. Strict eligibility for reference-relative musical-feature analysis increases from 14 to 1,548 outputs, and Longest Common Subsequence analysis does not show a systematic increase in high-overlap sequences relative to held-out baselines under the specified protocol. The results support the technical feasibility of persistent, operation-aware symbolic editing with a general-purpose instruction LLM. They do not establish superior musical quality or human-AI co-creativity, which remain questions for musician-centered evaluation.
|
| 1863 |
Simulation-Based Quantum System Inference with Neural Posterior Estimation
2609.34995
|
cs.LG
|
Hang Zou, Anton Frisk Kockum, Martin Rahm, Simon Olsson |
Models of quantum systems faithfully map system parameters to observations, but the inverse problem of parameter inference from measurement data presents a fundamental challenge: computationally intractable likelihoods due to an exponentially large Hilbert spa...Models of quantum systems faithfully map system parameters to observations, but the inverse problem of parameter inference from measurement data presents a fundamental challenge: computationally intractable likelihoods due to an exponentially large Hilbert space. Here, we introduce simulation-based quantum system inference, a unified, likelihood-free framework that learns parameter posteriors directly from classical simulation data. The central idea is to pair polynomial-cost classical simulators, such as Pauli propagation and tensor networks, with normalizing flows or other neural density estimators for accurate, reusable inference. A single model, trained once, maps any new measurement record to its posterior in one forward pass---turning per-experiment inference into a fixed, up-front cost. We numerically demonstrate the framework's versatility across Pauli noise learning, quantum error mitigation, quantum state tomography, and Hamiltonian learning, with examples involving 81-qubit shallow circuits and 735-parameter inference. In each case, the approach yields accurate estimates of identifiable parameters, while posterior uncertainty provides additional diagnostics of non-identifiability and indicates where further characterization is needed. Our framework reduces data-acquisition requirements in quantum experiments and accelerates parameter inference, providing a practical route to characterizing and improving large-scale quantum systems.
|
| 1864 |
Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device
2609.35005
|
cs.LG
|
Pawe{\l} Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefa\'nski, Szymon Klimaszewski |
Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables r...Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.
|
| 1865 |
Recommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance Decomposition
2609.35034
|
cs.LG
|
Riki Okamura, Toshiharu Sugawara |
Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the as...Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the assumed examination structures fail. We propose two estimators based on the decomposition of clicks into examination and relevance. First, the latent-examination independent inverse propensity score (LE-IIPS) estimator corrects the IIPS bias using policy examination probability ratios. Second, the examination-decomposed doubly robust (ED-DR) estimator extends LE-IIPS to a doubly robust framework. ED-DR is unbiased if the examination probabilities are correct regardless of relevance accuracy, or under ranking-independent examination, even if both model estimates are inaccurate. Experiments show that ED-DR achieves a lower MSE than existing methods with large sample sizes, especially when the examination depends on ranking. We also highlight its limitations under small samples or cascade user behavior conditions.
|
| 1866 |
GUIDE-FBO: Guidance via Uncertainty Intervention and Distributional Exchange for Federated Bayesian Optimization
2609.35038
|
cs.LG
|
Jintao Wei, Chenxi Li, Songhao Wang |
Federated Bayesian Optimization (FBO) enables distributed agents to collaboratively optimize expensive black-box objectives without sharing raw local observations. However, effective knowledge transfer remains challenging under communication constraints and ta...Federated Bayesian Optimization (FBO) enables distributed agents to collaboratively optimize expensive black-box objectives without sharing raw local observations. However, effective knowledge transfer remains challenging under communication constraints and task heterogeneity. We propose GUIDE-FBO, in which agents exchange compact distributions over the locations of their respective optima inferred from local Gaussian process (GP) posteriors, rather than raw observations, query points, or surrogate parameters. The server merges and reweights these distributional components before returning a subset to each agent. Each agent then constructs a Federated Interventional GP (FI-GP), which preserves the local posterior mean and spatially rescales its covariance for local decision making. For the upper confidence bound (UCB) instantiation, GUIDE-UCB, we prove that any bounded FI-GP uncertainty intervention preserves the leading-order cumulative regret rate of standard GP-UCB. When the transferred distributions place greater support near an optimum than in a suboptimal region, selecting the latter requires greater local posterior uncertainty. Experiments on 12 synthetic benchmarks and three real-world optimization tasks show that GUIDE-FBO remains effective across settings ranging from homogeneous to severely heterogeneous. Ablation results highlight the importance of spatially localized uncertainty intervention, while the communication analysis shows that GUIDE-FBO exchanges only compact distributional messages.
|
| 1867 |
Graph-Based Learning for Multi-Horizon Martian Atmospheric Forecasting
2609.35042
|
cs.LG
|
Gary Myler, James Holmes, Manish Patel, Amel Bennaceur |
Martian weather forecasting is important for future exploration, but atmospheric behaviour on Mars combines spatial, temporal, vertical, and dust-driven processes in ways that challenge current modelling and forecasting approaches. This paper introduces MaGMA ...Martian weather forecasting is important for future exploration, but atmospheric behaviour on Mars combines spatial, temporal, vertical, and dust-driven processes in ways that challenge current modelling and forecasting approaches. This paper introduces MaGMA (Martian Graph-based Multi-horizon Atmospheric Forecasting), a graph-based data engineering framework that transforms OpenMARS reanalysis fields into structured learning objects for Martian atmospheric forecasting. Local atmospheric patches are represented as graph nodes and linked through spatial neighbourhoods, temporal continuity, longer temporal dependencies, and dynamically similar atmospheric states. The model integrates recent atmospheric history, engineered physical descriptors, and vertical atmospheric information to support forecasting across multiple horizons. We evaluate MaGMA across five unseen Martian years, including regular years and a global dust storm year. In regular years, the model achieves overall R^2 values of approximately 0.73-0.85. For dust-column forecasting, it outperforms classical and deep temporal baselines in most year-horizon comparisons. During the global dust storm year, dust-column prediction remains strong at shorter horizons, with R^2 above 0.8 for the first two horizons, while broader multivariate performance declines. The results show that graph-based data engineering can create reusable and diagnostically useful representations for planetary atmospheric forecasting, while highlighting the need for better learning under rare extreme regimes and improved use of vertical atmospheric structure.
|
| 1868 |
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
2609.35047
|
cs.LG
|
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim |
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue ...A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, and Bayesian inference estimates their parameters and states from noisy observations. The resulting model lets the agent predict the outcomes of actions, choose informative experiments, and revise its hypotheses when predictions fail. Across five simulated domains, EMPIRIC learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines. On a physical robot, it learns wind forces and domino masses to solve a manipulation task. Website and code: https://yichao-liang.github.io/empiric
|
| 1869 |
Perceptual Quality Loss or Loss of Perceptual Quality?
2609.35054
|
cs.LG
|
Danilo de Oliveira, Tal Peer, Maur\'icio do V. M. da Costa, Timo Gerkmann |
Contemporary deep speech enhancement (SE) models are often trained with specific auxiliary terms in the loss function as a way to improve their performance in terms of perceptual metrics. Nevertheless, a higher score on a perceptual metric does not necessarily...Contemporary deep speech enhancement (SE) models are often trained with specific auxiliary terms in the loss function as a way to improve their performance in terms of perceptual metrics. Nevertheless, a higher score on a perceptual metric does not necessarily correlate with an improved listening experience. Through objective and subjective experiments, we assess the performance of SE models trained with two different types of auxiliary PESQ loss terms. The numerical evaluation on a suite of standard metrics suggests that, while models optimized for PESQ naturally obtain higher PESQ scores in the test set, for most other metrics the scores do not significantly change. In some cases, the PESQ loss even results in worse PESQ scores on mismatched data. A formal listening experiment reveals that the models without a PESQ loss were generally preferred over models that include it, across all settings. Finally, we analyze the relative importance of PESQ in the composite metrics CSIG, CBAK and COVL, and find that PESQ dominates all of them. Our study highlights the perils of over-reliance on PESQ and stresses the importance of a complete evaluation procedure for SE.
|
| 1870 |
TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash
2609.35065
|
cs.LG
|
Jay H. Park, Hyungjun Kim, Dong Kim |
Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical KV hit is not necessarily ready for GPU retrieval. Demand s...Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical KV hit is not necessarily ready for GPU retrieval. Demand staging exposes SSD latency, whereas immediate staging can reserve fast-tier capacity long before retrieval begins. We present TempoKV, a timing-aware resource-commitment layer that separates early knowledge of reuse from the acquisition of staging resources. It records reusable-KV hits as metadata-only claims and requests commitment when the runtime-estimated time until retrieval falls to the storage-estimated time needed to make KV resident and protected against eviction. These estimates adapt to runtime progress and staging state, while commitment remains subject to available protected capacity. We implement TempoKV in vLLM and LMCache on an SSD-backed CXL memory device without changing request scheduling. Across two models and three prefix cache ratios, TempoKV reduces protected fast-tier byte-time per request by 63-91% versus immediate staging while retaining much of the serving benefit of advance staging. In a fast-tier capacity sweep, output throughput and p95 time to first token (TTFT) remain nearly unchanged as capacity decreases from 100 to 25 GiB. Compared with unmodified LMCache's Device-DAX L1 configuration, TempoKV reduces p95 TTFT by up to 48.0% and increases output throughput by up to 27.8%.
|
| 1871 |
Continuous Variational Synthesis
2609.35083
|
cs.LG
|
Alan N. Amin, Mattia G. Gollub, Andrei Slabodkin, Elizabeth B. Wood, Eli N. Weinstein |
Biological machine learning was long bottlenecked by the ability to synthesize designed DNA. Variational synthesis models control chemical reactions to physically manufacture quadrillions of designed sequences in DNA. However, training these generative models ...Biological machine learning was long bottlenecked by the ability to synthesize designed DNA. Variational synthesis models control chemical reactions to physically manufacture quadrillions of designed sequences in DNA. However, training these generative models is challenging: constraints on chemical synthesis can force many parameters into a discrete space, limiting the ability to pre-train and fine-tune. In this article we train ``free'' variational synthesis models using stochastic gradient descent in continuous space, and then discretize with post-training quantization to impose hardware and wetware constraints. This enables variational synthesis models to satisfy stringent reward criteria, while still synthesizing diverse designs, achieving a strictly dominating quality-diversity Pareto frontier. We demonstrate by training variational synthesis models of enzymes, peptides, antibody CDRH3s, and regulatory DNA elements. In silico performance is maintained in vitro.
|
| 1872 |
DF-CBM: Region-Aware Concept Bottleneck Models for Deepfake Detection
2609.35096
|
cs.LG
|
Georgios Tsoumplekas, Vazgken Vanian, Alexandros Doumanoglou, Panos K. Papadopoulos, Yannis Spyridis |
Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear....Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only partially address this need since localization-based approaches lack semantic descriptions while language-based explanation methods are only weakly grounded in visual evidence. In this work, we propose DF-CBM, a region-aware concept bottleneck model for explainable deepfake detection. DF-CBM builds a compact vocabulary of manipulation-related concepts from textual artifact annotations and links each concept to plausible facial and boundary regions. It then predicts these concepts from visual features using a concept-specific masked attention mechanism guided by parsed facial masks and the final real/fake decision is made from the predicted concept bottleneck. Our experiments show that DF-CBM outperforms concept-based baselines in concept prediction and deepfake classification while remaining competitive with state-of-the-art black-box detectors. Finally, qualitative results and intervention analyses demonstrate that DF-CBM provides spatially grounded concept evidence and enables counterfactual explanations of how individual manipulation concepts influence the final prediction. Our code is available at: https://github.com/GeorgeTsoumplekas/DF-CBM.
|
| 1873 |
Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
2609.35110
|
cs.LG
|
Yitong Li, Jincheng Yu, Junsong Chen, Haopeng Li, Shuchen Xue |
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterat...Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.
|
| 1874 |
Coordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning Approach
2609.35152
|
cs.LG
|
Guodong Ma, Baofeng Sun, Wenyu Yang, Zhihong Yao |
Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimi...Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimization with the resilience and transferability of model-based control. We propose a hybrid framework to coordinate lane-level variable speed limits (VSLs) and ramp metering across SWSs. First, we reconstruct L-METANET, a lane-level macroscopic traffic flow model that captures free and forced lane changes. Second, we combine XGBoost-SHAP with a random-parameters binary logit (RPBL) model to derive analytical equations for merging and diverging collision risks and formulate system cost and reward functions. Third, we develop MPC-STMAPPO, a hierarchical controller integrating model predictive control (MPC) and multi-agent reinforcement learning (MARL). Its upper MPC layer uses L-METANET for long-horizon rolling optimization and generates baseline commands; its lower spatiotemporal MAPPO (ST-MAPPO) layer, enhanced with Mamba cells and graph attention, produces residual actions for short-horizon adjustment. Real-world experiments on the 18-km Eastern Expressway in Changchun, China, show that L-METANET accurately reproduces lane-changing-induced flow redistribution and capacity drops, with state evolution aligned with ground truth. XGBoost-SHAP-RPBL achieves AUCs above 0.80 in most tasks, outperforming conventional logit models. MPC-STMAPPO converges faster and performs better across multiple metrics than MPC- and MARL-based baselines. Under randomly fluctuating demand, it also significantly outperforms pure MARL in generalization, demonstrating strong potential for industrial deployment.
|
| 1875 |
G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA
2609.35189
|
cs.LG
|
Jia Song (The Hong Kong University of Science and Technology), Wenhow Li (The Hong Kong University of Science and Technology), Lichen Bai (The Hong Kong University of Science and Technology), Bada Ye (Tencent), Zeke Xie (The Hong Kong University of Science and Technology) |
Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We stud...Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2.0 dimensions define data buckets and an external multimodal reward pipeline assigns sample weights. We propose G$^3$-LoRA (Gradient-Guided Grouped LoRA), a data organization procedure that probes category-level gradients induced by reward-weighted video samples, removes the shared global update direction, clusters categories by residual gradient compatibility, trains group-specific LoRA experts, and consolidates them into one adapter by weight merging followed by on-policy distillation from the experts. We motivate this procedure by viewing reward-weighted flow matching as velocity-field regression: incompatible reward dimensions may prefer different denoising directions in overlapping noisy latent regions, causing shared LoRA training to average capabilities. On Wan2.1-T2V-1.3B-Diffusers, the merged grouped adapter improves the matched VBench2.0 evaluation over the base model, a joint reward-weighted LoRA baseline, and random, semantic, and raw-gradient partitions trained with the same pipeline; an independent evaluator agrees, and on CogVideoX-2B grouping avoids the negative transfer of joint training. The gain is not uniform: merging compresses the largest specialist gains, distillation recovers part of this loss, and camera motion and several local-quality dimensions remain challenging. Together, these results suggest that gradient compatibility can serve as a practical diagnostic for organizing reward-weighted video post-training data.
|
| 1876 |
ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation
2609.35200
|
cs.LG
|
Pankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj, Zhuoyue Li |
Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We pr...Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3\% average success on LIBERO and 62.4\% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7\% average success, against 8.3\% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT
|
| 1877 |
Uncertainty Quantification in Cardiac Model Personalisation from Ultrafast Ultrasound
2609.35214
|
cs.LG
|
Camilla Ferrario (CHU Bordeaux), Maelys Venet (CHU Bordeaux), Olivier Villemain (CHU Bordeaux), Maxime Sermesant (EPIONE) |
Cardiac model personalisation requires inferring mechanical parameters that are not directly measurable in vivo. Ultrafast ultrasound shear wave elastography (SWE) enables non-invasive tracking of myocardial stiffness dynamics over the cardiac cycle, providing...Cardiac model personalisation requires inferring mechanical parameters that are not directly measurable in vivo. Ultrafast ultrasound shear wave elastography (SWE) enables non-invasive tracking of myocardial stiffness dynamics over the cardiac cycle, providing a target for personalisation. However, mapping these observations to subject specific model parameters remains ill-posed, as multiple parameter sets can reproduce the same stiffness dynamics. We formulate SWE-informed personalisation as a statistical inference problem using simulation-based inference (SBI). Using a subject-adapted 0D cardiovascular model and neural posterior estimation, we estimate model-conditional posterior distributions over active stiffness scale k0, contraction rate kATP, and relaxation rate kSR, conditioned on SWE-derived curve features and subject specific context. Among six healthy volunteers, four passed objective prior-support diagnostics and were retained for quantitative posterior analysis. Curve-level RMSE against the observed SWE target decreased from 12.61 $\pm$ 5.55 kPa for the prior predictive median to 1.14 $\pm$ 0.38 kPa for the posterior predictive median, an 89.7 $\pm$ 4.2% reduction. Posterior analysis revealed parameter-specific uncertainty, k0-kATP compensation, weaker constraint of kSR, and the importance of prior-predictive diagnostics for assessing whether each subject is represented within the modelled SWE feature space. These results support SBI for uncertainty aware SWE-based personalisation, while identifying prior support and forward-model adequacy as key diagnostics.
|
| 1878 |
A Hierarchy of Entropy-Shapley Games for Multivariate Predictive Uncertainty
2609.35217
|
cs.LG
|
Niklas Koenen, Claudia Battistin, Jeriek Van den Abeele, Martin Jullum |
Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is i...Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is important for risk-aware decisions, model diagnostics, and deciding whether the uncertainty should be mitigated or hedged against. This attribution problem requires a choice of how dependencies between output components are treated. Existing approaches reduce the output to a scalar through aggregation or projection before attribution, thereby obscuring whether features affect marginal uncertainty, dependence structure, or both, while component-wise analyses can miss dependence effects entirely. We close this gap by introducing a hierarchy of three entropy-based Shapley games that make this output-side choice explicit for any ordered multivariate outcome, ranging from per-component marginal entropy to fully joint entropy. The hierarchy isolates a cross-component attribution term that captures how each feature shifts the dependence between output components, a quantity invisible to component-wise methods. We establish a chain-rule decomposition of the joint attribution and characterize the cross-component term through conditional total correlation, providing both closed-form and sample-based estimators. Finally, we demonstrate how the framework captures differences in learned joint structure across probabilistic models from distributional regression to a zero-shot time series foundation model.
|
| 1879 |
AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
2609.35247
|
cs.LG
|
Mohamed Eltahir, Fardows Adam, Duaa M. Tahir, Lama Alamoudi, Sana Ammar |
When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce ...When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual rationale constructed from the output head. The image is cut into K row and K column bands, each shown alone to the frozen model along with the query in the format of a yes/no relevance question. The outer product of the row and column ``yes'' posteriors gives the query-conditioned spatial map. Crucially, by defining a fixed read-out R (e.g., expectation, maximum) on top of AnswerMap, we can derive continuous outputs like location natively. This bypasses the reliance on discrete text tokens for continuous-output tasks and guarantees an image-dependent answer by construction. However, a rationale can be confabulated, so we validate AnswerMap across four models and three query distributions with two tests: (a) agreement with the model's own generated point and (b) deletion of the map's region. The map lands where the model points (AUC 0.85 against 0.38 for attention), and deleting its region flips 53% of correct answers (against 19% for attention's). Beyond establishing faithfulness, we demonstrate the map's task-agnostic utility through three distinct read-outs: its maximum flags hallucinated objects without generation, its expectation localizes correctly when the model's own pointing fails, and its top-mass region, fed back as a crop, fixes half of the model's wrong answers. AnswerMap thus offers a new lens on VLM interpretability and, through its read-outs, a new output interface for visual tasks beyond text tokens.
|
| 1880 |
Scaffold Then Internalize: Representation Injection for Diffusion Transformers
2609.35292
|
cs.LG
|
Han Fu, Jiacheng Chen, Baoquan Zhao, Weidong Chen, Wei Liu |
Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direc...Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce \textit{REPresentation Injection} (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over $43.5\times$. Code will be available at https://jeneveuxpas.github.io/REPI
|
| 1881 |
Simulation-Based Inference for Plate Reverb System Identification
2609.35295
|
cs.LG
|
Dylan Sechet, Marc Evrard, Matthieu Kowalski |
We address Task A of the 1st DAFx Parameter Estimation Challenge, which aims to retrieve the physical parameters of a plate model from an impulse response. To do so, we use the Simulation-Based Inference (SBI) framework, in which we train a neural network to e...We address Task A of the 1st DAFx Parameter Estimation Challenge, which aims to retrieve the physical parameters of a plate model from an impulse response. To do so, we use the Simulation-Based Inference (SBI) framework, in which we train a neural network to estimate a density over plate parameters given an impulse response, using a dataset generated by the simulator. Inference for a new impulse response then requires only a forward pass through the network, without involving the simulator. For each test observation, we fine-tune a specific network: additional simulation rounds are performed by sampling parameters from the current estimated distribution, simulating the corresponding impulse responses, and fine-tuning to produce the specialized network.
|
| 1882 |
Convex Optimization Is Free When Accuracy Is Expensive
2609.35418
|
cs.LG
|
Arthur Paing, Arthur Jacot |
This paper studies convex optimization when the gradient cannot be evaluated exactly, but only approximated by a hierarchy of algorithms whose compute grows like $\delta^{-\gamma}$ in the accuracy $\delta$. When $\gamma>2$, falling into the Harder-Than-Mont...This paper studies convex optimization when the gradient cannot be evaluated exactly, but only approximated by a hierarchy of algorithms whose compute grows like $\delta^{-\gamma}$ in the accuracy $\delta$. When $\gamma>2$, falling into the Harder-Than-Monte-Carlo (HTMC) regime, the price of accuracy outruns the variance reduction that Monte Carlo would buy and we show that minimizing a loss function costs no more, up to a factor depending only on $\gamma$, than a single evaluation of its gradient at the accuracy the problem demands. A randomized multilevel oracle replaces the deterministic approximation of accuracy $\delta$ by an unbiased estimator of it, whose variance $\sigma^2$ becomes a second, independently priced dial: the cost of one call drops from $\delta^{-\gamma}$ to $\delta^{2-\gamma}\sigma^{-2}$. Plain inexact gradient descent driven by that oracle reaches loss $\varepsilon$ at expected compute $\Theta(\varepsilon^{-\gamma})$ in the convex case, against $\Theta(\varepsilon^{-(\gamma+1)})$ for the same method run at a fixed accuracy: randomization buys a full power of $\varepsilon$. Under $\mu$-strong convexity the exponent halves, to $\varepsilon^{-\gamma/2}$, because the iterates settle at a noise floor and the bias budget relaxes accordingly. Both bounds are independent of the step size, and hence of the smoothness constant, and we show that the cost is a functional of the underlying gradient flow rather than of any discretization of it.
|
| 1883 |
Multi-Task Learning of Conditional Mean Operators: applications to dynamical systems and uncertainty quantification
2609.35429
|
cs.LG
|
Sami Chemlal, Thibaut Germain, R\'emi Flamary, Vladimir R. Kostic, Karim Lounici |
Estimating conditional statistics and learning representations of a population of conditional distributions are central problems in many data-driven applications, including uncertainty quantification and dynamical systems analysis. Conditional mean operators (...Estimating conditional statistics and learning representations of a population of conditional distributions are central problems in many data-driven applications, including uncertainty quantification and dynamical systems analysis. Conditional mean operators (CMOs), a class of linear operators between function spaces, resolve these objectives by providing access to a broad class of conditional statistics. However, existing methods typically estimate each CMO independently or constrain it to prespecified function spaces, thereby preventing the exploitation of shared structure across related distributions. In this work, we posit that related CMOs share finite-dimensional input and output function spaces, and are specialized for each task with a linear operator mapping these spaces. Based on this hypothesis, we introduce MTL-CMO, a multi-task framework that jointly learns shared function spaces and task-specific operators across multiple datasets. We further introduce T-CMO, a transfer learning method that reuses the shared spaces to estimate, in closed form, the operator of a new conditional distribution. We establish statistical guarantees quantifying the benefits of jointly learning the shared function spaces. Our experiments demonstrate that learning shared function spaces improves uncertainty quantification across a broad range of conditional distributions and, when applied to Langevin and plasma dynamics, yields compact representations of complex dynamics that retain physically meaningful information and enable parameter identification.
|
| 1884 |
Building Transformation Layers for Riemannian Neural Networks
2609.35436
|
cs.LGcs.AI
|
Ziheng Chen |
Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications. One recent focus is the generalization of Euclidean fully connected (FC) and convolutional layers to non-Euclide...Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications. One recent focus is the generalization of Euclidean fully connected (FC) and convolutional layers to non-Euclidean geometries. However, previous approaches typically focus on a few selected manifolds and rely on specific properties of the target manifold. In contrast, this work proposes a framework for constructing FC and convolutional layers over computationally tractable Riemannian spaces. This framework incorporates several previous FC layers across different geometries as special cases and is instantiated on ten representative manifolds, including three hyperbolic models, five geometries of the symmetric positive definite (SPD) manifold, and two Grassmannian perspectives. Experiments on different manifolds demonstrate the effectiveness and applicability of our approach. Code can be found at https://github.com/GitZH-Chen/RieTrans.
|
| 1885 |
From internal representations to model improvement through prediction errors
2609.35449
|
cs.LG
|
Yushi Nakaya, Kenichi Higuchi, Shuichi Ishida |
With limited annotation budgets, choosing which images to label determines how much a model improves. Data-selection methods that use features from a separately trained model, or scene descriptions written by vision-language models, have been successful, but t...With limited annotation budgets, choosing which images to label determines how much a model improves. Data-selection methods that use features from a separately trained model, or scene descriptions written by vision-language models, have been successful, but those signals do not directly capture changes in the model being improved. The target model's own internal features reflect what it has learned so far and change with retraining, making them a natural cue for choosing the next training data. However, feature rarity alone does not reveal the errors that matter for performance. Here we link internal features to prediction errors and their expected impact on performance and select images for labeling and retraining without using labels for candidate images. We evaluated the method with an object detector on two datasets and two pairs of random seeds. Adding internal features improved the identification of prediction errors in 15 of 16 conditions. When performance was averaged over successive labeling rounds, the method outperformed selection based only on feature rarity in all four evaluation settings and ranked among the top two of six methods. With other conditions held fixed, performance after retraining was again higher than with rarity-based selection, even though the latter collected more errors. With longer retraining, the proposed method ranked first among six methods. These results suggest that linking a model's internal features to its errors and their effects on performance may help select training images that improve performance, thereby allowing the model's current state to guide which images are labeled next.
|
| 1886 |
Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching
2609.35469
|
cs.LG
|
Chenyu Zhang, Yuhang Cao, Daru Du, Yingxi Lu, Jing Shao |
Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone....Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.
|
| 1887 |
Handwritten Text Recognition Lives in the High-Pixel Variance Subspace
2609.35473
|
cs.LG
|
Carlos Garrido-Munoz, Jorge Calvo-Zaragoza |
In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel spa...In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.
|
| 1888 |
TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training
2609.35481
|
cs.LG
|
Jiacheng Zhu, Xie Zhao, Gongming Zhao, Hongli Xu, Yao Fei |
Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage...Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage and reduce overall training efficiency. Existing expert-parallelism load-balancing (EPLB) systems commonly compute load-balancing plans on the CPU, incurring device--host data transfers and cross-rank synchronization that make scheduling at every layer and microbatch expensive. Their planning formulations also overlook the hierarchical communication costs of modern scale-up and scale-out GPU clusters. We present \textit{TopoEP}, a GPU-native, topology-aware load-balancing system for large-scale MoE training. At each MoE layer and training microbatch, \textit{TopoEP} converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead. To generate these decisions, \textit{TopoEP} uses a deterministic GPU solver that performs inter-node placement followed by intra-node refinement, allowing all ranks to independently produce bitwise-identical plans. On a 32-GPU NVIDIA H800 cluster, integrating \textit{TopoEP} with Megatron-LM improves end-to-end training throughput by 6.2\%--11.4\% across three representative MoE models.
|
| 1889 |
SRHarness: A Harness for Agentic Symbolic Regression
2609.35501
|
cs.LG
|
Zihan Yu, Shixuan Zhou, Hao Huang, Jingtao Ding, Yong Li |
Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model a...Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around three mechanisms: composable scientific actions that provide a common interface over raw, transformed, and candidate-derived quantities; persistent scientific state that retains evaluated hypotheses and exposes compact model-facing views; and trajectory lifecycle management that coordinates continuation, branching, restart, and termination. On LLM-SRBench, SRHarness consistently improves both numerical generalization and symbolic recovery under matched LLM backbones. With DeepSeek-v4-flash-0731, it achieves 93.69% symbolic accuracy on LSR-Transform, compared with 62.16% for SR-Scientist, and retains 72.97% accuracy on an anonymized variant that removes scientific descriptions and variable semantics, versus 39.64% for SR-Scientist. Under the same DeepSeek-v4-flash-0731 backbone, SRHarness also substantially outperforms Codex (72.97% vs. 20.72%) and reaches performance comparable to Codex with GPT-5.5, while simply providing Codex with the same scientific tools does not reproduce this advantage. These results show that effective agentic symbolic regression depends not only on models or tools, but also on structured runtime support for organizing scientific actions, accumulated hypotheses, and long-horizon search.
|
| 1890 |
MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?
2609.35515
|
cs.LG
|
Zihan Yu, Jiadong Zhang, Jialin Cheng, Jingtao Ding, Yong Li |
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law r...Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model, a structured set of scientifically meaningful relations whose joint consequences entail an observable phenomenal law, while agents receive only observational data and scientific context. We evaluate mechanism recovery through mechanism probes, which query internal scientific consequences that cannot be inferred from the phenomenal law alone. To reduce reliance on memorized textbook mechanisms, we construct unfamiliar variants through controlled, scientifically interpretable mutations of canonical mechanisms, and screen for mechanistic indistinguishability to exclude ambiguous instances admitting comparable competing mechanisms. Experiments across representative scientific agents reveal a substantial phenomenal--mechanism recovery gap: for Codex with GPT-5.6-sol, phenomenal-law accuracy reaches 35.00% on the Core-set while mechanism accuracy is only 13.75%, with mechanism recovery failing in 64.29% of cases where the phenomenal law is correctly recovered. The gap widens as mechanisms become increasingly mutated, and even providing the correct phenomenal law leaves mechanism recovery below 50%. These results reveal a substantial generalization gap in mechanistic reasoning and establish mechanism discovery as a distinct challenge beyond recovering observable scientific laws.
|
| 1891 |
Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions
2609.35569
|
cs.LG
|
Tianyao Shi, Xipeng Shen, Yi Ding |
Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. Yet these dimensions are largely evaluated in isolation, leaving it unclear when and how they lead to different o...Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. Yet these dimensions are largely evaluated in isolation, leaving it unclear when and how they lead to different optimization decisions. We present PRISM, a unified framework for characterizing and optimizing LLM serving across energy, carbon, water, and biodiversity impacts. Our analysis reveals a fundamental distinction: computing configurations determine energy consumption, whereas where and when LLM serving is deployed determine its carbon, water, and biodiversity impacts. Under a fixed deployment choice and operational-only accounting, all dimensions preserve the same energy-based configuration ranking. Deployment rankings can diverge across dimensions, while embodied impacts can break configuration invariance when they exceed a lifecycle crossover boundary. PRISM identifies these conditions, quantifies cross-dimensional regrets, and balances the four dimensions. In regional-routing experiments, PRISM reduces median worst-case regret by 50.2% relative to the strongest baseline.
|
| 1892 |
QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks
2609.35581
|
cs.LG
|
Pranav Gupta |
We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x ...We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.
|
| 1893 |
Learning Conditional Expectation Operators via Functional Newton Updates
2609.35598
|
cs.LG
|
Thiago Ramos, Alek Fr\"ohlich, Daniel Perazzo, Massimiliano Pontil |
We introduce the Functional Spectral-Newton Method (FSNM) for learning the leading singular structure of a conditional expectation operator without fixing a basis or reproducing kernel Hilbert space. FSNM fits a low-rank representation of the centered joint-to...We introduce the Functional Spectral-Newton Method (FSNM) for learning the leading singular structure of a conditional expectation operator without fixing a basis or reproducing kernel Hilbert space. FSNM fits a low-rank representation of the centered joint-to-product density ratio kernel by alternating functional Newton updates. Each update reduces to a preconditioned regression, which we approximate with vector-valued regression trees in a stagewise boosting procedure. At the population level, we establish descent and an $O(1/T)$ best-iterate block-stationarity rate under a relative weak-learner accuracy condition, and show that every nondegenerate local minimum over the full centered $L^2$ spaces is a globally optimal rank-$d$ approximation. Synthetic experiments show that FSNM recovers a low-rank density ratio and its leading spectral structure, and that the same learned kernel can answer multiple conditional queries without refitting.
|
| 1894 |
On-Policy Self-Distillation for Multi-Turn Image Editing
2609.35611
|
cs.LG
|
Liangbing Zhao, Le Zhuo, Mohamed Elhoseiny |
Instruction-based image editing has achieved strong performance in single-turn settings, yet practical editing is often iterative, with each instruction applied to the output of the previous turn. We find that existing editing models degrade rapidly under recu...Instruction-based image editing has achieved strong performance in single-turn settings, yet practical editing is often iterative, with each instruction applied to the output of the previous turn. We find that existing editing models degrade rapidly under recursive editing and attribute this failure to a train-test mismatch in the conditioning distribution: models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference time. To address this, we propose MT-OPSD, an on-policy self-distillation framework that trains the model on self-generated conditioning states with editing supervision from a clean-conditioned teacher, without requiring multi-turn annotations. We further introduce LME-Bench, a benchmark of 100 ten-turn editing sessions for evaluating long-horizon robustness. Experiments across three editing backbones show that MT-OPSD substantially improves long-horizon editing success and reduces multi-turn collapse while largely preserving single-turn editing quality.
|
| 1895 |
Elicitation and Decision Geometry in Single-Index Bandits
2609.35622
|
cs.LG
|
Sakshi Arya, Cheng Soon Ong |
We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. ...We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (NBL), a greedy procedure that uses a sequential Stein contrast to learn the optimal boundary directly, without estimating the reward functions or the common link. We characterize the local Riemannian dynamics of NBL through a decision stability coefficient balancing arm separation, link geometry, and the context distribution. We show that this stability is connected to the elicitation geometry of the underlying convex potential. Under local decision stability, NBL contracts toward the optimal boundary and achieves $O(\log n)$ expected regret. Numerical experiments illustrate the predicted stability regimes and compare NBL with a parametric greedy benchmark under link misspecification.
|
| 1896 |
RIDE: Reference-Anchored Inference-Time Diffusion Editing for Scaffold Hopping
2609.35623
|
cs.LG
|
Ruoxi Gao, Frazier N. Baker, Trieu Nguyen, Xia Ning |
Scaffold hopping is a critical task in drug discovery, which seeks to discover new, structurally distinct molecules that share key functional groups and similar 3D shape with a reference binding ligand. Existing diffusion-based scaffold hopping methods formula...Scaffold hopping is a critical task in drug discovery, which seeks to discover new, structurally distinct molecules that share key functional groups and similar 3D shape with a reference binding ligand. Existing diffusion-based scaffold hopping methods formulate the problem as conditional generation of scaffolds given the functional groups. However, they lack a principled mechanism to jointly enforce 2D structural novelty and preserve the 3D shape of the reference ligand. Here, we introduce RIDE, a Reference-anchored Inference-time Diffusion Editing framework for scaffold hopping. RIDE recovers the reference diffusion noise trajectory conditioned on the binding pocket and functional groups, selects an optimal trajectory segment for editing via noise perturbation, and conducts a value-guided scaffold sampling to generate new scaffolds. Extensive experimental results demonstrate that, compared to baselines, RIDE consistently generates scaffolds with lower 2D similarity and higher 3D similarity to the reference, with an average improvements of 11.7% and 7.3%, respectively. Further analysis reveals that RIDE can accommodate various reward functions, and can preserve 3D similarity even when this is not explicitly included in the reward. Two case studies illustrate RIDE's ability to generate distinct scaffolds with different structures and properties, and its ability to introduce substantial 2D variation while maintaining very high 3D similarity. RIDE is publicly available at https://anonymous.4open.science/r/RIDE-C8A0.
|
| 1897 |
Learned Preconditioning for a Primal-Dual Interior-Point Method
2609.35665
|
cs.LG
|
Abhinav Madabhushi, Jialin Liu, Minxin Zhang |
Interior-point methods (IPMs) are among the most widely used algorithms for constrained optimization, yet their Newton-based search directions require costly second-order information and large linear-system solves. Learning to optimize offers cheaper updates l...Interior-point methods (IPMs) are among the most widely used algorithms for constrained optimization, yet their Newton-based search directions require costly second-order information and large linear-system solves. Learning to optimize offers cheaper updates learned from data, but the singular behavior of logarithmic barriers near constraint boundaries makes IPMs highly sensitive to perturbations, complicating both warm starting and learning reliable updates. We introduce pdLIP, an IPM for smooth nonlinear programs that integrates learned preconditioning with pdProj, an all-shifted primal-dual projected-search IPM. A shared coordinate-wise recurrent network predicts a positive diagonal preconditioner that scales the right-hand side of the reduced Newton system for the primal step, and the remaining slack and multiplier directions are recovered analytically. The learned iterations avoid Hessian evaluations and Newton-system solves, using only first-order and coordinate-wise operations amenable to GPU parallelization. Training is self-supervised, with a loss based on a penalty-barrier merit function and the residual of perturbed optimality conditions, requiring neither target directions nor precomputed solutions. Primal and dual shifts mitigate the barrier's sensitivity to perturbations near constraint boundaries, enabling effective warm starting. Across four classes of 200-dimensional convex and nonconvex constrained problems, pdLIP warm starts reduce pdProj refinement iterations by 63-67% compared with cold starts at the same KKT residual tolerance of $10^{-8}$, with negligible warm-start generation cost relative to the subsequent pdProj solve. Improvements persist on box-constrained QPs with 1000 variables and extend to applications including portfolio optimization, support vector machines, and a nonlinear control example.
|
| 1898 |
The Hidden Perception Constraint in Task-Aware Compression
2609.35684
|
cs.LG
|
Sahan Liyanaarachchi, Semih Akkoc, Sennur Ulukus, Aylin Yener |
With the recent advancements of neural compressors, explicitly incorporating perception constraints into the design of compression schemes has gained significant attention. Traditionally, these perception constraints ensure that the distribution of the reconst...With the recent advancements of neural compressors, explicitly incorporating perception constraints into the design of compression schemes has gained significant attention. Traditionally, these perception constraints ensure that the distribution of the reconstruction does not significantly deviate from the distribution of the source, thus attesting to the perceptual quality of the reconstruction. In this work, we uncover several perception constraints that are naturally present in task-aware compression. In particular, we consider a problem where the primary task is reconstruction and the secondary task is classification (i.e., a statistical test). We study this problem at varying levels of domain information available to us and discuss how to utilize the naturally emerging perception constraints to design rate-minimal compression schemes that also maximize the utility of our secondary task. We show that in this setting, if the decision boundaries of the classifier are ill-defined (mismatch) for our source distribution, then matching onto a target distribution enhances our classification accuracy.
|
| 1899 |
Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
2609.35745
|
cs.LG
|
Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Shiqi Huang, Min Kass Chong |
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax v...Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.
|
| 1900 |
Statistical Learning of Contractive Dynamical Representations for Composite Adaptive Control
2609.35758
|
cs.LG
|
Min Kim, Jos\'e Leonardo Brenes, Fred Hadaegh, Soon-Jo Chung |
We present a representation-learning framework for composite adaptive tracking control under dynamically coupled disturbances. The framework connects classical disturbance-accommodating control (DAC) to recent last-layer adaptive disturbance-rejection methods....We present a representation-learning framework for composite adaptive tracking control under dynamically coupled disturbances. The framework connects classical disturbance-accommodating control (DAC) to recent last-layer adaptive disturbance-rejection methods. Specifically, we introduce a statistically principled hard expectation-maximization (hard-EM) procedure, with a Kalman smoother in the hard E-step, to identify dynamical representations of disturbance whose latent evolution is uniformly contractive. The learned representation evolves a latent disturbance-excitation state from measured plant features and control inputs and decodes that state into the time-varying disturbance acting on the nominal plant, thereby extending prior "fixed-decay" last-layer adaptive methods to a learned, predictive DAC-style formulation. Combined with Bayesian filtering of the learned latent state, this representation yields a composite adaptive tracking controller with predictive capability and provable exponential convergence to a bounded neighborhood. We validate our approach experimentally on a slippery ground vehicle carrying a liquid-sloshing tank and a pendulum load, and we further assess its robustness on a system of coupled Duffing oscillators. Across both settings, the method achieves accurate disturbance prediction and improved overall tracking performance relative to fixed-decay representation-learning ablations, LTI disturbance-accommodating baselines, and model-based PD baselines.
|
| 1901 |
PDMD: Projected Distribution Matching Distillation for Video Diffusion Models
2609.35768
|
cs.LG
|
Zimo Wang, Junkun Yuan, Angtian Wang, Haotian Yang, Canyu Zhang |
Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during train...Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.
|
| 1902 |
Advanced Policies: A First-Principles Path from Policy Gradient to Q-Learning
2005.08844
|
cs.LG
|
Donghoon Lee |
We apply entropy to reinforcement learning in two ways: entropy-augmented reward defines the soft objective, while relative-entropy regularization controls the size of a policy-improvement step without changing that objective. Together they yield a one-paramet...We apply entropy to reinforcement learning in two ways: entropy-augmented reward defines the soft objective, while relative-entropy regularization controls the size of a policy-improvement step without changing that objective. Together they yield a one-parameter family of advanced policies connecting a base policy to its soft-greedy policy. Under exact evaluation, every nontrivial member improves upon the common base, although the improvement need not be monotone along the family. Under policy-value consistency, this family forms an entropic mirror-descent path. We then relax this consistency, treating the policy and action-value function as independent coordinates of the advanced policy. Differentiating the objective of the advanced policy yields Advanced Actor-Critic (AAC), whose endpoint gradients recover soft policy gradient and an action-centered Q-learning-like update. We develop implementations for discrete and continuous actions. Experiments demonstrate effective learning at both endpoints and intermediate parameter values.
|
| 1903 |
Improving Causal Effect Estimation of Weighted RegressionBased Estimator using Neural Networks
2110.15075
|
cs.LG
|
Plabon Shaha, Talha Islam Zadid, Ismat Rahman, Md. Mosaddek Khan |
Estimating causal effects from observational data informs us about which factors are important in an autonomous system, and enables us to take better decisions. This is important because it has applications in selecting a treatment in medical systems or making...Estimating causal effects from observational data informs us about which factors are important in an autonomous system, and enables us to take better decisions. This is important because it has applications in selecting a treatment in medical systems or making better strategies in industries or making better policies for our government or even the society. Unavailability of complete data, coupled with high cardinality of data, makes this estimation task computationally intractable. Recently, a regression-based weighted estimator has been introduced that is capable of producing solution using bounded samples of a given problem. However, as the data dimension increases, the solution produced by the regression-based method degrades. Against this background, we introduce a neural network based estimator that improves the solution quality in case of non-linear and finitude of samples. Finally, our empirical evaluation illustrates a significant improvement of solution quality, up to around $55\%$, compared to the state-of-the-art estimators.
|
| 1904 |
FedSysID: A Federated Approach to Sample-Efficient System Identification
2211.14393
|
cs.LG
|
Han Wang, Leonardo F. Toso, James Anderson |
We study the problem of learning a linear system model from the observations of $M$ clients. The catch: Each client is observing data from a different dynamical system. This work addresses the question of how multiple clients collaboratively learn dynamical mo...We study the problem of learning a linear system model from the observations of $M$ clients. The catch: Each client is observing data from a different dynamical system. This work addresses the question of how multiple clients collaboratively learn dynamical models in the presence of heterogeneity. We pose this problem as a federated learning problem and characterize the tension between achievable performance and system heterogeneity. Furthermore, our federated sample complexity result provides a constant factor improvement over the single agent setting. Finally, we describe a meta federated learning algorithm, FedSysID, that leverages existing federated algorithms at the client level.
|
| 1905 |
Training-Free Uncertainty Estimation for Embedding Models
2306.00206
|
cs.LG
|
Young-Jin Park, Addison Kristanto Julistiono, Hao Wang, Shervin Ardeshir, Navid Azizan |
Embedding models, often obtained via self-supervised learning, extract general-purpose representations from data. Quantifying the reliability of these representations is crucial, as many downstream models rely on them as input for their own tasks. To this end,...Embedding models, often obtained via self-supervised learning, extract general-purpose representations from data. Quantifying the reliability of these representations is crucial, as many downstream models rely on them as input for their own tasks. To this end, we introduce a formal definition of representation reliability: the representation for a given test point is considered to be reliable if the downstream models built on top of that representation can, on average, consistently generate accurate predictions for that test point across various downstream tasks. However, accessing the downstream data to quantify the representation reliability is often limited or restricted for various reasons. We propose training-free methods for estimating the representation reliability without access to the downstream data. Our method is based on the concept of neighborhood consistency (NC) across distinct pre-trained representation spaces. The key insight is to find shared neighboring points as anchors to align these representation spaces before comparing them. We provide theoretical justifications for NC and develop two practical approaches: (1) directly computing NC when multiple pre-trained models are available, and (2) a perturbation-based NC (PNC), which creates synthetic ensembles from a single model through isotropic Gaussian noise, avoiding the computational cost of training deep ensembles. We further propose PNC-spread tuning, which systematically determines the perturbation magnitude by maximizing the spread of the PNC scores on a reference set. We demonstrate through comprehensive numerical experiments that our methods effectively capture the representation reliability with a high degree of correlation, achieving robust and favorable performance compared with baseline methods.
|
| 1906 |
Classifier-pruned Bayesian optimization for particle accelerator tuning: Exploring temporally structured manifold of 6D beam phase space
2412.01748
|
cs.LG
|
Mahindra Rautela, Alan Williams, Alexander Scheinker |
Complex dynamical systems, such as particle accelerators, require tuning over high-dimensional, nonlinear state and parameter spaces while experimental measurements and high-fidelity simulations can be expensive. Learned latent representations provide a compac...Complex dynamical systems, such as particle accelerators, require tuning over high-dimensional, nonlinear state and parameter spaces while experimental measurements and high-fidelity simulations can be expensive. Learned latent representations provide a compact domain for such optimization, but poorly supported regions of the learned manifold may decode into unrealistic physical states and yield deceptively favorable objectives. To address this challenge, we propose the Classifier-pruned Bayesian Optimization-based Latent-space Tuner (CBOL-Tuner), which performs Bayesian optimization over a temporally structured latent representation of 6D beam phase-space dynamics. The CBOL-Tuner integrates a conditional variational autoencoder for latent space representation, a long short-term memory network for temporal dynamics, a lightweight neural network for parameter estimation, and a classifier-pruned Bayesian optimizer to adaptively search and filter the latent space for optimal solutions. This framework enables feasibility-aware exploration of learned scientific representations for accelerator tuning.
|
| 1907 |
Goal-Conditioned Supervised Learning for Multi-Objective Recommendation
2412.08911
|
cs.LG
|
Shijun Li, Hilaf Hasson, Jing Hu, Joydeep Ghosh |
Multi-objective learning endeavors to concurrently optimize multiple objectives using a single model, aiming to achieve high and balanced performance across diverse objectives. However, this often entails a complex optimization problem to balance the learning ...Multi-objective learning endeavors to concurrently optimize multiple objectives using a single model, aiming to achieve high and balanced performance across diverse objectives. However, this often entails a complex optimization problem to balance the learning of potentially conflicting objectives, leading to solutions with higher memory requirements and computational complexity. This paper introduces a Multi-Objective Goal-Conditioned Supervised Learning (MOGCSL) framework for automatically learning to achieve multiple objectives from offline sequential data. MOGCSL extends the conventional GCSL method to multi-objective scenarios by redefining goals from one-dimensional scalars to multi-dimensional vectors. It benefits from naturally eliminating the need for complex architectures and optimization constraints. Moreover, MOGCSL inherently disentangles uninformative or noisy training instances that fail to achieve desirable long-term rewards across multiple objectives. We also introduce a novel goal-selection algorithm for MOGCSL to model and identify desired and achievable goals for inference. In this paper, we focus on its application to the next action prediction problem in commercial-grade recommender systems. In this context, any viable solution needs to be reasonably scalable and also be robust to large amounts of noisy data that is characteristic of this application space. We show that MOGCSL performs admirably on both counts by extensive experiments. Also, analysis and experiments are included to explain its strength in discounting the noisier portions of training data in recommender systems with multiple objectives.
|
| 1908 |
Stochastic Engrams for Efficient Continual Learning
2503.21436
|
cs.LG
|
Isabelle Aguilar, Luis Fernando Herbozo Contreras, Omid Kavehei |
The ability to learn continuously in artificial neural networks (ANNs) is often limited by catastrophic forgetting, a phenomenon in which new knowledge becomes dominant. By taking mechanisms of memory encoding in neuroscience (i.e., engrams) as inspiration, we...The ability to learn continuously in artificial neural networks (ANNs) is often limited by catastrophic forgetting, a phenomenon in which new knowledge becomes dominant. By taking mechanisms of memory encoding in neuroscience (i.e., engrams) as inspiration, we propose a novel approach that integrates stochastically-activated engrams as a gating mechanism for metaplastic binarized neural networks (mBNNs). This method leverages the computational efficiency of mBNNs combined with the robustness of probabilistic memory traces to mitigate forgetting and maintain the model's reliability. Previously validated metaplastic optimization techniques have been incorporated to further enhance synaptic stability. Compared to baseline binarized models and benchmark fully connected continual learning approaches, our method is the only strategy capable of achieving average accuracies over 70% in both class-incremental and domain-incremental MNIST benchmarks, matching full-precision state-of-the-art methods. Furthermore, we achieve a significant reduction in peak GPU and RAM usage, under 5% and 20%, respectively, as well as an ~8x reduction in memory footprint compared to full precision counterparts. Our findings demonstrate (A) an improved stability vs. plasticity trade-off, (B) reduced memory intensiveness, and (C) enhanced performance in binarized architectures. By uniting principles of neuroscience and efficient computing, we offer new insights into the design of scalable and robust deep learning systems.
|
| 1909 |
DRAN: A Distribution and Relation Adaptive Network for Spatio-temporal Forecasting
2504.01531
|
cs.LG
|
Xiaobei Zou, Luolin Xiong, Kexuan Zhang, Cesare Alippi, Yang Tang |
Spatio-temporal forecasting remains challenging under non-stationary environments because both data distributions and spatial relations evolve over time. Temporal normalization and de-normalization are widely used to mitigate distribution shifts, but they may ...Spatio-temporal forecasting remains challenging under non-stationary environments because both data distributions and spatial relations evolve over time. Temporal normalization and de-normalization are widely used to mitigate distribution shifts, but they may distort inter-node relationships and thereby impair spatial dependency modeling. To address these issues, we propose the Distribution and Relation Adaptive Network (DRAN) for spatio-temporal forecasting. DRAN incorporates a Spatial Factor Learner (SFL) module, which enables effective normalization and de-normalization while preserving spatial dependencies in spatio-temporal systems. To model evolving spatial interactions, DRAN further proposes the Dynamic-Static Fusion Learner (DSFL) module. DSFL decomposes features into static and dynamic components and adaptively fuses them according to input variability. Experiments on six benchmark datasets show that DRAN outperforms state-of-the-art baselines. Additional analyses demonstrate that SFL consistently reduces spatial-relation distortion across multiple normalization schemes, whereas DSFL captures complementary static and dynamic dependencies and adjusts their contributions according to temporal variability.
|
| 1910 |
AYLA: Architecting a loss landscape in shallow neural networks to accelerate feature recovery
2504.01875
|
cs.LG
|
Behnam Gheshlaghi, Shahin Atakishiyev |
Feature learning in shallow neural networks exhibits rich yet fragile dynamics, including prolonged plateaus, abrupt phase transitions, and sensitivity to optimization hyperparameters. While recent theoretical work has characterized these behaviors through the...Feature learning in shallow neural networks exhibits rich yet fragile dynamics, including prolonged plateaus, abrupt phase transitions, and sensitivity to optimization hyperparameters. While recent theoretical work has characterized these behaviors through the geometry of loss landscapes, saddle escape mechanisms, and emergent scaling laws, practical methods for actively shaping these dynamics remain limited. In this paper, we introduce AYLA, a principled loss reparameterization framework that dynamically modulates gradient magnitudes during training without altering the location of stationary points or optimal solutions. AYLA applies a smooth, sigmoid-controlled power-law transformation to empirical loss, yielding a state-dependent effective learning rate that accelerates descent in flat or saddle-dominated regions while stabilizing late-stage optimization. Crucially, AYLA preserves all critical points of the original objective, acting solely as a monotone transformation that reshapes optimization trajectories rather than objectives. We evaluate AYLA in controlled teacher student settings using two-layer tanh networks trained on synthetic Gaussian data. Across stochastic gradient descent and multiple loss-exponent schedules, AYLA consistently improves feature recovery. This evidence is observed in terms of weight alignment, per-neuron cosine similarity, hidden-activation correlation, and spectral properties of learned representations, while AYLA maintains competitive or faster loss convergence. Spectral analyses further demonstrate that AYLA mitigates rank collapse and promotes richer internal representations, signaling a transition from lazy to active feature-learning regimes. AYLA offers a lightweight, theoretically grounded way to improve shallow-network optimization, especially in resource-limited or noise-sensitive settings.
|
| 1911 |
Focus on Likely Classes for Test-Time Prediction
2505.03819
|
cs.LG
|
Johannes Schneider |
We ask: Can focusing on likely classes of a single, in-domain sample improve accuracy? Prior work argued "no", we answer "yes" on average. Standard entropy minimization yields largest gains, and we further dissect why by looking at its two core mechanisms: inc...We ask: Can focusing on likely classes of a single, in-domain sample improve accuracy? Prior work argued "no", we answer "yes" on average. Standard entropy minimization yields largest gains, and we further dissect why by looking at its two core mechanisms: increasing likely and decreasing unlikely class predictions. Simply maximizing logits of the two most likely classes explains most of its benefits and, in some cases, even outperforms entropy minimization. Thus, decreasing classes is a secondary concern. Our controlled experiment and theory investigate the most puzzling case, why simply optimizing the logits of the two most likely classes leads to accuracy gains. Our small scale networks suggest that the degree of uncertainty of a prediction, followed by model choice, and to a lesser extent dataset choice moderate the success chances of the optimization. We show that the true class tends to have larger gradients independent of whether the prediction is correct or not. Our theory thoroughly investigates one of multiple possible explanations from the angle of shared features. Our evaluation focuses on 26 large language models and 12 text datasets, and 18 image recognition models trained on ImageNet. It demonstrates gains of 0.6 percent on average for entropy minimization on uncertain samples using no extra data aside from the given test sample using a fixed learning rate. We also suggest gradient-ray ensembling. When traversing input/embedding space by applying the same gradient with differing learning rates and aggregating probabilities for the same sample further improves accuracy significantly to 0.75 percent relying on a coarse range for learning rates, thus improving accuracy and eliminating the need to find a precise learning rate at the expense of computation.
|
| 1912 |
LOD: Latent Objective Discovery in Heterogeneous Multi-Agent Reinforcement Learning
2505.09756
|
cs.LG
|
Ke Sun, Zhaoyang Shi |
In multi-agent reinforcement learning (MARL) settings, agents are often heterogeneous with diverse intrinsic utility functions. In this study, we introduce Latent Objective Discovery (LOD) as a framework for decomposing agents' distinct utilities into mixed un...In multi-agent reinforcement learning (MARL) settings, agents are often heterogeneous with diverse intrinsic utility functions. In this study, we introduce Latent Objective Discovery (LOD) as a framework for decomposing agents' distinct utilities into mixed underlying objectives in a low-dimensional latent space. Latent objective discovery aims to capture flexible and abstract coordination patterns among agents and each abstract objective is associated with a shared value function. During learning, all the learned values are aggregated by each agent according to their personalized weights learned in an adaptive manner. Therefore, agents inherit objective-level estimates for policy updates and value learning, enabling structured information sharing without requiring access to other agents' policies. Algorithmically, we incorporate our approach into Multi-Agent Proximal Policy Optimization (MAPPO) to exploit this structure. Theoretically, we establish convergence guarantees under linear function approximation within the actor-critic framework. Empirically, we extensively validate the advantage of introducing latent objective discovery in popular multi-agent testbeds with heterogeneous agents.
|
| 1913 |
Bidirectional Information Flow (BIF) - A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization
2505.11294
|
cs.LG
|
Juan D. Guerra, Thomas Garbay, Numa Dancause, Guillaume Lajoie, Marco Bonizzato |
Hierarchical Gaussian Process (H-GP) models divide problems into different subtasks, allowing different components to address each part, making them well-suited for problems with inherent compositional structure. However, existing H-GP frameworks typically emp...Hierarchical Gaussian Process (H-GP) models divide problems into different subtasks, allowing different components to address each part, making them well-suited for problems with inherent compositional structure. However, existing H-GP frameworks typically employ one-way information sharing - either top-down or bottom-up - which limits sample efficiency and slows convergence. We propose Bidirectional Information Flow (BIF), which establishes continuous two-way communication. BIF retains the modular structure of hierarchical models-the parent conditions its own posterior on child summaries, treating them as structured priors-while introducing top-down feedback to softly decompose environment observations from the parent into sub-responses. This mutual exchange improves sample efficiency, enables robust training, and allows modular reuse of learned subtask models. We prove analytically that the regret of a GP with a learned kernel scales linearly with the mismatch to the true kernel, tightening in the hierarchical case to the sum of child-level errors. Ablation shows that removing the downward pathway collapses child $R^2$ by up to 58%. Across synthetic, neurostimulation, and HPO benchmarks, BIF achieves up to $4\times$ higher parent $R^2$ and $\sim 100\%$ AUC improvement over vanilla GPBO, and outscores all hierarchical state-of-the-art methods on child $R^2$ given the correct acquisition function, while supporting modular child transfer to novel composite tasks.
|
| 1914 |
Understanding and Mitigating Under-Confidence in GNNs from the Final Layer
2505.11335
|
cs.LG
|
Jincheng Huang, Jie Xu, Xiaoshuang Shi, Ping Hu, Lei Feng |
Graph Neural Networks (GNNs) have demonstrated remarkable effectiveness on graph-based tasks. However, their predictive confidence is often miscalibrated, typically exhibiting under-confidence, which harms the reliability of their decisions. Existing calibrati...Graph Neural Networks (GNNs) have demonstrated remarkable effectiveness on graph-based tasks. However, their predictive confidence is often miscalibrated, typically exhibiting under-confidence, which harms the reliability of their decisions. Existing calibration methods for GNNs normally introduce additional calibration components, which fail to capture the intrinsic relationship between the model and the prediction confidence, resulting in limited theoretical guarantees and increased computational overhead. To address this issue, we propose a simple yet efficient graph calibration method. We establish a unified theoretical framework revealing that model confidence is jointly governed by class-centroid-level and node-level calibration at the final layer. Based on this insight, we theoretically show that reducing the weight decay of the final-layer parameters alleviates GNN under-confidence by acting on the class-centroid level, while node-level calibration acts as a finer-grained complement to class-centroid-level calibration, which encourages each test node to be closer to its predicted class prototype in the final-layer representations. Extensive experiments validate the superiority of our method. The code is released at https://github.com/huangJC0429/SCAR.
|
| 1915 |
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
2505.11840
|
cs.LG
|
Huan Li, Yiming Dong, Zhouchen Lin |
As the default optimizer for training large language models, AdamW has achieved remarkable success in deep learning. However, its convergence behavior is not theoretically well-understood. This paper establishes the convergence rate $\frac{1}{K}\sum_{k=1}^KE\l...As the default optimizer for training large language models, AdamW has achieved remarkable success in deep learning. However, its convergence behavior is not theoretically well-understood. This paper establishes the convergence rate $\frac{1}{K}\sum_{k=1}^KE\left[||\nabla f(x^k)||_1\right]\leq O(\frac{\sqrt{d}C}{K^{1/4}})$ for AdamW measured by $\ell_1$ norm, where $K$ represents the iteration number, $d$ denotes the model dimension, and $C$ matches the constant in the optimal convergence rate of SGD. Theoretically, we have $||\nabla f(x)||_2\ll ||\nabla f(x)||_1\leq \sqrt{d}||\nabla f(x)||_2$ for any high-dimensional vector $x$ and $E\left[||\nabla f(x)||_1\right]\geq\sqrt{\frac{2d}{\pi}}E\left[||\nabla f(x)||_2\right]$ when each element of $\nabla f(x)$ is generated from Gaussian distribution $\mathcal N(0,1)$. Empirically, our experimental results on real-world deep learning tasks reveal $||\nabla f(x)||_1=\varTheta(\sqrt{d})||\nabla f(x)||_2$. Both support that our convergence rate can be considered to be analogous to the optimal $\frac{1}{K}\sum_{k=1}^KE\left[||\nabla f(x^k)||_2\right]\leq O(\frac{C}{K^{1/4}})$ convergence rate of SGD in the ideal case. We also extend our result to NAdamW, an AdamW variant that employs a double-momentum mechanism, and demonstrate that it maintains the same convergence rate.
|
| 1916 |
Relation-Aware Graph Foundation Model
2505.12027
|
cs.LG
|
Jianxiang Yu, Jiapeng Zhu, Yibo Zhao, Hao Qian, Ziqi Liu |
In recent years, large language models (LLMs) have demonstrated remarkable capability to generalize across diverse natural language processing tasks, inspiring the development of graph foundation models (GFMs) for large-scale pre-training. However, unlike lang...In recent years, large language models (LLMs) have demonstrated remarkable capability to generalize across diverse natural language processing tasks, inspiring the development of graph foundation models (GFMs) for large-scale pre-training. However, unlike language models with explicit token units, graphs lack a well-defined unit for generalization, making it challenging to design effective pre-training strategies. In this work, we propose REEF, a novel GFM framework that leverages relation tokens as the fundamental units. We construct a vocabulary of relation tokens to encode relational information within graphs. To accommodate diverse relations, we introduce two hypernetworks that adaptively generate the parameters of aggregators and classifiers in graph neural networks based on relation tokens. In addition, we design another hypernetwork to construct dataset-specific projectors and incorporate a dataset-level feature bias into the initial node representations, enhancing flexibility across different datasets with the same relation. Extensive experiments demonstrate that REEF consistently outperforms existing methods in both pre-training and transfer learning, highlighting its potential as a general-purpose graph foundation model.
|
| 1917 |
Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning
2505.15064
|
cs.LG
|
Sho Sonoda, Yuka Hashimoto, Isao Ishikawa, Masahiro Ikeda |
We ask when adding hidden layers improves generalization, in a model that keeps the layers fixed and varies only their number. A hidden layer is a self-map of a state space, a depth-$k$ network composes at most $k$ hidden layers with an output layer, and depth...We ask when adding hidden layers improves generalization, in a model that keeps the layers fixed and varies only their number. A hidden layer is a self-map of a state space, a depth-$k$ network composes at most $k$ hidden layers with an output layer, and depth is compared within the nested family $H_0\subset H_1\subset\cdots$ built from one class $F$ of hidden layers; this compares a deep network with shallower networks built from the same layers, not with wider ones. Our message is that the statistical cost of depth is the metric entropy of the set $B(k,F)$ of compositions. The estimation error is bounded by an entropy integral over $B(k,F)$ with constants that do not depend on the depth, and the bound is matched from below when the output layer can see the hidden states. For Lipschitz layers on a bounded state space this entropy grows at most polynomially in $k$, and it stays bounded, or grows only like $\log k$, under contraction, equicontinuity, or nilpotent structure. Balanced against the approximation error, this gives depths $k^\ast(n)$ that grow with the sample size, and it separates the models whose estimation error is independent of the depth from those whose estimation error grows with it. Deep ReLU networks, unrolled solvers, and chain-of-thought computation are worked out; for the last two the number of steps is derived rather than assumed. All statements are machine-checked in Lean~4; the formalization and its blueprint are available at https://shosonoda.github.io/lean-deepgen/ .
|
| 1918 |
HERO: Preserving Structure and Semantics in Heterogeneous Continual Graph Learning
2505.17458
|
cs.LG
|
Guiquan Sun, Xikun Zhang, Jingchao Ni, Dongjin Song |
Machine learning on heterogeneous graphs has advanced rapidly, driven by the diverse entities, relations, and semantics inherent in real-world data. However, most existing studies assume static graphs, whereas real-world graph data are often continuously updat...Machine learning on heterogeneous graphs has advanced rapidly, driven by the diverse entities, relations, and semantics inherent in real-world data. However, most existing studies assume static graphs, whereas real-world graph data are often continuously updated. Continual learning in this setting is particularly challenging because historical knowledge is encoded not only in target-node representations, but also in their heterogeneous structural contexts and relation-dependent semantic responses. To this end, we introduce HERO, a heterogeneous continual graph learning framework that preserves historical knowledge at both structural and semantic levels while adapting to incoming tasks. We analyze experience replay as an approximation to the unavailable historical gradient and decompose its error into target-selection and context-reconstruction terms. This analysis shows that replaying target nodes alone is generally insufficient for relation-sensitive HGNNs, as removing typed neighborhood context can alter both representations and historical gradients. Motivated by this, HERO employs DiSCo to select representative target nodes and reconstruct compact heterogeneous contexts through relation-aware multi-type neighbor expansion. To further preserve knowledge that cannot be captured by replayed subgraphs alone, HERO aligns task-conditioned predictions and heterogeneous semantic responses through multi-level knowledge distillation. A lightweight look-ahead adaptation step is additionally used to improve plasticity on incoming tasks. Experiments on four datasets with three HGNN backbones show that HERO achieves the best or tied-best average performance among state-of-the-art baselines, while maintaining competitive forgetting. Our code is available at https://github.com/gqBond/HCGL.
|
| 1919 |
Federated Independent Component Analysis via Spectral Alignment and Robust Aggregation
2505.20532
|
cs.LG
|
Xin Bing, Dian Jin, Yuqian Zhang |
This paper studies robust estimation for Independent Component Analysis (ICA) in the federated learning setting, where data are locally distributed across clients and may exhibit substantial heterogeneity. The goal is to recover a common global mixing matrix b...This paper studies robust estimation for Independent Component Analysis (ICA) in the federated learning setting, where data are locally distributed across clients and may exhibit substantial heterogeneity. The goal is to recover a common global mixing matrix by aggregating local estimators computed by individual clients. The main difficulty is that local ICA estimators are identifiable only up to sign flips and column permutations and may have highly heterogeneous estimation quality. We propose a novel three-step aggregation method that first aligns the permutations across all local estimators via a particular spectral clustering approach, then aligns the signs within each estimated cluster via another tailored spectral approach, and finally applies the geometric median for robust aggregation. The proposed estimator is shown to remain accurate even when a substantial fraction of local estimators are of low quality or inconsistent, as long as each cluster contains a majority of accurate estimators. This contrasts with its analogue based on simple averaging, whose performance is determined by the worst local estimator. In the homogeneous setting, the robust estimator is also shown to be minimax optimal, up to a logarithmic factor, both when the number of clients remains fixed and when it diverges. The theoretical findings are corroborated by simulation studies and a real-data analysis.
|
| 1920 |
Tags for DAGs: Graph Refinement with Meta-Informed Relations
2506.19459
|
cs.LG
|
Florian Peter Busch, Moritz Willig, Florian Guldan, Kristian Kersting, Devendra Singh Dhami |
Causal discovery has shifted from data-centric methods to hybrid strategies that integrate semantic knowledge from experts or large language models (LLMs). Such external information is vital for identifying causal structures beyond the Markov Equivalence Class...Causal discovery has shifted from data-centric methods to hybrid strategies that integrate semantic knowledge from experts or large language models (LLMs). Such external information is vital for identifying causal structures beyond the Markov Equivalence Class (MEC), which data alone cannot resolve. However, expert availability is often limited, and LLMs frequently misidentify causal directions in specialized domains. To overcome such shortcomings, we propose a tag-based approach that leverages semantically meaningful labels while deriving causal directionality directly from data. Using variable-level tag assignments from available sources (e.g., LLMs), our tags for DAGs method learns from identifiable data structures to extract higher-level causal relations. These are then used to orient undirected edges, enabling causal discovery to move beyond the MEC without reliance on fallible external knowledge.
|
| 1921 |
On the Interaction of Compressibility and Adversarial Robustness
2507.17725
|
cs.LG
|
Melih Barsbey, Ant\^onio H. Ribeiro, Umut \c{S}im\c{s}ekli, Tolga Birdal |
As demands for resource efficiency and safety in modern neural networks intensify, substantial research effort has gone into model compression and adversarial robustness. Yet despite progress on each in isolation, a systematic understanding of how compressibil...As demands for resource efficiency and safety in modern neural networks intensify, substantial research effort has gone into model compression and adversarial robustness. Yet despite progress on each in isolation, a systematic understanding of how compressibility shapes robustness remains elusive. In this paper, we develop a principled framework to analyze how different forms of structured compressibility - such as neuron-level and spectral compressibility - affect adversarial robustness. We show that structured compressibility can induce a small number of highly sensitive directions in the representation space, which adversaries can exploit to construct effective perturbations. Our analysis yields a robustness bound that reveals how neuron and spectral compressibility impact $\ell_\infty$ and $\ell_2$ robustness via their effects on the learned representations. Crucially, the vulnerabilities we identify arise irrespective of how compressibility is achieved - whether via regularization, architectural bias, or learning dynamics. Through empirical evaluations across synthetic and realistic tasks, we confirm our theoretical predictions, and further demonstrate that these vulnerabilities persist under adversarial training and transfer learning, and contribute to the emergence of universal adversarial examples. Our findings show a fundamental tension between structured compressibility and robustness and highlight new pathways for designing models that are efficient and safe.
|
| 1922 |
Edge Selection for the Effective use of Piecewise-Constant Distributions as Neural Network Outputs for Event Prediction
2507.21616
|
cs.LG
|
Kevin Doran, Tom Baden |
We study the output representation of a neural network used for next event prediction. We propose partitioning the time axis into a fixed set of intervals and having a neural network output a categorical distribution over them, which we map to a (mostly) piece...We study the output representation of a neural network used for next event prediction. We propose partitioning the time axis into a fixed set of intervals and having a neural network output a categorical distribution over them, which we map to a (mostly) piecewise-constant probability density. We present an optimization procedure that selects interval edges in order to maximize data likelihood under the representation. The representation is well suited to processes whose inter-event distribution is a mixture of smooth and sharply peaked components$\unicode{x2013}$a pattern we find common in event data recorded from real-world processes.
|
| 1923 |
SHEFL: Sparse Heterogeneous Ensemble Federated Learning with Group-Balanced Aggregation
2508.08552
|
cs.LG
|
Keumseo Ryum, Jinu Gong, Joonhyuk Kang |
System heterogeneity in federated learning (FL) requires training workloads and communication demands to be adapted to client resource capacities. Federated ensemble learning distributes multiple predictors across clients, but assigning one ensemble member to ...System heterogeneity in federated learning (FL) requires training workloads and communication demands to be adapted to client resource capacities. Federated ensemble learning distributes multiple predictors across clients, but assigning one ensemble member to each client reduces the number of contributors to each member as the ensemble grows. We propose SHEFL, a global ensemble-based FL framework that uses the additional computation of high-power clients (HPCs) to train all ensemble members while low-power clients (LPCs) train one assigned member. SHEFL addresses increased HPC uplink demand and biased aggregation toward HPC updates through resource-aware sparse update allocation and group-balanced aggregation. Across four vision datasets, SHEFL improves over Fed-Ensemble (FedEns) in all data heterogeneity settings. Under a matched total uplink budget, SHEFL reaches the target accuracy in fewer rounds than FedEns and improves final accuracy by up to 14.65 points. Ablations further show that group-balanced aggregation improves accuracy and convergence under different resource allocations.
|
| 1924 |
Graph Structure Learning with Temporal Graph Information Bottleneck for Inductive Representation Learning
2508.14859
|
cs.LG
|
Jiafeng Xiong, Rizos Sakellariou |
Temporal graph learning is crucial for dynamic networks where nodes and edges evolve over time and new nodes continuously join the system. Inductive representation learning in such settings faces two major challenges: effectively representing unseen nodes and ...Temporal graph learning is crucial for dynamic networks where nodes and edges evolve over time and new nodes continuously join the system. Inductive representation learning in such settings faces two major challenges: effectively representing unseen nodes and mitigating noisy or redundant graph information. We propose GTGIB, a versatile framework that integrates Graph Structure Learning (GSL) with Temporal Graph Information Bottleneck (TGIB). We design a novel two-step GSL-based structural enhancer to enrich and optimize node neighborhoods and demonstrate its effectiveness and efficiency through theoretical proofs and experiments. The TGIB refines the optimized graph by extending the information bottleneck principle to temporal graphs, regularizing both edges and features based on our derived tractable TGIB objective function via variational approximation, enabling stable and efficient optimization. GTGIB-based models are evaluated to predict links on four real-world datasets; they outperform existing methods in all datasets under the inductive setting, with significant and consistent improvement in the transductive setting.
|
| 1925 |
Learning Domain- and Class-Disentangled Prototypes for Domain-Generalized EEG Emotion Recognition
2509.01135
|
cs.LG
|
Guangli Li, Canbiao Wu, Zhehao Zhou, Na Tian, Li Zhang |
Electroencephalography (EEG)-based emotion recognition plays a critical role in affective Brain-Computer Interfaces (aBCIs), yet its practical deployment remains limited by inter-subject variability, reliance on target-domain data, and unavoidable label noise....Electroencephalography (EEG)-based emotion recognition plays a critical role in affective Brain-Computer Interfaces (aBCIs), yet its practical deployment remains limited by inter-subject variability, reliance on target-domain data, and unavoidable label noise. To address these challenges, we propose a Multi-domain Aggregation Transfer learning framework with domain-class prototypes (MAT) for emotion recognition under completely unseen target domains. MAT introduces a feature decoupling module to disentangle class-invariant domain features from domain-invariant class features, enabling more robust and interpretable EEG representations. A Hierarchical-Domain Aggregation (HDA) mechanism based on Maximum Mean Discrepancy (MMD) constructs superdomains to model shared distributional structures across subjects, while adaptive prototype updating refines domain and class prototypes to capture stable intrinsic representations. Moreover, a pairwise learning strategy reformulates classification as similarity estimation between sample pairs, effectively mitigating the effect of label noise. Extensive experiments on three public EEG emotion datasets (SEED, SEED-IV, and SEED-V) show that the accuracy of MAT is improved by 2.87%, 3.84%, and 2.05% compared with the state-of-the-art (SOTA) model for unseen target domains. Our results provide a promising direction for emotion recognition under real-world unseen-subject scenarios.The source code is available at https://github.com/WuCB-BCI/MAT.
|
| 1926 |
An Efficient Subspace Algorithm for Federated Learning on Heterogeneous Data
2509.05213
|
cs.LG
|
Jiaojiao Zhang, Yizhao Fan, Yuqi Xu, Kun Yuan |
This work addresses the key challenges of applying federated learning to large-scale deep neural networks, particularly the issue of client drift due to data heterogeneity across clients and the high costs of communication, computation, and memory. We propose ...This work addresses the key challenges of applying federated learning to large-scale deep neural networks, particularly the issue of client drift due to data heterogeneity across clients and the high costs of communication, computation, and memory. We propose FedSub, an efficient subspace algorithm for federated learning on heterogeneous data. Specifically, FedSub utilizes subspace projection to guarantee local updates of each client within low-dimensional subspaces, thereby reducing communication, computation, and memory costs. Additionally, it incorporates low-dimensional dual variables to mitigate client drift. We provide convergence analysis that reveals the impact of key factors such as step size and subspace projection matrices on convergence. Experimental results demonstrate its efficiency.
|
| 1927 |
Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
2509.22161
|
cs.LG
|
Takashi Morita |
Vector quantization, which discretizes a continuous vector space into a finite set of representative vectors (a codebook), has been widely adopted in modern machine learning. Despite its effectiveness, vector quantization poses a fundamental challenge: the non...Vector quantization, which discretizes a continuous vector space into a finite set of representative vectors (a codebook), has been widely adopted in modern machine learning. Despite its effectiveness, vector quantization poses a fundamental challenge: the non-differentiable quantization step blocks gradient backpropagation. Smoothed vector quantization addresses this issue by relaxing the discrete selection of a codebook vector into a weighted combination of codebook entries, represented as the matrix product of a simplex vector and the codebook. Effective smoothing requires two properties: (1) smoothed code-selection vectors should remain close to a onehot vector, ensuring tight approximation, and (2) all codebook entries should be utilized, preventing code collapse. Existing methods typically address these desiderata separately. By contrast, the present study introduces a simple and intuitive regularization that promotes both simultaneously by minimizing the distance between each simplex vertex and its $K$-nearest smoothed code-selectors. Representative benchmarks on image encoding demonstrate that the proposed method achieves more effective codebook utilization and improves performance over prior approaches.
|
| 1928 |
WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning
2509.23219
|
cs.LG
|
Xin Li, Mengbing Liu, Yiyang Zhu, Wenhe Zhang, Li Wei |
Technical-domain benchmarks constructed from arXiv papers can overlap the same public text used in LLM pretraining. Auditing this risk at training-corpus scale requires searching billions of corpus n-grams while retaining per-item evidence that users can inspe...Technical-domain benchmarks constructed from arXiv papers can overlap the same public text used in LLM pretraining. Auditing this risk at training-corpus scale requires searching billions of corpus n-grams while retaining per-item evidence that users can inspect and recompute. We contribute a reverse-probe audit at a fixed 13-gram threshold: it indexes benchmark prompts, streams public pretraining corpora, and emits per-problem prompt-surface lexical-overlap metadata with memory that scales with the benchmark. We instantiate the protocol in WirelessMathBench-XL, a 4,027-problem wireless mathematical-reasoning benchmark built from 836 retained arXiv papers across 20 subfields. Against 12.9B streamed 13-grams from RedPajama-arXiv, the audit identifies a strict zero-hit view S0 covering 3,853 problems (95.7%). Filtering to S0 changes accuracy by less than 1 pp for every evaluated model; frontier calibration rows form one high-accuracy cluster between 86.5% and 91.3%, not a resolved rank order. Only 30/800 test items carry detected overlap. Under an all-flagged-correct counterfactual, their largest possible positive score inflation is 0.31-0.51 pp for the frontier rows, so full-versus-S0 is a bounded, structurally underpowered stability summary rather than a contamination-effect test or cleanliness claim. The audit channel does not cover paraphrase, target-answer, post-training, or closed-corpus exposure. The release includes source-paper identifiers, verifier-facing ground truths, audit and threshold metadata, filtered views, a paper-disjoint sensitivity view, evaluation traces, paired-bootstrap scripts, training recipes, Croissant metadata, and a Datasheet for Datasets.
|
| 1929 |
Graph Your Own Prompt
2509.23373
|
cs.LG
|
Xi Ding, Lei Wang, Piotr Koniusz, Yongsheng Gao |
We propose Graph Consistency Regularization (GCR), a novel framework that injects relational graph structures, derived from model predictions, into the learning process to promote class-aware, semantically meaningful feature representations. Functioning as a f...We propose Graph Consistency Regularization (GCR), a novel framework that injects relational graph structures, derived from model predictions, into the learning process to promote class-aware, semantically meaningful feature representations. Functioning as a form of self-prompting, GCR enables the model to refine its internal structure using its own outputs. While deep networks learn rich representations, these often capture noisy inter-class similarities that contradict the model's predicted semantics. GCR addresses this issue by introducing parameter-free Graph Consistency Layers (GCLs) at arbitrary depths. Each GCL builds a batch-level feature similarity graph and aligns it with a global, class-aware masked prediction graph, derived by modulating softmax prediction similarities with intra-class indicators. This alignment enforces that feature-level relationships reflect class-consistent prediction behavior, acting as a semantic regularizer throughout the network. Unlike prior work, GCR introduces a multi-layer, cross-space graph alignment mechanism with adaptive weighting, where layer importance is learned from graph discrepancy magnitudes. This allows the model to prioritize semantically reliable layers and suppress noisy ones, enhancing feature quality without modifying the architecture or training procedure. GCR is model-agnostic, lightweight, and improves semantic structure across various networks and datasets. Experiments show that GCR promotes cleaner feature structure, stronger intra-class cohesion, and improved generalization, offering a new perspective on learning from prediction structure. [Project website](https://darcyddx.github.io/gcr/) [Code](https://github.com/Darcyddx/graph-prompt)
|
| 1930 |
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
2509.25637
|
cs.LG
|
Kotaro Yoshida, Atsushi Nitanda |
Preconditioning is widely used in machine learning to accelerate convergence on the empirical risk, yet its role on the expected risk remains underexplored. In this work, we investigate how preconditioning affects feature learning and generalization performanc...Preconditioning is widely used in machine learning to accelerate convergence on the empirical risk, yet its role on the expected risk remains underexplored. In this work, we investigate how preconditioning affects feature learning and generalization performance. We first show that the input information available to the model is conveyed solely through the Gram matrix defined by the preconditioner's metric, thereby inducing a controllable spectral bias on feature learning. Concretely, instantiating the preconditioner as the $p$-th power of the input covariance matrix, we prove that, for a single ReLU neuron, increasing $p$ monotonically shifts the learned feature's sensitivity toward higher-variance eigenspaces, even though all choices of $p$ attain the same limiting population loss. We then empirically investigate how this spectral bias interacts with task-relevant features in three settings: robustness to noise, out-of-distribution generalization, and forward knowledge transfer. Our experiments show that learned representations favor the spectral components emphasized by preconditioning, and that generalization improves when this bias aligns with task-relevant features.
|
| 1931 |
Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models
2509.25826
|
cs.LG
|
Kun Feng, Shaocheng Lan, Yuchen Fang, Wenchao He, Sihan Lu |
Inherent temporal heterogeneity, such as varying sampling densities and periodic structures, has posed substantial challenges in zero-shot generalization for Time Series Foundation Models (TSFMs). Existing TSFMs predominantly rely on massive parameterization t...Inherent temporal heterogeneity, such as varying sampling densities and periodic structures, has posed substantial challenges in zero-shot generalization for Time Series Foundation Models (TSFMs). Existing TSFMs predominantly rely on massive parameterization to absorb such heterogeneity, as their static tokenization and positional encoding schemes entangle diverse temporal patterns into a fixed representation space, encouraging memorization rather than adaptation. To address this limitation, we propose Kairos, a flexible and parameter-efficient TSFM dedicated to forecasting tasks, which decouples temporal heterogeneity from model capacity through a novel tokenization perspective. Kairos introduces a dynamic patching tokenizer and a mixture-of-size encoding that adapt observational granularity to local information density, enabling fine-grained temporal abstraction without increasing model width or depth. In addition, we design a multi-granularity positional embedding based on dynamic rotary encodings, which conditions on instance-level spectral features and temporal structure induced by dynamic patching tokenization, allowing robust modeling of diverse temporal dependencies. Trained on a novel Predictability-Stratified Time-Series (PreSTS) corpus, Kairos achieves superior zero-shot performance with substantially fewer parameters on two mainstream benchmarks, GIFT-Eval and Time-Series-Library. The project page is at https://foundation-model-research.github.io/Kairos .
|
| 1932 |
Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
2510.04525
|
cs.LG
|
Satoshi Hayakawa, Yuhta Takida, Masaaki Imaizumi, Hiromi Wakaki, Yuki Mitsufuji |
Masked diffusion models have shown promising performance in generating high-quality samples in a wide range of domains, but accelerating their sampling process remains relatively underexplored. To investigate efficient samplers for masked diffusion, this paper...Masked diffusion models have shown promising performance in generating high-quality samples in a wide range of domains, but accelerating their sampling process remains relatively underexplored. To investigate efficient samplers for masked diffusion, this paper theoretically analyzes the MaskGIT sampler for image modeling, revealing its implicit temperature sampling mechanism. Through this analysis, we show that MaskGIT is asymptotically equivalent to a choose-then-sample (CTS) formulation, instantiated as the "moment sampler," which explicitly separates index selection from token sampling. This CTS reformulation is essential: it yields unbiased token sampling and exposes an algorithmic design space for index selection, both of which are inaccessible in MaskGIT's original formulation. Regarding token sampling, we reveal that MaskGIT implicitly adopts a low-temperature sampler, which explains why MaskGIT often degrades with more sampling steps. The CTS reformulation of MaskGIT allows us to fix the temperature sampling to ensure unbiasedness. We also improve the index selection in CTS through two key innovations: a partial caching technique for transformers that approximates longer sampling trajectories without proportional computational cost, and a hybrid approach formalizing the exploration-exploitation trade-off in adaptive unmasking. Experiments in image and text domains demonstrate our theory as well as the efficiency of our proposed methods, advancing both theoretical understanding and practical implementation of masked diffusion samplers.
|
| 1933 |
Federated Computation of ROC and PR Curves
2510.04979
|
cs.LG
|
Xuefeng Xu, Graham Cormode |
Receiver Operating Characteristic (ROC) and Precision-Recall (PR) curves are fundamental analytical tools for evaluating classification models, providing a comprehensive view of performance across decision thresholds. In modern data management settings, such e...Receiver Operating Characteristic (ROC) and Precision-Recall (PR) curves are fundamental analytical tools for evaluating classification models, providing a comprehensive view of performance across decision thresholds. In modern data management settings, such evaluation must often be performed over data that is distributed across multiple parties and subject to strict privacy constraints. This setting arises naturally in federated learning (FL) and collaborative data platforms, where raw prediction scores and labels cannot be centrally collected. We study the problem of computing ROC and PR curves over distributed data under formal privacy guarantees. The key challenge is that exact computation requires access to all prediction scores and labels, leading to linear communication costs and privacy risks. To address this, we propose a novel framework for approximating ROC and PR curves using quantiles of the score distribution, which can be computed efficiently under secure aggregation and distributed differential privacy. Our approach requires only $O(Q)$ communication, where $Q$ is the number of quantiles, and avoids sharing raw data entirely. We provide theoretical guarantees on the approximation quality by bounding the Area Error (AE) between the true and estimated curves, demonstrating principled trade-offs among approximation accuracy, privacy protection, and communication cost. Our method is robust to data heterogeneity and skew, making it suitable for real-world distributed data management scenarios. Experiments on real-world datasets show that our approach achieves high accuracy with low communication overhead under strong privacy constraints. Overall, this work provides a practical and theoretically grounded solution for privacy-preserving evaluation of classification models over distributed data.
|
| 1934 |
Auditing Information Disclosure During Large-Scale Gradient-Based Training via Gradient Uniqueness
2510.10902
|
cs.LG
|
Sleem Abdelghafar, Maryam Aliakbarpour, Christopher Jermaine |
Auditing information disclosure across every datapoint during the training of LLMs is challenging. We propose a principled, attack-agnostic approach that uses mutual information to measure what the final model reveals about a datapoint's training membership. W...Auditing information disclosure across every datapoint during the training of LLMs is challenging. We propose a principled, attack-agnostic approach that uses mutual information to measure what the final model reveals about a datapoint's training membership. We show that this final-model disclosure is upper bounded by the sum of per-iteration gradient disclosures and that, under a reasonable set of assumptions, these gradient disclosures increase with Gradient Uniqueness (GNQ), which measures how distinguishable a datapoint's gradient is relative to other gradients in the batch. While naively computing GNQ requires forming and inverting a $P\times P$ matrix for every datapoint (for a model with $P$ parameters), we introduce Batch-Space Ghost (BS-Ghost). This efficient algorithm performs all computations in a much smaller batch space and uses ghost kernels to compute GNQ "in-run" for every datapoint in the training corpus, with minimal computational and memory overhead. Our experiments show the following: (i) GNQ predicts MIA vulnerability without the need for shadow models. (ii) Beyond membership disclosure, GNQ predicts the success of reconstruction attacks. (iii) GNQ-guided removal and retraining identify datapoints that causally contribute to disclosure. (iv) GNQ outperforms counterfactual memorization in text extraction and common-knowledge discrimination without the need for additional model training. (v) For data attribution, GNQ-guided filtering reduces emergent misalignment in Qwen2.5-7B more than baselines. Further, GNQ attributes 1000 datapoints in 17 seconds---roughly $70\times$ faster than the baselines. (vi) GNQ explains how training choices affect training-set disclosure and captures how per-datapoint disclosure emerges during training.
|
| 1935 |
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
2510.13680
|
cs.LG
|
Bingbin Liu, Rachit Bansal, Depen Morwani, Nikhil Vyas, David Alvarez-Melis |
Preconditioned methods are central to deep learning optimization. Predominant approaches include computationally light diagonal preconditioners such as Adam, which rely on gradient statistics, and second-order methods such as Gauss-Newton (GN), which capture r...Preconditioned methods are central to deep learning optimization. Predominant approaches include computationally light diagonal preconditioners such as Adam, which rely on gradient statistics, and second-order methods such as Gauss-Newton (GN), which capture richer curvature information. Seeking the best of both worlds, we disentangle the preconditioner design space into several factors, separating the choice of diagonal scaling (Adam-style versus GN-style) from 1) the basis choices under which the diagonal scaling operates, and 2) the gradient noises from mini-batching. Our theoretical results show that GN's optimality for linear regression no longer holds under a poorly chosen basis under both population and stochastic updates, or a move from linear regression to a non-convex variant of logistic regression. In these settings, Adam-style methods offer genuine advantages over curvature-inspired preconditioning, rather than serving merely as a tractable proxy. Empirical results on synthetic problems and CIFAR-10 support these findings.
|
| 1936 |
Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
2510.15508
|
cs.LG
|
Naoki Yoshida, Satoshi Hayakawa, Yuhta Takida, Toshimitsu Uesaka, Hiromi Wakaki |
In this study, we propose an enhancement to the similarity computation mechanism in multimodal contrastive pretraining frameworks such as CLIP. Prior theoretical research has demonstrated that the optimal similarity metrics between paired modalities should cor...In this study, we propose an enhancement to the similarity computation mechanism in multimodal contrastive pretraining frameworks such as CLIP. Prior theoretical research has demonstrated that the optimal similarity metrics between paired modalities should correspond to the pointwise mutual information (PMI) between the two modalities. However, the current implementations of CLIP and its variants fail to fully utilize the underlying linear structure of PMI. We therefore propose KME-CLIP, which leverages this structure through the inner product in a reproducing kernel Hilbert space (RKHS). We theoretically prove that, under our assumptions, the KME-CLIP similarity can bring the contrastive loss arbitrarily close to its optimal value, which is attained by PMI, as the size of the point set grows, and we empirically evaluate KME-CLIP against CLIP and its kernel-based variants across several retrieval and classification tasks.
|
| 1937 |
SCORENF: Score-based Normalizing Flows for Sampling Unnormalized distributions
2510.21330
|
cs.LG
|
Vikas Kanaujia, Vipul Arora |
Unnormalized probability distributions are central to modeling complex physical systems across various scientific domains. Traditional sampling methods, such as Markov Chain Monte Carlo (MCMC), often suffer from slow convergence, critical slowing down, poor mo...Unnormalized probability distributions are central to modeling complex physical systems across various scientific domains. Traditional sampling methods, such as Markov Chain Monte Carlo (MCMC), often suffer from slow convergence, critical slowing down, poor mode mixing, and high autocorrelation. In contrast, likelihood-based and adversarial machine learning models, though effective, are heavily data-driven, requiring large datasets and often encountering mode covering and mode collapse. In this work, we propose ScoreNF, a score-based learning framework built on the Normalizing Flow (NF) architecture, integrated with an Independent Metropolis-Hastings (IMH) module, enabling efficient and unbiased sampling from unnormalized target distributions. We show that ScoreNF maintains high performance even with small training ensembles, thereby reducing reliance on computationally expensive MCMC-generated training data. We also present a method for assessing mode-covering and mode-collapse behaviours. We validate our method on synthetic 2D distributions (MOG-4 and MOG-8) and the high-dimensional $\phi^4$ lattice field theory distribution, demonstrating its effectiveness for sampling tasks.
|
| 1938 |
Managing Self-Learning Experts under Per-Round Budget Constraints
2510.22654
|
cs.LG
|
Ilgam Latypov, Alexandra Suvorikova, Alexey Kroshnin, Alexander Gasnikov, Yuriy Dorn |
This paper addresses the problem of sequential decision-making under learning budget constraints. Such settings naturally arise in applications like managing a portfolio of bandit or reinforcement learning (RL) algorithms. We propose a novel UCB-type algorithm...This paper addresses the problem of sequential decision-making under learning budget constraints. Such settings naturally arise in applications like managing a portfolio of bandit or reinforcement learning (RL) algorithms. We propose a novel UCB-type algorithm, M-LCB, designed to manage a pool of $K$ self-learning experts in a stochastic environment while accounting for a limited per-round learning budget $M$. At each round, M-LCB selects one expert to make a decision and at most $M \le K$ experts to learn. For selection, M-LCB uses confidence bounds constructed from limited prior knowledge about the experts (i.e., mild assumptions) and their observed training losses. We derive anytime regret bounds for M-LCB that scale with the individual regrets of the experts. In particular, if each expert has regret $\tilde O(T^\alpha)$ by round $T$, then M-LCB guarantees an overall regret of $\tilde O\left(\sqrt{KT/M} + (K/M)^{1-\alpha}T^\alpha\right)$ relative to the best expert in hindsight. Finally, we demonstrate the applicability of M-LCB using self-learning experts instantiated as (i) parametric models and (ii) bandit algorithms.
|
| 1939 |
A Game-Theoretic Spatio-Temporal Reinforcement Learning Framework for Collaborative Public Resource Allocation
2510.26184
|
cs.LG
|
Songxin Lei, Qiongyan Wang, Yanchen Zhu, Hanyu Yao, Sijie Ruan |
Public resource allocation involves distributing resources, including urban infrastructure, energy, and transportation, which are typically limited in capacity, to meet social demands. In real-world scenarios, resources are typically limited in capacity, which...Public resource allocation involves distributing resources, including urban infrastructure, energy, and transportation, which are typically limited in capacity, to meet social demands. In real-world scenarios, resources are typically limited in capacity, which makes coordination among multiple resources essential. However, existing methods often optimize resource movements in an isolated manner and do not explicitly account for capacity-aware collaboration under spatio-temporal dynamics. To address this limitation, we introduce the Collaborative Public Resource Allocation (CPRA) problem, and propose a Game-Theoretic Spatio-Temporal Reinforcement Learning (GSTRL) framework to solve it. Our contributions are twofold: 1) We formulate CPRA as a potential game and construct the potential function based on the objective function of CPRA, laying a theoretical foundation for approximating the Nash equilibrium of this NP-hard problem; and 2) Our GSTRL framework effectively captures the spatio-temporal dynamics of the overall system. We evaluate GSTRL on two real-world datasets, where experiments show its superior performance. Our source codes are available at https://github.com/thunderlrr/GSTRL.
|
| 1940 |
Wavelet-Based Parity Detection Revisited: Representation Dependence, Generalization, and Mechanistic Analysis
2511.00071
|
cs.LG
|
Ertu\u{g}rul Mutlu |
Parity is exactly determined by the least significant bit (LSB), so machine learning is unnecessary for solving the task. This work instead uses parity as a controlled diagnostic for studying how a classical signal-processing pipeline preserves, amplifies, or ...Parity is exactly determined by the least significant bit (LSB), so machine learning is unnecessary for solving the task. This work instead uses parity as a controlled diagnostic for studying how a classical signal-processing pipeline preserves, amplifies, or suppresses access to symbolic information embedded in a numerical representation. Integers from 0 to 10,000 are encoded as fixed-width 32-bit binary signals, decomposed with a level-3 Daubechies-2 (db2) discrete wavelet transform, summarized by mean absolute coefficient magnitude, and clustered independently per subband with k-means. Clustering is unsupervised, while cluster-to-parity calibration uses training labels only under a stratified 60/20/20 train/validation/test protocol. The frozen configuration reaches 84.26% accuracy on the held-out test set (95% Wilson CI: 82.60-85.79%), and 84.20% +/- 0.57% across 20 stratified resplits. Raw binary summary statistics reach 61.30%, while masking the LSB reduces validation accuracy to 48.15%, consistent with chance. The A3 approximation band alone retains 83.20%, whereas detail bands remain near chance. Performance is strongly representation-dependent: moving the parity bit changes validation accuracy up to 98.60%, and changing the wavelet boundary mode ranges from 54.45% to 83.20%. Cross-magnitude extrapolation degrades from 79.98% on 10,001-20,000 to 59.69% on 100,001-1,000,000. These results do not show that wavelets discover the arithmetic rule of parity. Rather, they show that information already present in the binary encoding becomes more or less recoverable depending on multiscale filtering, spatial alignment, and boundary handling.
|
| 1941 |
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
2511.03095
|
cs.LG
|
Gaia Grosso, Sai Sumedh R. Hindupur, Thomas Fel, Samuel Bright-Thonney, Philip Harris |
Modern artificial intelligence has revolutionized how we extract representations from scientific data, yet the statistical properties of these representations remain poorly controlled, causing misspecified anomaly detection methods to falter. The hardest anoma...Modern artificial intelligence has revolutionized how we extract representations from scientific data, yet the statistical properties of these representations remain poorly controlled, causing misspecified anomaly detection methods to falter. The hardest anomalies to detect are the rare, weakly separable ones hiding within the nominal distribution-a regime that grows in importance as models mature and easily separable signals are exhausted. We identify structural desiderata for detection in this regime under minimal prior information: sparsity, to enforce parsimony; locality, to preserve geometric sensitivity; and competition, to promote efficient allocation of model capacity. These principles define a class of self-organizing local kernels that adaptively partition the representation space around regions of statistical imbalance. As an instantiation, we introduce SparKer, a sparse ensemble of Gaussian kernels trained in a semi-supervised Neyman-Pearson framework to locally model the likelihood ratio between a sample that may contain anomalies and an anomaly-free reference. We provide theoretical insights into the mechanisms driving detection and self-organization, and demonstrate the approach on realistic high-dimensional problems in scientific discovery, open-world novelty detection, intrusion detection, and generative-model validation. Ensembles of only a handful of kernels identify statistically significant anomalies in representation spaces of thousands of dimensions while remaining sensitive across regimes, underscoring the interpretability, efficiency, and scalability of the approach.
|
| 1942 |
COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing
2511.06894
|
cs.LG
|
Wenlong Shang, Shihao Tian, Xutong Wan, Peng Chang |
Reconstruction-based methods are a dominant paradigm in time series anomaly detection (TSAD), however, their near-universal reliance on Mean Squared Error (MSE) loss results in statistically flawed reconstruction residuals. This fundamental weakness leads to n...Reconstruction-based methods are a dominant paradigm in time series anomaly detection (TSAD), however, their near-universal reliance on Mean Squared Error (MSE) loss results in statistically flawed reconstruction residuals. This fundamental weakness leads to noisy, unstable anomaly scores, hindering reliable detection. To address this, we propose Constrained Gaussian-Noise Optimization and Smoothing (COGNOS), a universal, model-agnostic enhancement framework that tackles this issue at its source. COGNOS introduces a novel Gaussian-White Noise Regularization strategy during training, which directly constrains the model's output residuals to conform to a Gaussian white noise distribution. This engineered statistical property creates the ideal precondition for our second contribution: Adaptive Residual Kalman Smoother that operates as a statistically robust estimator to denoise the raw anomaly scores. Extensive experiments on multiple benchmarks demonstrate that COGNOS consistently enhances the performance of state-of-the-art backbones significantly, validating the efficacy of coupling statistical regularization with adaptive filtering.
|
| 1943 |
Neural Tractability via Structure: Learning-Augmented Algorithms for Graph Combinatorial Optimization
2511.19573
|
cs.LG
|
Jialiang Li, Weitong Chen, Mingyu Guo |
Neural solvers provide fast solutions to graph combinatorial optimization problems, but training alone does not guarantee solution quality. Exact methods guarantee optimality but can be prohibitively expensive. We propose Neural Fixed-Parameter Tractable (N-FP...Neural solvers provide fast solutions to graph combinatorial optimization problems, but training alone does not guarantee solution quality. Exact methods guarantee optimality but can be prohibitively expensive. We propose Neural Fixed-Parameter Tractable (N-FPT), a neural-model-agnostic framework that uses parameterized algorithms to improve neural solutions without retraining and guide learning. It restricts neural advice to a treewidth modulator and completes the bounded-treewidth remainder exactly. It returns and certifies the best achievable solution under any valid advice, improving the neural solution whenever a better compatible completion exists. We prove this guarantee and global optimality under perfect advice. Its incremental-confidence N-FPT (NIC-FPT) and randomized-deferral N-FPT (NRD-FPT) variants consolidate neural advice. Beyond inference, N-FPT feedback becomes available as neural decisions make exact completion tractable. From that state onward, it supplies what final solution rewards lack: the first decision that rules out every optimal completion and alternatives that preserve one. One exact computation supplies reusable guidance for continuations from that state, even when neural samples miss its optimum. Experiments across graph combinatorial optimization problems demonstrate solution-quality gains with different neural models and robustness to graph-size and distribution shifts. The learned completion model reaches optimal completions more often than reinforcement learning (RL) alone, without an exact solver at inference. N-FPT thus uses graph structure to improve neural solutions and guide learning toward optimal completion.
|
| 1944 |
Time Series Foundation Models for Process Model Forecasting
2512.07624
|
cs.LG
|
Yongbo Yu, Jari Peeperkorn, Johannes De Smedt, Jochen De Weerdt |
Process Model Forecasting (PMF) aims to predict how the control-flow structure of a process evolves over time by modeling the temporal dynamics of directly-follows (DF) relations, complementing predictive process monitoring that focuses on single-case prefixes...Process Model Forecasting (PMF) aims to predict how the control-flow structure of a process evolves over time by modeling the temporal dynamics of directly-follows (DF) relations, complementing predictive process monitoring that focuses on single-case prefixes. Prior benchmarks show that machine learning and deep learning models provide only modest gains over statistical baselines, mainly due to the sparsity and heterogeneity of the DF time series. We investigate Time Series Foundation Models (TSFMs), large pre-trained models for generic time series, as an alternative for PMF. Using DF time series derived from real-life event logs, we compare zero-shot use of TSFMs, without additional training, with fine-tuned variants adapted on PMF-specific data. TSFMs generally achieve lower forecasting errors (MAE and RMSE) than traditional and specialized models trained from scratch on the same logs, indicating effective transfer of temporal structure from non-process domains. While fine-tuning can further improve accuracy, the gains are often small and may disappear on smaller or more complex datasets, so zero-shot use remains a strong default. Our study highlights the generalization capability and data efficiency of TSFMs for process-related time series and, to the best of our knowledge, provides the first systematic evaluation of temporal foundation models for PMF.
|
| 1945 |
CAT: Can Trust be Predicted with Context-Awareness in Dynamic Heterogeneous Networks?
2512.11352
|
cs.LG
|
Jie Wang, Zheng Yan, Jiahe Lan, Xuyan Li, Elisa Bertino |
Trust prediction provides valuable support for decision-making, risk mitigation, and system security enhancement. Recently, Graph Neural Networks (GNNs) have emerged as a promising approach for trust prediction, owing to their ability to learn expressive node ...Trust prediction provides valuable support for decision-making, risk mitigation, and system security enhancement. Recently, Graph Neural Networks (GNNs) have emerged as a promising approach for trust prediction, owing to their ability to learn expressive node representations that capture intricate trust relationships within a network. However, current GNN-based trust prediction models face several limitations: (i) Most of them fail to capture trust dynamicity, leading to questionable inferences. (ii) They rarely consider the heterogeneous nature of real-world networks, resulting in a loss of rich semantics. (iii) None of them support context-awareness, a basic property of trust, making prediction results coarse-grained. To this end, we propose CAT, the first Context-Aware GNN-based Trust prediction model that supports trust dynamicity and accurately represents real-world heterogeneity. CAT consists of a graph construction layer, an embedding layer, a heterogeneous attention layer, and a prediction layer. It handles dynamic graphs using continuous-time representations and captures temporal information through a time encoding function. To model graph heterogeneity and leverage semantic information, CAT employs a dual attention mechanism that identifies the importance of different node types and nodes within each type. For context-awareness, we introduce a new notion of meta-paths to extract contextual features. By constructing context embeddings and integrating a context-aware aggregator, CAT can predict both context-aware trust and overall trust. Extensive experiments on three real-world datasets demonstrate that CAT outperforms five groups of baselines in trust prediction, while exhibiting strong scalability to large-scale graphs and robustness against both trust-oriented and GNN-oriented attacks.
|
| 1946 |
Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
2512.12602
|
cs.LG
|
Jingdi Lei, Di Zhang, Soujanya Poria |
In this paper, we introduce Exact Flow Linear Attention~(EFLA), an exact-flow formulation of delta-rule linear attention. We show that the delta-rule update can be interpreted as an explicit Euler discretization of an underlying continuous-time system. EFLA re...In this paper, we introduce Exact Flow Linear Attention~(EFLA), an exact-flow formulation of delta-rule linear attention. We show that the delta-rule update can be interpreted as an explicit Euler discretization of an underlying continuous-time system. EFLA replaces this first-order update with the exact closed-form flow. By exploiting the rank-1 structure of the dynamics matrix, both the matrix exponential and the input integral collapse to a simple update that preserves delta-rule linear attention's algebraic structure, parameter count, linear-time complexity, and chunkwise parallelism. This attention mechanism removes the Euler discretization error of the delta-rule dynamics without introducing additional parameters. Experiments on robustness tests, language modeling benchmarks, and the MAD synthetic benchmark show that EFLA improves stability under corrupted and high-energy inputs, reduces perplexity, and achieves stronger downstream performance compared to SSM and Euler-style baselines. These results establish exact-flow integration as a principled and scalable update mechanism for delta-rule linear attention.
|
| 1947 |
Soft Geometric Inductive Bias for Object Centric Dynamics
2512.15493
|
cs.LG
|
Hampus Linander, Conor Heins, Alexander Tschantz, Marco Perin, Christopher Buckley |
Physical systems are naturally parameterized in terms of geometric entities and their transformations. While exact equivariance can be a powerful inductive bias, real-world applications frequently exhibit approximate or broken symmetries in practice. We introd...Physical systems are naturally parameterized in terms of geometric entities and their transformations. While exact equivariance can be a powerful inductive bias, real-world applications frequently exhibit approximate or broken symmetries in practice. We introduce object-centric world models that provide a soft geometric inductive bias without enforcing exact symmetries. By embedding object states as Clifford multivectors, our models encourage physically meaningful transformations while retaining the expressivity needed to handle asymmetric dynamics. We evaluate our approach on 2D rigid-body dynamics, 3D charged-particle systems, and real-world driving trajectories. Compared to both unstructured baselines and strictly equivariant models, our soft Clifford transformer achieves better long-horizon fidelity, particularly in regimes with broken symmetries. These results suggest that geometric algebra offers an effective middle ground, delivering sample-efficient dynamics without the inflexibility of hard mathematical constraints.
|
| 1948 |
Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
2512.15765
|
cs.LG
|
M\'elissa Tamine, Otmane Sakhi, Benjamin Heymann, Maxime Vono, Patrick Loiseau |
Data valuation is a natural framework for understanding which data sources matter most when aligning a Large Language Model (LLM) from multiple sources. The standard game-theoretic approach treats each source, or equivalently each preference dataset, as a play...Data valuation is a natural framework for understanding which data sources matter most when aligning a Large Language Model (LLM) from multiple sources. The standard game-theoretic approach treats each source, or equivalently each preference dataset, as a player in a cooperative game and assigns it a contribution score through the Shapley value. In practice, however, Shapley-based valuation is computationally prohibitive because it requires aligning a separate model for every possible coalition of sources, i.e., an exponential number of alignments. We address this challenge for Direct Alignment Algorithms (DAAs), including IPO, which learn through log-policy ratios with respect to a reference policy. We show that, when a model is aligned sequentially source by source, exact optimization makes each stage contribute additively to the log-probability of a full response, up to a prompt-dependent normalization constant. This allows the log-probability assigned by any coalition to a fixed response to be reconstructed from the base policy and the policies trained on each source individually. This reduces the alignment cost of Shapley-based valuation from exponential to linear, since only one model per source needs to be trained to evaluate coalition scores. We test whether this theoretical property remains approximately valid under finite training across several base models and real-world data sources. We finally compute the Shapley values of these sources under multiple reward models, showing how their estimated contributions vary across evaluation criteria.
|
| 1949 |
Role Support in Knowledge-Graph Error Ranking: Predictor Regimes and Evaluation Policies
2512.22318
|
cs.LG
|
Chorok Lee |
When do role-specific training counts improve knowledge-graph error ranking? We evaluate fifteen from-scratch predictors, nine development-selected continuation stages, and six fixed ensembles across three static graphs. A positive penalty for missing training...When do role-specific training counts improve knowledge-graph error ranking? We evaluate fifteen from-scratch predictors, nine development-selected continuation stages, and six fixed ensembles across three static graphs. A positive penalty for missing training support has no stable error-ranking orientation: stronger ComplEx training reverses its pooled error AUROC from 0.689 to 0.364 on FB15k-237, and from 0.729 to 0.430 in a prospectively specified CoDEx-S check. Pair decomposition shows that reversal also survives pair-weighted within-relation evaluation, but not equal-relation averaging. The statement depends on which relations and query pairs receive weight. Learned continuous counts remain useful: they improve AURC and Brier beyond explicit ensemble disagreement in twelve original ensemble/head comparisons and satisfy all four predefined external-dataset AURC contrasts. Binary indicators add no information beyond their defining counts and have less stable fitting effects. Candidate, target, population and acceptance policies change the comparisons; fixed development thresholds do not preserve a common test answer rate. The external study was fixed before inspection of its test outcomes, whereas subsequent explanations and reused-graph controls are explicitly exploratory. These measurements support relation-aware count baselines and policy-specific support diagnostics, not a new uncertainty decomposition, causal training mechanism, calibrated factual truth, or graph-population guarantee. Saved checkpoints, ledgers and separately authored numerical implementation replays make the scoped findings inspectable.
|
| 1950 |
SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints
2512.23770
|
cs.LG
|
Dominik Wagner, Ankit Kanwar, Luke Ong |
In safety-critical domains, reinforcement learning (RL) systems must satisfy strict, zero-cost safety constraints while achieving meaningful task performance. Existing model-free methods can struggle to achieve high safety without substantially compromising ta...In safety-critical domains, reinforcement learning (RL) systems must satisfy strict, zero-cost safety constraints while achieving meaningful task performance. Existing model-free methods can struggle to achieve high safety without substantially compromising task performance. We introduce \emph{Safety-Biased Trust Region Policy Optimisation (SB-TRPO)}, a principled approach to RL with zero-cost constraints, which requires only a fixed fraction of the maximal cost reduction achievable within the trust region, thus retaining flexibility for reward optimisation. We show that the idealised update nevertheless converges to zero cost and maximal reward amongst zero-cost policies in finite MDPs. A practical gradient-based approximation provides local improvements in both safety and reward under suitable gradient alignment. Experiments on \emph{Safety Gymnasium} demonstrate high safety alongside strong task performance.
|
| 1951 |
Understanding and inverse design of implicit bias in stochastic learning: a geometric perspective
2601.06597
|
cs.LG
|
Nicola Aladrah, Emanuele Ballarin, Matteo Biagetti, Alessio Ansuini, Alberto d'Onofrio |
Can we design a model such that its stochastic training favours a desired class of solutions without enforcing an explicit penalty? Under suitable conditions, the interplay between symmetries of a model's weight parametrization and stochastic training favours ...Can we design a model such that its stochastic training favours a desired class of solutions without enforcing an explicit penalty? Under suitable conditions, the interplay between symmetries of a model's weight parametrization and stochastic training favours particular solutions, inducing an implicit bias. Building on this mechanism, we develop a framework for inverse-designing such biases by constructing novel parametrizations and their associated symmetries. We show how holomorphic functions make this construction and calculation simple and explicit. Specifically, we introduce a new parametrization that biases learned weights toward the binary values $\{-1,+1\}$. Numerical experiments confirm the theoretical predictions. They also show that our parametrization reproduces the same preference induced by an explicitly regularized model without adding a penalty to the training loss.
|
| 1952 |
Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning
2601.06641
|
cs.LG
|
Quan Minh Nguyen, Min-Seon Kim, Hoang M. Ngo, Trong Nghia Hoang, Hyuk-Yoon Kwon |
Membership inference attacks (MIAs) pose a serious privacy threat in federated learning (FL). While MIAs have been extensively studied in standard FL, the recent shift toward federated fine-tuning introduces new and largely unexplored attack surfaces. In this ...Membership inference attacks (MIAs) pose a serious privacy threat in federated learning (FL). While MIAs have been extensively studied in standard FL, the recent shift toward federated fine-tuning introduces new and largely unexplored attack surfaces. In this work, we show that federated prompt-tuning, which adapts pre-trained foundation models using lightweight input prefixes, exposes a novel and effective vector for membership inference. We propose PromptMIA, a membership inference attack tailored to federated prompt-tuning, in which a malicious server introduces adversarially crafted prompts and exploits their updates during collaborative training to determine whether a target data point belongs to a client's private dataset. We formalize this threat via a security game and demonstrate that PromptMIA achieves consistently high attack advantage across diverse benchmark datasets, substantially outperforming current SOTA federated MIAs. We also provide a theoretical lower bound on the attack advantage that explains the observed empirical behavior. Finally, we show that existing MIA defenses are often ineffective against PromptMIA, highlighting the need for defense mechanisms specifically tailored to prompt-tuning in federated settings.
|
| 1953 |
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
2601.18510
|
cs.LG
|
Yibo Li, Zijie Lin, Ailin Deng, Xuan Zhang, Yufei He |
While Large Language Model (LLM) agents excel at general tasks, they inherently struggle with continual adaptation due to the frozen weights after deployment. Conventional reinforcement learning (RL) offers a solution but incurs prohibitive computational costs...While Large Language Model (LLM) agents excel at general tasks, they inherently struggle with continual adaptation due to the frozen weights after deployment. Conventional reinforcement learning (RL) offers a solution but incurs prohibitive computational costs and the risk of catastrophic forgetting. We introduce Just-In-Time Reinforcement Learning (JitRL), a training-free framework that enables test-time policy optimization without any gradient updates. JitRL maintains a dynamic, non-parametric memory of experiences and retrieves relevant trajectories to estimate action advantages on-the-fly. These estimates are then used to directly modulate the LLM's output logits. We theoretically prove that this additive update rule is the exact closed-form solution to the KL-constrained policy optimization objective. Extensive experiments on WebArena and Jericho demonstrate that JitRL establishes a new state-of-the-art among training-free methods. Crucially, JitRL outperforms the performance of computationally expensive fine-tuning methods (e.g., WebRL) while reducing monetary costs by over 30 times, offering a scalable path for continual learning agents. The code is available at https://github.com/liushiliushi/JitRL.
|
| 1954 |
Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference
2601.22132
|
cs.LG
|
Ziming Dong, Hardik Sharma, Evan O'Toole, Jaya Prakash Champati, Kui Wu |
Large Language Models (LLMs) deliver state-of-the-art performance on complex reasoning tasks, but their inference costs limit deployment at scale. Small Language Models (SLMs) offer dramatic cost savings yet lag substantially in accuracy. Existing approaches -...Large Language Models (LLMs) deliver state-of-the-art performance on complex reasoning tasks, but their inference costs limit deployment at scale. Small Language Models (SLMs) offer dramatic cost savings yet lag substantially in accuracy. Existing approaches - routing and cascading - treat the LLM as an all-or-nothing resource: either the query bypasses the LLM entirely, or the LLM generates a complete response at full cost. We introduce LLM Shepherding, a framework that requests only a short prefix (a hint) from the LLM and provides it to SLM. This simple mechanism is surprisingly effective for math and coding tasks: even hints comprising 10-30% of the full LLM response improve SLM accuracy significantly. Shepherding generalizes both routing and cascading, and it achieves lower cost under oracle decision-making. We develop a two-stage predictor that jointly determines whether a hint is needed and how many tokens to request. On the widely-used mathematical reasoning (GSM8K, CNK12) and code generation (HumanEval, MBPP) benchmarks, Shepherding reduces costs by 42-94% relative to LLM-only inference. Compared to state-of-the-art routing and cascading baselines, shepherding delivers up to 2.8x cost reduction while matching accuracy. To our knowledge, this is the first work to exploit token-level budget control for SLM-LLM collaboration.
|
| 1955 |
DP-{\lambda}CGD: Efficient Noise Correlation for Differentially Private Model Training
2601.22334
|
cs.LG
|
Nikita P. Kalinin, Ryan McKenna, Rasmus Pagh, Christoph H. Lampert |
Differentially private stochastic gradient descent (DP-SGD) is the gold standard for training machine learning models with formal differential privacy guarantees. Several recent extensions improve its accuracy by introducing correlated noise across training it...Differentially private stochastic gradient descent (DP-SGD) is the gold standard for training machine learning models with formal differential privacy guarantees. Several recent extensions improve its accuracy by introducing correlated noise across training iterations. Matrix factorization mechanisms are a prominent example, but they can require storing previously added noise vectors, leading to substantial memory overhead. In this work, we propose a new algorithm for correlated-noise private training that correlates noise only with the immediately preceding iteration and cancels a tunable portion of it. By regenerating the previous noise using a pseudorandom number generator rather than storing it, our method has essentially the same memory footprint as DP-SGD. We show that the computational overhead is minimal and empirically demonstrate improved accuracy over DP-SGD.
|
| 1956 |
Environment-Conditioned Tail Reweighting for Invariant Learning under Mixed Shifts
2601.22944
|
cs.LG
|
Yuanchao Wang, Tianqi Zhong, Fengnan Li, Zhao-Rong Lai |
Out-of-distribution generalization becomes challenging when spurious correlations vary across environments while difficult or underrepresented samples remain insufficiently emphasized within them. Invariant learning and reweighting address complementary aspect...Out-of-distribution generalization becomes challenging when spurious correlations vary across environments while difficult or underrepresented samples remain insufficiently emphasized within them. Invariant learning and reweighting address complementary aspects of this mixed-shift problem, yet applying reweighting only to prediction leaves the invariance constraint evaluated on a different training risk. We introduce Environment-Conditioned Tail Reweighting (ECTR), which couples cross-environment TV invariance with within-environment adversarial tail weighting through a shared weighted risk. Sample weights are normalized within each environment, and the resulting risk is used consistently for both predictive learning and the total-variation (TV) stationarity penalty, while an environment-wise KL term controls adversarial concentration. Our analysis characterizes the stationarity-dependent weighting signal and its KL-regularized distributional interpretation, alongside conditional-risk and optimization properties under the stated assumptions. Controlled mixed-shift ablations consistently support the shared-risk coupling across three difficulty settings. Across synthetic and real-world benchmarks, ECTR achieves favorable results under both observed- and inferred-environment settings, including improvements over TV-based parent methods and strong performance against additional invariant-learning baselines. Training-corruption experiments further show that ECTR assigns less excess weight to corrupted samples than the tested fixed-tail rules. Together, these results provide a unified theoretical and empirical account of environment-conditioned tail reweighting for invariant learning under mixed shifts.
|
| 1957 |
Forest-Guided Semantic Transport for Label-Supervised Manifold Alignment
2602.00974
|
cs.LG
|
Adrien Aumon, Myriam Lizotte, Guy Wolf, Kevin R. Moon, Jake S. Rhodes |
Label-supervised manifold alignment bridges the gap between unsupervised and correspondence-based paradigms by leveraging shared label information to align multimodal datasets. However, existing methods either rely on label-independent intra-domain geometry or...Label-supervised manifold alignment bridges the gap between unsupervised and correspondence-based paradigms by leveraging shared label information to align multimodal datasets. However, existing methods either rely on label-independent intra-domain geometry or incorporate supervision primarily through the alignment objective, rather than directly constructing task-aware intra-domain affinities. To address this limitation, we introduce FoSTA (Forest-guided Semantic Transport Alignment), which learns supervised forest geometries independently within each domain and uses their shared class-semantic structure to infer cross-domain correspondences without predefined anchors. FoSTA extends RF-GAP affinities to partially labeled data and aligns the resulting semantic representations through scalable hierarchical transport. Extensive comparisons with established baselines demonstrate strong correspondence recovery and semantic structure preservation on synthetic and real-world benchmarks, including multimodal data integration and single-cell batch correction.
|
| 1958 |
Plain Transformers are Surprisingly Powerful Link Predictors
2602.01553
|
cs.LG
|
Quang Truong, Yu Song, Donald Loveland, Mingxuan Ju, Tong Zhao |
Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies. While Graph Neural Networks (GNNs) are the standard solution, state-of-the-art pipelines often rely on explicit structural h...Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies. While Graph Neural Networks (GNNs) are the standard solution, state-of-the-art pipelines often rely on explicit structural heuristics or memory-intensive node embeddings -- approaches that struggle to generalize or scale to massive graphs. Emerging Graph Transformers (GTs) offer a potential alternative but often incur significant overhead due to complex structural encodings, hindering their applications to large-scale link prediction. We challenge these sophisticated paradigms with PENCIL, an encoder-only plain Transformer that replaces hand-crafted priors with attention over sampled local subgraphs, retaining the scalability and hardware efficiency of standard Transformers. Through experimental and theoretical analysis, we show that PENCIL extracts richer structural signals than GNNs, implicitly generalizing a broad class of heuristics and subgraph-based expressivity. Empirically, PENCIL outperforms heuristic-informed GNNs and is far more parameter-efficient than ID-embedding--based alternatives, while remaining competitive across diverse benchmarks -- even without node features. Our results challenge the prevailing reliance on complex engineering techniques, demonstrating that simple design choices are potentially sufficient to achieve the same capabilities. Our code is publicly available at https://github.com/quang-truong/pencil.
|
| 1959 |
VLM-Guided Experience Replay
2602.01915
|
cs.LG
|
Elad Sharony, Tom Jurgenson, Orr Krupnik, Dotan Di Castro, Shie Mannor |
Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have enabled powerful semantic and multimodal reasoning capabilities, creating new opportunities to enhance sample efficiency, high-level planning, and interpretability in reinfo...Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have enabled powerful semantic and multimodal reasoning capabilities, creating new opportunities to enhance sample efficiency, high-level planning, and interpretability in reinforcement learning (RL). While prior work has integrated LLMs and VLMs into various components of RL, the replay buffer, a core component for storing and reusing experiences, remains unexplored. We propose addressing this gap by leveraging VLMs to guide the prioritization of experiences in the replay buffer. Our key idea is to use a frozen, pre-trained VLM as an automated evaluator to identify and prioritize promising sub-trajectories from the agent's experiences. Across scenarios, including game-playing and robotics, spanning both discrete and continuous domains, agents trained with our proposed prioritization method achieve 15-57% higher average success rates and improve sample efficiency by 35-55% compared to previous approaches. Project page: https://esharony.me/projects/vlm-rb/
|
| 1960 |
Poly-attention: a general scheme for higher-order self-attention
2602.02422
|
cs.LG
|
Sayak Chakrabarti, Toniann Pitassi, Josh Alman |
The self-attention mechanism, at the heart of the Transformer model, is able to effectively model pairwise interactions between tokens. However, numerous recent works have shown that it is unable to perform basic tasks involving detecting triples of correlated...The self-attention mechanism, at the heart of the Transformer model, is able to effectively model pairwise interactions between tokens. However, numerous recent works have shown that it is unable to perform basic tasks involving detecting triples of correlated tokens, or compositional tasks where multiple input tokens need to be referenced to generate a result. Some higher-dimensional alternatives to self-attention have been proposed to address this, including higher-order attention and Strassen attention, which can perform some of these polyadic tasks in exchange for slower, superquadratic running times. In this work, we define a vast class of generalizations of self-attention, which we call poly-attention mechanisms. Our mechanisms can incorporate arbitrary higher-order (tensor) computations as well as arbitrary relationship structures between the input tokens, and they include the aforementioned alternatives as special cases. We then systematically study their computational complexity and representational strength, including giving new algorithms and matching complexity-theoretic lower bounds on the time complexity of computing the attention matrix exactly as well as approximately, and tightly determining which polyadic tasks they can each perform. Our results give interesting trade-offs between different desiderata for these mechanisms, including a tight relationship between how expressive a mechanism is, and how large the coefficients in the model may be so that the mechanism can be approximated in almost-linear time. Notably, we give a new attention mechanism which can be computed exactly in quadratic time, and which can perform function composition for any fixed number of functions. Prior mechanisms, even for just composing two functions, could only be computed in superquadratic time, and our new lower bounds show that faster algorithms for them are not possible.
|
| 1961 |
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
2602.05548
|
cs.LG
|
Zhiqi Yu, Zhangquan Chen, Mengting Liu, Heye Zhang, Liangqiong Qu |
Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO, has become the standard for eliciting LLM reasoning. However, its efficiency in exploration and difficulty adaptation remains an open challenge. In this work, we identify an implicit adv...Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO, has become the standard for eliciting LLM reasoning. However, its efficiency in exploration and difficulty adaptation remains an open challenge. In this work, we identify an implicit advantage symmetry inherent in Group Relative Advantage Estimation (GRAE) as a structural property that provides a new perspective for understanding these bottlenecks. This symmetry induces two critical limitations: (i) at the group level, strict symmetry in weights between correct and incorrect trajectories leaves unsampled action logits unchanged, thereby hindering the exploration of novel correct solution. (ii) at the sample level, the algorithm implicitly prioritizes medium-difficulty samples, remaining agnostic to the non-stationary demands of difficulty focus. Through controlled experiments, we reveal that this symmetric property is sub-optimal, yielding two pivotal insights: (i) asymmetrically down-weighting the advantages of correct trajectories encourages essential exploration but risks instability; (ii) learning efficiency can be boosted by a curriculum-like transition-prioritizing simpler samples initially before gradually shifting to complex ones. Motivated by these findings, we propose Asymmetric GRAE (A-GRAE), which dynamically modulates exploration incentives and sample-difficulty focus. Experiments across seven benchmarks demonstrate that A-GRAE consistently improves GRPO and its variants across both LLMs and MLLMs. Code is available at https://github.com/HKU-HealthAI/A-GRAE
|
| 1962 |
$f$-FUM: Federated Unlearning via min--max and $f$-divergence
2602.06187
|
cs.LG
|
Radmehr Karimian, Amirhossein Bagheri, Meghdad Kurmanji, Nicholas D. Lane, Gholamali Aminian |
Federated learning (FL) enables collaborative training while keeping raw data on clients. Deletion requests and the discovery of poisoned data create a need to remove selected contributions from an already trained model. This is challenging in FL because clien...Federated learning (FL) enables collaborative training while keeping raw data on clients. Deletion requests and the discovery of poisoned data create a need to remove selected contributions from an already trained model. This is challenging in FL because client data are decentralized and their influence is entangled through repeated aggregation. We present f-FUM, an active federated unlearning method that builds on teacher-student forget/retain optimization. Starting from a pretrained global model, the method keeps that model fixed as a teacher. Clients maximize an $f$-divergence between student and teacher predictive distributions on forget examples, then minimize KL-based teacher-student disagreement and supervised loss on retained examples. The updates are aggregated in separate, sample-weighted synchronization phases without moving raw data to the server or changing the model architecture. We evaluate forget-side divergence choices in client-level and data-level deletion settings involving backdoors, label confusion, and clean-data deletion. Under equal synchronization-phase budgets, divergence choice changes the forgetting-utility trade-off, yielding improvements in some settings and mixed results in others. In the evaluated configurations, f-FUM uses up to 8 times fewer synchronization phases than full federated retraining.
|
| 1963 |
SOCKET: SOft Collision Kernel EsTimator for Sparse Attention
2602.06283
|
cs.LG
|
Sahil Joshi, Agniva Chowdhury, Wyatt Bellinger, Amar Kanakamedala, Ekam Singh |
Exploiting sparsity is key to efficient long-context inference, as attention dominates the cost of autoregressive decoding. Sparse attention reduces this cost by restricting computation to a subset of tokens, but its effectiveness hinges on fast and accurate t...Exploiting sparsity is key to efficient long-context inference, as attention dominates the cost of autoregressive decoding. Sparse attention reduces this cost by restricting computation to a subset of tokens, but its effectiveness hinges on fast and accurate token scoring and selection at inference time. Data-agnostic approaches offer an attractive way to perform this selection, but often incur substantial memory overhead to maintain high recall. We revisit Locality-Sensitive Hashing (LSH) and introduce SOCKET, a SOft Collision Kernel EsTimator that replaces hard bucket matches with probabilistic, similarity-aware aggregation. Traditional LSH relies on binary collision signals, providing limited information for ranking tokens and necessitating many hash tables for accurate retrieval. In contrast, soft LSH accumulates graded collision evidence across hash tables, closely preserving the true top-$k$ ordering with significantly less memory. This reframes LSH from a candidate-generation mechanism into a principled scoring kernel for sparse attention. Building on this insight, SOCKET enables efficient token selection without ad hoc voting and matches or outperforms existing sparse attention methods across multiple long-context benchmarks and diverse language models. With a custom set of CUDA/Triton kernels for scoring, selection, and attention, SOCKET achieves up to approximately $1.5\times$ higher throughput than FlashAttention. Code is open-sourced at https://github.com/amarka8/SOCKET.
|
| 1964 |
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
2602.08847
|
cs.LG
|
Lang Feng, Longtao Zheng, Shuo He, Fuxiang Zhang, Bo An |
Multi-agent LLM systems enable advanced reasoning and tool use via role specialization, yet reliable reinforcement learning (RL) post-training for such systems remains difficult. In this work, we theoretically pinpoint a key reason for training instability whe...Multi-agent LLM systems enable advanced reasoning and tool use via role specialization, yet reliable reinforcement learning (RL) post-training for such systems remains difficult. In this work, we theoretically pinpoint a key reason for training instability when extending group-based RL to multi-agent LLM systems. We show that under GRPO-style optimization, a global normalization baseline may deviate from diverse agents' reward distributions, which ultimately leads to gradient-norm instability. Based on this finding, we propose Dr. MAS, a simple and stable RL training recipe for multi-agent LLM systems. Dr. MAS uses an agent-wise remedy: normalizing advantages per agent using each agent's own reward statistics, which calibrates gradient scales and dramatically stabilizes training, both theoretically and empirically. Beyond the algorithm, Dr. MAS provides an end-to-end RL training framework for multi-agent LLM systems, supporting scalable orchestration, flexible per-agent LLM serving and optimization configs, and shared resource scheduling of LLM actor backends. We evaluate Dr. MAS on multi-agent math reasoning and multi-turn search benchmarks using Qwen2.5 and Qwen3 series models. Dr. MAS achieves clear gains over vanilla GRPO (e.g., +5.6\% avg@16 and +4.6\% pass@16 on math, and +15.2\% avg@16 and +13.1\% pass@16 on search) while largely eliminating gradient spikes. Moreover, it remains highly effective under heterogeneous agent-model assignments while improving efficiency.
|
| 1965 |
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
2602.12029
|
cs.LG
|
Sunghyeon Woo, Hoseung Kim, Sunghwan Shim, Minjung Jo, Hyunjoon Jeong |
Multi-agent systems increasingly orchestrate multiple specialized language models to solve complex real-world problems, often invoking them over a shared context. This execution pattern repeatedly processes the same prompt prefix across models. Consequently, e...Multi-agent systems increasingly orchestrate multiple specialized language models to solve complex real-world problems, often invoking them over a shared context. This execution pattern repeatedly processes the same prompt prefix across models. Consequently, each model redundantly executes the prefill stage and maintains its own key-value (KV) cache, increasing aggregate prefill load and worsening tail latency by intensifying prefill-decode interference in existing LLM serving stacks. Disaggregated serving reduces such interference by placing prefill and decode on separate GPUs, but disaggregation does not fundamentally eliminate inter-model redundancy in computation and KV storage for the same prompt. To address this issue, we propose PrefillShare, a novel algorithm that enables sharing the prefill stage across multiple fine-tuned models in a disaggregated setting. PrefillShare factorizes the model into prefill and decode modules, freezes the prefill module, and fine-tunes only the decode module. This design allows multiple task-specific models to share a prefill module and the KV cache generated for the same prompt. We further introduce a routing mechanism that enables effective prefill sharing in a vLLM-based disaggregated system. PrefillShare not only matches full fine-tuning accuracy on a broad range of tasks and models, but also delivers 4.5x lower p95 latency and 2.0x higher throughput in multi-model agent workloads.
|
| 1966 |
Preventing Rank Collapse in Federated Low-Rank Adaptation with Client Heterogeneity
2602.13486
|
cs.LG
|
Fei Wu, Jia Hu, Geyong Min, Shiqiang Wang |
Federated low-rank adaptation (FedLoRA) has facilitated communication-efficient and privacy-preserving fine-tuning of foundation models for downstream tasks. In practical federated learning scenarios, client heterogeneity in system resources and data distribut...Federated low-rank adaptation (FedLoRA) has facilitated communication-efficient and privacy-preserving fine-tuning of foundation models for downstream tasks. In practical federated learning scenarios, client heterogeneity in system resources and data distributions motivates the use of heterogeneous LoRA ranks across clients. However, we identify a previously overlooked phenomenon in heterogeneous FedLoRA with SVD-based allocation, termed rank collapse, where the energy of the global update becomes concentrated in the minimum shared rank, resulting in suboptimal performance and high sensitivity to rank configurations. Through theoretical analysis, we reveal the underlying mechanism of rank collapse: a mismatch between rank-agnostic aggregation weights and rank-dependent client contributions, which systematically suppresses higher-rank updates at a geometric rate over rounds. Motivated by this insight, we propose raFLoRA, a rank-partitioned aggregation method that decomposes local updates into rank partitions and then aggregates each partition weighted by its effective client contributions. Extensive experiments across vision, language, and reasoning tasks show that raFLoRA prevents rank collapse, improves model performance, and enhances robustness across diverse heterogeneous configurations compared with strong FedLoRA baselines.
|
| 1967 |
Pseudo-differential-enhanced physics-informed neural networks
2602.14663
|
cs.LG
|
Andrew Gracyk |
We present pseudo-differential enhanced physics-informed neural networks (PINNs), an extension of gradient enhancement but in Fourier space. Gradient enhancement of PINNs dictates that the PDE residual is taken to a higher differential order than prescribed by...We present pseudo-differential enhanced physics-informed neural networks (PINNs), an extension of gradient enhancement but in Fourier space. Gradient enhancement of PINNs dictates that the PDE residual is taken to a higher differential order than prescribed by the PDE, added to the objective as an augmented term in order to improve training and overall learning fidelity. We propose the same procedure after application via Fourier transforms, since differentiating in Fourier space is multiplication with the Fourier wavenumber under suitable decay. Our methods are fast and efficient. Our methods oftentimes achieve superior PINN versus numerical error in fewer training iterations, potentially pair well with few/fixed samples in collocation, and can achieve breakthrough rather than gradual effects on the error. Moreover, our methods are suitable for fractional derivatives. We establish that our methods, due to the dynamical effects, improve spectral eigenvalue decay of the neural tangent kernel (NTK), and so our methods contribute towards the learning of high frequencies in early training, mitigating the effects of frequency bias up to the polynomial order and possibly greater with smooth activations. In particular, our primary contribution can be characterized as gradient enhancement affects spectral bias, and we specialize to Fourier space, although our theory does not prove the ambient case exactly due to a result with the Plancherel theorem on the descent paths. Our methods accommodate advanced techniques in PINNs, such as Fourier feature embeddings. A pitfall of discrete Fourier transforms via the Fast Fourier Transform (FFT) is mesh subjugation, and so we demonstrate compatibility of our methods for greater mesh flexibility and invariance on alternative Euclidean and non-Euclidean domains via Monte Carlo methods and otherwise, although possessing trade-offs of their own.
|
| 1968 |
Use What You Know: Causal Foundation Models with Partial Graphs
2602.14972
|
cs.LG
|
Arik Reuter, Anish Dhir, Cristiana Diaconu, Jake Robertson, Ole Ossen |
Estimating causal quantities traditionally relies on bespoke estimators tailored to specific assumptions. Recently proposed Causal Foundation Models (CFMs) promise a more unified approach by amortising causal discovery and inference in a single step. However, ...Estimating causal quantities traditionally relies on bespoke estimators tailored to specific assumptions. Recently proposed Causal Foundation Models (CFMs) promise a more unified approach by amortising causal discovery and inference in a single step. However, in their current state, they do not allow for the incorporation of any domain knowledge, which can lead to suboptimal predictions. We bridge this gap by introducing methods to condition CFMs on causal information, such as the causal graph or more readily available ancestral information. When access to complete causal graph information is too strict a requirement, our approach also effectively leverages partial causal information. We systematically evaluate conditioning strategies and find that injecting learnable biases into the attention mechanism, together with a graph-convolutional encoder, is a highly effective method to utilise full and partial causal information. Our experiments show that this conditioning allows a general-purpose CFM to match the performance of specialised models trained on specific causal structures. Overall, our approach addresses a central hurdle on the path towards all-in-one causal foundation models: the capability to answer causal queries in a data-driven manner while effectively leveraging any amount of domain expertise.
|
| 1969 |
Neural Proposals, Symbolic Guarantees: Neuro-Symbolic Graph Generative Modeling
2602.16954
|
cs.LG
|
Chuqin Geng, Li Zhang, Mark Zhang, Zhaoyue Wang, Haolin Ye |
While deep generative models excel at capturing graph data distributions, they struggle to satisfy complex, hard constraints. In unconstrained settings, these models typically produce valid topologies; yet imposing strict compositional rules, like those in dru...While deep generative models excel at capturing graph data distributions, they struggle to satisfy complex, hard constraints. In unconstrained settings, these models typically produce valid topologies; yet imposing strict compositional rules, like those in drug discovery, creates an out-of-distribution (OOD) setting where purely neural methods frequently fail. Because these neural approaches rely on soft conditioning and post-hoc filtering on such tasks, they cannot provide the formal guarantees needed for high-stakes domains. To address this, we introduce Neuro-Symbolic Graph Generative Modeling (NSGGM), a framework built on the principle of Neural Proposals, Symbolic Guarantees. NSGGM decouples generation: an autoregressive model proposes structural scaffolds, and a Satisfiability Modulo Theories (SMT) solver handles the final discrete assembly of the proposed substructures. Empirically, NSGGM is competitive with state-of-the-art methods on unconstrained tasks. To evaluate logical-constraint satisfaction inspired by real drug discovery workflows, we introduce MolSAT, a benchmark for hard compositional rules. On MolSAT, purely neural baselines completely fail OOD (0% satisfaction with zero training support), while NSGGM achieves >95% satisfaction in-distribution and 64-86% with zero training support.
|
| 1970 |
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
2602.17684
|
cs.LG
|
Xiao Zhu, Xinyu Zhou, Boyu Zhu, Hanxu Hu, Mingzhe Du |
Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and reliability of high-...Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and reliability of high-quality test cases. We propose CodeScaler, a reward model designed to scale both reinforcement learning training and test-time inference for code generation. CodeScaler is trained on carefully curated preference data derived from verified code problems and incorporates syntax-aware code extraction and validity-preserving reward shaping to ensure stable and robust optimization. Across four coding benchmarks, CodeScaler consistently outperforms execution-based RL by +1.55 points on Qwen3-8B-Base and +4.23 points on Qwen3-14B-Base. By further scaling to 44K problems with additional synthetic data, CodeScaler yields +14.64 points improvement over the base model without requiring any test cases. At inference time, CodeScaler serves as an effective test-time scaling method, achieving performance comparable to unit test approaches while providing a 10-fold reduction in latency. Moreover, CodeScaler surpasses existing reward models on RM-Bench not only in the code domain (+3.3 points), but also in general and reasoning domains (+2.7 points on average).
|
| 1971 |
Don't stop me now: How Validation Criteria Affect Checkpoint Selection and Early Stopping
2602.22107
|
cs.LG
|
Andrea Apicella, Francesco Isgr\`o, Andrea Pollastro, Roberto Prevete |
Checkpoint selection is a standard component of neural network training, yet the validation criterion used to select a checkpoint is often chosen heuristically. Moreover, the same criterion may be used either only to rank checkpoints after completion of a pred...Checkpoint selection is a standard component of neural network training, yet the validation criterion used to select a checkpoint is often chosen heuristically. Moreover, the same criterion may be used either only to rank checkpoints after completion of a predefined training run or also to determine when training should stop, thereby affecting both the selected checkpoint and the set of checkpoints available for selection. In this work, we systematically investigate the role of validation criteria under these two settings. We separately vary the training loss, the validation criterion, and the target evaluation metric, and compare post-hoc checkpoint selection, in which training proceeds for all predefined epochs, with patience-based early stopping, in which the validation criterion also controls training termination. We consider three Cross-Entropy, C-Loss, and PolyLoss as training losses, and accuracy, macro-F1, and Matthews correlation coefficient as target metrics. Selection quality is assessed through the relative gap between the test performance of the validation-selected checkpoint and the best-observed test performance for the same target metric over the complete predefined training run.
|
| 1972 |
Forecasting Bacterial Antimicrobial Resistance Trends Using Machine Learning on WHO GLASS Surveillance Data: A Retrieval-Augmented Generation Approach for Policy Decision Support
2602.22673
|
cs.LG
|
Md Tanvir Hasan Turja |
Background: Antimicrobial resistance (AMR) is a global health threat. While the WHO Global Antimicrobial Resistance and Use Surveillance System (GLASS) provides standardized data, population-level machine learning forecasting of resistance trends remains limit...Background: Antimicrobial resistance (AMR) is a global health threat. While the WHO Global Antimicrobial Resistance and Use Surveillance System (GLASS) provides standardized data, population-level machine learning forecasting of resistance trends remains limited. Translating forecasts into policy also requires transparent interpretation. Methods: Surveillance data (2021-2023) comprising 5,909 observations across 44 countries and five WHO regions were processed. Six models (Naive, Linear, Ridge, XGBoost, LightGBM, LSTM) were benchmarked to forecast one-year-ahead resistance rates using prior-year resistance, antibiotic consumption, and related features. MAE, RMSE, and sMAPE were computed with 95% bootstrap confidence intervals for MAE. A local Retrieval-Augmented Generation (RAG) system (Gemma 4) translated forecast findings into policy guidance grounded in WHO documents. Results: XGBoost achieved the best predictive performance (test MAE = 6.13% [95% CI: 5.83-6.44]), an 85.3% error reduction versus the naive baseline (MAE = 41.79%). SHAP analysis identified prior-year resistance as the dominant predictor (50.5% gain). Regional forecast error tracked surveillance coverage, from 3.65% in the European Region to 8.61% in South-East Asia. The RAG pipeline generated source-attributed policy responses without fabricated citations. Conclusion: Short-term AMR resistance rates exhibit strong temporal autocorrelation that can be accurately forecasted using gradient boosting. Coupling these forecasts with a hallucination-resistant RAG system provides a scalable, evidence-based decision-support framework for AMR governance. v3 correction: a leakage audit found feature-construction defects; on the corrected panel a persistence forecast (MAE 5.79) significantly outperforms XGBoost (6.55), and the 85.3% claim is withdrawn (see Correction).
|
| 1973 |
LFPO: Likelihood-Free Policy Optimization for Masked Diffusion Models
2603.01563
|
cs.LG
|
Chenxing Wei, Jiazhen Kang, Hong Wang, Jianqing Zhang, Hao Jiang |
Reinforcement Learning with Verifiable Rewards (RLVR) has achieved remarkable success in improving autoregressive models, especially in domains requiring correctness like mathematical reasoning and code generation. However, directly applying such paradigms to ...Reinforcement Learning with Verifiable Rewards (RLVR) has achieved remarkable success in improving autoregressive models, especially in domains requiring correctness like mathematical reasoning and code generation. However, directly applying such paradigms to Diffusion Large Language Models (dLLMs) is fundamentally hindered by the intractability of exact likelihood computation, which forces existing methods to rely on high-variance approximations. To bridge this gap, we propose Likelihood-Free Policy Optimization (LFPO), a native framework that maps the concept of vector field flow matching to the discrete token space. Specifically, LFPO formulates alignment as geometric velocity rectification, which directly optimizes denoising logits via contrastive updates. This design effectively bypasses the errors inherent in likelihood approximation, yielding the precise gradient estimation. Furthermore, LFPO enforce consistency by predicting final solutions from intermediate steps, effectively straightening the probability flow to enable high-quality generation with significantly fewer iterations. Extensive experiments demonstrate that LFPO not only outperforms state-of-the-art baselines on code and reasoning benchmarks but also accelerates inference by approximately 20% through reduced diffusion steps.
|
| 1974 |
Physics-Informed Neural Networks with Architectural Physics Embedding for Large-Scale Wave Field Reconstruction
2603.02231
|
cs.LG
|
Huiwen Zhang, Feng Ye, Chu Ma |
Large-scale wave field reconstruction requires precise solutions but faces challenges with computational efficiency and accuracy. The physics-based numerical methods like Finite Element Method (FEM) provide high accuracy but struggle with large-scale or high-f...Large-scale wave field reconstruction requires precise solutions but faces challenges with computational efficiency and accuracy. The physics-based numerical methods like Finite Element Method (FEM) provide high accuracy but struggle with large-scale or high-frequency problems due to prohibitive computational costs. Pure data-driven approaches excel in speed but often lack sufficient labeled data for complex scenarios. Physics-informed neural networks (PINNs) integrate physical principles into machine learning models, offering a promising solution by bridging these gaps. However, standard PINNs embed physical principles only in loss functions, leading to slow convergence, optimization instability, and spectral bias, limiting their ability for large-scale wave field reconstruction. This work introduces architectural physics embedded (PE)-PINN, which integrates additional physical guidance directly into the neural network architecture beyond Helmholtz equations and boundary conditions in loss functions. Specifically, a new envelope transformation layer is designed to mitigate spectral bias with kernels parameterized by source properties, material interfaces, and wave physics. Experiments demonstrate that PE-PINN achieves more than 10 times speedup in convergence compared to standard PINNs and several orders of magnitude reduction in memory usage compared to FEM. This breakthrough enables high-fidelity modeling for large-scale 2D/3D electromagnetic wave reconstruction involving reflections, refractions, and diffractions in room-scale domains, readily applicable to wireless communications, sensing, room acoustics, and other fields requiring large-scale wave field analysis.
|
| 1975 |
The Trace Is the State: Exact Credit Assignment for LLM Agent Teams
2603.06859
|
cs.LG
|
Yanjun Chen, Yirong Sun, Hanlin Wang, Jinghan Wang, Xinming Zhang |
Credit assignment for a team of LLM agents, what each message was worth, has mostly been treated as prediction: a learned critic, a trajectory-level score, or an agent ablation stands in for the counterfactual that defines credit, which is rarely run. Teams th...Credit assignment for a team of LLM agents, what each message was worth, has mostly been treated as prediction: a learned critic, a trajectory-level score, or an agent ablation stands in for the counterfactual that defines credit, which is rarely run. Teams that communicate through a shared context are different: when everything a downstream agent reads is written into the trace, the trace is the state, and that counterfactual can be executed. A credit signal can then be judged as any estimator is: by bias, variance, and agreement with an independent reference. C3, credit assignment by counterfactual continuation, substitutes one message at a decision point and continues the run to the terminal reward, so its credit is unbiased, exact up to Monte Carlo error. Given the sampled alternatives, that error's variance follows a derived law with no term for the number of agents, and the observed noise follows the law on 6 workflows of 2 to 10 decision points. On the 2-agent chain, from 4 continuations, C3 ranks alternatives at 0.69 rank correlation against a 16-continuation reference, near the 0.73 at which 2 such references agree; a critic trained on the same rollouts reaches 0.29. Used as the advantage in training a 2-agent team, C3 beats MAGRPO, the stronger baseline in our comparison, and spends 37% fewer training tokens than MAPPO, since only the messages downstream of a decision point are regenerated. When the trace is the state, credit need not be predicted; it can be exact.
|
| 1976 |
HO-SFL: Hybrid-Order Split Federated Learning with Backprop-Free Clients and Dimension-Free Aggregation
2603.14773
|
cs.LG
|
Qiyuan Chen, Xian Wu, Yi Wang, Xianhao Chen |
Fine-tuning large models on edge devices is severely hindered by the memory-intensive backpropagation (BP) in standard frameworks like federated learning and split learning. While substituting BP with zeroth-order optimization can significantly reduce memory f...Fine-tuning large models on edge devices is severely hindered by the memory-intensive backpropagation (BP) in standard frameworks like federated learning and split learning. While substituting BP with zeroth-order optimization can significantly reduce memory footprints, it typically suffers from prohibitively degraded convergence speed. To resolve this dilemma, we propose Hybrid-Order Split Federated Learning (HO-SFL). By reformulating the split learning process within a Lagrangian framework, HO-SFL decouples the optimization landscape: The server performs precise first-order updates (i.e., BP), whereas clients conduct memory-efficient zeroth-order optimization. This hybrid design not only eliminates the need for client-side BP but also enables dimension-free model aggregation, drastically lowering communication costs. Crucially, we provide a theoretical convergence analysis, demonstrating that HO-SFL mitigates the dimension-dependent convergence slowdown of zeroth-order optimization, achieving a convergence rate comparable to first-order methods. Extensive experiments on tasks across vision and language modalities validate that HO-SFL achieves convergence speeds comparable to first-order baselines while significantly reducing communication costs and client memory footprints.
|
| 1977 |
Models Designed to Forget: Machine Unlearning via Key Deletion
2603.15033
|
cs.LG
|
Sonia Laguna, Jorge da Silva Goncalves, Moritz Vandenhirtz, Alain Ryser, Irene Cannistraci |
Machine unlearning for vision models is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training images. Despite this, most existing approximate unlearning methods tackle the pro...Machine unlearning for vision models is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training images. Despite this, most existing approximate unlearning methods tackle the problem from a post-hoc perspective. They attempt to erase the influence of targeted samples through parameter updates that typically require access to the full training data. This creates a mismatch with real deployment scenarios where unlearning requests can be anticipated, revealing a fundamental limitation of post-hoc approaches. We motivate unlearning by design, a novel paradigm for approximate methods in which models are directly trained to support forgetting as an inherent architectural capability. We instantiate this idea with Machine UNlearning via KEY deletion (MUNKEY), a memory-augmented transformer that decouples instance-specific memorization from model weights. Here, unlearning corresponds to removing the instance-identifying key, enabling zero-shot forgetting without weight updates or access to the original samples or labels. Across natural image benchmarks, fine-grained visual recognition, and medical datasets, MUNKEY outperforms all post-hoc baselines. Our results establish that unlearning by design enables fast, deployment-oriented unlearning while preserving predictive performance.
|
| 1978 |
Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
2603.16731
|
cs.LG
|
Kristi Topollai, Anna Choromanska |
Quantizing optimizer states is becoming an important ingredient of memory-efficient large-scale pre-training, but the resulting optimizer dynamics remain only partially understood. We study low-precision exponential moving average (EMA) optimizer states and sh...Quantizing optimizer states is becoming an important ingredient of memory-efficient large-scale pre-training, but the resulting optimizer dynamics remain only partially understood. We study low-precision exponential moving average (EMA) optimizer states and show how quantization can cause many nominal updates to round back to the same stored value, making the state effectively stale and slowing adaptation beyond what the nominal decay would suggest. We then develop a simple predictive model of stalling that estimates one-step stalling probabilities and characterizes how stalling builds up over time after the initialization. This perspective provides a mechanistic explanation for why optimizer-state resets help in low precision: once a quantized EMA becomes effectively stale, resetting it can temporarily restore responsiveness. Motivated by this picture, we derive a simple theory-guided method for choosing useful reset periods, showing that in low precision the key question is not only whether resets help, but when they should be applied. Experiments in controlled simulations and LLM pre-training show that suitable reset schedules recover the performance lost to low-precision state storage while substantially reducing optimizer-state memory.
|
| 1979 |
Discovering What You Can Control: Interventional Boundary Discovery for Reinforcement Learning
2603.18257
|
cs.LG
|
Jiaxin Liu, Anzhe Cheng, Paul Bogdan |
When an RL agent's observations contain distractors driven by the same confounders as its true state, observational data alone cannot identify which dimensions the agent controls. In our benchmarks, even state-conditioned observational selectors can collapse w...When an RL agent's observations contain distractors driven by the same confounders as its true state, observational data alone cannot identify which dimensions the agent controls. In our benchmarks, even state-conditioned observational selectors can collapse when distractors mimic controllable state variables. We propose Interventional Boundary Discovery (IBD), which treats the agent's own action channel as a source of randomized interventions: randomizing actions implements an interventional contrast, and per-dimension two-sample tests with FDR correction produce a binary mask over observation dimensions. Across 12 continuous-control settings with up to 100 distractors, IBD matches oracle return in 11 of 12 settings, while observational baselines including mutual information, state-conditioned forward models, and gradient-based sensitivity often underperform simply passing the full observation to SAC. Code is available at https://github.com/jiaxin26/IBD-RL
|
| 1980 |
Binary Classification from Coupled Pairwise Labels
2603.19713
|
cs.LG
|
Tomoya Tate, Kosuke Sugiyama, Masato Uchida |
Even when it is difficult to assign absolute class labels to individual instances, relational information may still be available, such as whether two instances belong to the same class or which instance is more likely to belong to the positive class. In this s...Even when it is difficult to assign absolute class labels to individual instances, relational information may still be available, such as whether two instances belong to the same class or which instance is more likely to belong to the positive class. In this study, we refer to these two types of information as Similarity/Dissimilarity (SD) labels and Pairwise Comparison (Pcomp) labels, respectively, and consider binary classification that uses both types of relational information from the same instance pairs. SD learning uses the distinction between similar and dissimilar pairs but does not use the ordering within each pair, whereas Pcomp learning uses the ordering within each pair but does not distinguish between similar and dissimilar pairs. We therefore propose SD-Pcomp learning, whose objective function simultaneously preserves the structures of both SD learning and Pcomp learning. The proposed objective function admits two decompositions: one consists of an SD estimator plus a term that represents ordering information from Pcomp labels, and the other consists of a Pcomp estimator plus a term that represents pair-type information from SD labels. These decompositions clarify how the complementary information provided by SD and Pcomp labels is integrated into the proposed objective function. Experiments on eight datasets compare the proposed method with SD learning, Pcomp learning, and a method that takes a convex combination of their objective functions. We evaluate the effect of using both types of relational information on classification performance in terms of classification accuracy and AUC.
|
| 1981 |
Joint Surrogate Learning of Objectives, Constraints, and Sensitivities for Efficient Multi-objective Optimization of Neural Dynamical Systems
2603.20984
|
cs.LG
|
Frithjof Gressmann, Ivan Georgiev Raikov, Seung Hyun Kim, Mattia Gazzola, Lawrence Rauchwerger |
Gaussian process surrogates dominate constrained multi-objective optimization because they are effective in data-scarce regimes, but their cubic scaling in training samples limits their ability to capture shared structure between objectives and constraints as ...Gaussian process surrogates dominate constrained multi-objective optimization because they are effective in data-scarce regimes, but their cubic scaling in training samples limits their ability to capture shared structure between objectives and constraints as problems grow in dimensionality. We show that deterministic neural network surrogates, equipped with feature tokenization and adaptive output normalization, match or exceed Gaussian process accuracy, while scaling to high-dimensional output spaces and training on all data including infeasible samples. Jointly training a single Feature Tokenizer Transformer to predict objectives, constraint satisfaction, and parameter sensitivities yields a unified gradient that simultaneously improves objective values, steers toward feasibility, and identifies the most influential parameters: a coherent search signal that disjoint per-output models cannot provide. We validate this on biophysical neural optimization problems of increasing complexity. In the hardest regime, with wide, uninformed parameter bounds where random sampling finds zero feasible solutions, descending the surrogate's learned constraint gradient steers the search into the feasible region and recovers near-optimal solutions where standard surrogate optimization and constrained Bayesian optimization find none.
|
| 1982 |
P^2O: Joint Policy and Prompt Optimization
2603.21877
|
cs.LG
|
Xinyu Lu, Kaiqi Zhang, Jinglin Yang, Boxi Cao, Yaojie Lu |
Reinforcement Learning with Verifiable Rewards (RLVR) enhances Large Language Model (LLM) reasoning but is suffer from advantage collapse: when all rollouts of a query receive identical rewards, the group variance vanishes, most damagingly on hard samples, whe...Reinforcement Learning with Verifiable Rewards (RLVR) enhances Large Language Model (LLM) reasoning but is suffer from advantage collapse: when all rollouts of a query receive identical rewards, the group variance vanishes, most damagingly on hard samples, where scaling rollout budgets yields little. We introduce Joint Policy and Prompt Optimization (P O) to mitigate this collapse by alternating continuous policy updates with discrete prompt evolution. P O mines hard samples with a success-rate threshold, evolves reasoning prompts for them with GEPA, and internalizes the elicited trajectories via context distillation, which optimizes each trajectory under the original query and thus removes inference-time prompting, with a Context Ratio Mask (CRM) filtering out extreme likelihood ratios. P O restores critical advantage signals and surpasses the GRPO baseline by up to 8.2 points in average accuracy on six held-out benchmarks across all training datasets and backbones, while also outperforming DAPO and other baselines. The gains are especially pronounced on hard benchmarks, reaching up to 16.3 points above GRPO on average across AIME24 and AIME25. Our findings expose the limits of standard exploration in sparse-reward environments, illuminating the potential of unifying evolutionary algorithms with reinforcement learning. This integration of discrete semantic search and continuous parameter updates provides a self-reinforcing framework that facilitates more effective LLM alignment.
|
| 1983 |
Do Papers Tell the Whole Story? A Benchmark and Framework for Uncovering Hidden Implementation Gaps in Bioinformatics
2603.22018
|
cs.LG
|
Tianxiang Xu, Xiaoyan Zhu, Xin Lai, Xin Lian, Sizhe Dang |
As bioinformatics software is increasingly applied across a broader range of scenarios and the rapid development of large language models (LLMs) further lowers the barriers to software use and development, the composition of the bioinformatics research communi...As bioinformatics software is increasingly applied across a broader range of scenarios and the rapid development of large language models (LLMs) further lowers the barriers to software use and development, the composition of the bioinformatics research community is undergoing substantial change. Consequently, a growing number of researchers require a deeper understanding of methodological details and software behavior. In this context, systematically analyzing the relationship between paper descriptions and code implementations is emerging as an important new challenge in the field. To address this challenge, we introduce paper-code consistency analysis as a new research perspective and construct BioCon, the first benchmark dataset for paper-code consistency analysis in bioinformatics. Furthermore, we develop a unified cross-modal analysis framework to systematically investigate this problem from three perspectives: sentence-level detection, cross-modal retrieval, and project-level assessment. Experimental results demonstrate that the proposed framework can effectively model the semantic relationships between scientific publications and software implementations. Further case studies reveal that paper-code inconsistency is not a single phenomenon but arises from multiple underlying causes, among which Author-Perceived Non-Essential Details represents the most prevalent category. These findings suggest that paper-code consistency analysis is not merely a technical problem but also raises broader discussions regarding knowledge dissemination, the boundaries of code disclosure, and community norms. We hope that this work will encourage the bioinformatics community to re-examine the relationship between scientific publications and software implementations while providing a foundation for future research in paper-code consistency analysis.
|
| 1984 |
LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models
2603.28301
|
cs.LG
|
Chanyoung Kim, Minwoo Kim, Minseok Kang, Hyunwoo Kim, Dahuin Jung |
Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in downstream robotic settings, they are typically fine-tuned with limited data, leading to overfitting to spec...Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in downstream robotic settings, they are typically fine-tuned with limited data, leading to overfitting to specific instruction formulations and leaving robustness to paraphrased instructions underexplored. To study this gap, we introduce LIBERO-Para, a controlled benchmark that independently varies action expressions and object references for fine-grained analysis of linguistic generalization. Across seven VLA configurations (0.6B-7.5B), we observe consistent performance degradation of 22-52 pp under paraphrasing. This degradation is primarily driven by object-level lexical variation: even simple synonym substitutions cause large drops, indicating reliance on surface-level matching rather than semantic grounding. Moreover, 80-96% of failures arise from planning-level trajectory divergence rather than execution errors, showing that paraphrasing disrupts task identification. Binary success rate treats all paraphrases equally, obscuring whether models perform consistently across difficulty levels or rely on easier cases. To address this, we propose PRIDE, a metric that quantifies paraphrase difficulty using semantic and syntactic factors. Our benchmark and corresponding code are available at: https://github.com/cau-hai-lab/LIBERO-Para
|
| 1985 |
Meta-TTL: Meta-Learning Self-Improvement Policies for Language Agents
2604.00830
|
cs.LG
|
Zhanzhi Lou, Hui Chen, Yibo Li, Qian Wang, Bryan Hooi |
Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at inference time. At the core of TTL is a self-improvement policy that updates the actor policy based on experience fro...Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at inference time. At the core of TTL is a self-improvement policy that updates the actor policy based on experience from previous episodes, thereby improving future behavior. Existing methods rely on hand-crafted self-improvement rather than optimizing them for downstream improvement. We argue that optimal self-improvement policies should be learned from task environments, not hand-engineered based on human intuition. To achieve this, we introduce \textbf{Meta-TTL}, a framework that formulates the discovery of effective self-improvement policies as a bi-level optimization problem. Within this framework, the inner loop executes the standard TTL process, measuring how effectively a candidate self-improvement policy helps an agent correct errors across sequential episodes. Guided by the agent's performance, the outer loop performs reflective meta-training across diverse training tasks, using a balanced improvement score (BIS) to balance task contributions during candidate selection. We evaluate Meta-TTL on Jericho, WebArena-Lite, and -bench across both in-distribution (ID) and out-of-distribution (OOD) settings. Meta-TTL consistently outperforms existing baselines, improving TTL over the strongest baseline by up to 23% on ID tasks and 27% on OOD tasks. These results suggest that the optimized self-improvement policy encodes transferable meta-strategies that generalize beyond the training task distribution.
|
| 1986 |
Reasoning Shift: How Context Silently Shortens LLM Reasoning
2604.01161
|
cs.LG
|
Gleb Rodionov, Roman Garipov, George Yakushev |
Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remarkable performance on complex, long-term reasoning tasks. However, the robustness of these reasoning behaviors re...Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remarkable performance on complex, long-term reasoning tasks. However, the robustness of these reasoning behaviors remains underexplored. To investigate this, we conduct a systematic evaluation of multiple reasoning models across three scenarios: (1) problems augmented with lengthy, irrelevant context; (2) multi-turn conversational settings with independent tasks; and (3) problems presented as a subtask within a complex task. We observe an interesting phenomenon: reasoning models tend to produce much shorter reasoning traces (up to 74%) for the same problem under different context conditions compared to the traces produced when the problem is presented in isolation. A finer-grained analysis reveals that this compression is associated with a decrease in self-verification and uncertainty management behaviors, such as double-checking. Importantly, we show that even when additional self-checks are forced, their efficiency depends not only on the content of the reasoning traces, but also on the presence of redundant context. We hope our findings draw additional attention to both the robustness of reasoning models and the problem of context management for LLMs.
|
| 1987 |
Contrast Matters: Understanding Robustness of In-Context Fine-Tuning to Target-Context Relatedness
2604.01601
|
cs.LG
|
Deeptanshu Malu, Deevyanshu Malu, Aditya Nemiwal, Sunita Sarawagi |
In-context fine-tuning (IC-Train), training an LLM with labeled examples in-context, is increasingly used in place of standard fine-tuning for domain adaptation and continual absorption of labeled data. We study the robustness of the in-context learning abilit...In-context fine-tuning (IC-Train), training an LLM with labeled examples in-context, is increasingly used in place of standard fine-tuning for domain adaptation and continual absorption of labeled data. We study the robustness of the in-context learning ability that emerges from such training: does the fine-tuned model perform well across test inputs whose in-context examples range from unrelated to nearly identical? Across 32 configurations spanning four open-source LLMs and eight test sets over machine translation, Text-to-SQL, and multilingual semantic parsing, we show that robustness hinges on an overlooked design choice: how in-context examples are selected relative to the target during training. The two prevailing strategies turn out to be accurate over complementary parts of this spectrum: random contexts yield a model that gains little from related examples even when they are placed in its context, while retrieved similar contexts weaken accuracy on targets lacking close neighbors and raise the propensity to copy labels from context. Probes tracking in-weights learning, in-context learning, and copying trace these failures to distinct training dynamics, and show that introducing contrast in target-context similarity both within a context and across batches, restores robustness across the entire spectrum.
|
| 1988 |
Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability
2604.02653
|
cs.LG
|
Eric Gan |
Empirically, modern deep learning training often occurs at the Edge of Stability (EoS), where the sharpness of the loss exceeds the threshold below which classical convergence analysis applies. Despite recent progress, existing theoretical explanations of EoS ...Empirically, modern deep learning training often occurs at the Edge of Stability (EoS), where the sharpness of the loss exceeds the threshold below which classical convergence analysis applies. Despite recent progress, existing theoretical explanations of EoS either rely on restrictive assumptions or focus on specific squared-loss-type objectives. In this work, we introduce and study a structural property of loss functions that we term product-stability. We show that for losses with product-stable minima, gradient descent applied to objectives of the form $(x,y) \mapsto l(xy)$ can provably converge to the local minimum even when training in the EoS regime. This framework substantially generalizes prior results and applies to a broad class of losses, including binary cross entropy. Using bifurcation diagrams, we characterize the resulting training dynamics, explain the emergence of stable oscillations, and precisely quantify the sharpness at convergence. Together, our results offer a principled explanation for stable EoS training for a wider class of loss functions.
|
| 1989 |
An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
2604.07666
|
cs.LG
|
Andreas Plesner, Francisco Guzm\'an, Anish Athalye |
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used for post-training Large Language Models, but practical verifiers can make errors. We study how the rate and structure of reward noise affect RLVR in code generation, with a preliminary scient...Reinforcement Learning with Verifiable Rewards (RLVR) is widely used for post-training Large Language Models, but practical verifiers can make errors. We study how the rate and structure of reward noise affect RLVR in code generation, with a preliminary scientific-reasoning check. In multi-seed Qwen3 8B experiments on MBPP, mean validation reward over two fixed late evaluations is within 1 percentage point of the clean baseline at the tested resampled group-rollout noise rates through 20%, and within about 2 points at 30%. Confidence intervals allow larger losses; these point estimates do not establish a general tolerance threshold. We also examine full-program pass@k, four controlled noise structures, two model-based verifiers, and policy models from three families spanning 4B-9B parameters. We derive conditional advantage distributions for symmetric and asymmetric group noise, including retained format penalties, and show why clipping limits simple gradient-scaling arguments. The analysis identifies information preserved by whole-group corruption and limits on interpreting our asymmetric sweep as a precision-recall comparison. Overall, the results indicate that imperfect verification can support effective RLVR in the tested settings, while aggregate error rates alone do not characterize the learning signal.
|
| 1990 |
EvoLen: Evolution-Guided Tokenization for DNA Language Model
2604.08698
|
cs.LG
|
Nan Huang, Xiaoxiao Zhou, Junxia Cui, Mario Tapia-Pacheco, Tiffany Amariuta |
Tokens serve as the basic units of representation in DNA language models (DNALMs), yet their design remains underexplored. Unlike natural language, DNA lacks inherent token boundaries or predefined compositional rules, making tokenization a fundamental modelin...Tokens serve as the basic units of representation in DNA language models (DNALMs), yet their design remains underexplored. Unlike natural language, DNA lacks inherent token boundaries or predefined compositional rules, making tokenization a fundamental modeling decision rather than a naturally specified one. While existing approaches like byte-pair encoding (BPE) excel at capturing token structures that reflect human-generated linguistic regularities, DNA is organized by biological function and evolutionary constraint rather than linguistic convention. We argue that DNA tokenization should prioritize functional sequence patterns like regulatory motifs-short, recurring segments under evolutionary constraint and typically preserved across species. We incorporate evolutionary information directly into the tokenization process through EvoLen, a tokenizer that combines evolutionary stratification with length-aware decoding to better preserve motif-scale functional sequence units. EvoLen uses cross-species evolutionary signals to group DNA sequences, trains separate BPE tokenizers on each group, merges the resulting vocabularies via a rule prioritizing preserved patterns, and applies length-aware decoding with dynamic programming. Through controlled experiments, EvoLen improves the preservation of functional sequence patterns, differentiation across genomic contexts, and alignment with evolutionary constraint, while matching or outperforming standard BPE across diverse DNALM benchmarks. These results demonstrate that tokenization introduces a critical inductive bias and that incorporating evolutionary information yields more biologically meaningful and interpretable sequence representations. Code, pretrained and fine-tuned checkpoints, and tokenizer files are available at https://github.com/HN020719/EvoLen and https://huggingface.co/EvoLenTokenizer.
|
| 1991 |
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
2604.13010
|
cs.LG
|
Yecheng Wu, Song Han, Han Cai |
On-policy distillation (OPD) is an effective post-training paradigm for large language models but requires a live teacher server throughout training, resulting in substantial infrastructure overhead. We investigate whether OPD can be performed offline by preco...On-policy distillation (OPD) is an effective post-training paradigm for large language models but requires a live teacher server throughout training, resulting in substantial infrastructure overhead. We investigate whether OPD can be performed offline by precomputing teacher log-probabilities once over SFT rollouts and reusing them during training. We find that naively doing so fails to reliably match standard OPD, and trace the root cause to a previously overlooked condition we term teacher consistency, requiring that the same teacher be used for both supervised fine-tuning and OPD. Violating this condition introduces a gradient bias that degrades performance for both offline and online OPD. Building on this insight, we propose Lightning OPD, an offline on-policy distillation framework that enforces teacher consistency and eliminates the need for a live teacher server entirely. We prove that, under teacher consistency, Lightning OPD shares the same optimum as standard OPD, with bounded gradient discrepancy and an implicit regularization effect that helps prevent policy drift. Experiments on math reasoning and code generation show that Lightning OPD achieves comparable performance to standard OPD while delivering 4.0x higher training efficiency. Starting from an SFT-initialized Qwen3-8B-Base model, Lightning OPD reaches 69.9% on AIME 2024 in just 30 GPU hours. Lightning OPD further scales to MoE architectures, training Qwen3-30B-A3B to 71.0% on AIME 2024 on a single 8xH100 node, substantially lowering the barrier for academic research on LLM post-training. Our code is released at https://github.com/jet-ai-projects/Lightning-OPD.
|
| 1992 |
MambaSL: Exploring Single-Layer Mamba for Time Series Classification
2604.15174
|
cs.LG
|
Yoo-Min Jung, Leekyung Kim |
Despite recent advances in state space models (SSMs) such as Mamba across various sequence domains, research on their standalone capacity for time series classification (TSC) has remained limited. We propose MambaSL, a framework that minimally redesigns the se...Despite recent advances in state space models (SSMs) such as Mamba across various sequence domains, research on their standalone capacity for time series classification (TSC) has remained limited. We propose MambaSL, a framework that minimally redesigns the selective SSM and projection layers of a single-layer Mamba, guided by four TSC-specific hypotheses. To address benchmarking limitations -- restricted configurations, partial University of East Anglia (UEA) dataset coverage, and insufficiently reproducible setups -- we re-evaluate 20 strong baselines across all 30 UEA datasets under a unified protocol. As a result, MambaSL achieves state-of-the-art performance with statistically significant average improvements, while ensuring reproducibility via public checkpoints for all evaluated models. Together with visualizations, these results demonstrate the potential of Mamba-based architectures as a TSC backbone.
|
| 1993 |
Model Compression with Exact Budget Constraints via Riemannian Manifolds
2605.00649
|
cs.LG
|
Michael Helcig, Dan Alistarh |
Assigning one of K options to each of N groups under a total cost budget is a recurring problem in efficient AI, arising in mixed-precision quantization, non-uniform pruning, and expert selection. The objective (model loss) depends on all assignments jointly a...Assigning one of K options to each of N groups under a total cost budget is a recurring problem in efficient AI, arising in mixed-precision quantization, non-uniform pruning, and expert selection. The objective (model loss) depends on all assignments jointly and does not decompose across groups, so combinatorial solvers can only optimize proxy objectives. Evolutionary search evaluates the actual loss but lacks gradients, while penalty-based methods enforce the budget only approximately and often require heavy hyperparameter tuning. We show that under softmax relaxation, the budget constraint defines a smooth Riemannian manifold in logit space with unusually clean geometry: the normal vector is available in closed form, shifting logits along the cost vector changes expected cost monotonically, and vector transport reduces to a single inner product. Building on this, we propose Riemannian Constrained Optimization (RCO), which wraps tangent projection, binary-search retraction, and momentum transport around a standard Adam step. Combined with Gumbel straight-through estimation and budget-constrained dynamic programming for discrete feasibility, RCO optimizes the actual loss with first-order methods, enforces the expected budget exactly at every iterate, and introduces no constraint-related hyperparameters. The same construction handles multiple simultaneous budgets with no additional coefficients. RCO matches or exceeds state-of-the-art methods on synthetic problems and realistic LLM compression settings, often at considerably lower wall-clock cost. Source code is available at https://github.com/IST-DASLab/RCO.
|
| 1994 |
A Theory of Saddle Escape in Deep Nonlinear Networks
2605.01288
|
cs.LG
|
Divit Rawal, Michael R. DeWeese |
In deep networks with small initialization, training can exhibit long plateaus separated by sharp feature-acquisition transitions. Whereas shallow nonlinear networks and deep linear networks are well studied, extending these analyses to deep nonlinear networks...In deep networks with small initialization, training can exhibit long plateaus separated by sharp feature-acquisition transitions. Whereas shallow nonlinear networks and deep linear networks are well studied, extending these analyses to deep nonlinear networks remains challenging. We derive exact scalar and matrix identities for the imbalance of layer weight norms, holding for any smooth activation and any differentiable loss, and use the resulting approximate balance law to control the full finite-width gradient flow through its first escape. This gives a scaling law in the escape time: with a prefix of $r$ layers initialized at the bottleneck scale $\varepsilon$, the escape time is $\Theta(\Gamma^{-1})$ for $r=1$, $\Theta(\Gamma^{-1}\log(1/\varepsilon))$ for $r=2$, and $\Theta(\Gamma^{-1}\varepsilon^{-(r-2)})$ for $r\geq3$, where $\Gamma=\|\mathbb{E}[yx]\|$. We further show that a multimode teacher leaves the saddle once rather than in stages: with scalar output, all mode overlaps grow through a common amplitude, with relative weights fixed in advance. We find close agreement between our theory and experiment.
|
| 1995 |
Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time
2605.01766
|
cs.LG
|
Itai Allouche, Joseph Keshet |
Multimodal large language models (MLLMs) achieve strong performance on vision- and audio-language tasks, yet can generate responses that conflict with the given visual or auditory inputs, a problem known as multimodal hallucinations. Prior work suggests that t...Multimodal large language models (MLLMs) achieve strong performance on vision- and audio-language tasks, yet can generate responses that conflict with the given visual or auditory inputs, a problem known as multimodal hallucinations. Prior work suggests that this occurs when models rely more on textual cues and learned language patterns than on evidence from the perceptual input. To obtain a more direct account of this imbalance, we apply Layer-wise Relevance Propagation (LRP), which attributes predictions to individual input tokens, and use the resulting relevance scores to analyze and mitigate hallucinations. First, we examine whether this imbalance leads to multimodal hallucinations. We find that hallucinations often arise when the model relies less on perceptual inputs, and that changing this reliance affects its predictions. We further leverage LRP and propose a training-free framework that shifts relevance toward perceptual tokens by optimizing key-value representations during decoding, without modifying model parameters or requiring training data. We call this method Learning Inference-time Modality Enhancement (LIME). Despite using no spatial or temporal supervision, LIME concentrates relevance on query-relevant regions. We evaluate LIME across multiple multimodal benchmarks in both vision and audio domains, demonstrating consistent reductions in hallucinations and enhanced grounding while preserving generation quality.
|
| 1996 |
Federated Semi-Supervised Graph Neural Networks with Prototype-Guided Pseudo-Labeling for Privacy-Preserving Gestational Diabetes Mellitus Prediction
2605.01810
|
cs.LG
|
G. Victor Daniel |
Gestational Diabetes Mellitus (GDM) is a high-prevalence pregnancy complication that requires accurate early risk stratification to reduce maternal and fetal morbidity. However, real-world clinical deployment of machine learning is hindered by two coupled cons...Gestational Diabetes Mellitus (GDM) is a high-prevalence pregnancy complication that requires accurate early risk stratification to reduce maternal and fetal morbidity. However, real-world clinical deployment of machine learning is hindered by two coupled constraints: (i) label scarcity, where a large fraction of electronic health records (EHR) lack confirmed diagnostic labels, and (ii) data privacy, which prevents sharing patient-level data across hospitals. This paper proposes FedTGNN-SS, a privacy-preserving federated semi-supervised framework for clinical tabular EHR. Each hospital builds a local k-nearest-neighbor patient similarity graph and trains a topology-adaptive GNN encoder. To robustly exploit unlabeled records, FedTGNN-SS combines (1) prototype-guided pseudo-labeling with neighborhood agreement, (2) adaptive graph refinement that periodically updates the k-NN graph using learned embeddings, (3) clinical-aware consistency augmentation applied only to continuous variables, and (4) privacy-safe prototype sharing that exchanges only class-level centroids. Across three diabetes-related datasets (GDM: N = 3,525; Pima: N = 768; Early Stage: N = 520) under 10\%-80\% missing labels per silo, FedTGNN-SS achieves 56 significant wins ($p < 0.05$) against 11 federated baselines and attains strong AUROC under extreme scarcity (Pima: 0.8037 at 80\% missing, Early Stage: 0.9634 at 80\% missing).
|
| 1997 |
Demystifying Manifold Constraints in LLM Pre-training
2605.04418
|
cs.LG
|
Kang An, Jiaxiang Li, Donald Goldfarb, Shiqian Ma |
The recent success of matrix optimizers (e.g., Muon) suggests that specific normalization of momentum, such as orthogonalization and row-wise normalization, benefits both the stability and acceleration of LLM training. Consequently, several recent studies have...The recent success of matrix optimizers (e.g., Muon) suggests that specific normalization of momentum, such as orthogonalization and row-wise normalization, benefits both the stability and acceleration of LLM training. Consequently, several recent studies have suggested that weights should also be normalized, leading to a Riemannian optimization problem. While such constrained training frameworks demonstrate superior performance, the effects of explicitly constraining weights, and their interaction with existing stabilization mechanisms, remain less understood. To bridge this gap, we study manifold constrained training dynamics through activation scales, rotational dynamics, and the update-to-weight ratio. We propose a Riemannian spectral steepest descent optimizer called MACRO, alongside a radius selection principle to serve as our testbed. Our analysis and numerical experiments reveal that RMSNorm and manifold constraints serve overlapping roles, and that weight decay can be completely eliminated when manifold constraints are applied. By controlling the update-to-weight ratio, constrained training significantly alleviates update cancellation, empirically demonstrating that MACRO is robust to low-precision computation and competitive with existing algorithms for standard LLM pre-training.
|
| 1998 |
Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies
2605.05373
|
cs.LG
|
David Leeftink, Max Hinne, Marcel van Gerven |
A key capability of intelligent agents is to act effectively under incomplete state observations. Recurrent policies address this by compressing observation histories into a hidden state. In this work, we show that the hidden state of a recurrent policy admits...A key capability of intelligent agents is to act effectively under incomplete state observations. Recurrent policies address this by compressing observation histories into a hidden state. In this work, we show that the hidden state of a recurrent policy admits a control-theoretic interpretation: it plays the role of the co-state in Pontryagin's minimum principle, and the readout that maps it to actions implements control-Hamiltonian minimization. The hidden state thus tracks the gradient of the value function, encoding the optimality structure of the underlying control problem. We formalize this correspondence through a class of policies we refer to as co-state policies (CPs) and show that several modern recurrent cells implicitly realize this structure. The correspondence also allows for a co-state loss for actor-critic training, in which the critic's gradient serves as a target for the actor's hidden state. Empirically, we find that hidden states trained with the co-state loss encode co-state information beyond what is linearly decodable from the environment state alone, and that the loss improves policies both within and outside this class on challenging locomotion tasks, including the H1 and Berkeley humanoids. By connecting the minimum principle to recurrent memory, we provide a control-theoretic account of what hidden states compute in continuous control and a mechanism for shaping them toward optimality.
|
| 1999 |
Two-Stage Learned Decomposition for Scalable Routing on Multigraphs
2605.05389
|
cs.LG
|
Filip Rydin, Morteza Haghir Chehreghani, Bal\'azs Kulcs\'ar |
Most neural methods for Vehicle Routing Problems (VRPs) are limited to Euclidean settings or simple graphs. In this work, we instead consider multigraphs, where parallel edges represent distinct travel options with varying trade-offs (e.g., distance vs. time)....Most neural methods for Vehicle Routing Problems (VRPs) are limited to Euclidean settings or simple graphs. In this work, we instead consider multigraphs, where parallel edges represent distinct travel options with varying trade-offs (e.g., distance vs. time). Multigraphs are highly relevant in practice, yet few neural methods are designed for them, and those that do exist face major scalability issues. We address these scalability issues with Node-Edge Policy Factorization (NEPF), which splits the routing policy into a node permutation stage and an edge selection stage. To enable the decomposition, we introduce a pre-encoding edge aggregation scheme and a non-autoregressive architecture for the edge stage, as well as a hierarchical reinforcement learning method to train the stages jointly. Our experiments across six VRP variants demonstrate that NEPF trains and runs up to orders of magnitude faster than prior neural multigraph methods and scales to considerably larger instances, while matching or improving on their solution quality.
|
| 2000 |
AeroJEPA: Learning Semantic Latent Representations for Scalable 3D Aerodynamic Field Modeling
2605.05586
|
cs.LG
|
Francisco Giral, Abhijeet Vishwasrao, Andrea Arroyo Ramo, Mahmoud Golestanian, Federica Tonti |
High-fidelity CFD is essential for aerodynamic design, but repeated simulations are computationally expensive, motivating surrogate models for rapid evaluation across geometries and operating conditions. Most existing surrogates are designed for direct field r...High-fidelity CFD is essential for aerodynamic design, but repeated simulations are computationally expensive, motivating surrogate models for rapid evaluation across geometries and operating conditions. Most existing surrogates are designed for direct field regression, requiring the evaluation of millions of field points even when only an aerodynamic quantity or a localized region is needed, while their internal representations are not intended for direct use in downstream tasks. We introduce AeroJEPA, a framework inspired by joint-embedding predictive architectures that represents the problem in two distinct latent spaces: context tokens encode geometry, while predicted tokens encode the aerodynamic state. Both representations remain directly accessible for downstream tasks, such as linear readouts of design variables and aerodynamic quantities without decoding and integrating the full field. When spatial detail is needed, a continuous implicit decoder evaluates the field only at the requested coordinates while reusing the encoded geometry. We evaluate AeroJEPA on HiLiftAeroML, with multi-million-point fields, and SuperWing, which spans a broad family of transonic wings. Compared with state-of-the-art direct-regression surrogates, AeroJEPA trades peak full-field accuracy for compact, reusable representations. In our selective-decoding experiment, however, AeroJEPA substantially outperforms the evaluated direct-regression surrogates while avoiding predictions over the remainder of the aircraft. The learned representations further support controlled interpolation, concept-vector arithmetic, and preliminary constrained latent-space optimization. These results show how predictive representations can support aerodynamic analysis with or without full-field reconstruction.
|
| 2001 |
From Dual Tracking to Clipping: Provably Faster Distributionally Robust Multi-Objective Optimization
2605.05660
|
cs.LG
|
Yufeng Yang, Fangning Zhuo, Ziyi Chen, Heng Huang, Yi Zhou |
Multi-objective optimization (MOO) has received growing attention in applications that require learning under multiple criteria. However, most existing MOO formulations do not explicitly account for distributional shifts in the data. We introduce distributiona...Multi-objective optimization (MOO) has received growing attention in applications that require learning under multiple criteria. However, most existing MOO formulations do not explicitly account for distributional shifts in the data. We introduce distributionally robust multi-objective optimization (DR-MOO), which minimizes multiple objectives under their respective worst-case distributions. We propose Pareto-type solution concepts for DR-MOO and develop multi-gradient descent algorithms (MGDA) with provable guarantees. Leveraging a Lagrangian dual reformulation, we first design a double-loop MGDA that uses an inner loop to estimate dual variables and achieves a total sample complexity $\mathcal{O}(\epsilon^{-8})$ for reaching an $\epsilon$-Pareto-stationary point. To further improve convergence, we combine large-batch sampling with gradient clipping to accommodate generalized smoothness and control bias in stochastic preference updates, eliminating the need for double sampling. This yields a single-loop double-clip MGDA with substantially improved sample complexity $\mathcal{O}(\epsilon^{-4})$. Our theory applies to nonconvex problems without requiring uniformly bounded gradients of the dual objectives. Experiments demonstrate that our methods are competitive with state-of-the-art MGDA baselines.
|
| 2002 |
Learning beyond Site Bias for OOD Generalization in Brain Networks
2605.06050
|
cs.LG
|
Yingxu Wang, Kunyu Zhang, Yanwu Yang, Mengzhu Wang, Thomas Wolfers |
Graph-based learning from functional magnetic resonance imaging (fMRI) has shown strong potential for brain network analysis. However, existing methods often degrade under cross-site out-of-distribution (OOD) settings, as site-dependent confounder effects can ...Graph-based learning from functional magnetic resonance imaging (fMRI) has shown strong potential for brain network analysis. However, existing methods often degrade under cross-site out-of-distribution (OOD) settings, as site-dependent confounder effects can obscure disease-related connectivity patterns and static functional connectivity (FC) does not explicitly capture informative within-scan variations. In this paper, we propose Cross-site OOD Robust brain nEtwork (CORE), a unified framework for brain network learning across unseen sites. First, CORE estimates site-specific confounder effects and aggregates the resulting deconfounders by cross-source reliability for unseen-site correction, while extracting a population scaffold of reproducible label-associated connections from site-wise residualized FC. It then summarizes temporal variations on scaffold edges into compact descriptors for line-graph modeling. Finally, prior-guided subject-adaptive gating modulates message passing, balancing population priors with individual variability. Extensive leave-one-site-out experiments on ABIDE, REST-meta-MDD, SRPBS, and ABCD demonstrate that CORE consistently outperforms competitive baselines, with up to a 10.3% relative improvement in accuracy. These gains also persist across different brain parcellation schemes on ABIDE.
|
| 2003 |
SymDrift: One-Shot Generative Modeling under Symmetries
2605.06140
|
cs.LGcs.AI
|
Samir Darouich, Vinh Tong, Llu\'is Pastor-P\'erez, Tanja Bien, Loay Mualem |
Generative modeling of physical systems, such as molecules, requires learning distributions that are invariant under global symmetries, such as rotations in three-dimensional space. Equivariant diffusion and flow matching models can incorporate such invariance...Generative modeling of physical systems, such as molecules, requires learning distributions that are invariant under global symmetries, such as rotations in three-dimensional space. Equivariant diffusion and flow matching models can incorporate such invariances effectively, even when trained on a non-invariant empirical distribution, but they typically rely on costly multi-step sampling. Recently, drifting models have emerged as an efficient alternative, enabling single-step generation and achieving state-of-the-art performance in generative modeling tasks. However, we show that drifting models face a symmetry-specific challenge, since an equivariant generator does not generally produce the same drifting field as the one obtained from the symmetrized target distribution. Addressing this issue would require expensive symmetrization of the empirical distribution. To avoid this cost, we propose SymDrift, a framework that makes the drifting field itself symmetry-aware. We introduce two complementary strategies: (i) a symmetrized drift in coordinate space based on optimal alignment, and (ii) a $G$-invariant embedding that removes symmetry ambiguity by construction. Empirically, SymDrift outperforms existing one-shot methods on standard benchmarks for conformer and transition state generation, while remaining competitive with significantly more expensive multi-step approaches. By enabling one-shot inference, SymDrift reduces computational overhead by up to 40$\times$ compared to existing baselines, making it promising for high-throughput applications such as virtual drug screening and large-scale reaction network exploration.
|
| 2004 |
LINC: Decoupling Local Consequence Scoring from Hidden Matching in Constructive Neural Routing
2605.06332
|
cs.LG
|
Shaofeng Qin, Li Wang |
Constructive neural routing solvers usually score the next action by matching a decoder context to candidate embeddings, leaving deterministic one-step consequences such as travel, waiting, slack, and capacity changes implicit. We propose LINC, a decoder-side ...Constructive neural routing solvers usually score the next action by matching a decoder context to candidate embeddings, leaving deterministic one-step consequences such as travel, waiting, slack, and capacity changes implicit. We propose LINC, a decoder-side candidate decision architecture that computes these consequences explicitly. LINC uses them according to their decision role: candidate-level consequences are scored by a state-conditioned shared linear comparator, while feasible-set summaries modulate the decoder context. This preserves standard global matching while reducing the burden on the hidden state to reconstruct transition arithmetic. The Capacitated Vehicle Routing Problem with Time Windows (CVRPTW) serves as the main constrained-routing testbed, and the same interface extends to the Capacitated Vehicle Routing Problem (CVRP) and Traveling Salesman Problem (TSP). Across external benchmarks and no-retraining scale-transfer settings, LINC consistently improves strong neural baselines, with the advantage becoming more pronounced as test size moves further beyond the training scale, especially on constrained routing problems.
|
| 2005 |
Diverse Sampling in Diffusion Models with Divergence-Free Particle Guidance
2605.06553
|
cs.LG
|
Gal Vinograd, Idan Achituve, Ethan Fetaya |
Modern generative models can produce high-quality samples, but independent samples under the same condition often yields highly similar outputs. Particle-based methods increase diversity by letting the samples in a batch interact, typically through a repulsive...Modern generative models can produce high-quality samples, but independent samples under the same condition often yields highly similar outputs. Particle-based methods increase diversity by letting the samples in a batch interact, typically through a repulsive force. These forces, however, also push each sample away from the data distribution and produce visible artifacts. We introduce EDDY, a training-free particle guidance method designed to avoid this failure mode. Instead of a repulsive gradient, EDDY couples particles through anti-symmetric matrix fields passed through a Stein operator, a family of drift perturbations that leaves the Fokker--Planck equation invariant. Unlike repulsive guidance, these interactions do not alter a particle's distribution while the batch is independent. For the same reason they cannot create diversity on their own, so EDDY pairs them with a negatively correlated initialization that keeps each particle's prior exact. To use EDDY with perceptual kernels such as DINOv2, we approximate its second-order terms with finite differences and Hutchinson estimates. Across FLUX.1-dev, FLUX.2-klein and SDXL, EDDY achieves higher image quality and prompt alignment than existing particle guidance methods at matched diversity.
|
| 2006 |
UniPool: Learning Expert-to-Layer Ownership from Brief Global Access
2605.06665
|
cs.LG
|
Minbin Huang, Han Shi, Chuanyang Zheng, Yimeng Wu, Guoxuan Chen |
Most Mixture-of-Experts (MoE) transformers preassign each expert to one layer for the whole of training. We study whether a brief global-access phase can learn a compatible expert-to-layer allocation that is then executed privately. UniPool-Lock gives every la...Most Mixture-of-Experts (MoE) transformers preassign each expert to one layer for the whole of training. We study whether a brief global-access phase can learn a compatible expert-to-layer allocation that is then executed privately. UniPool-Lock gives every layer its own router over one global pool, balances aggregate pool usage, and scores experts with a scale-stable NormRouter. After the short ownership-learning phase, it assigns each expert to one layer, locks the disjoint allocation, and continues training as a layer-private MoE. With the same expert-FFN budget and routed expert FLOPs, this recipe lowers held-out loss relative to vanilla MoE by 0.024-0.037 across five dense-equivalent scales from 182M to 1.5B, including -0.0247 at 1.5B after 60B tokens. The ownership-learning phase lasts 2K steps, about 3.3% of training, and post-lock step time is within -0.7% to +2.8% of vanilla MoE in our throughput measurements. Controls locate the gain in full-pool training and in the allocation it produces: a random disjoint allocation fixed at initialization, trained with the same router and losses, matches vanilla, whereas locking the allocation learned during the full-pool phase recovers nearly all of the persistent full-pool gain, and substituting a random allocation at the lock forfeits about half of it. Keeping full-pool access throughout training (UniPool-Full) also improves over vanilla MoE at four scales and outperforms it with only 66.7% (182M) to 50% (469M and 650M) of its expert parameters.
|
| 2007 |
Enabling Unsupervised Training of Deep EEG Denoisers With Intelligent Partitioning
2605.06724
|
cs.LG
|
Qiyu Rao, Haozhe Tian, Homayoun Hamedmoghadam, Danilo Mandic |
Denoising electroencephalogram (EEG) is an inherently challenging task, since neural activity is not only subtle but also inseparable from spectrally overlapping noise artifacts. Today, effective EEG denoising is more important than ever, given the rapid adopt...Denoising electroencephalogram (EEG) is an inherently challenging task, since neural activity is not only subtle but also inseparable from spectrally overlapping noise artifacts. Today, effective EEG denoising is more important than ever, given the rapid adoption of wearables across various applications. Deep learning methods have shown promising results in decomposition-free denoising that handles the time-varying pervasive EEG artifacts. However, training highly expressive neural networks requires artifact-free EEG, which is inherently unobtainable. To address this, we propose Intelligent Partitioning for Self-supervised Denoising (iPSD). Our method eliminates the need for clean references by learning to partition an input EEG segment into independent noisy realizations with the same underlying signal. This enables self-supervision of deep learning denoisers, even in zero-shot settings where only a single EEG segment to be denoised is available. We validate iPSD through extensive experiments, including validations on wearable EEG from in-ear sensors. The results show that iPSD achieves state-of-the-art performance, most notably under extremely low signal-to-noise ratios (down to -10 dB) and challenging artifacts (e.g., EMG), with spectral fidelity orders of magnitude higher than competitive baselines.
|
| 2008 |
Skip What You Can Predict: Predictive Repositioning for Policy Optimization for Efficient LLM Training
2605.06755
|
cs.LG
|
Ismam Nur Swapnil, Aranya Saha, Tasneea Zahra, Tanvir Ahmed Khan, Mohammad Ariful Haque |
Reinforcement learning with verifiable rewards (RLVR) can improve the reasoning ability of large language models, but repeatedly updating a policy on the same rollout batch is expensive: every additional update requires another backward pass, and multi-step me...Reinforcement learning with verifiable rewards (RLVR) can improve the reasoning ability of large language models, but repeatedly updating a policy on the same rollout batch is expensive: every additional update requires another backward pass, and multi-step methods pay for all intermediate optimization steps. We introduce Predictive Repositioning for Policy Optimization (PrePO), which uses two observed optimizer transitions to estimate a farther point along the same-batch optimization trajectory, moves partway toward that point, evaluates the original objective there, and applies a corrective update. This gives PrePO a fixed active cost that does not grow with the virtual optimization depth. Our analysis gives finite-horizon error bounds for AdamW and Muon and shows that sufficiently accurate endpoint estimates preserve the usual descent and convergence behavior of smooth gradient descent. In RLVR experiments, PrePO reaches matched performance targets in fewer optimization steps and lower wall-clock time than the corresponding baselines. We further evaluate the same update mechanism in supervised fine-tuning on a different dataset, showing that its use is not restricted to the original RLVR setting. Together, these results suggest that PrePO provides a practical mechanism for approximating repeated same-batch optimization while avoiding the cost of explicitly executing every intermediate update.
|
| 2009 |
On Privacy in Data-Space Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
2605.06835
|
cs.LG
|
Behnoosh Zamanlooy, Elaheh Bassak, Fatemeh Tavakoli, Sara Kodeiri, Marcelo Lotif |
Tabular data plays an important role in many fields and industries, including those with elevated privacy considerations and risks. As such, there is a rising interest in generating high-quality synthetic proxies for real tabular data as a means of reducing pr...Tabular data plays an important role in many fields and industries, including those with elevated privacy considerations and risks. As such, there is a rising interest in generating high-quality synthetic proxies for real tabular data as a means of reducing privacy risk and proprietary data exposure. With data-space tabular diffusion models (TDMs) demonstrating leading performance in synthesizing such data, understanding and measuring the privacy risks associated with these models is imperative. Leveraging state-of-the-art membership inference attacks for such TDMs in both black- and white-box settings, this work quantifies the impact of training setup, synthesis choices, and attacker knowledge on privacy leakage. Moreover, the results demonstrate that adversaries need not have perfect knowledge of the training setup, identical data distributions, or massive compute resources to construct successful attacks. Finally, the substantial pitfalls associated with heuristic privacy metrics, such as distance-to-closest record, are identified.
|
| 2010 |
Continuous First, Discrete Later: VQ-VAEs Without Dimensional Collapse
2605.06870
|
cs.LG
|
Xinyu Zhao, Nikita Karagodin, Hamed Hassani, Sinan Hersek, Paul Pu Liang |
While many approaches to improve VQ-VAE performance focus on codebook size and utilization, the effect of dimensional collapse, where trained VQ-VAE representations live in an extremely low-dimensional subspace (1-2% of full rank), remains unaddressed. We show...While many approaches to improve VQ-VAE performance focus on codebook size and utilization, the effect of dimensional collapse, where trained VQ-VAE representations live in an extremely low-dimensional subspace (1-2% of full rank), remains unaddressed. We show theoretically and empirically that dimension collapse causes a hard loss lower bound that various codebook improvement techniques fail to surpass. Our analytic framework extends the sequential learning effect of Saxe et al. [2014] by introducing ideas from rate-distortion theory and explains how the latent collapse is caused by the VQ suppressing lower-variance directions. Our theory justifies a simple solution: a "warm-up phase" that trains the model as an (unquantized) autoencoder before introducing VQ. On both synthetic experiments and large-scale image (VQGAN) and audio (WavTokenizer) VQ-VAEs, we show that AE Warm-Up successfully restores representation dimension, leading to lower reconstruction and perceptual loss at the same training budget. Across codebook sizes $K \in$ {$2^{10}, 2^{14}, 2^{16}$}, AE warm-up raises VQGAN codebook effective dimension from 3-5 to 17-19 and reduces rFID by 17-35%; on WavTokenizer at $K \in$ {$2^{13}, 2^{14}$}, it raises codebook dimension from 4 to 17-19 and improves PESQ by 11-14%. We empirically characterize how warm-up duration governs the achievable final loss. In agreement with experiment, our theoretical analysis predicts downstream performance as a function of warm-up length, enabling an adaptive criterion for switching from AE Warm-up to VQ-VAE training.
|
| 2011 |
Why Does Agentic Safety Fail to Generalize Across Tasks?
2605.06992
|
cs.LG
|
Yonatan Slutzky, Yotam Alexander, Tomer Slor, Yoav Nagel, Nadav Cohen |
AI agents are increasingly deployed in multi-task settings, where the task to perform is specified at test time, and the agent must generalize to unseen tasks. A major concern in such settings is safety: often, an agent must not only execute unseen tasks, but ...AI agents are increasingly deployed in multi-task settings, where the task to perform is specified at test time, and the agent must generalize to unseen tasks. A major concern in such settings is safety: often, an agent must not only execute unseen tasks, but do so while avoiding risks and handling ones that materialize. Empirical evidence suggests that even when the ability to execute generalizes to unseen tasks, the ability to do so safely frequently does not. This paper provides theory and experiments indicating that failures of agentic safety to generalize across tasks are not merely due to limitations of training methods, but can also reflect an inherent property of safety itself: the relationship between a task and its safe execution is more complex than the relationship between a task and its execution alone. Theoretically, we analyze linear-quadratic control with $H_{\infty}$-robustness, and prove that in many cases, the mapping from task specification to an optimal controller has higher Lipschitz constant with safety requirements than without, yielding a Lipschitz bound of independent interest. Empirically, we demonstrate our conclusions in simulated quadcopter navigation with a neural network agent and in CRM with an LLM agent, showing that imitating safe and unsafe teachers is similarly straightforward on tasks seen in training, yet generalizing across tasks is considerably more difficult with the safe teacher. Our findings suggest that current efforts to enhance agentic safety may be insufficient, and point to a need for fundamentally different approaches.
|
| 2012 |
CellScientist: From Execution Feedback to Auditable Model-Revision Trajectories for Cellular Perturbation Prediction
2605.07335
|
cs.LG
|
Mengran Li, Bo Li, Jiaying Wang, Wenbin Xing, Chengyang Zhang |
Cellular perturbation-response modeling requires coordinated choices of representations, fusion mechanisms, objectives, and training procedures. Large language models (LLMs) can propose executable candidates, but unconstrained revision can produce invalid impl...Cellular perturbation-response modeling requires coordinated choices of representations, fusion mechanisms, objectives, and training procedures. Large language models (LLMs) can propose executable candidates, but unconstrained revision can produce invalid implementations, change task semantics, or discard useful components. We present CellScientist, a protocol-constrained workflow that converts execution and validation feedback into auditable model-revision trajectories. It records design states and outcomes, routes discrepancies to specific components, and applies local revisions under a fixed task contract. A matched-budget study fixes the candidate language, predictor, fitting, and evaluator: structured revision finds better held-out predictors at small budgets under two LLM backbones. Operational audits link history to fewer repeated proposals, contract checks to contained violations, and discrepancy routing to targeted repairs. Refits of two frozen designs on an independently acquired cohort evaluate external predictive utility. Open-workflow trajectories retain improvements, regressions, and failures, while transcriptomic and single-cell searches extend application to additional response spaces. CellScientist produces both a selected predictor and an inspectable record of its development. Project page: https://limengran98.github.io/CellScientist/.
|
| 2013 |
On the Invariance and Generality of Neural Scaling Laws
2605.07546
|
cs.LG
|
Xing Han, Liu Ziyin, Suchi Saria, Paul Pu Liang |
Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocation in new domains and tasks. Yet such laws are most needed precisely where they are hardest to obtain: fittin...Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocation in new domains and tasks. Yet such laws are most needed precisely where they are hardest to obtain: fitting one for a new model task pair demands expensive sweeps that typically exhaust the very compute budget the law is meant to economize. This paper poses the research question of how to develop generalizable scaling laws: laws fit once on a well-resourced source domain and reliably transported to new domains where running a full sweep is infeasible, which requires a fundamental understanding of when and why scaling properties change. We address this by identifying the right invariants: scaling laws are preserved under bijective (information-preserving) transformations of the data and modified in predictable, information-theoretically grounded ways under non-bijective transformations that lower its information resolution $\rho$: a single axis along which a law fit in one domain can be transported to another. We validate this across language, vision, and speech, and demonstrate two cross-domain applications: predicting scaling for language models trained on electronic health records from laws fit on general text, and predicting time-series classification scaling under varying levels of noise injection, recovering the data-scaling exponents to within $3\%$ error.
|
| 2014 |
Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning
2605.07977
|
cs.LG
|
Seohyun Lee, Wenzhi Fang, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. Brinton |
Recent works have advanced feedback-based learning systems, whereby a foundation model is able to intake incoming feedback (e.g., from a user) to self-improve, creating a self-loop system of training. However, existing works are limited in needing to consider ...Recent works have advanced feedback-based learning systems, whereby a foundation model is able to intake incoming feedback (e.g., from a user) to self-improve, creating a self-loop system of training. However, existing works are limited in needing to consider an offline setup to allow for such feedback-based methods, requiring privileged ground-truth contexts for training. Moreover, there is limited consideration of federated learning (FL), which is particularly well-suited for incorporating external feedback across large networks of end users, for example, but requires methods to be efficient for training on resource-constrained edge devices. Therefore, we introduce SPEAR (Self-Play Enhancement via Advantage-Weighted Refinement), an efficient online learning algorithm for federated LLM fine-tuning. SPEAR utilizes a feedback-guided self-play loop to construct naturally contrastive pairs per prompt which are used to train the model with (i) standard maximum likelihood on satisfactory completions and (ii) confidence-weighted unlikelihood on tail tokens of unsatisfactory completions. SPEAR requires only an interaction-level binary satisfaction signal together with non-answer feedback, without requiring a ground-truth answer to be provided, and does not need expensive group generations, meaning training both online and in a resource-efficient manner is feasible. We validate SPEAR across various benchmark datasets, demonstrating its superior performance in comparison to relevant baselines.
|
| 2015 |
Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit
2605.08352
|
cs.LG
|
Konstantin Riedl, Konstantinos Spiliopoulos, Justin Sirignano |
A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. As the number of hidden units tends to infinity, the NN training dynamics converge in probability to the solution of a dete...A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. As the number of hidden units tends to infinity, the NN training dynamics converge in probability to the solution of a deterministic limit equation involving a "Newton neural tangent kernel" (NNTK). Explicit rates characterizing this convergence are provided and, in the infinite-width limit, we prove that the NN converges exponentially fast to the target data (i.e., a global minimizer with zero loss). We show that this convergence is uniform across the frequency spectrum, addressing the spectral bias inherent in gradient descent. The eigenvalues of the NTK for gradient descent accumulate at zero, leading to slow convergence for target data with high-frequency components. In contrast, the NNTK has uniformly lower bounded eigenvalues if the regularization parameter is selected appropriately, allowing Newton's method to converge more quickly for data with high-frequency components. Mathematical challenges that need to be addressed include the implicit parameter update of the Newton method with a potentially indefinite Hessian matrix and the fact that the dimension of this linear system of equations tends to infinity as the NN width grows. This substantially complicates deriving the training dynamics in the overparameterized limit as well as proving the convergence of the finite-width dynamics thereto. Our analysis identifies a scaling formula for selecting the regularization parameter, which we show can vanish at a suitable rate as the NN width becomes larger. In addition, we prove that, for sufficiently large numbers of hidden units, the regularized Hessian remains positive definite during training and the Newton updates for individual NN parameters converge to zero, demonstrating that the model behaves as a linearization around the initialization.
|
| 2016 |
Recovering Physical Dynamics from Discrete Observations via Intrinsic Differential Consistency
2605.08454
|
cs.LG
|
Yuxiang Luo, Andrew Perrault |
Recovering continuous-time dynamics from discrete observations is difficult because local supervision loses fidelity as the observation interval grows. We replace local supervision with a global structural constraint: any flow representing autonomous dynamics ...Recovering continuous-time dynamics from discrete observations is difficult because local supervision loses fidelity as the observation interval grows. We replace local supervision with a global structural constraint: any flow representing autonomous dynamics must satisfy the semi-group property under time translation. We train a time-conditioned secant velocity field whose deviation from this property---which we call Symmetry Rupture---serves two roles: as a training regularizer it confines the hypothesis space to flows that compose consistently across temporal scales; as an inference oracle it guides an adaptive solver to select the largest step size that preserves internal consistency, without relying on local truncation error estimates. On the diffusion-reaction and Navier-Stokes benchmarks, our method achieves the lowest rollout RMSE among all evaluated methods using $\leq 3$ function evaluations per step. On the near-conservative shallow water benchmark, our method matches or exceeds competing flow-based methods at far lower computational cost; the strongest Neural ODE baseline achieves lower absolute RMSE but at the expense of $>100$ function evaluations per step. In the direct auto-regressive setting, our solver maintains stable long-horizon rollouts where flow-matching baselines diverge and Neural ODE requires up to $12\times$ more function evaluations. The core contribution is a model that internalizes temporal consistency, removing the need for high-order external solvers to compensate for structural bias.
|
| 2017 |
Towards Effective Theory of LLMs: A Representation Learning Approach
2605.09294
|
cs.LG
|
Muhammed Ustaomeroglu, Guannan Qu |
We propose Representational Effective Theory (RET), a framework for describing large language model computation in terms of learned macrostates rather than microscopic details. RET learns these macrostates from hidden-state trajectories using a BYOL/JEPA-style...We propose Representational Effective Theory (RET), a framework for describing large language model computation in terms of learned macrostates rather than microscopic details. RET learns these macrostates from hidden-state trajectories using a BYOL/JEPA-style self-supervised objective, coarse-graining activations into macrovariables that preserve higher-level structure relevant for prediction and interpretation. We evaluate whether these macrovariables are practically relevant for interpretability: RET yields temporally consistent states that reveal ``mental-state'' trajectories of reasoning, capture high-level semantic structure, support early prediction of behavioral outcomes such as sycophancy and alignment faking, while providing causal handles for steering generations toward interpretable computational phases. Together, these results suggest that LLM computation admits useful effective descriptions via RET: high-level, dynamically meaningful variables for interpretation, prediction, and control.
|
| 2018 |
Neural Cluster First, Route Second: Capacitated Vehicle Routing via Differentiable Optimal Transport
2605.09301
|
cs.LG
|
Samuel J. K. Chin, Maximilian Schiffer |
The Capacitated Vehicle Routing Problem (CVRP) underpins modern last-mile logistics, where routing decisions recur over the same fixed service area, like a city. In this setting, routing problems share a fixed set of potential customer locations, while active ...The Capacitated Vehicle Routing Problem (CVRP) underpins modern last-mile logistics, where routing decisions recur over the same fixed service area, like a city. In this setting, routing problems share a fixed set of potential customer locations, while active customers and demands vary between instances. We study how this spatial support can be exploited through reusable learned representations and design our method around three symmetries of the symmetric Euclidean CVRP: $E(2)$ transformations, vehicle-route permutations, and tour reversal. We introduce Neural Cluster-First--Route-Second (CFRS), a neural extension of the Fisher--Jaikumar framework that predicts seed-selection scores and customer-to-cluster assignment costs non-autoregressively and respects the three symmetries. A differentiable entropic optimal transport layer provides capacity-aware supervision and guides discrete capacitated assignment, followed by independent traveling salesman subproblems for route recovery. Component ablations show consistent benefits from learned seed selection, while learned assignment costs perform best near the training size and classical FJ costs perform better at larger sizes under exact decoding. On the fixed-support distribution with constant capacity, a model trained on $N=100$ achieves a $3.77\%$ routing gap relative to HGS at $N=1000$ without retraining. A shallow variant with one attention layer in each transformer achieves a $5.08\%$ gap at this scale, with spatial embeddings consistently improving routing quality over raw coordinates. Embedding interpolation further accommodates entirely unseen customer locations without retraining. On standard CVRP benchmarks, a separately trained model achieves a $2.73\%$ routing gap relative to LKH-3 at $N=100$.
|
| 2019 |
FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning
2605.09355
|
cs.LG
|
Xing Han, Shravan Chaudhari, Tanvi Ranade, Rama Chellappa, Suchi Saria |
Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, tasks are co-available at design time where related tasks could borrow representational strength from one ano...Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, tasks are co-available at design time where related tasks could borrow representational strength from one another, (2) continual adaptation, in which new tasks emerge after deployment with previously unseen modality combinations. However, neither regime alone suffices: the pretraining task set is never exhaustive, while bypassing joint training forfeits the transfer gains and efficiency among co-trainable tasks. Sparse Mixture-of-Experts (MoE) is a natural fit for this dual requirement: sparse activation enables modular capacity expansion as new tasks arrive, while routing decouples modality-level computation from task-level composition. In this work, we propose a scalable MoE framework for multitask pretraining and continual learning across flexible modality combinations. The framework is designed to support training on multimodal tasks with diverse modality configurations by leveraging modality-specific routers that process tokens from each modality across tasks. Furthermore, it enables continual learning over sequential multimodal tasks within a fixed-capacity MoE by compressing accumulated expert knowledge into low-rank memory subspaces, while expanding only the lightweight routers. We validate the effectiveness of our method on multiple healthcare multimodal benchmarks. It demonstrates competitive multitask pretraining performance while alleviating catastrophic forgetting and improving parameter efficiency.
|
| 2020 |
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
2605.10770
|
cs.LG
|
Eleonora Gualdoni, Sonia Laguna, Louis Bethune, Joao Monteiro, Pierre Ablin |
Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving previous capabilities, such as general knowledge, instruction following, or safety evaluations. Existing data mixing strategies rely on fixed he...Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving previous capabilities, such as general knowledge, instruction following, or safety evaluations. Existing data mixing strategies rely on fixed heuristics or adaptive rules that cannot explicitly enforce preservation of such capabilities. We propose DynaMiCS, a dynamic mixture optimizer that casts multi-domain fine-tuning as a constrained optimization problem. At each update, DynaMiCS performs short domain-specific probing runs to estimate a slope matrix of local cross-domain effects, capturing how training on each fine-tuning dataset affects each evaluation domain. These estimates are then used to compute mixture weights through optimization over the probability simplex, with the objective of improving target-domain performance while keeping constrained-domain metrics within a specified tolerance of reference levels. Because these effects are measured by finite differences rather than gradients, targets and constraints need not be differentiable, or present in the fine-tuning data, and can be specified directly as benchmark accuracies. Across scenarios with varying numbers of target and constrained domains, and with loss- or accuracy-based objectives, DynaMiCS achieves stronger target-domain improvements and higher constraint satisfaction than static, dynamic, similarity-based and probing-based alternatives, without a reference model, per-example scoring, or manually tuned weights.
|
| 2021 |
When and How to Canonize: A Generalization Perspective
2605.11008
|
cs.LG
|
Yonatan Sverdlov, Benjamin Friedman, Snir Hordan, Nadav Dym |
While invariant architectures are standard for processing symmetric data, there is growing interest in achieving invariance by applying group averaging or canonization to non-invariant backbones. However, the theoretical generalization properties of these alte...While invariant architectures are standard for processing symmetric data, there is growing interest in achieving invariance by applying group averaging or canonization to non-invariant backbones. However, the theoretical generalization properties of these alternative strategies remain poorly understood. We introduce a theoretical framework to analyze the generalization error of these methods by bounding their covering numbers. We establish a rigorous generalization hierarchy: the error bounds of canonized models are at best equal to the error bounds of structurally invariant and group-averaged models, and at worst equal to the bounds of non-invariant baselines. Furthermore, we show that there exist optimal canonizations which attain the optimal error bounds, and poor canonizations which attain the non-invariant error bounds, and that this depends on the regularity of the canonization. Finally, applying this framework to permutation groups in point cloud processing, we rigorously prove that the covering number of lexicographical sorting grows exponentially with point cloud dimension, whereas Hilbert curve canonization guarantees polynomial growth. This provides the first formal theoretical justification for the empirical success of Hilbert curve serialization in state-of-the-art point cloud architectures. We conclude with experiments that support our theoretical claims. Code is available at https://github.com/yonatansverdlov/Canonization
|
| 2022 |
ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
2605.11009
|
cs.LG
|
Qian Chen, Junqiao Zhao, Hongtu Zhou, Hang Yu, Yanping Zhao |
Long-horizon, sparse-reward tasks pose a fundamental challenge for reinforcement learning, since single-step TD learning suffers from bootstrapping error accumulation across successive Bellman updates. Actor-critic methods with action chunking address this by ...Long-horizon, sparse-reward tasks pose a fundamental challenge for reinforcement learning, since single-step TD learning suffers from bootstrapping error accumulation across successive Bellman updates. Actor-critic methods with action chunking address this by operating over temporally extended actions, which reduce the effective horizon, enable fast value backups, and support temporally consistent exploration. However, existing methods rely on a fixed chunk size and therefore cannot adaptively balance reactivity against temporal consistency. A large fixed chunk size reduces responsiveness to new observations, while a small one produces incoherent motions, forcing task-specific tuning of the chunk size. To address this limitation, we propose Adaptive Chunk Size Actor-Critic (ACSAC). ACSAC leverages a causal Transformer critic to evaluate expected returns for action chunks of different sizes. At each chunk boundary, it adaptively selects the chunk size that maximizes the expected return, supporting flexible, state-dependent chunk sizes. We prove that the ACSAC Bellman operator is a $\gamma$-contraction whose unique fixed point is the action-value function of the adaptive policy. Experiments on OGBench demonstrate that ACSAC achieves strong performance on long-horizon, sparse-reward manipulation and navigation tasks across both offline RL and offline-to-online RL settings.
|
| 2023 |
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning
2605.11235
|
cs.LG
|
Han Zheng, Yining Ma, Karthick Gunasekaran, Bharathan Balaji, Zheng Du |
In LLM Reinforcement Fine-Tuning (RFT), curriculum learning drives both efficiency and performance. Yet, current methods externalize curriculum judgment via handcrafted heuristics or auxiliary models, risking misalignment with the policy's training dynamics. I...In LLM Reinforcement Fine-Tuning (RFT), curriculum learning drives both efficiency and performance. Yet, current methods externalize curriculum judgment via handcrafted heuristics or auxiliary models, risking misalignment with the policy's training dynamics. In this paper, we introduce METIS (METacognitive Internalized Self-judgment), a novel framework that internalizes curriculum judgment as a native capability. Leveraging a critical observation that within-prompt reward variance effectively gauges prompt informativeness, METIS predicts this metric based on recent training outcomes as lightweight in-context learning examples. This intrinsic self-judgment then dynamically dictates the training allocation. Moreover, METIS closes the loop between judgment and optimization by jointly optimizing the standard RFT rewards and a self-judgment reward. This allows the policy to learn what to learn next, as a form of metacognition. Across mathematical reasoning, code generation, and agentic function-calling benchmarks, METIS delivers superior performance while achieving up to a 2.1x training speedup, with controlled ablations and in-depth analysis further validating the benefits of internalized curriculum judgment. By bypassing handcrafted heuristics and auxiliary models, our work establishes a simple, closed-loop, and highly efficient curriculum internalization paradigm for LLM reinforcement fine-tuning.
|
| 2024 |
A Dual Representation of Influence Functions for Linearizable Models
2605.11239
|
cs.LG
|
Zhenhuan Sun, Shahrokh Valaee |
In this paper, we present a dual representation of influence functions, whose computational complexity scales with dataset size rather than model size. Both analytically and experimentally, we show that this representation can be an efficient alternative to th...In this paper, we present a dual representation of influence functions, whose computational complexity scales with dataset size rather than model size. Both analytically and experimentally, we show that this representation can be an efficient alternative to the original influence functions for estimating changes in parameters, model outputs and loss due to data point removal, when model size is large relative to dataset size, or when evaluating the original influence functions in parameter space is infeasible. The dual representation, however, is limited to linearizable models, which are models whose behavior can be approximated by their linearizations throughout training, and requires materializing a matrix, whose size grows with the product of model output dimension and dataset size.
|
| 2025 |
Deep Minds and Shallow Probes
2605.11448
|
cs.LG
|
Su Hyeong Lee, Risi Kondor |
Neural representations are not unique objects. Even when two systems realize the same downstream computation, their hidden coordinates may differ by reparameterization. A probe family intended to reveal structure already present in a representation should ther...Neural representations are not unique objects. Even when two systems realize the same downstream computation, their hidden coordinates may differ by reparameterization. A probe family intended to reveal structure already present in a representation should therefore be stable under the relevant representation symmetries rather than be tied to a particular basis. We prove that, under standard nondegeneracy assumptions, behaviorally equivalent LLMs have hidden representations related by an invertible affine transformation. The resulting symmetry principle singles out a unique hierarchy of shallow coordinate-stable probes, with linear probes as its degree-1 member. We also show that a natural object for cross-model probe transfer is a shared probe-visible quotient--the representation modulo directions invisible to the probe family--rather than the full hidden state. Experiments on synthetic and real-world tasks support both predictions, showing where degree-2 probes help beyond linear ones and how quotient-based transfer enables coverage-aware monitor portability across model families. These results point toward a broader geometric representation theory of neural probing, with coverage-aware monitor transfer as a concrete operational consequence.
|
| 2026 |
Hypernetworks for Dynamic Feature Selection
2605.12278
|
cs.LG
|
Javier Fumanal-Idocin, Raquel Fernandez-Peralta, Javier Andreu-Perez |
Dynamic feature selection (DFS) is a machine learning framework in which features are acquired sequentially for individual samples under budget constraints. The exponential growth in the number of possible feature acquisition paths forces a DFS model to balanc...Dynamic feature selection (DFS) is a machine learning framework in which features are acquired sequentially for individual samples under budget constraints. The exponential growth in the number of possible feature acquisition paths forces a DFS model to balance fitting specific scenarios against maintaining general performance, even when the feature space is moderate in size. In this paper, we study the structural limitations of existing DFS approaches to achieve an optimal solution. Then, we propose \textsc{Hyper-DFS}, a hypernetwork-based DFS approach that generates feature subset-specific classifier parameters on demand. We show that the use of hypernetworks compared to mask-embedding methods results in a smaller structural complexity bound. We also use a Set Transformer encoding to create a smooth conditioning space for the hypernetwork, so that functionally similar tasks are also geometrically close. In our benchmarks, \textsc{Hyper-DFS} performed best or second best compared to all state-of-the-art approaches on synthetic and real-life tabular data. It is also best or second best across all image datasets tested, and shows stronger zero-shot generalisation to feature subsets never seen during training than existing DFS approaches. Code available in: https://github.com/Fuminides/hyper_DFS
|
| 2027 |
PyroAdapt: Adapting Wildfire Prediction under Spatial Heterogeneity and Temporal Shift
2605.12435
|
cs.LG
|
Enyi Jiang, Wu Sun |
Prediction of wildfire occurrence is a rare-event problem compounded by spatial heterogeneity and temporal distribution shift, as fire occurrences are vastly outnumbered by non-occurrences, and predictor--fire relationship varies across space and time. Models ...Prediction of wildfire occurrence is a rare-event problem compounded by spatial heterogeneity and temporal distribution shift, as fire occurrences are vastly outnumbered by non-occurrences, and predictor--fire relationship varies across space and time. Models trained on historical fire data may perform poorly under new conditions and require adaptation to the target distribution before operational use. We propose PyroAdapt, a pretrain--retrieve--rank framework that adapts a pretrained model to target conditions by retrieving historical locations with similar conditions and fine-tuning on the retrievals through risk ranking. For spatial adaptation, we condition risk on terrain, ecoregion embeddings, and fire rates, accounting for spatial context in the retrieval, and learn risk ordering from same-day fire--nonfire cell pairs. We compare direct ranking, residual pairwise DPO (RDPO), and selective ranking through a unified score-gap formulation that characterizes their gradient allocation. Over California (discretized into 666 0.25x0.25 grid cells), these objectives raise daily average precision from 21.62% for continued focal fine-tuning to 24.35--24.57%, and Top5% recall from 18.70% to 22.11--22.79%. Under a fixed daily detection budget of 34 cells (5% area), selective ranking captures 344 additional positive cell--days. For fires in the top 5%/10%/20% of dry matter consumption, selective ranking raises recall by 39.70/28.18/20.50 percentage points, respectively. Furthermore, rolling evaluations over Yosemite show that the gains from ranking persist under temporal distribution shift. Together, these results show that PyroAdapt prioritizes the most fire-prone locations under a daily budget constraint and detects more extreme fire events.
|
| 2028 |
Hessian Matching for Machine-Learned Coarse-Grained Molecular Dynamics
2605.12823
|
cs.LG
|
Sanya Murdeshwar, Sanjit Shashi, Kevin Bachelor, William Noid, Razvan Marinescu |
Coarse-grained (CG) molecular dynamics enables simulations of atomic systems such as biomolecules at timescales inaccessible to all-atom (AA) methods, but existing CG neural potentials trained via force matching capture only the gradient of the free-energy sur...Coarse-grained (CG) molecular dynamics enables simulations of atomic systems such as biomolecules at timescales inaccessible to all-atom (AA) methods, but existing CG neural potentials trained via force matching capture only the gradient of the free-energy surface, leaving its curvature unsupervised. We introduce a framework that augments force matching with stochastic Hessian-vector product (HVP) matching, instilling second-order curvature information into CG potentials without constructing the full Hessian. We derive a decomposition of the target CG Hessian into a model-independent projected AA Hessian, precomputed once before training, and a model-dependent covariance correction computed online at negligible cost. We then construct an unbiased stochastic estimator of the Hessian-matching objective by using random probe vectors. We evaluate our method by training on a benchmark set of nine fast-folding proteins and gauging agreement of the learned free-energy landscapes with reference data along the directions of slowest collective motion using time-lagged independent component analysis. HVP matching improves slow-mode accuracy over force matching for all but one of our benchmark proteins. It also sharpens local structure, reducing bond-length error on eight of nine proteins, and in several cases by an order of magnitude. Our results demonstrate that higher-order physical supervision is a practical path to more accurate CG potentials for biomolecular simulation.
|
| 2029 |
CoRe-Gen: Robust Spectrum-to-Structure Generation under Imperfect Fingerprint Conditions
2605.12980
|
cs.LG
|
Tianbo Liu, Chixiang Lu, Jing Hao, Hengyu Zhang, Lifei Wang |
Molecular structure elucidation from tandem mass spectra (MS/MS) remains challenging, particularly for de novo generation beyond database coverage. A common approach decomposes the task into spectrum-to-fingerprint prediction followed by fingerprint-to-structu...Molecular structure elucidation from tandem mass spectra (MS/MS) remains challenging, particularly for de novo generation beyond database coverage. A common approach decomposes the task into spectrum-to-fingerprint prediction followed by fingerprint-to-structure decoding, enabling the use of large-scale molecular corpora. However, at deployment, the decoder relies on predicted rather than oracle fingerprints, introducing structured errors that propagate into generation. Moreover, common autoregressive decoders threshold fingerprint probabilities and serialize active bits as discrete input tokens, blocking molecular generation gradients from the spectrum encoder. We present \textit{CoRe-Gen}, which improves the intermediate condition through synthetic-spectrum pretraining, matches deployment-time noise through frequency-aware corruption, and introduces a differentiable fingerprint-to-memory bridge for joint encoder--decoder finetuning. The bridge preserves soft bit confidences and conditions a structure-aware autoregressive decoder without discrete fingerprint tokenization. CoRe-Gen achieves 21.49\%/35.42\% Top-1/Top-10 exact-match accuracy on NPLIB1 and, under encoder-aligned evaluation, 20.02\%/22.45\% on MassSpecGym, while retaining efficient autoregressive inference.
|
| 2030 |
Learning to Select Source Domains: Proxy-Rewarded Policy Optimization for Molecular OOD Generalization
2605.13932
|
cs.LG
|
Zhuohao Lin, Kun Li, Jiameng Chen, Jiajun Yu, Duanhua Cao |
Molecular property prediction under severe out-of-distribution (OOD) shifts remains challenging because conventional scaffold splits can retain local structural similarity between training and test molecules, while adaptation from heterogeneous source domains ...Molecular property prediction under severe out-of-distribution (OOD) shifts remains challenging because conventional scaffold splits can retain local structural similarity between training and test molecules, while adaptation from heterogeneous source domains may cause negative transfer. We introduce SCOPE-Bench, a scaffold-cluster benchmark constructed by partitioning Bemis-Murcko scaffolds in an explicit physicochemical descriptor space, and POMA, a retrieve-compose-adapt framework for source-domain selection when target property labels are unavailable. POMA uses unlabeled target structures to retrieve labeled source scaffolds as proxy targets. A policy is trained from reductions in proxy-task mean absolute error after adaptation, and the selected source subset is then used for target adaptation with covariance alignment at the whole-molecule and BRICS-derived substructure levels. We evaluate HOMO, LUMO, and HOMO-LUMO gap prediction on QM9 using three 3D molecular backbones and 15 target-scaffold tasks. Relative to conventional scaffold splitting, prediction errors on SCOPE-Bench increase by up to 8.0-fold, with a mean increase of 5.9-fold. Relative to the supervised baseline under the strict OOD split, POMA reduces mean absolute error by up to 11.2%, with an average relative reduction of 6.2% over the nine backbone-property combinations. These results support adaptive source selection as a useful strategy for molecular prediction under strong structural distribution shifts.
|
| 2031 |
Gaussian Relational Graph Transformer
2605.15575
|
cs.LG
|
Zezhong Ding, Jin Li, Xugang Wang, Xike Xie |
Relational graph learning enables predictive modeling over relational databases by directly capturing dependencies across interconnected tables. While relational graph Transformers extend the receptive field beyond local message passing, a larger receptive fie...Relational graph learning enables predictive modeling over relational databases by directly capturing dependencies across interconnected tables. While relational graph Transformers extend the receptive field beyond local message passing, a larger receptive field does not necessarily provide more useful information: sampling must preserve relational structure without introducing excessive semantically irrelevant nodes, while attention must account for the temporal relevance of the retained information. We propose GelGT, a Gaussian relational graph transformer that explicitly addresses these challenges. GelGT introduces a structure-semantic collaborative sampling strategy to preserve structural connectivity while filtering irrelevant semantic information, and incorporates a Gaussian graph attention mechanism with a learnable Gaussian bias on the sampled subgraphs to dynamically encode temporal dependencies. We provide theoretical analysis characterizing the structural preservation, semantic refinement, and temporal discrimination of these mechanisms. Experiments on \textbf{7} real-world relational datasets covering \textbf{21} prediction tasks show that GelGT consistently outperforms existing relational graph learning methods, with improvements of up to \textbf{13.8\%}.
|
| 2032 |
Goal-Conditioned Supervised Learning for LLM Fine-Tuning
2605.16345
|
cs.LG
|
Shijun Li, Kaiwen Dong, Xiang Gao, Joydeep Ghosh |
Large language models often require fine-tuning to better align their behavior with user intent at deployment. Existing approaches are commonly divided into online and offline paradigms. Online methods, such as RL-based alignment, can directly optimize outcome...Large language models often require fine-tuning to better align their behavior with user intent at deployment. Existing approaches are commonly divided into online and offline paradigms. Online methods, such as RL-based alignment, can directly optimize outcome quality but typically rely on external reward models and iterative rollouts, making them costly and difficult to deploy in many cases. Offline methods are more efficient, but prevailing approaches such as supervised fine-tuning (SFT) and direct preference optimization (DPO) remain limited: SFT typically collapses graded feedback into binary supervision, while DPO depends on paired preference data that is often unavailable or expensive to construct. In this paper, we propose goal-conditioned supervised learning (GCSL) as an offline fine-tuning framework for LLMs. Our core idea is to treat feedback signals directly as an explicit goal and train the model, purely through supervised learning, to generate responses that achieve that goal. To better exploit graded feedback, we further introduce a novel goal formulation that defines learning as consistently pursuing outcomes above a target quality threshold, rather than imitating samples from a selected high-quality subset. This design mitigates the bounded-learning effect by learning transferable patterns of meeting or exceeding quality thresholds. We also propose natural-language goal representations to further connect these patterns to the LLM's pretrained knowledge and generalization capabilities. We evaluate our method on three tasks: non-toxic generation, code generation, and LLM for recommendation. Results show that our approach consistently outperforms standard offline fine-tuning baselines while retaining the efficiency, scalability, and simple data requirements of supervised learning.
|
| 2033 |
DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition
2605.16441
|
cs.LG
|
Jiahui Li, Ruili Fang, Zishuai Liu, WenZhan Song, Jin Lu |
Beat-level Electrocardiography (ECG) arrhythmia detection aims to assign an arrhythmia class to each beat in a recording, yet many existing systems treat beats as isolated local instances. This is limiting because beat labels often depend on multi-beat rhythm ...Beat-level Electrocardiography (ECG) arrhythmia detection aims to assign an arrhythmia class to each beat in a recording, yet many existing systems treat beats as isolated local instances. This is limiting because beat labels often depend on multi-beat rhythm context, including timing, compensatory pauses, and beat-to-beat morphological consistency. We present DeepArrhythmia, a tool-grounded multimodal framework for segment-contextualized beat-level ECG arrhythmia classification. Given a multi-beat ECG segment, DeepArrhythmia combines the raw ECG signal and a rendered waveform image, localizes R peaks to identify beat instances, and produces structured beat-level predictions. The framework decouples physiological measurement from evidence integration using specialized tools for beat localization, numerical rhythm--morphology extraction, and morphology-focused textual analysis. DeepArrhythmia uses segment-level confidence to route between minimal and rich evidence states, since richer physiological evidence is not uniformly useful. This agentic design integrates rhythm context, explicit physiological grounding, and selective evidence acquisition for decision making.
|
| 2034 |
EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
2605.16692
|
cs.LG
|
Thomas Evers, Cristian Meo, Wendelin Bohmer, Justin Dauwels, Yaniv Oren |
We introduce EfficientTDMPC, a sample-efficient model-based reinforcement learning method for continuous control built on the TD-MPC family of algorithms. Central to this family is a planner that aims to find an action sequence that maximizes the estimated ret...We introduce EfficientTDMPC, a sample-efficient model-based reinforcement learning method for continuous control built on the TD-MPC family of algorithms. Central to this family is a planner that aims to find an action sequence that maximizes the estimated return. The return is estimated using a learned model and value networks, each of which can introduce error. EfficientTDMPC introduces three contributions that improve performance by aiming to reduce this error. First, we introduce a multi-horizon planning objective that evaluates the value at different rollout depths and averages them. Second, to our knowledge we are the first to train a value-equivalent dynamics ensemble. Our improved objective then averages over rollouts from multiple dynamics heads. Third, we add pessimistic reanalyze for tasks that can terminate early. Applying our contributions to a recent baseline (BMPC) yields EfficientTDMPC, which to our knowledge is the new state of the art in sample efficiency on HumanoidBench and the DeepMind Control Suite, reaching BMPC's final aggregated performance using 57\% fewer environment steps.
|
| 2035 |
Fidelity Probes for Specification--Code Alignment
2605.17246
|
cs.LG
|
Ferhat Erata, Hao Zhou, Jun Huan |
Software modernization often relies on natural-language specifications recovered from an existing implementation. Incorrect or missing requirements can lead to defects in the modernized system. We introduce fidelity probes to check a specification against the ...Software modernization often relies on natural-language specifications recovered from an existing implementation. Incorrect or missing requirements can lead to defects in the modernized system. We introduce fidelity probes to check a specification against the code and guide its revision. Each probe pairs a question about program behaviour with a reference answer derived from the code. A language model answers the question using only the specification, and the fraction of agreeing answers defines fidelity. Disagreements flag conflicting claims or missing information and guide proposed corrections and additions to the requirements. An LLM generates probes directly from code or phrases facts selected through control-flow, data-flow, and call-graph analysis. We use an observability rule to focus probe questions on behaviour that the modernized system is intended to preserve, including user-visible outputs, changes to stored business data, and interactions with other systems. We revise the specification using fresh probes and evaluate each version on a fixed held-out set. A repair-regression model describes how new failures can offset the gains from repairs. On CardDemo, held-out fidelity improves from 0.59 to 0.89, exceeding free-form revision from source code. We also evaluate the method on industrial systems and public documentation, and assess probe quality and reported defects through human reviews and comparison with an independent requirements audit.
|
| 2036 |
AMO: Operator-level Adaptive Muon Orthogonalization
2605.17806
|
cs.LG
|
Xinlin Zhuang, Panyi Ouyang, Yichen Li, Jiangming Shi, Yizhang Chen |
Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iteration as its core operation. Standard Muon applies a uniform NS schedule to all parameter matrices, overlooking poss...Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iteration as its core operation. Standard Muon applies a uniform NS schedule to all parameter matrices, overlooking possible differences in orthogonalization difficulty and its impact on performance. Through a systematic empirical study, we show that this per-matrix heterogeneity is pervasive and strongly associated with matrix geometry, which evolves dynamically across operator types, training stages, and network depths. Therefore, uniform NS schedules can lead to uneven orthogonalization quality across the model. Motivated by these findings, we propose Operator-level Adaptive Muon Orthogonalization (AMO), an observe-then-commit method that measures weight geometry by operator type early in training and then uses these signals to allocate the NS budget for the remainder of training. AMO delivers consistent improvements over uniform-schedule Muon across standard, prolonged, and continual pre-training, surpassing the strongest baseline by +0.76 on Llama3.1-1.4B and +0.51 on Qwen3-1.7B in average downstream performance of 12 evaluation tasks, with gains persisting at Llama3.1-4B scale.
|
| 2037 |
DCFold: Efficient Protein Structure Generation with Single Forward Pass
2605.17899
|
cs.LG
|
Zhe Zhang, Yuanning Feng, Yuxuan Song, Keyue Qiu, Hao Zhou |
AlphaFold3 introduces a diffusion-based architecture that elevates protein structure prediction to all-atom resolution with improved accuracy. This state-of-the-art performance has established AlphaFold3 as a foundation model for diverse generation and design ...AlphaFold3 introduces a diffusion-based architecture that elevates protein structure prediction to all-atom resolution with improved accuracy. This state-of-the-art performance has established AlphaFold3 as a foundation model for diverse generation and design tasks. However, its iterative design substantially increases inference time, limiting practical deployment in downstream settings such as virtual screening and protein design. We propose DCFold, a single-step generative model that attains AlphaFold3-level accuracy. Our Dual Consistency training framework, which incorporates a novel Temporal Geodesic Matching (TGM) scheduler, enables DCFold to achieve a 15x acceleration in inference while maintaining predictive fidelity. We validate its effectiveness across both structure prediction and binder design benchmarks.
|
| 2038 |
Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models
2605.18202
|
cs.LG
|
Samuele Bortolotti, Emanuele Marconato, Andrea Pugnana, Andrea Passerini, Stefano Teso |
Neuro-Symbolic Concept-based Models (NeSy-CBMs) are a family of architectures that integrate neural networks with symbolic reasoning for enhanced reliability in high-stakes applications. They work by first extracting high-level concepts from the input and then...Neuro-Symbolic Concept-based Models (NeSy-CBMs) are a family of architectures that integrate neural networks with symbolic reasoning for enhanced reliability in high-stakes applications. They work by first extracting high-level concepts from the input and then inferring a task label from these compatibly with given logical constraints. Yet, their label and concept predictions can be overconfident, making it difficult for stakeholders to gauge when the model's decisions can be trusted. We address this issue by integrating ideas from Conformal Prediction (CP), a framework providing rigorous, distribution-free coverage guarantees. We formalize three desiderata -- consistency, coverage, and conciseness -- that any conformal method for NeSy-CBMs should satisfy, and show that existing approaches fall short of at least one. We then introduce COCOCO, a post-hoc framework that conformalizes concepts and labels jointly and reconciles them via a single deduction-abduction revision step. COCOCO satisfies all three desiderata, retains distribution-free coverage, is robust to imperfect knowledge and supports user-specified size budgets. Our experiments on 8 data sets highlight how COCOCO compares favorably against competitors and natural baselines in terms of performance and set size.
|
| 2039 |
Beyond Square Roots: A Memory-Efficient Explicit Factorization for Multi-Epoch Private Learning
2605.18379
|
cs.LG
|
Nikita P. Kalinin, Aki Rehn, Joel Daniel Andersson, Antti Honkela, Christoph H. Lampert |
Correlated-noise mechanisms are among the most promising approaches for improving the utility of differentially private model training, but rigorous guarantees require explicit, analyzable factorizations, and practical deployment requires memory efficiency. Re...Correlated-noise mechanisms are among the most promising approaches for improving the utility of differentially private model training, but rigorous guarantees require explicit, analyzable factorizations, and practical deployment requires memory efficiency. Recent works have developed banded inverse factorizations, which address both requirements by exploiting a banded structure in the correlation matrix. The bandwidth controls the size of the noise buffer used to correlate noise across iterations, and thus governs the tradeoff between utility and memory cost. At the two ends of this tradeoff, DP-$\lambda$CGD achieves high memory efficiency by using only a one-step noise buffer, but this limits its utility gains, while the banded inverse square root (BISR) factorization exploits larger correlation windows and is asymptotically optimal for large bandwidths but performs poorly at low bandwidths. To bridge this gap, we introduce $\gamma$-BIFR, a unified generalization of both factorizations. In the low-memory, low-bandwidth regime, $\gamma$-BIFR improves RMSE, amplified RMSE, and private training performance, while yielding tighter theoretical guarantees for multi-participation error in multi-epoch training.
|
| 2040 |
PACE-FNO: Physics-Aligned Canonical Equivariance for Fourier Neural Operators
2605.18606
|
cs.LG
|
Jiaxiao Xu, Changhong Mou, Yeyu Zhang, Fengxiang He |
Neural operators are often tested on states that differ physically from training data. A distinct failure occurs when the physical dynamics are unchanged but the observed coordinate frame differs from training. PACE-FNO addresses this case by estimating the fr...Neural operators are often tested on states that differ physically from training data. A distinct failure occurs when the physical dynamics are unchanged but the observed coordinate frame differs from training. PACE-FNO addresses this case by estimating the frame, predicting after pulling the field to a canonical representative, and restoring the requested terminal frame. The default inference path uses one forward prediction; optional test-time adaptation (TTA) updates only the low-dimensional coordinate. On translated and Galilean-shifted Burgers and shallow-water systems, PACE-FNO lowers out-of-distribution (OOD) relative error by up to $12\times$ relative to a data-augmented Fourier Neural Operator (FNO+Aug). A matched-estimator control supports attributing the gain to prediction in the canonical frame rather than added estimator capacity. Our analysis separates the canonical approximation error from the two alignment residuals. We also test regional WeatherBench2 time-series forecasting with the fifth-generation ECMWF reanalysis (ERA5) under natural time OOD, where PACE-FNO and FNO achieve nearly identical performance. Experiments with approximate rotation, other backbones, irregular domains, rollouts, and image data delineate the conditions under which the mechanism is effective.
|
| 2041 |
When Does Equivariance Help? Canonical Alignment in Neural Fluid Surrogates
2605.18816
|
cs.LG
|
Patryk Rygiel, Julian Suk, Kak Khee Yeung, Christoph Brune, Jelmer M. Wolterink |
Neural surrogates can accelerate computational fluid dynamics (CFD) simulations by orders of magnitude, but practical deployment in engineering and healthcare applications requires architectures that scale to high-resolution meshes and learn effectively from l...Neural surrogates can accelerate computational fluid dynamics (CFD) simulations by orders of magnitude, but practical deployment in engineering and healthcare applications requires architectures that scale to high-resolution meshes and learn effectively from limited data. Explicit equivariance offers a principled inductive bias, yet its accuracy benefits may depend on the prediction task and the distribution of anatomical orientations. We investigate this dependence across three hemodynamic benchmarks with different degrees of natural canonical alignment. To support this study, we introduce the Anchored-Branched Geometric Algebra Transformer (AB-GATr), an $E(3)$-equivariant surrogate that efficiently predicts coupled surface and volume quantities. Across these benchmarks, AB-GATr consistently outperforms the evaluated non-equivariant models, including variants trained with rotational augmentation, while achieving accuracy competitive with $E(3)$-equivariant LaB-GATr at substantially lower training cost. In comparison, rotational augmentation provides inconsistent benefits across architectures and can reduce accuracy. A controlled experiment on ShapeNet-Car shows that strong canonical alignment can favor non-equivariant models, but their accuracy generally deteriorates as training orientations broaden and can decline sharply under broader test rotations. We further investigate these patterns using extended symmetry-breaking diagnostics and probes of the predictive information associated with canonical alignment across all benchmarks. Together, these results support explicit equivariance for the evaluated hemodynamic tasks with natural orientation variation, while showing that its accuracy benefits depend on the task and orientation distribution.
|
| 2042 |
Learning Robust Recommenders from Noisy Implicit Feedback via GMM-Weighted Bayesian Transition Matrix
2605.20721
|
cs.LG
|
Zongyu Li, Xuanyu Liu, Gongce Cao, Shirui Sun, Yaqi Fang |
Label noise is a central challenge in learning from implicit feedback for recommendation. Conventional approaches discard noisy examples for robustness, but this sacrifices data efficiency. Unlike filtering approaches, Bayes-label transition matrix (BLTM) base...Label noise is a central challenge in learning from implicit feedback for recommendation. Conventional approaches discard noisy examples for robustness, but this sacrifices data efficiency. Unlike filtering approaches, Bayes-label transition matrix (BLTM) based methods keep all data, but their transition matrix estimates are skewed in practice. To reduce this skew, we introduce GMM-weighted Bayes-label Transition Matrix (RGBT), which augments BLTM with GMM-based instance weights. A GMM assigns each instance a reliability score, and these scores calibrate the BLTM to reduce bias. We show theoretically that RGBT retains all samples, yields consistent estimates, and provably reduces variance compared to CLTM, generalizing to any method with deterministic labels and non-negative weights. Experiments on real and synthetic datasets show that RGBT handles noisy samples more effectively than sample-selection methods, and calibrates the transition matrix more accurately than existing approaches.
|
| 2043 |
Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning
2605.21127
|
cs.LG
|
Lukas Twist, Helen Yannakoudakis, Jie M. Zhang |
Explicit reasoning models are trained to produce intermediate reasoning traces before final answers, but downstream fine-tuning is often performed on ordinary instruction--response data that contains no such traces. We show that this mismatch can induce reason...Explicit reasoning models are trained to produce intermediate reasoning traces before final answers, but downstream fine-tuning is often performed on ordinary instruction--response data that contains no such traces. We show that this mismatch can induce reasoning-trace collapse: a fine-tuned model continues to produce plausible final answers while losing the structurally valid explicit reasoning traces that made it a reasoning model in the first place. We introduce a structural evaluation framework that separates answer correctness from reasoning-trace validity, measuring valid, empty, missing, and truncated reasoning alongside reasoning-conditioned task performance. Using this framework, we study four open-weight reasoning models and find that standard supervised fine-tuning can rapidly suppress valid reasoning traces, and that answer-only metrics can substantially obscure this failure: in several settings, performance conditional on valid reasoning remains high while the rate of valid reasoning falls sharply. We further show that simple loss-masking strategies can substantially mitigate collapse without requiring teacher-generated reasoning traces. These results suggest that evaluations of fine-tuned reasoning models should report structural reasoning reliability metrics in addition to final-answer performance, especially when adaptation data does not contain explicit reasoning traces.
|
| 2044 |
Factored Diffusion Policies:Compositionally Generalized Robot Control with a Single Score Network
2605.22596
|
cs.LG
|
Sayan Mitra, Ege Yuceel, Noah Giles, Abhishek Pai |
Robotic tasks are typically specified by a tuple of factors, such as the object to be grasped, the obstacles to be avoided, the color of the target, and so on. Collecting expert demonstrations for every combination of factor values grows combinatorially. We pr...Robotic tasks are typically specified by a tuple of factors, such as the object to be grasped, the obstacles to be avoided, the color of the target, and so on. Collecting expert demonstrations for every combination of factor values grows combinatorially. We present factored diffusion policies: a single shared diffusion network trained with per-factor null-token dropout, whose score decomposes additively across factors at inference. Under approximate conditional independence between factors given the action-observation pair, this composition approximates the true joint score with a bounded uniform error, reducing the training-task budget from a product of factor cardinalities to a sum. A trajectory-tube certificate chains this score-level bound through the reverse-time sampling ODE and a contracting tracking controller into a closed-loop state-trajectory tube whose radius factors into an ODE-sensitivity constant and a per-factor score-error budget. Unlike compositional-diffusion methods for control that combine separately trained networks, we use one shared network. Drone racing experiments confirm both the generalization bound and the certificate. On state-based multi-gate racing, the factored policy passes 90% of held-out gates -- matching an oracle -- while a K-network composition baseline collapses to 3%; on vision-based single-gate traversal, it transfers zero-shot to an unseen venue with +11.7pp success-rate gain and 2.4X crash-rate reduction.
|
| 2045 |
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
2605.22967
|
cs.LG
|
Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose |
When Masked Diffusion Models (MDMs) generate sequences through iterative refinement, the rich internal computation over masked positions is discarded, forcing every subsequent refinement step to recompute the valuable internal information stored as model repre...When Masked Diffusion Models (MDMs) generate sequences through iterative refinement, the rich internal computation over masked positions is discarded, forcing every subsequent refinement step to recompute the valuable internal information stored as model representations. To avoid a hard reset between denoising rounds, we propose Learned Relay Representations (Relay), a method that allows MDMs to be forward-thinking when denoising by explicitly learning how to propagate latent information for the benefit of future denoising steps. Relay introduces a differentiable per-token channel that passes information between forward passes and is trained via truncated backpropagation through time (BPTT). We show that this framework can be scaled to state-of-the-art Diffusion Language Models (DLMs), and is seamlessly compatible with techniques like block diffusion and KV caching. We first provide a thorough justification of the design choices in Relay on a challenging Sudoku-based planning task. We then scale Relay to Fast-dLLM v2, a state-of-the-art DLM, outperforming standard supervised finetuning on coding tasks while reducing inference latency by up to 32%. Our empirical results demonstrate that state-of-the-art DLMs can be explicitly trained to relay latent information forward across decoding steps, advancing the performance-latency Pareto frontier. We provide code for all our experiments.
|
| 2046 |
Anytime Training with Schedule-Free Spectral Optimization
2605.23061
|
cs.LG
|
Anuj Apte, Pranav Deshpande, Niraj Kumar, Shouvanik Chakrabarti, Junhyung Lyle Kim |
Standard neural network training relies on learning-rate schedules tied to a fixed horizon, leading to strong path dependence and costly re-tuning as data availability changes. Schedule-Free (SF) methods address this by removing explicit schedules, yet SF-Adam...Standard neural network training relies on learning-rate schedules tied to a fixed horizon, leading to strong path dependence and costly re-tuning as data availability changes. Schedule-Free (SF) methods address this by removing explicit schedules, yet SF-AdamW, the current state-of-the-art anytime optimizer, consistently underperforms well-tuned AdamW baselines. We propose SF-NorMuon, a schedule-free spectral optimizer that closes this gap: with a single hyperparameter configuration, SF-NorMuon matches or exceeds tuned AdamW on 125M and 772M parameter language models across $1$--$8\times$ Chinchilla horizons, and trails horizon-aware, cosine-scheduled NorMuon by only $\sim 0.03$ nats, substantially narrowing the gap between anytime and scheduled spectral optimizers. On the theoretical side, we prove a stationarity guarantee for schedule-free spectral dynamics and identify weight decay at the fast iterate as essential for long-horizon stability. SF-NorMuon enables practitioners to obtain high-quality checkpoints at any point during training without committing to a horizon in advance. Fully matching horizon-aware spectral optimizers without a schedule remains an open problem, which we highlight as a direction for future work. By closing the performance gap with tuned AdamW baselines, SF-NorMuon makes horizon-free optimization more practical, taking a step towards truly open-ended, continual learning.
|
| 2047 |
Complete-muE: Optimal Hyperparameter Transfer and Scaling for MoE Models
2605.23893
|
cs.LG
|
Hongwu Peng, Ohiremen Dibua, Yuanjun Xiong, Yifan Gong, Jin Huang |
We propose Complete-muE, a framework which targets hyperparameter transfer across dense FFN and any Mixture-of-Experts (MoE) setups in transformer blocks. Existing tools such as $\mu$P (requires fixed architectue) or SDE (requires fixed per-step token count) c...We propose Complete-muE, a framework which targets hyperparameter transfer across dense FFN and any Mixture-of-Experts (MoE) setups in transformer blocks. Existing tools such as $\mu$P (requires fixed architectue) or SDE (requires fixed per-step token count) cannot directly solve the hyperparameter transfer problem in MoE setups because Dense to MoE transfer or MoE total experts scaling changes both architecture and tokens per expert. Complete-muE solves this challenge with a two-bridge system: Bridge~I maps between dense FFN and Dense MoE by active-width $\mu$P with a normalized router scale. Bridge~II maps between Dense MoE and sparse MoE by activated-expert scaling, where the first-order SDE LR/WD correction cancels while a bounded residual $\sigma_0$ shift remains. The resulting transfer rule, which we term as Complete muE, covers changes in activated experts, total capacity, granularity, and shared/group-balanced hybrids for MoE models as well as network width/depth, batch size, and duration changes for general Transformer models. Extensive language model and diffusion model pretraining experiments confirm that complete-muE yields relatively stable hyperparameter optima across model architectures and parameter counts -- with only minor drift consistent with the non-strict SDE behavior of Bridge~II. In practice this drift is small enough that hyperparameters tuned on a single dense reference transfer near-optimally to all MoE configurations -- \emph{tune dense once, transfer to all} is the practical recipe at the core of Complete-muE. This enables MoE models to achieve accelerated convergence speedup over dense models when scaling model capacity without costly hyperparameter search.
|
| 2048 |
From Privacy to Generalization: Linear Max-Information Bounds for Differentially Private Learning Algorithms
2605.26222
|
cs.LG
|
Christoph H. Lampert, Max Cairney-Leeming, Hossein Zakerinia |
Understanding the relationship between generalization and privacy remains a challenge in modern machine learning theory, particularly for deep networks that are trained by variants of differentially private stochastic gradient descent (DP- SGD). In this work w...Understanding the relationship between generalization and privacy remains a challenge in modern machine learning theory, particularly for deep networks that are trained by variants of differentially private stochastic gradient descent (DP- SGD). In this work we make progress on this persistent open problem. First, we derive explicit upper bounds on the approximate max-information of any algorithm that fulfills $(\epsilon, \delta)$-differential privacy or R\'enyi differential privacy, thereby going beyond the classical results for pure $\epsilon$-differential privacy. Subsequently, we show even stronger guarantees for two common private learning algorithms, output perturbation with the Gaussian mechanism, and streaming DP-SGD, by exploiting the structure of their internal randomization. As an application of our results, we demonstrate how to obtain non-vacuous PAC-Bayes generalization bounds for deep networks, in which the prior distribution is learned by DP-SGD instead of the classical way of choosing it in a data-independent way.
|
| 2049 |
Inference-Native Zeroth-Order Optimization for LLMs
2605.28760
|
cs.LG
|
Zelin Li, Caiwen Ding |
Zeroth-order (ZO) methods train large language models using only forward passes, yet common ZO execution paths perform substantial work beyond what the algorithm itself requires. To remove this extra execution overhead, we present Infer-ZO, which separates the...Zeroth-order (ZO) methods train large language models using only forward passes, yet common ZO execution paths perform substantial work beyond what the algorithm itself requires. To remove this extra execution overhead, we present Infer-ZO, which separates the evaluations required by the algorithm from how they are executed, allowing them to reuse existing inference-engine optimizations. After removing inherited execution work, a complete Infer-ZO step on Qwen3-14B adds only 1.16% wall-clock time over matched inference, bringing ZO execution close to its fundamental inference workload. With Infer-ZO, ZO evaluations run at near-inference cost with frozen base weights and can share GPU batches with ordinary inference requests. In co-serving experiments, foreground and background throughput remain within 1-3% of the corresponding background-request baseline. Across 15 Qwen3, Llama, and OPT models, Infer-ZO achieves 2.10x-9.94x end-to-end step speedups over released LoZO. Code is available at https://github.com/playeriv65/zo-vllm.
|
| 2050 |
Digitally enriching a high-risk population for pancreatic cancer using routine blood-based measures and clinical histories
2605.30275
|
cs.LG
|
Chris Varghese, Leo Y. Li-Han, Richa Bisht, Ellen Larson, Frank Lee |
Earlier detection of pancreatic cancer is key to enabling wider access to curative treatment and reducing cancer deaths; however, screening is presently not viable. Latent digital indicators of pathology are evident in an individual's disease and blood test tr...Earlier detection of pancreatic cancer is key to enabling wider access to curative treatment and reducing cancer deaths; however, screening is presently not viable. Latent digital indicators of pathology are evident in an individual's disease and blood test trajectories and may predict the development of pancreatic cancer. Longitudinal sequences of coded diagnoses and blood test values accrued by patients throughout their clinical interactions were used to train a custom Transformer-based neural network with a multi-head attention mechanism to predict risk of pancreatic cancer with a multi-year lead time and risk-stratify populations for targeted screening. Mayo Clinic Platform with trained model from Mayo Clinic Rochester validated at Mayo Clinic Arizona, Mayo Clinic Florida, and Mayo Clinic Health Systems. The cohort comprised 6,017 adults with pancreatic cancer and 177,081 controls (median age 75, 45% female) with median 12 years (interquartile range 6.9-16.2) of medical history prior to pancreatic cancer diagnosis. External validation via leave-one-site-out, out-of-sample testing predicting pancreatic cancer 1-, 2-, and 3-years prior to diagnosis demonstrated mean area under the receiver operating characteristic of 0.837 (95% confidence interval 0.827-0.848), 0.797 (95% confidence interval 0.782-0.813), and 0.760 (95% confidence interval 0.745-0.776), respectively. Estimated pancreatic cancer risks were well-calibrated (calibration plot slope 1.08, intercept of -0.077; Brier score 0.025), and a Bayesian population pancreatic cancer prevalence update allows estimated cancer risk outputs to be transportable across settings. At testing, a screening threshold of >3.3% risk of pancreatic cancer in 1-year offered a diagnostic odds ratio of 18.2. Our work therefore lays the foundation for a digital population-level risk enrichment tool that could widen access to curative-intent management.
|
| 2051 |
Effective Biological Representation Learning by Masking Gene Expression
2605.31562
|
cs.LG
|
Kian Kenyon-Dean, Alina Selega, Ihab Bendidi, Jordan M. Sorokin, Luca Bertinetto |
RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery. Modeling such data is challenging due to inherent technical noise and experimenta...RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery. Modeling such data is challenging due to inherent technical noise and experimental batch effects, as evidenced by many existing transcriptomic foundation models (FMs) underperforming relative to linear baselines. Such results raise the question of whether deep representation learning provides a distinct advantage over the direct use of raw transcript counts. Our work explores this by developing a new self-supervised model, TxFM, with a focus on inductive representation learning evaluations. TxFM employs a masked autoencoding approach tailored to diverse RNA-seq count data, and our ablation study empirically identifies crucial architecture configurations required for strong transfer performance. Additionally, we curate a public training corpus, DiverseRNA-1.4M, and find that TxFM trained on this curated dataset yields high-fidelity gene representations that outperform FMs trained on atlas-scale corpora over 100x larger. Overall, our results indicate that inductive self-supervised learning is a viable modeling approach for transcriptomics representation, provided a careful synthesis of model architecture and training data curation.
|
| 2052 |
BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization
2606.00079
|
cs.LG
|
Jiayu Zhao, Zihan Teng, Minhao Fan, Tianrui Ma, Wentao Ren |
Mixture-of-Experts (MoE) large language models incur substantial memory costs due to their large expert parameter counts. Mixed-precision quantization reduces these costs by allocating different bit-widths to experts or linear blocks according to their importa...Mixture-of-Experts (MoE) large language models incur substantial memory costs due to their large expert parameter counts. Mixed-precision quantization reduces these costs by allocating different bit-widths to experts or linear blocks according to their importance. However, assigning a single precision within each expert or linear block overlooks its internal structural heterogeneity. This limitation motivates two key questions: (1) how to define a fine-grained unit for quantization within a linear transformation; and (2) how to characterize the quantization cost of each unit under actual activation patterns and different bit-widths. To address these two questions, we propose BitsMoE, a cost-aware mixed-precision quantization framework built on two complementary techniques: (1) Shared-basis Spectral Decomposition (SSD) separates expert weights into a shared basis and expert-specific spectral components, defining structural quantization units while exploiting cross-expert redundancy. (2) Factorized Quantization Cost Modeling (FQCM) estimates component-wise costs from output reconstruction loss by combining intrinsic spectral importance, activation-dependent importance, and bit-width-dependent distortion. Using these component-wise costs, we formulate bit allocation as an integer linear program (ILP) that minimizes total modeled quantization cost under a fixed memory budget. On Qwen3-30B-A3B at 2-bit, BitsMoE achieves 64.29% average accuracy over seven downstream tasks, outperforming the evaluated MoE-specific methods, including those using ILP-based bit allocation, and exceeding GEMQ by 2.80 percentage points. Under the same setting, it achieves a $16.47\times$ end-to-end offline quantization speedup over GEMQ. It also achieves up to $6.46\times$ the decode throughput of GPTQ.
|
| 2053 |
Context-aware tokenization for Cross-subject Emotion Decoding from EEG
2606.00884
|
cs.LG
|
Jiaxin Qing, Lexin Li |
In cross-subject EEG emotion decoding, neighboring windows can inform the representation of a target window, but they contain both background variation and sustained task signals. This creates a challenge for contextual representation learning because subtract...In cross-subject EEG emotion decoding, neighboring windows can inform the representation of a target window, but they contain both background variation and sustained task signals. This creates a challenge for contextual representation learning because subtracting shared activity can also remove useful information. We introduce the Morlet Spectral Transformer (MST), which conditions target-window tokens on a structured spectral summary before attention. MST averages the Morlet log-amplitude spectra of unlabeled same-trial neighbors and uses frequency-specific electrode projections to combine the target spectrum with its reference-relative residual. This retains access to absolute activity while incorporating context into a fixed attended target-token grid. Without external pretraining, MST achieves 66.5\%, 40.7\%, 36.1\%, and 27.9\% accuracy on SEED, SEED-IV, SEED-V, and SEED-VII under leave-one-subject-out evaluation, averaged over three random seeds, outperforming the evaluated pretrained and from-scratch baselines. On FACED, MST achieves 17.52\% accuracy in nine-class cross-subject evaluation, exceeding the strongest evaluated baseline by 1.15 percentage points. We also performed the ablation studies across all 15 SEED subjects to evaluate the contribution of different components of MST.
|
| 2054 |
ConTraIRL: Factorized Contrastive Abstractions for Transferable IRL
2606.03017
|
cs.LG
|
Yikang Gui, Bikramjit Banerjee, Prashant Doshi |
Reward transfer in Inverse Reinforcement Learning (IRL) is unreliable when policies must generalize to unseen combinations of environment dynamics and task goals. We propose Factorized Contrastive Abstractions for Transferable IRL (ConTraIRL), a framework that...Reward transfer in Inverse Reinforcement Learning (IRL) is unreliable when policies must generalize to unseen combinations of environment dynamics and task goals. We propose Factorized Contrastive Abstractions for Transferable IRL (ConTraIRL), a framework that enables compositional reward transfer by learning decoupled latent representations of these two factors. ConTraIRL uses a dual-encoder architecture that maps observations into separate dynamics and goal latent spaces, trained with a dual contrastive objective. Temporal alignment encourages the dynamics encoder to learn goal-invariant structure, while the goal encoder captures dynamics-invariant features. This factorization supports reward inference under recombined dynamics-goal settings. Experiments on continuous control benchmarks demonstrate effective few-shot transfer to unseen dynamics-goal pairings, improving sample efficiency and reward recovery over transfer IRL baselines.
|
| 2055 |
SPR: Toward a Graph Foundation Model for Transferable Graph Cognition via Spectral Patterns and Relational Geometry
2606.03315
|
cs.LG
|
Ankang Yang, Jitao Zhao, Dongxiao He, Liang Yang, Di Jin |
Recently, Graph Foundation Models (GFMs) have attracted increasing attention for their potential to learn unified and generalizable knowledge across diverse graphs, thereby supporting a wide range of graph scenarios. However, unlike natural language and images...Recently, Graph Foundation Models (GFMs) have attracted increasing attention for their potential to learn unified and generalizable knowledge across diverse graphs, thereby supporting a wide range of graph scenarios. However, unlike natural language and images, graphs lack an intuitive and unified form for perceiving and organizing transferable knowledge. Therefore, despite many initial explorations, a key problem remains unresolved: the transferable cognitive mechanism for graphs. Existing GFMs usually rely on intuitively defined mechanisms to encode transferable graph patterns, such as handcrafted structural templates (e.g., cycles and trees), fixed propagation mechanisms, and predefined random-walk patterns. These mechanisms are often prescriptive and rigid, limiting their ability to flexibly characterize diverse graph patterns across domains. This motivates the exploration of a more flexible, graph-native, and transferable cognitive mechanism. To this end, we analyze graphs from the spectral perspective and propose SPR. SPR decomposes graph information into Chebyshev polynomial bases and learns shared spectral responses, enabling unified yet adaptive cognition of continuously varying spectral patterns across graphs. We further model cross-graph relational geometry to distill recurring pairwise organization from multiple graphs, providing a shared relational reference for transferring such graph cognition across domains. Extensive experiments across diverse graphs and downstream scenarios demonstrate that SPR enables effective and transferable graph cognition.
|
| 2056 |
Dual Advantage Fields
2606.04188
|
cs.LG
|
Alexey Zemtsov, Maxim Bobrin, Alexander Nikulin, Dmitry V. Dylov, Fakhri Karray |
Offline goal-conditioned reinforcement learning requires both long-horizon reachability estimates and local action comparisons. Dual goal representations provide value fields that capture global goal reachability, but they do not directly specify which action ...Offline goal-conditioned reinforcement learning requires both long-horizon reachability estimates and local action comparisons. Dual goal representations provide value fields that capture global goal reachability, but they do not directly specify which action should be preferred at a given state. We propose Dual Advantage Fields (DAF), a policy-extraction method that turns a bilinear dual value model into a local advantage signal. Under bilinear dual parameterization, the goal embedding is the gradient of the value field with respect to the state representation, which directly specifies which action should be preferred. DAF introduces an action-effect model that predicts the discounted direction of change in state representation space induced by actions and scores them by the similarity between this change and the goal direction. In the realizable case, this score equals the goal-conditioned Bellman advantage, yielding a standard local policy-improvement guarantee. On OGBench locomotion, manipulation, and puzzle tasks, DAF improves aggregate RLiable metrics and performs strongly in settings where locally correct actions differ from direct movement toward the final goal.
|
| 2057 |
Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks
2606.06833
|
cs.LG
|
Jiani Xie, Andrew C. Cullen, Paul Montague, Benjamin I. P. Rubinstein |
Automatic Speech Recognition (ASR) systems operating in real-time settings must process acoustic input under strict temporal constraints, where transcription decisions are inherently made on incomplete information. This causal constraint serves as an informati...Automatic Speech Recognition (ASR) systems operating in real-time settings must process acoustic input under strict temporal constraints, where transcription decisions are inherently made on incomplete information. This causal constraint serves as an information bottleneck on attackers, significantly limiting attack performance. Our new Semantic Gambit attack overcomes this causal limitation by augmenting the adversary with predictive context derived from a Large Language Model in real-time. Our experiments show that this form of augmentation can elevate the corpus-level Word Error Rate to 35.6%-a three-fold increase over the current state-of-the-art streaming attack. Ultimately, this work reveals how common, low-latency LLM tooling can be exploited to systematically subvert real-time ASR pipelines.
|
| 2058 |
The Routing Plateau: Understanding the Accuracy Limits of LLM Routers
2606.07587
|
cs.LG
|
Yifan Lu, Qiyue Zhang, Shenrun Zhang, Zhibo Yu, Zhuang Wang |
LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by adaptively selecting a model for each query. Recent work has explored a broad range of routing methods, including clustering-based routers, learned classifiers, ...LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by adaptively selecting a model for each query. Recent work has explored a broad range of routing methods, including clustering-based routers, learned classifiers, pairwise ranking, and confidence-based approaches. Our extensive study of 21 routing methods across five benchmarks reveals a consistent phenomenon that we call the routing plateau (Fig. 1): many methods, including kNN, achieve very similar accuracy and converge to a narrow performance range that remains far below the oracle router. Our analysis supports a correctness-prediction bottleneck hypothesis: current routers primarily learn global-average model performance trends rather than fine-grained, query-specific routing signals. As a result, they collectively fail on queries that require instance-specific routing decisions. Moreover, to understand whether the plateau can be alleviated with a better training setup, we construct a 300K-query benchmark (Nine-by-300k). More data, larger encoders, and end-to-end fine-tuning improve eight routers by 1.24 pp on average, but leave the plateau largely intact. These findings suggest that further progress may require inputs beyond the query itself, such as partial output trajectories that reveal how models attempt the task.
|
| 2059 |
Agentic Search for Counterfactual Recourse under Fixed LLM Budgets
2606.08696
|
cs.LG
|
Yasuo Tabei |
Counterfactual recourse aims to provide actionable feature changes that would alter an unfavorable decision made by a predictive model. In practice, affected individuals often benefit from multiple feasible alternatives rather than a single optimal explanation...Counterfactual recourse aims to provide actionable feature changes that would alter an unfavorable decision made by a predictive model. In practice, affected individuals often benefit from multiple feasible alternatives rather than a single optimal explanation. A natural way to produce such alternatives is to prompt large language models (LLMs). However, prompting incurs a practical constraint: the number of LLM calls is often the dominant computational and economic cost. Together, the need for multiple alternatives and this cost constraint shift the problem from finding a single high-quality counterfactual to efficiently generating a set of oracle-validated counterfactuals under a fixed LLM-call budget. In this work, we study counterfactual recourse generation in the LLM-agentic setting as a fixed-budget search problem and propose Recourse Monte Carlo Tree Search (ReCo-MCTS), an agentic tree-search framework that aims to increase the yield of unique, oracle-validated counterfactuals under this budget while accounting for the extent of the required changes. ReCo-MCTS combines LLM-based multi-candidate generation, constraint checking, black-box oracle evaluation, and UCT-guided tree search to accumulate valid counterfactuals under a fixed LLM-call budget. On four real-world tabular datasets, ReCo-MCTS returns more counterfactuals than the evaluated baselines in our main comparison under the specified resource limits, with trade-offs in the extent of the required changes.
|
| 2060 |
nCMD: Benign-Anchored Feature Selection for Imbalanced Network Intrusion Detection
2606.09934
|
cs.LG
|
Abu Fuad Ahmad, Istiaque Ahmed |
Feature selection is critical for network intrusion detection systems (NIDS) operating under high-dimensional, highly imbalanced traffic, as found in operational and defense networks. Traditional filter methods rank features using global statistics computed sy...Feature selection is critical for network intrusion detection systems (NIDS) operating under high-dimensional, highly imbalanced traffic, as found in operational and defense networks. Traditional filter methods rank features using global statistics computed symmetrically across classes and thus fail to capture the asymmetry of intrusion detection, where attacks are best characterized as deviations from dominant benign traffic. We propose benign-anchored Classwise Mean Deviation (nCMD), a lightweight and interpretable method that scores feature relevance based on the deviation of attack-class distributions from the benign-class mean, rather than a globally biased reference. This approach aligns feature selection with the operational semantics of NIDS at no additional computational cost. Across four benchmark datasets (CICIDS2017, CICDDoS2019, NSL-KDD, and UNSW-NB15), multiple feature budgets, and three downstream classifiers, nCMD matches or exceeds classical filter baselines in macro-averaged F1-score. It achieves the best result on three of the four datasets and under every classifier, with the strongest improvements observed under tight feature budgets and severe class imbalance. These results support benign-anchored ranking as a scalable and interpretable preprocessing component for resource-constrained NIDS.
|
| 2061 |
AugRelNet: Relation-Augmented Dynamics for Structure Discovery and Forecasting from Limited Data
2606.11251
|
cs.LG
|
Xingji Cui |
Complex dynamics often arise from persistent interactions whose effects vary with the system state. We introduce AugRelNet for short- and medium-horizon trajectory forecasting from limited data. A shared learnable relation matrix explicitly routes source signa...Complex dynamics often arise from persistent interactions whose effects vary with the system state. We introduce AugRelNet for short- and medium-horizon trajectory forecasting from limited data. A shared learnable relation matrix explicitly routes source signals to targets, while evolving local features \(u\) and a persistent response state \(\eta\) transform them into nonlinear, history-dependent responses. Feeding these responses back into the dynamics separates shared relational organization from its changing realization. The learned relations reflect dense coupling in Lotka--Volterra systems and develop structured, neighbour-preferential organization in locally interacting systems. On 40-dimensional Lorenz--96, AugRelNet uses about 4,000 training states, roughly an order of magnitude fewer than representative prior Lorenz--96 forecasting studies. Across five seeds, every true interaction ranks above every nonedge, with a mean absolute-weight contrast of approximately \(146{:}1\). Under the same data budget, eight-step RMSE is about \(23\)--\(59\times\) lower than MNO, DySLIM, DeepSkip, and AL-RNN, \(14\)--\(55\times\) lower than NRI and dense GraphODE, and 37\% lower than PySINDy with a matched quadratic dictionary. The same relational representation extends to grids by treating sites as variables and sharing relations across spatial offsets: in discrete (Game of Life), continuous (FitzHugh--Nagumo), and stochastic (Ising) dynamics, learned weights rank physical neighbours above other candidate offsets while supporting multi-step prediction. AugRelNet learns these relations through prediction losses and regularization, without structural labels or system-specific physical losses.
|
| 2062 |
Running the Gauntlet: Hard Agentic Tasks
2606.14397
|
cs.LG
|
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel, Michal Zakrzewski, Sebastian Montagna |
As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and...As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we introduce GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), across five less-covered professional applications (Video Editor, Workflow Builder, 3D Modeller, Flight Analyser, and Circuit Designer), each with 27 vision-intensive tasks (135 in total). Our benchmark provides a modular pipeline that comprises an environment compatible with both open- and closed-source agent frameworks, a controlled web-based application, a well-structured task suite, and an automated evaluation engine with diverse metrics. Contrary to widespread expectations, our empirical results reveal that frontier agentic systems remain far from achieving human-level performance. Even the state-of-the-art agent achieves only a 28.2% success rate on our GauntletBench, highlighting the limitations in these overlooked capabilities and generalisation. By comparison, non-expert human annotators achieve over 80% success on our challenging yet feasible tasks, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.
|
| 2063 |
StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling
2606.15197
|
cs.LG
|
Jiajun Li, Yu Ding, Shisi Guan, Ran Hou, Wanyuan Wang |
Optimization modeling is inherently hierarchical, requiring a precise sequence of symbolic commitments. Traditional learning-based automated optimization modeling methods improve modeling policies through large-scale annotated or curated training data, but are...Optimization modeling is inherently hierarchical, requiring a precise sequence of symbolic commitments. Traditional learning-based automated optimization modeling methods improve modeling policies through large-scale annotated or curated training data, but are costly to adapt to new problem distributions. Meanwhile, one-shot generation remains brittle in hierarchical modeling, where early symbolic errors can propagate into invalid formulations. Test-time scaling offers a promising alternative by enabling structural exploration with additional instance-level computation; however, existing search-based methods typically rely on a fixed policy, causing repeated rollouts to inherit similar modeling biases and providing limited credit assignment for intermediate decisions. To address these limitations, we propose \textbf{StarOR}, a synergistic search-and-adaptation framework that couples MCTS with Test-Time Reinforcement Learning for optimization modeling. StarOR decomposes the modeling process into four stages and updates a transient LoRA adapter via GRPO at each non-terminal node. By using MCTS-generated siblings as local comparison sets, StarOR transforms search-time exploration into instance-specific policy refinement. Moreover, an unsupervised multi-faceted reward system provides fine-grained feedback for intermediate formulation decisions without ground-truth labels. Across five optimization benchmarks, StarOR attains a 65.0\% average accuracy with a 4B backbone, matching the highest average and outperforming the evaluated same-backbone test-time baselines. Code is available at \href{https://github.com/Liwow/StarOR}{StarOR}.
|
| 2064 |
Wasserstein Convergence of ODE-Based Samplers in Decentralized Diffusion Model via Velocity Field Decomposition
2606.15835
|
cs.LG
|
Chencheng Tang, Xuanyu Xue, Fangyikang Wang, Chao Zhang, Hubery Yin |
Diffusion models have achieved impressive empirical success in generative tasks, and their convergence theory is now relatively well understood. Motivated by privacy and scalability, recent decentralized diffusion architectures replace a single global velocity...Diffusion models have achieved impressive empirical success in generative tasks, and their convergence theory is now relatively well understood. Motivated by privacy and scalability, recent decentralized diffusion architectures replace a single global velocity field with multiple local experts and a routing mechanism, yielding a sampling dynamics with stochastic expert switching that falls outside standard diffusion convergence analyses. In this work, We study a decentralized diffusion framework with stochastic velocity fields and ODE-based sampling. We establish a convergence guarantee in Wasserstein-2 distance, showing that the distribution of the $N$-step discretization converges to the analytical solution at rate $\mathcal{O}(N^{-1/2}+\varepsilon)$ in $W_2$, where $\varepsilon$ captures the neural approximation errors. To our knowledge, this is the first $W_2$ convergence result for decentralized diffusion models with an ODE-based sampling scheme.
|
| 2065 |
RepNN: Tackling spectral bias in deep neural networks for regression and PDE problems via parameter reparameterization
2606.16575
|
cs.LG
|
Yong Wang, Tao Zhou, Xuhui Meng |
Deep neural networks (DNNs) have achieved remarkable success in scientific computing, yet they often suffer from spectral bias in capturing oscillatory and multiscale behaviors. In this study, we investigate this limitation by examining the failure of shallow ...Deep neural networks (DNNs) have achieved remarkable success in scientific computing, yet they often suffer from spectral bias in capturing oscillatory and multiscale behaviors. In this study, we investigate this limitation by examining the failure of shallow ReLU neural networks in fitting high-frequency functions. This observation identifies two important factors in resolving rapid oscillations: the initial slope scale and the distribution of partition points induced by the networks. Motivated by this analysis, we propose RepNN, a reparameterized neural network model with ReLU or tanh activations designed for high-frequency and multiscale problems. The key idea is to reparameterize the weights and biases in the first hidden layer, which enables effective control of the initial slope scale and provides an appropriate distribution of the initial partition points. Furthermore, treating the reparameterized weights and biases as trainable parameters allows the DNN to achieve adaptive frequency scaling during training. In addition, we derive quantitative estimates for the output and slope magnitudes of the reparameterized DNN to guide the initialization of the proposed method. Numerical experiments, including multiscale one-, two-, and four-dimensional function approximations, forward and inverse PDE problems in combination with physics-informed neural networks (PINNs), and operator learning for an earthquake problem using real data, demonstrate that RepNN improves the predicted accuracy of vanilla DNNs in capturing highly oscillatory features. These results indicate that RepNN provides an effective and flexible approach for overcoming spectral bias and applying DNNs to multiscale problems.
|
| 2066 |
When Does Depth Matter For In-Context Learning? Adaptive Inference in Deep Transformers
2606.16694
|
cs.LG
|
Ravin Raj, Gautam Reddy |
Transformers perform computations through many successive attention and feedforward blocks, allowing them to learn complex correlations between a large collection of coupled variables. When does stacking successive attention-feedforward computations over many ...Transformers perform computations through many successive attention and feedforward blocks, allowing them to learn complex correlations between a large collection of coupled variables. When does stacking successive attention-feedforward computations over many layers provide a computational advantage over a single transformer block? We address this question by examining in-context learning in generalized linear attention transformers. We first introduce a general theory of distributed inference in such transformers, subject to constraints on communication and depth. We show that such systems can exploit internal representations (`function vectors') to infer a latent context variable at increasingly finer scales over its layers. For an in-context linear regression task, the theory predicts that while one-layer transformers without feedforward blocks are optimal for Gaussian priors over the context variable, multi-layer transformers are superior for non-Gaussian, tree-like priors. Trained linear attention transformers reproduce quantitative predictions from the theory. Using causal key-patching experiments, we verify that function vectors in intermediate layers mediate adaptive routing of information. Our results suggest that depth and feedforward blocks enable transformers to implement adaptive inference, and this is advantageous when the distribution over latent variables has hierarchical structure.
|
| 2067 |
Filtered Conformal Ellipsoids for Graph-Native Time Series
2606.17014
|
cs.LG
|
Yannick Limmer |
Joint prediction sets for multivariate time series should control a single event while adapting to cross-coordinate dependence. We study filtered conformal ellipsoids: a frozen state-space filter emits a one-step predictive mean and covariance, and split-confo...Joint prediction sets for multivariate time series should control a single event while adapting to cross-coordinate dependence. We study filtered conformal ellipsoids: a frozen state-space filter emits a one-step predictive mean and covariance, and split-conformal calibration is applied to the resulting Mahalanobis scores. The filter is used to choose the ellipsoid shape; conformal calibration chooses the scalar radius, so the construction benefits from a learned predictive covariance without relying on Gaussian tail probabilities for coverage. The main difficulty is that filtered scores are dependent and learned recurrent filters need not contract in their raw hidden state; we therefore analyse contraction in an observable predictive-law quotient that identifies hidden states producing the same future sequence of emitted Gaussian laws. Under a stable Bayes Gaussian-projection filter, covariance bounds, and a finite-horizon observability Fisher condition, small excess Gaussian negative log-likelihood implies contraction of the learned emitted laws. Combined with a threshold-autocovariance envelope this yields a Chebyshev-type approximate coverage bound for filtered split-conformal prediction under dependence; a sharper Bernstein-type bound requires an additional geometric-mixing concentration assumption. Under Gaussian oracle realisability we also obtain a near-oracle log-volume comparison within the class of conditionally valid Gaussian ellipsoid rules. We instantiate the framework with a GCN-GRU filter with diagonal-plus-low-rank covariance. On moderate-size graph-native traffic benchmarks (METRLA-$20$ and PEMSBAY-$50$), the learned filter gives sharper at-target ellipsoids than static-covariance and non-filter baselines; at full-graph scale and on non-graph-native datasets, factor and copula baselines can be stronger.
|
| 2068 |
Recursive Scaling in Masked Diffusion Models
2606.18022
|
cs.LG
|
Alba Carballo-Castro, Julianna Piskorz, Paulius Rauba, Mihaela van der Schaar, Pascal Frossard |
Masked diffusion models (MDMs) generate sequences by iteratively refining a partially masked state and committing tokens in parallel. We introduce recursion in MDMs and propose new Recursive Masked Diffusion Models (R-MDMs), which apply a shared denoising tran...Masked diffusion models (MDMs) generate sequences by iteratively refining a partially masked state and committing tokens in parallel. We introduce recursion in MDMs and propose new Recursive Masked Diffusion Models (R-MDMs), which apply a shared denoising transformer $L$ times within each denoising step, adding recursive depth as an additional compute axis without increasing parameter count. Across structured generation tasks, recursive depth improves quality at fixed parameter budget, matches substantially larger non-recursive models at matched FLOPs, and can reduce the number of denoising steps needed to reach a target quality. We interpret these gains with a dependence--fidelity decomposition of parallel decoding error: recursion refines model marginals at a fixed masked state, whereas denoising steps change that state by committing tokens. Building on this analysis, we propose to treat decoding as a two-axis decision (how many loops to run and which tokens to commit) and show that entropy-guided adaptive rules improve the quality--compute frontier over fixed schedules, transferring across various tasks on Sudoku, Countdown, RNA, and executable math generation. Together, these results establish recursive depth as a practical, complementary test-time scaling mechanism for MDMs.
|
| 2069 |
CODEBLOCK: Learning to Supervise Code at the Right Granularity
2606.18286
|
cs.LG
|
Zhijie Deng, Ling Li, Jinlong Pang, Kaiqin Hu, Qi Xuan |
Supervised fine-tuning of code LLMs typically applies uniform cross-entropy loss to all response tokens, implicitly assuming that every token provides equally useful learning signals. Recent token-level selection methods challenge this assumption in natural-la...Supervised fine-tuning of code LLMs typically applies uniform cross-entropy loss to all response tokens, implicitly assuming that every token provides equally useful learning signals. Recent token-level selection methods challenge this assumption in natural-language SFT by supervising only high-value tokens. However, such pointwise selection can fragment the syntactic structures and program dependencies of code, leaving supervision scattered across incomplete code units. In experiments on 30K code instruction-response pairs, under the same 10% token budget, supervising complete coding blocks improves average performance by 11.7 points over prior approaches that supervise isolated tokens. Motivated by this observation, we propose CodeBlock, a structure-aware sparse supervision framework that uses complete, parser-aligned coding items as the basic units of supervision. CodeBlock constructs coding items from high-quality code instruction data, estimates their supervision utility using GCE, which is more robust to low-probability tokens, and further adjusts their supervision priority using data-flow reach and bridge signals. During training, the full response is retained as context, while loss is applied only to the selected supervision units. Experiments show that partial supervision over only about 6.9% of response tokens consistently outperforms full-token SFT across all five model settings, suggesting that effective code SFT depends not only on identifying high-value supervision, but also on allocating it at the appropriate structural granularity.
|
| 2070 |
DRIFT: Data Selection for LLM Instruction Tuning via On-Policy Attribution
2606.18307
|
cs.LG
|
Zefan Wang, Lincheng Li, Tianyu Yu, Yuan Yao |
Data selection for instruction tuning is usually studied for efficiency, reaching strong performance with a fraction of the training data. We ask whether data selection can instead raise the performance ceiling, producing a better model than training on all th...Data selection for instruction tuning is usually studied for efficiency, reaching strong performance with a fraction of the training data. We ask whether data selection can instead raise the performance ceiling, producing a better model than training on all the available data. To test this, we take a large language model (LLM) that has already completed supervised fine-tuning (SFT) on its full corpus and train it further on examples selected from that same corpus. Any gain therefore lifts the model above the ceiling set by full-data training. Existing selection methods yield little or no gain in this setting. We propose DRIFT (Data Refinement via On-Policy Influence Functions for Supervised Fine-Tuning), a data attribution method built on influence functions. DRIFT samples the model's own responses to validation queries and weights them by correctness to form a validation objective. It ranks the examples of the original corpus by their estimated influence on this objective, and the model then continues SFT on the top-ranked subset. This design is motivated by the proximity gap in influence function theory: we hypothesize that these on-policy responses are better validation targets than external references in terms of locality. DRIFT also corrects influence scores for their dependence on gradient norm within each validation task. DRIFT outperforms all evaluated data selection baselines on Olmo3-7B-Instruct-SFT and OpenR1-Distill-7B, raising average accuracy across eight benchmarks from 39.17 to 40.14 and from 57.23 to 58.66, respectively, with gains on benchmarks not used for attribution.
|
| 2071 |
When to Commit and When to Defer: Maturing Markov Decision Processes under Refining Information and Expiring Opportunities
2606.18820
|
cs.LG
|
Jiaxi Liu, Aiping Yang, Yuhang Yang, Shuqi Zhang, Zewei Dong |
Sequential decisions often become better informed as action opportunities expire. We introduce Maturing Markov Decision Processes (MMDPs), a structured subclass of augmented finite-horizon Markov decision processes, for organizing when to commit and when to de...Sequential decisions often become better informed as action opportunities expire. We introduce Maturing Markov Decision Processes (MMDPs), a structured subclass of augmented finite-horizon Markov decision processes, for organizing when to commit and when to defer. The Expiring-Actions-First Principle provides a sufficient condition for deferral based on information gain and delay loss. This structure motivates stage-local policies, action abstraction, and policy-guided planning. Replenishment experiments show improved learning and test performance under the MMDP interface and, in the primary setting, a larger test-performance gain from forecast refinement. In synthetic cash management, MMDP improves PPO test returns across both network sizes; five-account equal-size mask controls show that action count alone does not explain its learned gain, and the advantage persists with action abstraction and policy-guided search. A production-scale cash-management case study adds complementary evidence. Together, these results support commitment timing as a useful inductive bias for learning and planning.
|
| 2072 |
CRAX: Fast Safe Reinforcement Learning Benchmarking
2606.20376
|
cs.LGcs.AI
|
Tristan Tomilin, Mourad Boustani, Mickey Beurskens, Khoi H. B. Nguyen, Thiago D. Sim\~ao |
Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. While benchmarks have been central to progress in RL, existing 3D physics-based safety benchmarks remain computationally sl...Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. While benchmarks have been central to progress in RL, existing 3D physics-based safety benchmarks remain computationally slow, limiting large-scale experimentation and rapid prototyping. To address this gap, we propose CRAX (Constrained RL Accelerated with JAX). Built on top of the MuJoCo XLA (MJX) physics engine, CRAX leverages vectorized operations and hardware acceleration, yielding up to 200x faster training over comparable CPU-based safety benchmarks. The benchmark features eight tasks spanning three difficulty levels and multiple agent morphologies. Evaluating seven popular safe RL methods, we find that none dominates across tasks, and that learning safe policies from pixels remains largely unsolved.
|
| 2073 |
Real vs. Complex Spectral Bases for Neural Operators: The Role of Green's Function Alignment
2606.24851
|
cs.LG
|
Jason Sulskis, Sathya Ravi |
Fourier Neural Operators (FNO) learn solution operators of partial differential equations by parameterizing global convolutions in the complex Fourier domain. For real-valued PDE solutions, the complex FFT carries representational redundancy through conjugate ...Fourier Neural Operators (FNO) learn solution operators of partial differential equations by parameterizing global convolutions in the complex Fourier domain. For real-valued PDE solutions, the complex FFT carries representational redundancy through conjugate symmetry. We introduce the Hartley Neural Operator (HNO), the exact real-valued mirror of FNO: it replaces the FFT with the purely real Discrete Hartley Transform and learns a single real multiplier per retained spectral mode, with no complex arithmetic. Because the real Hartley spectrum is not halved by conjugate symmetry, HNO retains twice as many frequency corners as FNO but one real weight where FNO carries a complex pair, so the two operators are iso-parametric at equal width and differ only in spectral basis. Our central thesis is that the best basis is a property of the operator. Self-adjoint elliptic operators (Poisson, biharmonic) have real, symmetric Green's functions that the real Hartley multiplier diagonalizes exactly, and HNO is favored there. Time-dependent operators carry phase, from oscillation in the wave equation to transport in advection, Burgers, and Navier-Stokes, which a real diagonal multiplier cannot represent, so FNO is favored there, and increasingly so with the operator's phase content, leaving the phaseless heat equation as the borderline case. Training both operators identically and benchmarking across PDE classes, initial-condition families, and boundary conditions, we find an elliptic-versus-time-dependent split that is monotone in operator phase content and matches the Green's-function theory we develop. Rather than a universal winner, our findings give a predictive rule: match the spectral basis to the symmetry of the solution operator.
|
| 2074 |
Fast LeWorldModel
2606.26217
|
cs.LG
|
Yuntian Gao, Xiangyu Xu |
Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candidate action sequences by repeatedly applyi...Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candidate action sequences by repeatedly applying a local one-step latent transition model. This autoregressive rollout makes planning computationally expensive and exposes the predicted trajectory to accumulated latent errors as the horizon grows. We propose Fast LeWorldModel (Fast-LeWM), a fast latent world model that replaces repeated local rollout with action-prefix prediction. Given the current latent and a candidate action sequence, Fast-LeWM encodes its prefixes and predicts the future latents reached after executing those prefixes in parallel. By making action prefixes the basic prediction unit, Fast-LeWM directly models action effects accumulated to different extents over multiple horizons. This prefix-level supervision forces the model to learn how states continuously evolve under different action prefixes, rather than only fitting one-step state transitions. During planning, the predictor can use the prefix token from the encoded action sequence to evaluate the corresponding future latent without explicitly rolling through each intermediate imagined state. Across multiple tasks, Fast-LeWM improves average success over LeWM while substantially reducing planning time, achieving lower open-loop latent loss whose growth becomes significantly slower as the rollout horizon increases.
|
| 2075 |
The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment
2606.27739
|
cs.LG
|
Tianyu Jia, Yue Fang, Hongxin Ding, Rihong Qiu, Zhibang Yang |
Process reward models (PRMs) enhance the reasoning capabilities of large language models (LLMs) by providing fine-grained feedback, yet training PRMs typically requires expensive stepwise annotations. Outcome-supervised PRMs offer a scalable alternative by lea...Process reward models (PRMs) enhance the reasoning capabilities of large language models (LLMs) by providing fine-grained feedback, yet training PRMs typically requires expensive stepwise annotations. Outcome-supervised PRMs offer a scalable alternative by learning from final-answer correctness alone, but this introduces a fundamental credit assignment challenge, i.e., attributing outcomes to responsible reasoning steps. Existing approaches rely on either uniform or causal assignment, both of which fail to anchor credit in step correctness and thus hinder process error identification. In this work, we propose Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment (LCA), an outcome-supervised PRM framework that jointly learns credit assignment and reward modeling under the principle of Weakest Link Assignment: a reasoning chain is as strong as its weakest link. To address mutual dependence between credit assignment and reward modeling, we formalize outcome-supervised PRM as a Multiple Instance Learning (MIL) problem and introduce Softmax-Weighted-Sum (SWS) pooling, an MIL pooling technique tailored for strong dependence and redundancy among reasoning states. We establish instance-level Bayes consistency of the resulting learning objective under mild assumptions. Extensive experiments demonstrate that LCA outperforms state-of-the-art outcome-supervised PRMs across multiple tasks and backbones. Code is available at https://github.com/JiaTianyu20031016/MIL.git.
|
| 2076 |
A Gravitational Interpretation of Safety Reversion under Fine-Tuning
2606.28525
|
cs.LG
|
Samuele Poppi, Nils Lukas |
Safety alignment in large language models can degrade during post-training even when neither the data nor the objective is intentionally adversarial. Alignment rebound and reverse dynamics suggest that this degradation may reactivate behavior suppressed during...Safety alignment in large language models can degrade during post-training even when neither the data nor the objective is intentionally adversarial. Alignment rebound and reverse dynamics suggest that this degradation may reactivate behavior suppressed during safety alignment. Building on these ideas, we hypothesize that ordinary non-adversarial post-training follows a reversion direction: the activation-space displacement from the safety-aligned model toward a more permissive, earlier helpful-only state. We see that for Llama, every tested trajectory across references, tasks, and seeds exceeds a matched empirical null, while at aligned Llama and Qwen checkpoints, a vocabulary readout shows that the direction locally favors task-engaging over fixed refusal-like openings. Its geometric expression is behaviorally informative: as post-training proceeds, alignment with the direction and harmfulness increase together, yielding a strong descriptive correlation (Spearman r=0.958). To move beyond correlation, we test causal relevance during adaptation using objectives constructed from this coordinate. Across all tested Llama, Qwen, and Gemma settings from 3B to 14B, an optimizer-matched objective opposing positive motion reduces geometric alignment and harmfulness relative to ordinary fine-tuning, whereas a separately stabilized objective reinforcing that motion increases both. Every model and scale exhibits the same mean block-baseline-push ordering, showing that the causal relevance of the reversion direction is not tied to one architecture or model size. Finally, we show that a standard safety-rehearsal objective, built without access to the direction, independently opposes it and cuts cumulative reversion by about 30% in Llama and Qwen.
|
| 2077 |
Mind the Residual Gap: Probabilistic Downscaling under Real-World Bias
2606.30821
|
cs.LG
|
Yujin Kim, Nidhi Soma, Sarah Dean |
Probabilistic downscaling models the conditional distribution of high-resolution fields given coarse inputs and is a core challenge in atmospheric science, climate modeling, and other multiscale physical systems. A widely used paradigm decomposes the problem i...Probabilistic downscaling models the conditional distribution of high-resolution fields given coarse inputs and is a core challenge in atmospheric science, climate modeling, and other multiscale physical systems. A widely used paradigm decomposes the problem into a deterministic mean predictor followed by a stochastic residual generator. While effective in idealized settings, this mean-residual approach frequently produces biased and underdispersive ensembles in real-world applications. We identify a fundamental source of this failure: residual target misspecification, where the residual distribution induced during training systematically differs from the correction distribution required at test time, a mismatch amplified by downscaling bias. We theoretically connect this misspecification to underdispersion through the spread-skill ratio and show that conditional residual distribution matching controls both residual-energy and conditional-mean mismatch. Motivated by this analysis, we introduce ReMatch (Residual Distribution Matching), which aligns training residual targets toward a held-out calibration regime via optimal transport. On a controlled synthetic benchmark with varying bias levels and a real-world HRRR-ERA5 wind-field downscaling task, ReMatch substantially reduces underdispersion, improves calibration, and outperforms strong mean-residual and super-resolution baselines. Our code is available at https://github.com/sdean-group/ReMatch.git.
|
| 2078 |
Revisiting the Volume Hypothesis
2606.31282
|
cs.LG
|
Ari Pakman, Lior Kreimer, Yakir Berchenko |
Modern deep neural networks often contain far more parameters than needed to fit their training data, yet they achieve impressive generalization. A common explanation for this success is the implicit bias of stochastic gradient descent (SGD). An alternative vo...Modern deep neural networks often contain far more parameters than needed to fit their training data, yet they achieve impressive generalization. A common explanation for this success is the implicit bias of stochastic gradient descent (SGD). An alternative volume hypothesis posits that, within low training-loss regions, loss-landscape basins leading to strong generalization occupy much larger regions of weight space than basins that generalize poorly, and therefore SGD is simply more likely to land in the former. Recent experimental explorations of this idea present seemingly contradictory results. While in one set of experiments randomly sampling the network weights until achieving zero training error yielded poor generalization, molecular dynamics density estimates supported the volume hypothesis. We observe that these experiments were performed at different dataset size regimes, and explore an intermediate regime using the Replica Exchange Wang-Landau algorithm to estimate the joint density of states over training and test accuracies in binary networks. Across several architectures and datasets, we show that the generalization advantage of gradient learning over random sampling training generally diminishes as the training data size grows, suggesting a resolution of the paradox.
|
| 2079 |
Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition
2607.00535
|
cs.LG
|
Zhiqi Li, Wen Zhang, Bo Zhu |
Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by learning long-range transport maps between noise and data. However, their deterministic transitions do not directly provide the stochastic trajectories and tractable ...Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by learning long-range transport maps between noise and data. However, their deterministic transitions do not directly provide the stochastic trajectories and tractable likelihood ratios required by reinforcement learning (RL) post-training. Existing SDE-based stochasticization techniques target velocity-based samplers and do not directly extend to long-range flow-map transitions. We propose Flow-Map GRPO, an online RL post-training framework for deterministic few-step flow-map generators. Its key component, Anchored Stochastic Flow Map Composition (ASFMC), combines deterministic transport with anchor-based conditional resampling. We establish the conditions under which this construction preserves the marginal probability path and develop tractable local- and endpoint-anchor policies for two-time and single-time flow maps. These policies enable a unified GRPO training procedure. Experiments on FLUX-based MeanFlow and sCM generators demonstrate substantial improvements in OCR, PickScore, and GenEval at different numbers of inference steps, including joint OCR--PickScore gains with mixed rewards. Controlled ablations show that the stochastic transition design is essential for translating training rewards into generation quality. Flow-Map GRPO enables effective RL alignment of pretrained deterministic flow-map generators while retaining their original parameterization, without retraining them as native stochastic models.
|
| 2080 |
Scaling Laws for Collapse in Asynchronous GRPO
2607.01083
|
cs.LG
|
Jingwei Song, Haofeng Xu, Jie Xiao, Chengke Bao, Jingwei Shi |
Asynchronous reinforcement learning improves the throughput of large language model post-training by decoupling rollout generation from policy optimization, but introduces a mismatch between the behavior and learner policies. How the resulting policy staleness...Asynchronous reinforcement learning improves the throughput of large language model post-training by decoupling rollout generation from policy optimization, but introduces a mismatch between the behavior and learner policies. How the resulting policy staleness couples with the learning rate to govern training stability and collapse time remains poorly understood. We investigate this coupling in vanilla GRPO through controlled sweeps of the synchronization interval $S$ and constant learning rate $\eta$ on Llama-3.2-1B/3B, complemented by experiments on Qwen3-8B. We identify two empirical scaling laws: (i) Stability-boundary scaling: the largest stable learning rate scales approximately as $S^{-1}$, yielding a stability boundary characterized by an approximately constant product $S\eta$. (ii) Collapse-time scaling: among collapsing runs, estimated collapse times scale approximately as $\eta^{-1}$, corresponding to a model- and setup-dependent cumulative learning-rate budget that aligns across synchronization intervals in the Llama sweeps. We interpret these laws through a local analysis of the behavior-dependent GRPO surrogate and a complementary mean-field model. Under local regularity conditions, the analysis yields an $O(S\eta)$ upper bound on the staleness-induced update bias that resets at synchronization. The mean-field model shows how sufficiently strong positive feedback can sustain directional drift when update directions persist across synchronization cycles. When drift speed saturates under optimizer normalization, this mechanism predicts exit from a local surrogate-validity region after an approximately fixed cumulative learning rate. Together, these findings motivate a practical calibration rule: estimate the stability threshold and collapse budget from a coarse sweep, then jointly select $S$ and $\eta$ for the intended training horizon.
|
| 2081 |
Sampling Meets Interaction: Sequentially-Controlled Multi-Particle Flow-Maps for Efficient Inference-time Search
2607.01144
|
cs.LG
|
Binglin Ji, Anindya Sarkar, Hengchang Lu, Jens Sj\"olund, Yevgeniy Vorobeychik |
While generative models have enabled training-free reward alignment, existing particle-based methods are fundamentally constrained by their propensity to local exploration within narrow regions of the underlying distribution, severely restricting sample divers...While generative models have enabled training-free reward alignment, existing particle-based methods are fundamentally constrained by their propensity to local exploration within narrow regions of the underlying distribution, severely restricting sample diversity. This limitation becomes especially acute under tight reward-feedback budgets, where effective search demands broad, strategic exploration to uncover high-utility regions. To address this, we propose Sequentially-Controlled Interactive Multi-Particle Flow-Maps (IMPFM), a framework for feedback-efficient search. IMPFM progressively transports a group of interactive particles toward the target distribution, maintaining the broad coverage essential for heterogeneous preference alignment. IMPFM leverages a principled and efficient posterior sample-sharing mechanism across particles powered by flow maps. By correcting individual particle drift with the collective value gradient from the entire ensemble's posterior samples at each correction step, the framework maximizes sample utility to enable global exploration while actively mitigating reward over-optimization, typical of standard control frameworks. Paired with a principled exploration-exploitation reweighting mechanism involving multi-particle interaction, this sequentially corrected multi-particle dynamics explicitly preserves structural diversity and overcomes the weight degeneracy inherent to standard Sequential Monte Carlo (SMC) samplers. Crucially, we prove that the resulting sampling framework yields a multi-particle interaction-aware Feynman-Kac corrector that progressively steers the multi-particle system toward a KL-tilted target distribution, facilitating global exploration and preventing mode collapse. Extensive empirical evaluations and ablations across diverse search and alignment tasks confirm the efficacy of IMPFM over existing baselines.
|
| 2082 |
BFMT: Enhancing Search Capabilities of Tree Sampler via Bootstrap Flow-Map Tree
2607.02915
|
cs.LG
|
Binglin Ji, Anindya Sarkar, Hengchang Lu, Kunyu Wang, Alessandro Tibo |
The ability to efficiently navigate high-dimensional spaces to identify optimal candidates-whether designing targeted molecular structures or generating images with preferred attributes-stands as a central pillar of modern scientific progress. Recent tree-base...The ability to efficiently navigate high-dimensional spaces to identify optimal candidates-whether designing targeted molecular structures or generating images with preferred attributes-stands as a central pillar of modern scientific progress. Recent tree-based samplers enable highly scalable inference-time search by eliminating the reward-gradient bottleneck inherent to particle-based methods. Despite their potential, existing tree-based samplers are fundamentally constrained by the massive number of function evaluations (NFEs) required for node valuation, crippling their utility in exploration-heavy tasks. Furthermore, their reliance on small, uniform transitions at each depth precludes dynamic, adaptive search capabilities. To address this, we introduce Bootstrap Flow-Map-Tree (a.k.a BFMT), a computationally efficient sampling framework that enables full tree-path construction from any tree depth using a single function evaluation, drastically reducing computational overhead while providing critical foresight for sequential sampling. By enabling dynamic transition time-step scheduling, BFMT efficiently allocates its sampling budget, smoothly transitioning from broad global exploration to fine-grained local refinement of high-utility modes discovered through exploration. Extensive experiments and ablation studies across diverse domains-from high-dimensional images to large-scale molecules-demonstrate BFMT's superiority over baseline approaches in both search and alignment tasks.
|
| 2083 |
What Does a Routing Oracle Measure Under Stochastic Decoding? Coupling, Scorer Choice, and Single-Commit Ceilings
2607.03436
|
cs.LG
|
Teng-Ruei Chen |
Routing benchmarks often compare a policy that commits to one model before seeing its response with a hindsight oracle that credits any correct recorded output. Under stochastic decoding these are different decision classes. For success marginals $p_{im}$ unde...Routing benchmarks often compare a policy that commits to one model before seeing its response with a hindsight oracle that credits any correct recorded output. Under stochastic decoding these are different decision classes. For success marginals $p_{im}$ under a frozen query--model--decoder--scorer protocol, we distinguish the clairvoyant single-commit ceiling $R_i=\max_m p_{im}$, product-coupling union $U_i^\perp=1-\prod_m(1-p_{im})$, and premium $\Delta_i^\perp=U_i^\perp-R_i$. The marginals alone identify exactly the union interval $[R_i,\min\{1,\sum_m p_{im}\}]$: zero premium is attainable over compatible couplings, not established for an actual deployment. If generation is independent across models, product is the specified protocol's union probability. We audit frozen correctness tensors from 11 open models and 30 archived responses per query--model cell on GSM8K, MATH-500, and GPQA-Diamond. Full-pool display-channel product premiums are 0.371, 3.542, and 5.100 percentage points; the premium intervals from empirical marginals are $[0,0.473]$, $[0,4.787]$, and $[0,7.744]$. Their zero lower endpoints are algebraic. All retained scorers, eight finite-draw paths, and all 2,047 nonempty subpools expose scorer, estimator, and pool sensitivity, not confidence intervals. The GSM8K and GPQA display scorers were developed after limited output inspection. A separate retrospective held-out policy illustration instantiates the policy-specific gap decomposition without establishing new-data generalization. A limited reference-based human check supports scorer agreement only on definite-consensus subsets. Hash-bound evidence supports number checks, not end-to-end reproduction. The contribution is a measurement contract for interpreting oracle gaps conditional on coupling, decision class, scorer, pool, and finite draws, not a population effect or an equal-cost routing gain.
|
| 2084 |
Transformers with Physics-Informed Encodings and Simulation-Based Inference for Robust Detection of Eccentric Binary Black Holes in Pulsar Timing Array Data
2607.03904
|
cs.LG
|
Subhajit Dandapat, Alvin J. K. Chua |
Pulsar timing arrays (PTAs) provide a unique window into nanohertz gravitational waves (GWs), but extracting astrophysical parameters from noisy, long-baseline timing residuals remains computationally challenging with traditional Bayesian techniques due to the...Pulsar timing arrays (PTAs) provide a unique window into nanohertz gravitational waves (GWs), but extracting astrophysical parameters from noisy, long-baseline timing residuals remains computationally challenging with traditional Bayesian techniques due to the high dimensionality of the parameter space, complex and correlated noise models, and the cost of repeated likelihood evaluations. We introduce a Transformer with a physics-informed positional-encoding framework for the efficient inference of eccentric binary black holes in relativistic orbits from PTA data. Our approach embeds analytical GW phase evolution directly into the model through structured positional encodings, enabling the network to learn physically meaningful representations from raw PTA timing residuals. We then use generative models, including discrete and continuous conditional normalizing flows, to infer posterior distributions within a simulation-based inference framework. Across a range of signal-to-noise ratios, the proposed method achieves improved accuracy, sharper posteriors, and faster inference compared to physics-agnostic baselines. While presented for deterministic white-noise signals, the modular framework readily generalizes to realistic PTA analyses incorporating red noise and additional components. This work highlights the potential of physics-aware deep learning models as scalable alternatives to conventional inference pipelines for next-generation PTA datasets.
|
| 2085 |
RL Forgets! Towards Continual Policy Optimization
2607.04364
|
cs.LG
|
Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou, Bo Ye, Jian Zhao |
Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks. Recent studies report that reinforcement learning is less prone to forgetting than supervised fine-tuning, motivating the view that RL is inherently r...Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks. Recent studies report that reinforcement learning is less prone to forgetting than supervised fine-tuning, motivating the view that RL is inherently resistant to forgetting. However, this view remains insufficiently validated, as existing evidence is largely drawn from outdated or homogeneous benchmarks. We revisit this assumption by introducing MRCL, a Multimodal Reasoning Continual Learning benchmark built from recent and diverse multimodal reasoning tasks. Experiments on MRCL show that standard reinforcement learning still suffers from catastrophic forgetting during continual post-training. We trace the failure to an objective mismatch. The KL regularization used in common policy optimization methods is evaluated on current-task data, whereas forgetting is caused by behavioral drift on prior-task distributions. To address this problem, we propose Continual Policy Optimization (CPO), a replay-free method grounded in a prior-task behavioral KL objective. CPO derives a local Fisher surrogate from the historical KL objective and uses parameter movement as a gradient-free proxy for Fisher sensitivity, enabling sparse regularization with negligible additional overhead. Experiments on three model scales and comparisons with multiple RL baselines show that CPO consistently reduces forgetting while maintaining effective adaptation and preserving broader pretrained capabilities. The implementation code is available at https://github.com/MaolinLuo/CPO.
|
| 2086 |
Higher-Order Geometric Updates for Levenberg-Marquardt Method via Riemann Normal Coordinates
2607.07623
|
cs.LG
|
Jianing Liu, Dong H. Zhang |
Nonlinear least-squares objectives form the foundation of scientific machine-learning tasks, yet even curvature-aware optimizers remain geometrically inconsistent at finite step sizes: Levenberg-Marquardt (LM) derives its direction from local Riemannian metric...Nonlinear least-squares objectives form the foundation of scientific machine-learning tasks, yet even curvature-aware optimizers remain geometrically inconsistent at finite step sizes: Levenberg-Marquardt (LM) derives its direction from local Riemannian metrics but realizes it as a straight parameter update. Here we introduce RNC-LM, which carries the LM direction along a locally constructed curved trajectory in Riemann normal coordinates. A recursive reformulation of the geodesic equation generates arbitrary finite-order corrections while reusing the same damped Gauss-Newton matrix factorization, and curve length is controlled separately from damping. On a reaction-diffusion physics-informed neural network benchmark, RNC-LM reduces relative $L^2$ errors below $8\times10^{-3}$, whereas L-BFGS, LM and LM with geodesic acceleration remain near one. On a large-scale machine-learning potential fitting task with 985,160 configurations, fourth-order RNC-LM reaches a fixed training-error target with a $34\times$ wall-clock speedup over LM. These results establish finite-step geometric realization as a distinct optimizer design principle.
|
| 2087 |
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
2607.07740
|
cs.LGcs.AI
|
Haozhan Tang, Zerui Wang, Yuxian Gu, Song Han, Han Cai |
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretra...Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. The dominant zero-shot methods (YaRN, Self-Extend, DCA) fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity while a conservative one breaks down at long contexts; recent length-aware variants adapt the mapping, but with a fitted or distance-dependent schedule. We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically to the current sequence length via a parameter-free analytic schedule, recovering the base model exactly at short inputs while extrapolating cleanly at long ones. An inclusion-exclusion attention merge and on-the-fly RoPE correction enable a fused CuTe implementation. On H100 at 64K-128K, prefill retains 83-88% of FlashAttention-3 throughput across the evaluated Qwen3 sizes and 88-93% of a matched CuTe control; Qwen3-8B single-batch generation reaches 1.04-1.08 times FlashAttention-3 throughput. On Qwen3-1.7B/4B/8B up to 128K context, Jet-Long leads RULER by +4.79/+2.18/+2.03 percentage points over the strongest baseline at 1.7B/4B/8B, achieves the best overall accuracy on HELMET-RAG (a benchmark identified by HELMET as the most efficient predictor of downstream long-context performance) and attains the lowest PG-19 perplexity. Additional evaluations cover Meta-Llama-3-8B, post-trained Qwen3 checkpoints, and the hybrid Jet-Nemotron architecture, supporting broader applicability without retraining. The local-window hyperparameter remains robust across the tested settings.
|
| 2088 |
Hyperbolic Manifold Constrained Tabular Neural Network
2607.09710
|
cs.LG
|
Tian Li, Huizhi Liang, Lucy Robinson, Varun Ojha |
Tabular prediction is central to a wide range of real-world applications. Tabular data typically contain heterogeneous features as well as rich and complex relational information that can imply a latent structural manifold. Hyperbolic geometry can help capture...Tabular prediction is central to a wide range of real-world applications. Tabular data typically contain heterogeneous features as well as rich and complex relational information that can imply a latent structural manifold. Hyperbolic geometry can help capture complex structural relations in data. However, most existing tabular prediction models are constructed and optimized in Euclidean space. How to incorporate hyperbolic geometry into supervised tabular learning remains underexplored. We propose \textbf{HTNN}, a supervised hyperbolic manifold constrained tabular neural network for tabular prediction. HTNN consists of a hyperbolic feature-value representation layer for heterogeneous categorical and numerical features, followed by a conventional MLP predictor. HTNN employs a \emph{geometry-aware training} and \emph{geometry-free inference} optimization framework. The \emph{geometry-aware training} allows hyperbolic geometry to shape the latent representation learning of heterogeneous feature values. After training, the latent hyperbolic representations can be converted into ordinary Euclidean space for efficient \emph{geometry-free inference}. We conducted extensive experiments on the TALENT benchmark. HTNN ranks first among 36 methods on 200 classification datasets and third among 34 methods on 100 regression datasets. Experimental results show that the proposed hyperbolic manifold constrained tabular neural network is effective.
|
| 2089 |
A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models
2607.14522
|
cs.LG
|
Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang |
We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC). We consider...We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC). We consider policy optimization problems and derive corresponding policy gradient methods, leading to continuous-time variants of proximal policy optimization (PPO) and group relative policy optimization (GRPO). As a primary application, we develop a continuous-time RL framework for fine-tuning score-based discrete diffusion models, which enables reward-driven optimization without requiring differentiability on the reward signals. In contrast to the existing GRPO-based approaches that only rely on terminal rewards, our formulation allows intermediate reward or advantage signals to be incorporated throughout the denoising trajectory. Importantly, when specialized to masked diffusion models (MDMs), our framework encompasses a rich class of policy parameterizations over the vocabulary simplex with analytically tractable probability ratios, providing a unified perspective on exploration and policy optimization in MDMs. For diffusion large language models (dLLMs), we further propose trajectory subsampling techniques to efficiently estimate computationally prohibitive trajectory likelihoods, reducing the computational cost of computing per-position probability ratios. We showcase the effectiveness of our methods on both low-dimensional entropy-regularized optimization problems and RL post-training of dLLMs on reasoning and coding tasks.
|
| 2090 |
What Does Model Growth Really Add? Functional Capacity Beyond Parameter Count
2607.14571
|
cs.LG
|
Dante Lok |
Model growth adds parameters to a trained checkpoint and continues training, aiming to increase capacity while reusing the existing model instead of training from scratch. But does growing a trained model actually create new functional capacity that survives s...Model growth adds parameters to a trained checkpoint and continues training, aiming to increase capacity while reusing the existing model instead of training from scratch. But does growing a trained model actually create new functional capacity that survives subsequent training? Parameter count alone cannot answer this question, and ``capacity'' is itself ambiguous; it can refer to parameter count, effective dimensionality, or the ability to represent new functions. We measure it directly, as the number of new, functionally independent directions a growth step adds, formalized as the rank increment of the model's functional Jacobian and estimated matrix-free at scale. For zero-gated function-preserving growth, this increment is exactly $K$ under a transversality condition, i.e. $\operatorname{rank}(J_{\mathrm{grown}})=\operatorname{rank}(J_{\mathrm{old}})+K$. Continually training a Transformer grown from $253\mathrm{M}$ to $857\mathrm{M}$ parameters across three domains, we find these added directions persist, with all $K=36$ remaining independent after continual learning, even when the model catastrophically forgets (perplexity $28\to481$) or we deliberately destroy the old function ($25.9\to127$). Geometric capacity and retained function are therefore decoupled. The functional capacity introduced by growth persists, while the previously learned function is not retained. A second, geometrically distinct mechanism, Net2Net, reaches the same $K$-dimensional capacity from directions that are dormant at initialization ($0 \to K$) and then retains it, so the decoupling is not specific to the zero-gate parameterization. The implication is precise. Catastrophic forgetting need not reflect a loss of functional capacity, and preserving dimensionality alone is insufficient to preserve the learned function.
|
| 2091 |
Robust Peak-cost Constrained Reinforcement Learning
2607.15457
|
cs.LG
|
Shilpa Mukhopadhyay, Sourav Ganguly, Santosh Mohan Rajkumar, Honghao Wei, Debdipta Goswami |
We study robust peak-cost constrained reinforcement learning (\ours), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in which a sin...We study robust peak-cost constrained reinforcement learning (\ours), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in which a single large violation can be catastrophic and therefore cannot be adequately captured by the standard CMDP framework based on expected cumulative cost. Existing reachability-constrained RL methods adopt Lagrangian-based approaches, yet the underlying duality properties of peak-cost constrained MDPs remain unclear. We show that, unlike standard CMDPs, peak-cost constrained MDPs may not admit zero duality gap. We further consider a robust formulation to address simulator-to-real-world mismatch in the transition dynamics. To solve this problem, we develop a surrogate optimization framework and a robust value estimation method based on integral probability metrics. We prove that, with appropriate hyperparameter choices, the surrogate solution attains the same robust reward value as the original problem while violating the constraint by at most \(\epsilon\). Experiments show that the proposed method effectively enforces safety under dynamics perturbations while retaining strong reward performance across diverse environments.
|
| 2092 |
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
2607.16242
|
cs.LG
|
Changyue Li, Jiaming He, Youliang Yuan, Jialin Wu, Boxi Yu |
Fine-Tuning-as-a-Service (FTaaS) platforms let users perform supervised fine-tuning (SFT) on customized data, but this pipeline can erode model safety alignment. To recover safety without re-running full alignment, existing realignment methods focus on calibra...Fine-Tuning-as-a-Service (FTaaS) platforms let users perform supervised fine-tuning (SFT) on customized data, but this pipeline can erode model safety alignment. To recover safety without re-running full alignment, existing realignment methods focus on calibrating the integration of safety patches into fine-tuned models. These methods exhibit a persistent safety-utility trade-off: weak repair leaves harmful behavior intact, while stronger repair increasingly damages the benign task. This paper shifts the focus from online calibration to offline patch learning and aims to learn a safety patch that restores safety while preserving task-specific capabilities. To this end, we propose TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states. TRACE trains the safety patch during the offline stage, and reuses it across all user checkpoints without per-user calibration. We evaluate two representative models using three harmful SFT datasets, together with three utility benchmarks. Across six benchmarks and two models, TRACE consistently dominates the safety-utility frontier. TRACE improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended fine-tuned model.
|
| 2093 |
Lightweight Wrappers for Adapting Time Series Foundation Models to Regional Drought Forecasting
2607.17511
|
cs.LG
|
Wentao Gao, Jiuyong Li, Lin Liu, Thuc Duy Le, Jixue Liu |
Large \emph{Time Series Foundation Models} (TSFMs) demonstrate strong zero-shot forecasting capabilities across diverse domains. However, their application to regional climate forecasting faces practical challenges: model weights are often proprietary, local t...Large \emph{Time Series Foundation Models} (TSFMs) demonstrate strong zero-shot forecasting capabilities across diverse domains. However, their application to regional climate forecasting faces practical challenges: model weights are often proprietary, local training records are limited, and computational budgets are constrained, making traditional fine-tuning approaches infeasible. To address these constraints, we introduce a lightweight, black-box adaptation framework (requiring no access to backbone parameters and no backbone fine-tuning) that enhances frozen TSFMs at inference time through two plug-and-play wrappers: \textbf{SMR\textsuperscript{2}} (Stationarity aware multi-resolution Residual), which decomposes the input into multi-resolution temporal views, learns stride specific residual corrections that capture regional dynamics, then adaptively ensembles them into a single forecast, and \textbf{MBB} (Moving Block Bootstrap), which preserves temporal dependencies through block resampling and ensembles over temporally coherent residual perturbations to stabilize the point forecast. Both wrappers instantiate the same bagging style principle: they build diverse views of the input or its residuals, forecast each with the same frozen backbone, and aggregate, so all adaptation comes from inference time ensembling rather than any weight update. Evaluated on one month ahead Standardized Precipitation Evapotranspiration Index (SPEI) prediction across multiple sites in South Australia, our framework consistently improves forecasting performance across several backbone models, demonstrating up to 26\% mean squared error (MSE) reduction over the corresponding frozen backbone while enabling practical deployment in resource constrained regional forecasting systems.
|
| 2094 |
CoCurve: Cross-Module Co-Pruning Curvature for Structured LLM Pruning
2607.17568
|
cs.LG
|
Zhiren Gong, Zihao Zeng, Tiantong Wang, Yixin Wang, Honoka Anada |
Resource-constrained deployment requires sustaining large language model (LLM) capabilities as models scale under fixed memory and computation budgets. Structured pruning advances this deployment frontier with smaller dense checkpoints, yet deciding what to pr...Resource-constrained deployment requires sustaining large language model (LLM) capabilities as models scale under fixed memory and computation budgets. Structured pruning advances this deployment frontier with smaller dense checkpoints, yet deciding what to prune remains bottlenecked by interactions among joint removals. We introduce Cross-Module Co-Pruning Curvature (CoCurve), which formulates structured pruning as set-dependent predictive risk over a unified inventory of attention heads and feed-forward groups. A co-pruning graph built from single-unit forward ablations assigns individual risk to nodes and reinforcement or cancellation to interaction edges, conditioning each decision on the units already removed. We evaluate 6 LLMs (3B--70B) and 3 vision--language models (VLMs) across 3 perplexity corpora, 12 language tasks, and 7 multimodal benchmarks. Across the five-model 20--40% grid and the 70B 10--50% sweep, CoCurve ranks first in 53/60 corpus comparisons (15/15 at 70B); matched-quality interpolation permits 2.2--6.6 more pruning points in 9/10 cases, while CoCurve leads all 6 VLM Avg$_7$ blocks. After the same lightweight recovery, its retained structures remain strongest through 50% pruning, where the 8B checkpoint regains 10.8 Avg$_{12}$ points; physical slicing delivers $1.58\times$ dense prefill throughput with 41% lower peak memory. Mechanism analysis across 10 LLMs and 7 VLMs finds organized within- and cross-module edge structure; matched low-saliency, high-coupling removals degrade 19/20 capability groups by up to 23.7 points.
|
| 2095 |
CriPO: Enhancing Rubric-based RL via Self-Distillation
2607.18082
|
cs.LG
|
Mingxuan Xia, Yuhang Yang, Chao Ye, Shuai Zhu, Shenzhi Yang |
Rubric-based Reinforcement Learning (RL) has recently shown promise in improving Large Language Models (LLMs) on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored...Rubric-based Reinforcement Learning (RL) has recently shown promise in improving Large Language Models (LLMs) on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria) receive no optimization signal. Recent methods address this by incorporating rubric information as external guidance during rollout generation, yet they introduce a train-inference mismatch: the policy is optimized on rollouts produced under external guidance while this guidance is absent at inference time, causing error accumulation through autoregressive decoding. Moreover, these exploration-focused approaches overlook a fundamentally different failure mode that we term Suppressed Criteria---criteria that are satisfied by some rollouts yet whose learning signals might be lost during optimization because scalar reward aggregation assigns them non-positive aggregate advantages. Our analysis reveals that suppressed criteria constitute a persistent and non-negligible failure mode---over 25% of samples exhibit this issue throughout training. To simultaneously address both unexplored and suppressed criteria without introducing training-inference mismatch, we propose Criterion-Distilled Policy Optimization (CriPO), which enhances rubric-based RL via on-policy self-distillation. For unexplored criteria, CriPO constructs a behavior-injection teacher and computes a filtered forward-KL loss to inject missing behaviors into the policy. For suppressed criteria, CriPO uses a counterfactual teacher to locate criterion-relevant tokens in negative-advantage rollouts, and corrects their advantages in GRPO to preserve useful patterns. Experiments on medicine and science benchmarks demonstrate that CriPO outperforms existing rubric-based RL methods, e.g., achieving an average gain of 3.3 points over GRPO on Qwen3-4B.
|
| 2096 |
Conditioned Direct Feedback Alignment via Activity and Error Geometry
2607.18574
|
cs.LG
|
Houman Safaai, Varun Reddy, Bernardo L. Sabatini |
Direct feedback alignment (DFA) trains hidden layers with fixed random projections of the output error. Even when this feedback provides useful credit, unequal scales across activity or error directions can distort the local update. We study conditioned DFA (n...Direct feedback alignment (DFA) trains hidden layers with fixed random projections of the output error. Even when this feedback provides useful credit, unequal scales across activity or error directions can distort the local update. We study conditioned DFA (nDFA), a family that adapts established inverse-moment preconditioning to either side of this update. An aligned linear analysis describes how activity conditioning changes spectral learning rates and early-stopping risk. Synthetic experiments show gains from both stabilization and changes in update direction. Confirmatory CIFAR-10 experiments show that activity conditioning, including its FOOF formulation, improves tuned DFA at matched measured training time. Controls support a role for centered correlations beyond the tested mean and diagonal alternatives, while early-only conditioning retains most of the benefit with less training work. Conditioning also benefits backpropagation, and timing interventions do not establish a mechanism specific to random feedback. Error conditioning improves short-budget models and changes update directions, but its small additional improvement under stable training does not survive correction for multiple comparisons. Predicted benefits on image-background benchmarks are not confirmed. These findings establish practical benefits and important limits of conditioning learning with fixed random feedback.
|
| 2097 |
Interval and fuzzy physics-augmented neural networks (iPANN and fPANN) for uncertainty quantification and propagation in constitutive modeling
2607.20339
|
cs.LG
|
Somesh Pratap Singh, Govinda Anantha Padmanabha, Jingye Tan, Steven Yang, Reese E. Jones |
Constitutive modeling under uncertainty remains a central challenge for reliable mechanics simulations, particularly when the available stress-deformation data are sparse, noisy, or heterogeneous. We propose interval and fuzzy physics-augmented neural networks...Constitutive modeling under uncertainty remains a central challenge for reliable mechanics simulations, particularly when the available stress-deformation data are sparse, noisy, or heterogeneous. We propose interval and fuzzy physics-augmented neural networks (iPANNs and fPANNs) for uncertainty-aware hyperelastic constitutive modeling. iPANNs learn sparse lower, mean, and upper free energy density branches whose stresses, obtained by automatic differentiation, ultimately enclose noisy stress observations. In contrast to this deterministic interval description, fPANNs embed the learned iPANN branches into a fuzzy-set representation through membership-level (alpha-cut) interpolation, yielding a nested family of admissible responses. iPANNs and fPANNs encode mechanistic constraints by preserving objectivity, consistency and promoting polyconvexity and, smoothed L0 regularization promoting interpretable energy representations. The bound models are trained through a two-stage transfer-learning procedure in which a sparse mean constitutive response is learned first and then fine-tuned into lower and upper energy branches. We evaluate the framework on synthetic isotropic hyperelastic data with heteroscedastic noise, varying random realizations, shifted noise means, and varying noise magnitudes. The results show that the learned bounds enclose noisy stress observations while generalizing to the test set. Further, we examine the propagation of uncertainty through the mean, upper and lower bound predictions of the learned iPANN models in a finite element setting. The proposed framework provides a compact, physics-consistent route for distribution-free aleatoric uncertainty quantification in hyperelastic constitutive modeling, and propagation in downstream finite element simulations.
|
| 2098 |
On the Convergence of Stochastic Low-Rank Adaptation
2607.21975
|
cs.LG
|
Ru Wang, Chengchang Liu, John C. S. Lui |
Low-rank adaptation (LoRA) optimizes $J(B,A)=\mathcal L(W_\mathrm{base}+sBA)$ over two adapters $B \in \mathbb{R}^{m \times r}$ and $A \in \mathbb{R}^{r \times n}$ that form a low-rank update to a frozen pretrained weight matrix $W_\mathrm{base} \in \mathbb{R}...Low-rank adaptation (LoRA) optimizes $J(B,A)=\mathcal L(W_\mathrm{base}+sBA)$ over two adapters $B \in \mathbb{R}^{m \times r}$ and $A \in \mathbb{R}^{r \times n}$ that form a low-rank update to a frozen pretrained weight matrix $W_\mathrm{base} \in \mathbb{R}^{m \times n}$. The prior analysis shows LoRA-GD takes $\exp\{\mathcal{O}(\epsilon^{-2})\}$ oracle calls to find an $\epsilon$-stationary point such that $\|\nabla J(B,A)\|\leq \epsilon$ in the deterministic setting. We sharpen the analysis and show that $\mathcal{O}(\epsilon^{-4})$ full-gradient evaluations suffice for the same first-order criterion. We further study stochastic LoRA under unbiased gradient estimates and finite variance. We propose LoRA-NSGDM, which finds an $\epsilon$-stationary point with $\mathcal{O}(\epsilon^{-8})$ stochastic oracle complexity. Under the additional mean-square smoothness condition, we use variance reduction strategy and propose LoRA-STORM, which improves the stochastic oracle complexity to $\mathcal{O}(\epsilon^{-6})$.
|
| 2099 |
Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram
2607.22980
|
cs.LG
|
Ethan Davis |
Brain-computer interfaces (BCIs) have long sought calibration-free operation, yet classifiers are typically benchmarked on discrimination alone. Discrimination is blind to calibration, a meaningful gap given that electroencephalogram (EEG) signals are nonstati...Brain-computer interfaces (BCIs) have long sought calibration-free operation, yet classifiers are typically benchmarked on discrimination alone. Discrimination is blind to calibration, a meaningful gap given that electroencephalogram (EEG) signals are nonstationary and point-estimate classifiers can become overconfident under distribution shift. We conducted a large-scale study contrasting Bayesian complete-pooling models against frequentist baselines for cross-subject, left-hand versus right-hand motor imagery EEG classification across 20 datasets. Each of six frequentist pipelines was paired with an analogous Bayesian pipeline sharing identical feature engineering, fit via Markov chain Monte Carlo posterior sampling. Our primary metric was the Brier score, decomposed into reliability and resolution, alongside the area under the receiver operating characteristic curve for discrimination and Shannon entropy for sharpness. For each metric we fit a random-effects meta-analysis with Knapp-Hartung adjustment, verified by leave-one-out influence analysis. Bayesian complete-pooling produced statistically significant improvements in reliability and increases in predictive uncertainty (lower sharpness), but 95% confidence intervals bounded both effects near zero, and neither held significance when datasets with flagged posterior sampling were excluded. Brier score, resolution, and discrimination showed no significant differences. Between-study heterogeneity was low across all metrics. Bayesian pipelines consumed roughly thirteen times more energy than their frequentist counterparts, a cost that remains modest relative to common household appliances. Bayesian complete-pooling alone offers limited benefit for cross-subject motor imagery classification, but its modest cost makes partial-pooling across subjects and sessions a feasible next step.
|
| 2100 |
Variational Boosting for Physics-Informed Neural Networks
2607.23940
|
cs.LG
|
Kaylee Vo, Pavlos Protopapas |
Physics-Informed Neural Networks (PINNs) solve differential equations by minimizing the residual of a nonlinear operator over a neural parameterization of the solution. However, monolithic PINNs often suffer from ill-conditioning, spectral bias, and optimizati...Physics-Informed Neural Networks (PINNs) solve differential equations by minimizing the residual of a nonlinear operator over a neural parameterization of the solution. However, monolithic PINNs often suffer from ill-conditioning, spectral bias, and optimization instability. We introduce a variational boosting framework in which solutions are constructed additively in function space. Each stage trains a weak learner whose converged correction satisfies a local orthogonality condition, equivalent to a projected functional gradient descent step onto the tangent space of the network's function manifold. Because each correction network is deliberately small, the restricted minimization admits full Newton or conjugate gradient updates, which are typically infeasible in large PINNs. The resulting method separates global nonlinear refinement into a sequence of well-conditioned subproblems while preserving the full variational structure of the operator. This framework provides a geometric interpretation of multi-stage PINNs as projected functional gradient descent and enables stable second-order optimization for nonlinear differential equations.
|
| 2101 |
One Analyst Is Not Ground Truth: Grading Agent-Built Financial Models Against Observed Professional Practice
2607.24889
|
cs.LG
|
Jiacheng Lu, Yiming Li, Sinuo Wang, Wentao Zhao, Rui Sun |
Language model agents can now construct complete financial models, but it remains unclear whether they have learned the financial judgment that gives those models meaning. Existing benchmarks often anchor correctness to an expert-authored solution (e.g., numer...Language model agents can now construct complete financial models, but it remains unclear whether they have learned the financial judgment that gives those models meaning. Existing benchmarks often anchor correctness to an expert-authored solution (e.g., numerical targets or detailed rubrics) which introduces a implicit assumption: \emph{one expert solution can serve as ground truth.} This is appropriate when finance provides a unique answer, but not when judgment is required. Evidence from professional practice challenges this assumption: When financial models built by different analysts for the same company are graded against one another, the median score is only 0.33 under standard tolerances, \emph{revealing that single-reference grading confounds professional disagreement with error.} We therefore introduce GAUGE, a benchmark that decomposes financial-model evaluation into deterministic checks, rubric judgments and numerical rules. Built from 1,001 professional valuation spanning 922 companies and all 25 GICS industry groups, GAUGE checks mechanical properties deterministically and evaluates judgment-bearing quantities against ranges in professional practice. We then validate GAUGE as a measurement instrument by testing expertise ordering, held-out professional values, and robustness to LM judgements. Our findings show that current LM agents are far better at constructing financial models than at deriving the company-specific assumptions that drive valuation, a gap that even persists after fine-tuning.
|
| 2102 |
Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
2607.26358
|
cs.LG
|
Keegan Harris, Brian W. Lee, Ian Waudby-Smith, Philip Amortila, Nika Haghtalab |
Reinforcement learning (RL) fine-tuning is widely used in language model training to improve performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this...Reinforcement learning (RL) fine-tuning is widely used in language model training to improve performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. In practice, the coefficient is typically chosen heuristically or via hyperparameter search, which can lead to unnecessary overhead in training cost or undesirable reward-retention trade-offs. We instead propose a game-theoretic framework that gives this trade-off an explicit statistical interpretation. Specifically, we study a sequential game in which an agent chooses a policy to maximize cumulative reward while a monitor observes policy outputs over time and tests for deviations from the reference policy. Although not originating from the same perspective, we show that the resulting equilibrium policy can nonetheless be expressed as the solution to a KL-regularized RL problem for an optimal regularization parameter that can be viewed as maximizing reward per unit of statistical distinguishability. Drawing on classical results from concave-convex fractional programming, we provide a principled method for learning this equilibrium coefficient via reduction to the KL-regularized RL objective, thus allowing for flexible integration into standard fine-tuning pipelines. In experiments with Qwen3-8B and Llama-3.2-1B, we show that our methods result in competitive reward-retention trade-offs in a continual learning setting, and illustrate how our framework may be used to audit API providers serving open-source models.
|
| 2103 |
Existence-Field Diffusion Model for Spatial Point Processes with Variable Cardinality
2607.26428
|
cs.LG
|
Xiaoyin Pan, Christian R. Shelton, Rakshith Mahishi, Chengkuan Hong |
We study generative modeling of spatial point processes (SPP), where both the number of points and their spatial configuration are governed by a joint distribution. While diffusion models have achieved strong performance in modeling complex distributions, exte...We study generative modeling of spatial point processes (SPP), where both the number of points and their spatial configuration are governed by a joint distribution. While diffusion models have achieved strong performance in modeling complex distributions, extending them to variable-cardinality SPP remains challenging. Existing approaches either sample cardinality before generating locations conditionally, or introduce specialized discrete operations to change the number of points during generation. We propose the existence-field diffusion model (EFDM), which associates each potential point with a continuous variable representing its degree of presence. EFDM uses Gaussian diffusion to jointly model point locations and cardinality through a simple continuous representation. Existence probabilities can increase or decrease in both forward and reverse diffusion, without explicit point-addition or point-deletion steps. The same construction naturally accommodates categorical attributes, allowing locations, existence, and attributes to be modeled together. Experiments demonstrate improved molecular stability and validity over the compared baselines on the molecular dataset, alongside competitive performance on real-world spatial and synthetic datasets.
|
| 2104 |
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
2607.26828
|
cs.LG
|
Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen, Xiaokun Zhang |
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and...Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to incur different token costs. Under a fixed search budget, the controller must decide which frontier to develop and whether its gain justifies the realized cost before the budget is exhausted. We prove that, in the worst case, cost-blind control can forfeit all but a vanishing fraction of attainable quality as frontier count and cost heterogeneity grow. To address this limitation, we introduce \textbf{CostAda}, an adaptive controller built on \emph{cost-calibrated frontier utility}. By valuing progress relative to realized action cost and conditioning credit on the remaining budget, CostAda coordinates local exploration intensity, frontier allocation, and budgeted tactic intervention. Realized action cost and remaining budget therefore shape the search rather than serving only as accounting variables or a stopping rule. Across eight benchmarks with GLM-5 and GPT-5.4, CostAda reaches the strongest baseline's full-budget quality with at most half the budget on thirteen of sixteen benchmark--backbone pairs and achieves the strongest final quality on all sixteen pairs.
|
| 2105 |
What Can Latent World Models Know? Physical Information in Multimodal Predictive Representations
2607.27017
|
cs.LG
|
Kaizhen Tan, Sizhe Xu, Xin Xu, Siru Tao, Yixiao Li |
A central premise of latent world models is that predicting the future encourages representations to internalize the physics of their environment. We ask which physical quantities are accessible in learned latent states, how this depends on training, and how t...A central premise of latent world models is that predicting the future encourages representations to internalize the physics of their environment. We ask which physical quantities are accessible in learned latent states, how this depends on training, and how those quantities relate to the model's predictions. We present PokeWorld, a simulated environment in which a robot finger pushes objects whose mass, drag, and contact stiffness vary across episodes while remaining visually identical. We first measure parameter recovery from raw observation sequences, then train matched action-conditioned world models. Prediction targets can strongly shape latent content. For example, contact stiffness becomes decodable when touch is predicted, while providing touch only as an input does not. Longer-horizon prediction improves estimation of position and velocity from learned representations relative to single-step prediction. Drag is recoverable from raw observations, but has weak linear readout from the learned latent states. Yet the models' glide forecasts depend systematically on drag. Experiments on RH20T reproduce the same input-target dependence on real multimodal robot data. Together, these results show that observations, prediction targets, and prediction horizons shape both which physical quantities are accessible in latent states and how they affect future predictions.
|
| 2106 |
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
2607.28959
|
cs.LG
|
Weiyi He, Yuping Lin, Jiliang Tang, Yue Xing |
Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). While existing mitigation strategies, e.g., latent adversarial...Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). While existing mitigation strategies, e.g., latent adversarial training (LAT), have been developed, they still incur a high computational cost. In this work, focusing on LLM-based classifiers, we investigate computation-efficient strategies to speed up LAT from two complementary perspectives: (1) Defense-side optimization: We explore representation fine-tuning (ReFT) within LAT, and reveal a potential issue if there is a mismatch on which tokens to apply ReFT and the attack. (2) Attack-side optimization: When computing adversarial attacks in each LAT iteration, we extract only the relevant circuits from the LLM to construct a lightweight surrogate model, avoiding the computation in the forward-backward passes through the full model during the attack generation. For both perspectives, we provide theoretical justifications and numerical evidence to illustrate the effectiveness of the proposed strategies. Compared to standard LAT with full fine-tuning, our method on average reduces per-step adversarial-training FLOPs by 29.6% while requiring only 0.0118% trainable parameters, at a moderate cost in robustness.
|
| 2107 |
DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation
2607.29078
|
cs.LG
|
Yuchen Xia, Qianguo Sun, Chao Song, Junlong Wu, Yiyan Qi |
While on-policy distillation (OPD) reduces exposure bias by training student language models on their own rollouts, early student errors in long-horizon agentic scenarios can lead to contexts unfamiliar to the teacher. To improve trajectory quality, recent wor...While on-policy distillation (OPD) reduces exposure bias by training student language models on their own rollouts, early student errors in long-horizon agentic scenarios can lead to contexts unfamiliar to the teacher. To improve trajectory quality, recent work on agentic OPD introduces teacher intervention into training rollouts by switching the executor between the student and the teacher. However, existing methods determine how much teacher intervention is needed---but not when. To address this limitation, we propose DASH-OPD (Discrepancy-Aware Switching with Hysteresis for OPD), the first agentic OPD method to perform adaptive, bidirectional executor switching. At each turn, DASH-OPD measures teacher--student discrepancy using a mean log-likelihood ratio over action tokens. High student-to-teacher ratios on student turns serve as drift signals, while low teacher-to-student ratios on teacher turns serve as recovery signals. These signals are accumulated over multiple turns to form drift and recovery evidence, respectively. DASH-OPD switches executors when either type of evidence exceeds its corresponding switching threshold, introducing hysteresis that prevents rapid switching triggered by transient discrepancy fluctuations. Across three benchmarks and two student model scales, DASH-OPD outperforms five baselines in all 14 task-performance comparisons, while requiring the fewest interaction turns in nine of ten efficiency comparisons. Code, trained models, and training logs are available at https://github.com/Lucian1115/DASH-OPD
|
| 2108 |
Introspecting Alignment Shifts Beyond Behaviors Implanted Through Fine-Tuning
2608.04347
|
cs.LG
|
Kotaro Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto, Yuji Naraki, Ryotaro Shimizu |
Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source m...Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source model. Recent work has shown that large language models can be trained using LoRA-based modules known as introspection adapters (IAs) to describe behavioral changes induced by fine-tuning. However, existing studies primarily consider settings in which the model is fine-tuned on datasets explicitly designed to implant a specific behavior and is then asked to explain the implanted behavior. This differs from practical deployment scenarios, where the misalignment is not necessarily deliberately implanted, but arises only as a side effect. To bridge this gap, we formulate a novel problem setting, in which the target of introspection is not necessarily a behavior explicitly implanted through fine-tuning, but rather alignment shifts that may emerge as unintended side effects, and we construct a dataset for this setting. Furthermore, to enhance sensitivity to internal model changes, we propose the Delta-Aware Introspection Adapter (DAIA), a novel mechanism designed to explicitly process both base-model activations and activation differences induced by fine-tuning. Our empirical evaluation shows that introspection learning generalizes to unseen fine-tuned models and safety categories, and that DAIA generally outperforms existing IAs.
|
| 2109 |
Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
2608.04408
|
cs.LG
|
De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma |
On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through b...On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through budget-matched teacher-continuation and rollback branches. Based on their relative success, states are categorized as recoverable, irreversible-but-avoidable, or ambiguous, and these labels guide whether training retains, rolls back, or conventionally supervises the corresponding trajectory. On AIME branch diagnostics, the mean continuation-minus-rollback effect is 0.185 for recoverable states and -1.000 for irreversible-but-avoidable states, demonstrating opposite intervention preferences. A branch-derived recoverability proxy achieves an AUC of 1.000, substantially outperforming divergence alone at 0.392. Across frozen evaluations, recoverability-aware control achieves the strongest recorded performance, reaching 0.578 success on held-out AIME2025 compared with 0.517 for the best baseline. It also improves AIME2024-2025 average@32 from 0.2656 to 0.3125 and GPQA-Diamond average@32 from 0.2702 to 0.3070. Component ablations further show that retaining teacher-correctable prefixes provides the largest individual contribution. These findings establish recoverability as an outcome-grounded decision variable for selective supervision in OPD.
|
| 2110 |
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant
2608.05127
|
cs.LG
|
Adel Javanmard, David P. Woodruff, Vahab Mirrokni |
Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, rely on high-dimensional geometric constructions but incur unfavorable dimension...Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, rely on high-dimensional geometric constructions but incur unfavorable dimension-dependent variance. In this work, we propose Subsampled Stochastic TurboQuant (SSTQ), a framework that combines a bounded Kashin representation, data-independent coordinate subsampling, and privacy-aware one-dimensional quantization. SSTQ includes two variants: (1) a Flat Randomized Response variant that is unbiased and, for a fixed codebook bit-width, frame redundancy, and dimension-independent Kashin level, achieves reconstruction MSE that scales linearly with the ambient dimension $d$, while using only $\lceil \log_2 N \rceil + b$ bits per message. Here, $N = \Theta(d)$ denotes the frame size in the Kashin transform and $b$ is the codebook bit-width; and (2) a metric-aware truncated-Laplace variant that removes the exponential dependence on bit-width at the cost of a non-vanishing bias. We also derive a convex uniform-surrogate codebook objective whose worst-case codebook-dependent upper bound improves from $O(4^b)$ to $O(2^b)$. Experiments on synthetic regression, Fashion-MNIST, and CIFAR-10 compare the per-message privacy-utility and uplink-communication trade-offs of SSTQ with those of established baselines, demonstrating favorable utility and communication efficiency.
|
| 2111 |
Multivariate Time Series Forecasting needs Cross Variable Loss
2608.05742
|
cs.LGcs.AI
|
Kuiye Ding, Yifan Hu, Hanchen Wang, Hao Xue |
Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. While existing studies mainly focus on cross-variable dependencies in historical observations, dependencies among future valu...Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. While existing studies mainly focus on cross-variable dependencies in historical observations, dependencies among future values are much less explored. Specifically, modern forecasting models largely follow the Direct Forecasting (DF) paradigm, generating multi-step forecasts with point-wise objectives that do not explicitly constrain cross-variable structure. In this work, we show that the DF objective is mismatched in the presence of cross-variable and lagged dependencies, revealing an objective gap. To address this issue, we propose \textbf{C}ross-\textbf{V}ariable \textbf{Loss} (CvLoss), a plug-in structural regularizer that constrains forecast residuals on a cross-variable graph. CvLoss penalizes inconsistent edge-wise residual differences over forecast patches, encouraging consistency across both synchronous and asynchronous interactions. Our experiments show that CvLoss consistently improves competitive forecasting models, outperforms representative learning objectives, and is compatible with a variety of forecasting backbones.
|
| 2112 |
A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies
2608.05995
|
cs.LG
|
Frieder Wizgall, Georg Tirpitz, Moritz Seiler, Kerstin Ritter, B\'alint Mucs\'anyi |
Reliable uncertainty estimates are critical in safety-sensitive applications. For such estimates to be useful in practice, it is crucial to understand the sources underlying a model's uncertainty, motivating the disentanglement of total uncertainty into episte...Reliable uncertainty estimates are critical in safety-sensitive applications. For such estimates to be useful in practice, it is crucial to understand the sources underlying a model's uncertainty, motivating the disentanglement of total uncertainty into epistemic and aleatoric uncertainty. Existing notions of uncertainty differ in the sources they capture and, consequently, in their definitions of aleatoric and epistemic uncertainty, with no universally accepted definition. We define uncertainty through sample-conditional pointwise posterior risk, which is the expected loss of a predictor under the distribution of plausible ground-truth functions given the observed sample. This definition unifies probabilistic and risk-based concepts of uncertainty. To assess state-of-the-art uncertainty disentanglement methods, we develop a framework that directly compares their estimates against ground-truth uncertainty defined primarily by posterior risk, alongside commonly used alternative uncertainty definitions. We find that Spectral-normalized Neural Gaussian Processes and Variational Latent Gaussian Processes most closely recover the ground-truth uncertainty, while most methods track posterior variance more closely than posterior risk, missing the predictors' bias. Beyond method rankings, we investigate how strongly estimated aleatoric and epistemic uncertainty are entangled and how sensitive uncertainty quality is to modeling choices, yielding practical guidance for uncertainty disentanglement. To support further method development and validation, we release 13 semi-synthetic UCI/OpenML datasets with known posteriors, enabling the computation of ground-truth uncertainty. Code and data will be made publicly available upon acceptance.
|
| 2113 |
FastKron: Efficient Quantization with Kronecker-Factored Hessians
2608.06291
|
cs.LG
|
Johann Birnick, Rayan Saab |
We accelerate a family of algorithms for neural network quantization which utilize a two-sided version of the GPTQ/LDLQ algorithm. Standard GPTQ-style adaptive rounding uses one-sided correlation information derived from input activations. A natural two-sided ...We accelerate a family of algorithms for neural network quantization which utilize a two-sided version of the GPTQ/LDLQ algorithm. Standard GPTQ-style adaptive rounding uses one-sided correlation information derived from input activations. A natural two-sided extension can additionally capture correlations across output channels. It utilizes a general Kronecker-factored approximation of the weight matrix curvature. This approach has been used in BoA and YAQA. But making a concrete algorithmic implementation of this two-sided GPTQ variant is nontrivial. BoA uses a large number of sequential steps, while YAQA improves over the sequential depth but still has quartic total cost. We introduce FastKron, an efficient algorithmic implementation that combines anti-diagonal parallelism with a recursive divide-and-conquer construction. For an $m\times n$ weight matrix, FastKron uses $O(m+n)$ sequential steps while reducing the total work from $O(m^2n^2)$ to $O(mn(m+n))$. Thus, it matches the cubic scaling of GPTQ while exploiting richer curvature information. Moreover, FastKron is modular with respect to both the base quantizer and the Hessian estimator. We also provide practical benchmarks, consider a range of Hessian approximations that FastKron can be used with, and provide an efficient technique to compute these Hessians.
|
| 2114 |
Memory--Batch Tradeoffs in Lipschitz Bandits
2608.07922
|
cs.LG
|
Zicheng Lyu, Zengfeng Huang |
Lipschitz bandits admit near-optimal regret $\widetilde O_d(T^{(d+1)/(d+2)})$ with little memory under full adaptivity, or with few batches under unrestricted memory. We characterize the minimax expected pseudo-regret over $T$ rounds with $W$ bits of memory an...Lipschitz bandits admit near-optimal regret $\widetilde O_d(T^{(d+1)/(d+2)})$ with little memory under full adaptivity, or with few batches under unrestricted memory. We characterize the minimax expected pseudo-regret over $T$ rounds with $W$ bits of memory and at most $B$ batches in dimension $d$. For every memory budget $W$, we prove a lower bound of order $T^{\frac{d+2}{d+3}} (1+(B-1)W)^{-\frac{1}{d(d+3)}}$. When $W\gtrsim_d\log(eT)$, algorithms with fixed batch boundaries attain, up to logarithmic factors, the larger of this bound and the optimal $B$-batch regret with unrestricted memory. The analysis separates fine-scale comparisons from the regional information needed to allocate their samples at low regret. Attaining $\widetilde O_d(T^{(d+1)/(d+2)})$ regret requires both $B=\Omega_d(\log\log T)$ and $(B-1)W=\widetilde{\Omega}_d(T^{d/(d+2)})$. With one bit of memory, minimax regret is $\widetilde\Theta_d(TB^{-1/(d+2)})$ under adaptive batch boundaries while $\Theta_d(T)$ under fixed boundaries.
|
| 2115 |
HOPPER: Learnable Hop Extraction for Linearized Graph Sequence Models
2608.09031
|
cs.LG
|
Isuru Herath, Arin Gopakumar, Sharan Sahu |
Graph neural networks typically propagate information through repeated message-passing layers, coupling propagation distance with the number of nonlinear transformations applied. This coupling can make deep architectures difficult to optimize and lead to over-...Graph neural networks typically propagate information through repeated message-passing layers, coupling propagation distance with the number of nonlinear transformations applied. This coupling can make deep architectures difficult to optimize and lead to over-smoothing, over-squashing, and loss of long-range information. Linearized Graph Sequence Models (LGSMs) address this issue by separating propagation depth from processing depth and representing successive propagation states of each node as a sequence. However, existing LGSMs construct these sequences using fixed graph operators, limiting their ability to adapt propagation to the input graph, node features, and downstream task. We introduce HOPPER, an end-to-end learnable extension of LGSM that learns how hop sequences should be extracted before processing by a modern state-space model. HOPPER supports feature-conditioned, structure-aware, graph and hop-adaptive propagation while preserving permutation equivariance, with standard adjacency-based and non-backtracking LGSM sequences arising as special cases of the extractor family. HOPPER is state-of-the-art or competitive across ECHO-Synth and performs strongly on City-Networks. On the LRIM physics-based long-range dependency benchmark, varying the maximum neighborhood size used for message-backtracking cancellation, corresponding to the structural memory window, substantially affects performance. Ablations further isolate the contributions of the learnable extraction mechanism and its structural and feature-adaptive components, showing that adaptive hop-sequence construction provides gains beyond the downstream sequence model alone. Together, these results demonstrate that learnable sequence extraction is a flexible and effective framework for long-range graph representation learning across synthetic, physics-based and real-world graph benchmarks.
|
| 2116 |
From Objectives to What Models Learn: A Landau Theory of Invariant Learning
2608.09396
|
cs.LG
|
Pinli Wang, Yue He, Peng Cui |
Invariant-learning objectives pursue similar goals yet produce qualitatively different regularization paths, leaving unclear when shortcuts can be suppressed without damaging stable modes. Our starting point is simple: in the desired shortcut-suppressed regime...Invariant-learning objectives pursue similar goals yet produce qualitatively different regularization paths, leaving unclear when shortcuts can be suppressed without damaging stable modes. Our starting point is simple: in the desired shortcut-suppressed regime, shortcut loading is small, so the objective's low-order expansion governs local stability and residual amplitude. This brings the problem into the domain of Landau phase-transition theory. We establish a mathematical isomorphism between the near-critical normal form of predictive-mode learning and Landau theory, identifying the learning objective as an effective free energy and critical-mode amplitude as an order parameter. The resulting low-order objective signatures predict distinct regularization phenotypes: quadratic $R_2$ terms shift phase boundaries and enable finite-strength elimination, whereas quartic $R_4$ terms continuously attenuate acquired modes while leaving nonzero residual loading at finite strengths. Higher-order terms may further shape nonlinear tails. In a canonical bilinear model, we derive exact phase boundaries and equilibrium loadings, identifying conditions for a selective-retention window in which shortcut suppression preserves stable structure. Controlled bilinear and ReLU experiments, including an MNIST construction, support the predicted signature-phenotype relation, while coupled-feature experiments validate its extension to collective modes. The framework connects the mathematical structure of invariant-learning objectives to what models learn as regularization varies.
|
| 2117 |
Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
2608.11479
|
cs.LG
|
Siqiao Mu, Diego Klabjan |
We establish the first convergence guarantees of gradient descent for general feedforward neural networks of any width or depth, with any initialization or dataset. We only assume that the activation functions are linearly bounded, Lipschitz continuous, and Li...We establish the first convergence guarantees of gradient descent for general feedforward neural networks of any width or depth, with any initialization or dataset. We only assume that the activation functions are linearly bounded, Lipschitz continuous, and Lipschitz smooth---properties that hold for linear, tanh, softplus, sigmoid, and smoothed ReLU functions---and that the loss function is Lipschitz smooth in the model outputs, a mild condition satisfied by mean squared error and binary cross-entropy loss. By relating the Lipschitz properties of one layer to the next, we obtain a novel generalized Lipschitz smoothness condition for an $L$-layer neural network where the change in gradient is upper bounded by the change in the parameter space, multiplied by a polynomial of the parameter norms at both endpoints of degree $2L - 2$. This yields a descent lemma where the loss decreases as long as the learning rate is small enough with respect to the parameter norms. By ensuring that the leading term of the polynomial grows at a controlled rate, we prove that the minimum squared gradient norm converges to zero in $T$ iterations at rate $O(1/T^{1/L})$.
|
| 2118 |
FLARE++: Low-rank attention with dynamic attention routing
2608.11519
|
cs.LG
|
Vedant Puri, Yongjie Jessica Zhang, Levent Burak Kara |
Full self-attention is a strong token mixer for PDE surrogates on irregular domains, but its quadratic cost limits its use on large problems. Latent-space attention methods such as PerceiverIO, Transolver, and FLARE (Fast Low-rank Attention Routing Engine) avo...Full self-attention is a strong token mixer for PDE surrogates on irregular domains, but its quadratic cost limits its use on large problems. Latent-space attention methods such as PerceiverIO, Transolver, and FLARE (Fast Low-rank Attention Routing Engine) avoid that cost by routing attention among $N$ tokens through $M\ll N$ learned latents. They compress by dot-product matching of the input tokens against $M$ learned query tokens or projection weights: once trained, the same learned templates serve every input. We remove this restriction with FLARE++, a low-rank attention architecture with input-conditioned routing queries. FLARE++ uses FLARE's own encoder to map the $N$ input tokens to $M$ query tokens, which correct the learned queries. The adapted queries then determine how that same input is compressed and redistributed. This preserves FLARE's explicit low-rank factorization and linear $\mathcal O(NM)$ complexity, and expresses the complete routing operation with standard scaled dot-product attention (SDPA) calls alone. We also provide a multi-GPU context-parallel implementation that shards input tokens across devices without ever gathering the full token sequence on one of them. FLARE++ reduces FLARE's error by $25\%$ on average across five standard PDE benchmarks, achieving the lowest errors among the efficient models compared. The gains persist on industrial-scale DrivAerML aerodynamics and on Long Range Arena, where average accuracy rises by $3.5\%$ over fixed-query FLARE. Code is available at https://github.com/vpuri3/FLARE.py.
|
| 2119 |
A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers
2608.14689
|
cs.LG
|
Pratikkumar Babariya |
Residual connections are a fundamental component of transformer architectures, yet the roles of the attention and feed-forward residual pathways remain poorly understood when considered independently. This paper presents a reproducibility study of partial resi...Residual connections are a fundamental component of transformer architectures, yet the roles of the attention and feed-forward residual pathways remain poorly understood when considered independently. This paper presents a reproducibility study of partial residual ablations in Pre-LN GPT-style transformers trained at two scales (10M and 124M parameters). I compare four architectural configurations by selectively removing the attention residual connection, the feed-forward residual connection, or both. At 10M scale, a controlled 8-seed deterministic sweep shows a clear asymmetry: removing the attention residual (FFNOnly) reaches the No-Residual collapse regime (3.350 +/- 0.002), whereas removing the feed-forward residual (AttnOnly) remains well separated from it (1.580 +/- 0.003). A subsequent controlled 124M sweep establishes that the direction of this asymmetry persists under matched seeds: across five seeds each, AttnOnly reaches 6.037 +/- 0.355 versus 7.569 +/- 0.001 for FFNOnly, with no overlap between the observed ranges. AttnOnly nevertheless remains substantially degraded relative to FullResidual (4.577 +/- 0.012) and exhibits markedly greater seed sensitivity at 124M. During the investigation, I identified and corrected an experimental measurement confound in runtime gain scaling and retained an intermediate reproduction failure rather than excluding it. I propose cross-position routing through self-attention as a falsifiable hypothesis for the observed asymmetry; no clean causal test of this mechanism has yet been completed. To support reproducibility, I release the source code, experiment configurations, training logs, and experimental results, including intermediate non-reproducing runs.
|
| 2120 |
Online Convex Optimization with Dueling Feedback
2608.15050
|
cs.LG
|
Yiyang Lu, Hareshkumar Jadav, Mohammad Pedramfar, Ranveer Singh, Vaneet Aggarwal |
Noisy binary comparison between two candidates is a common interface between human and learning systems, especially in modern large language model (LLM) post-training alignment. We study online convex optimization with dueling (pairwise comparison) feedback, w...Noisy binary comparison between two candidates is a common interface between human and learning systems, especially in modern large language model (LLM) post-training alignment. We study online convex optimization with dueling (pairwise comparison) feedback, where the learner observes only a binary preference between two queried points. We consider adversarial sequences of convex losses and measure regret with the loss at both queried points, under a comparison link with a known nonzero slope at the origin. We propose a simple reduction that converts dueling feedback into approximate gradients, enabling the use of standard first-order methods. We show that regret guarantees transfer under this reduction, yielding $\mathcal O(T^{3/4})$ static and adaptive regret, and $\mathcal O(T^{3/4}\sqrt{1+P_T/D})$ dynamic regret with unknown comparator path length $P_T$. For strongly convex losses, the static and adaptive bounds improve to $\widetilde{\mathcal O}(T^{2/3})$. For smooth losses, we presents unified dueling ellipsoidal FTRL, and proves $\widetilde{\mathcal O}(T^{2/3})$ static regret, which improves to $\widetilde{\mathcal O}(\sqrt T)$ under additional strong convexity.
|
| 2121 |
SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences
2608.15429
|
cs.LG
|
Tsz Fung Pang, Po Jen Chen, Nimish Ronghe, Farhad Farahani, Bo Zhang |
Prior embedding models for sequential recommendation typically operate within a homogeneous action space, limiting their ability to capture cross-surface behavioral signals spanning distinct behavioral domains. We present SAGA, a generative action embedding mo...Prior embedding models for sequential recommendation typically operate within a homogeneous action space, limiting their ability to capture cross-surface behavioral signals spanning distinct behavioral domains. We present SAGA, a generative action embedding model that encodes multi-surface user interaction sequences across a Financial Service organization's ecosystems, from checkout, peer-to-peer (P2P) transactions, in-app engagement, email to account actions, into a unified user representation for downstream recommendation tasks. Central to SAGA is a per-field tokenization schema that decomposes each action event into multiple field-level tokens (e.g. product, interaction, surface), enabling field-level attention and per-field training objectives that fused single-token approaches cannot support. Through an offline ablation study on loss formulation, tokenization granularity and training data scope, we isolate the contribution of each design choice. A downstream model integrated with SAGA-generated user embeddings delivers the strongest overall click and conversion lift across diverse downstream touchpoints, compared to all ablated and alternative architectures.
|
| 2122 |
Spectral Saliency for Machine Unlearning
2608.15548
|
cs.LG
|
Cedar Site Bai, Amber Yijia Zheng, Raymond A. Yeh, Brian Bullins |
Machine unlearning (MU) aims to remove the influence of specific training data while preserving model utility. As the name suggests, MU can be viewed as the inverse of learning, using gradient-based updates to reduce the influence of a forget-set by counteract...Machine unlearning (MU) aims to remove the influence of specific training data while preserving model utility. As the name suggests, MU can be viewed as the inverse of learning, using gradient-based updates to reduce the influence of a forget-set by counteracting the previously learned behavior. Recently, Muon, a gradient descent variant, has been introduced. Muon applies spectral magnitude normalization to encourage exploration of rare directions and demonstrates promising performance. Inspired by Muon, we adopt the spectral view for unlearning and propose Spectral Saliency Unlearning (SSU). SSU thresholds weak singular components and updates only those directions supported by a confident unlearning signal. We further provide theoretical justification for this thresholding approach from the perspective of the forgetting-retention trade-off. Experiments across image classifiers, diffusion models, and LLMs demonstrate SSU's effectiveness.
|
| 2123 |
SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces
2608.15665
|
cs.LG
|
Ziming Yu, Shuyao Xiao, Xingyu Zhao, Sike Wang, Pan Zhou |
Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZer...Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization through a carefully designed dual low-dimensionality strategy: (i) multi-query forward-difference gradient estimation in periodically refreshed random subspaces to mitigate noise amplification in moment buffers, and (ii) Adam updates with periodic restarts performed directly in low-dimensional space rather than full-parameter space. In experiments, this dual design retains memory overhead comparable to momentum-free ZO methods while achieving stronger optimization performance than the evaluated ZO baselines. Theoretically, in the exact-directional limit, $K$-query averaging preserves conditional unbiasedness, while the coefficient estimator's covariance and mean-squared error, as well as query-induced second-moment inflation, scale exactly as $1/K$. Extensive experiments across SuperGLUE with models from 1.3B to 32B parameters under both full fine-tuning and LoRA schemes demonstrate consistent improvements over competing ZO methods. SubZero+ significantly narrows the performance gap with first-order optimization while preserving ZO's inference-time memory efficiency.
|
| 2124 |
Multinomial Subset Routing with Sum-Max Rewards and Operational Constraints
2608.16375
|
cs.LG
|
Quan Zhou, Yiyan Huang |
We study online routing to subsets of experts under aggregate bandit feedback and long-run operational constraints. Expert contributions can be complementary: for each task dimension, the best selected expert determines the contribution, and the total reward a...We study online routing to subsets of experts under aggregate bandit feedback and long-run operational constraints. Expert contributions can be complementary: for each task dimension, the best selected expert determines the contribution, and the total reward aggregates these dimension-wise maxima. At the same time, capacity, budget, and fairness requirements impose lower and upper bounds on long-run expert activation frequencies. A fixed deterministic subset cannot generally satisfy such heterogeneous requirements, motivating a stochastic routing policy. We formalize this problem as Multinomial Subset Routing (MSR). The learner maintains a distribution \(q\) over \(K\) experts, samples an expert independently \(M\) times from \(q\), and routes each task to the distinct sampled experts. We propose OMD-Approachability, which combines online mirror descent with Blackwell's approachability to optimize MSR under two-sided operational constraints. We establish \(O(T^{-1/2})\) average reward regret and expected constraint violation. We also quantify the approximation gap induced by a linear surrogate, and evaluate the approach on a real-world crowdsourcing dataset.
|
| 2125 |
Reinforced Planning with Latent World Models
2608.18669
|
cs.LG
|
Armin Sommer, Jannik Schilling |
Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has made substantial progress in learning world models that predict the consequences of action sequences, yet the p...Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has made substantial progress in learning world models that predict the consequences of action sequences, yet the procedures used to plan with these models remain largely hand-designed. Most planners rely on fixed search or optimization rules; approaches that learn aspects of search typically imitate a predefined optimizer or use planning to inform an amortized policy, rather improving multi-step plans. We introduce \textbf{Reinforced Planning}, a method that learns the plan-update itself by reinforcing update rules that produce better plans, using gradients propagated through a differentiable world model. We instantiate Reinforced Planning in RP1, which learns a critic over imagined outcomes via temporal-difference learning and a neural plan-improvement operator trained via imagined rollouts with a pretrained world model. RP1 can be trained fully offline without environment interaction; environment episodes are used only for checkpoint selection. Across visual navigation, arm reaching, and robotic manipulation on two world-model backbones, RP1 matches or exceeds existing planners, achieving near-perfect success in several settings while using $1{,}000\times$ fewer world-model rollouts than the strongest alternative (CEM) and planning up to $67\times$ faster under concurrent planners inference.
|
| 2126 |
Beyond Multimodal Alignment: Shared Physical Representations Across Sensors and Action Orders
2608.19492
|
cs.LG
|
Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong |
Multimodal world models are often evaluated by whether different sensors produce similar representations. However, similar representations do not necessarily imply that the models make the same physical predictions, or that those representations can be reused ...Multimodal world models are often evaluated by whether different sensors produce similar representations. However, similar representations do not necessarily imply that the models make the same physical predictions, or that those representations can be reused when actions are combined in a new order. We study both questions through the physical responses predicted by a model. We first use the Cluster Haptic dataset to ask whether audio and acceleration can independently recover the behavior of the same surface from different observations. Predictions from the two sensors are substantially closer for the same surface than for different surfaces, with a $4.5\times$ gap on average, while both also outperform an average-surface prediction. We then show that this agreement alone does not determine how familiar actions should compose. In a controlled elastoplastic system, shared step dynamics fit observed programs less accurately than a whole-program predictor but generalize better to unseen action orders, with the ranking reversing on both held-out transitions across three independent initializations. Fusing free-decay and hysteresis observations further improves prediction, with diagonal Gaussian beliefs yielding the lowest errors. Together, these results distinguish cross-sensor consistency, multimodal fusion, and generalization to new action orders as separate questions in evaluating multimodal physical representations.
|
| 2127 |
In Two Minds about Lifelong Learning: Exploring Hemispheric Redundancy and Specialisation in Neural Models
2608.19514
|
cs.LG
|
Benjamin Smith, Levin Kuhlmann, Kaushik Roy, Gideon Kowadlo |
Persistent intelligent systems require the ability to learn continually, but current machine learning approaches face significant challenges in this area compared to biological learning systems. Machine learning algorithms typically trade off retention of prev...Persistent intelligent systems require the ability to learn continually, but current machine learning approaches face significant challenges in this area compared to biological learning systems. Machine learning algorithms typically trade off retention of previously learned information and adaptation to new or changing data patterns. When continual learning capabilities are absent, algorithms must undergo retraining using the entire data set, an approach that becomes impractical when original training data are unavailable due to storage constraints, financial or computational costs, or privacy restrictions. However, biological animals can learn continually, without experiencing catastrophic forgetting. This paper attempts to build a high-level framework for how animals learn and preserve knowledge by modelling neural components and states that are known to be related to memory consolidation. We focus on three concepts: experience replay, REM sleep, and bilaterality. We propose 4MAS (4 Module Awake/Sleep), a novel macroarchitecture demonstrating how machine learning models might benefit from asymmetric hemispheres, each with their own long- and short-term memory mechanisms, and how a period of sleep between incremental learning tasks might benefit memory consolidation. Finally, we present results showing that our architecture achieves statistical improvements over established generative replay baselines on Split-MNIST (98.3%) and Split-Fashion-MNIST (84.9%), while reducing forgetting by more than half relative to the strongest replay baseline and demonstrating monotonic capacity scaling on Split-CIFAR-100 (29.29%).
|
| 2128 |
Learning Exact NVIDIA SASS Encoders with $\mathbb{F}_2$ Linear Algebra
2608.20532
|
cs.LG
|
Jiading Gai |
NVIDIA provides a SASS disassembler but no public SASS assembler for recent data-center GPUs, limiting controlled machine-code rewriting. We present F2Asm, which learns exact 128-bit SASS encoders from paired disassembly and original CUBIN instruction words. T...NVIDIA provides a SASS disassembler but no public SASS assembler for recent data-center GPUs, limiting controlled machine-code rewriting. We present F2Asm, which learns exact 128-bit SASS encoders from paired disassembly and original CUBIN instruction words. To our knowledge, F2Asm is the first system to learn SASS instruction encoders as vector-valued affine maps over $\mathbb{F}_2$ and the first open-source NVIDIA SASS assembler to support Rubin SM107. F2Asm uses Gaussian elimination over $\mathbb{F}_2$ to incrementally build a compact basis, detect inconsistencies, and reject inputs outside the learned span. F2Asm separates target-specific control bits, relocation rules, and CUBIN metadata from its learning algorithm. We train encoders for Hopper SM90/SM90a, Blackwell SM100, and Rubin SM107 using 3,225 CUBINs from pinned NVIDIA and third-party production libraries, CUDA 13.3 packages, and CUDA 13.4 Developer Preview archives. In round-trip tests, F2Asm reassembles each CUBIN's disassembled SASS, and all compared executable text sections match the originals exactly. Joint training with F2Asm yields one shared encoder for five Blackwell SM targets and another for three Rubin SM targets, providing strong evidence of a common SASS encoding scheme for instructions shared within each family. Continual training extends the Rubin encoder to all 17,159 previously unsupported cuTile and GROMACS queries with 1,504 additional basis rows, matching the derived lower bound.
|
| 2129 |
The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models
2608.22876
|
cs.LG
|
Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi, Jaewon Jang |
Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely verified. We formalize prefix invariance and give a lightweight audit, two forward passes, no training or gradients, yield...Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely verified. We formalize prefix invariance and give a lightweight audit, two forward passes, no training or gradients, yielding a per-layer score localizing where causality breaks. Attention-mask inspection, the field's default check, is incomplete: causality is a graph-level property, and leaks can occur via scans, aggregations, or normalization despite correct masks. Across 192 injected-fault trials on eight checkpoints, mask inspection detected none, while our audit localized all 192/192 to the exact layer. Static/dynamic analysis of chunked-scan code in transformers found the same defect in Zamba2 and Nemotron-H, an inter-chunk axis error fixed via the reference implementation. The method fits on one page and runs in seconds.
|
| 2130 |
Low-Latency Activation-Regularized Sparse Neural Operators with Distillation Assistance Towards Real-Time Neuromorphic Virtual Sensing
2608.23987
|
cs.LG
|
William Howes, Farid Ahmed, Syed Bahauddin Alam |
Virtual sensing enables digital twins and safety-critical systems to reconstruct and forecast spatial-temporal physics in real time. However, conventional computational and data-driven methods often face challenges in generalization, latency, and energy effici...Virtual sensing enables digital twins and safety-critical systems to reconstruct and forecast spatial-temporal physics in real time. However, conventional computational and data-driven methods often face challenges in generalization, latency, and energy efficiency for edge deployment. Neural operators offer a promising alternative but remain reliant on power-intensive hardware. Spiking neurons and neuromorphic computing can improve efficiency, yet surrogate-gradient training and multi-step spiking introduce convergence and latency challenges. We propose the Sparse-Activation-ReLU (SAR) layer, a single-step alternative that promotes activation sparsity without surrogate-gradient training while remaining compatible with event-based computing. Within a trunk-based NOMAD architecture, SAR achieves over a fivefold improvement in the combined Latency-Error-Energy (LEE) metric compared with Variable Spiking Neuron (VSN) and Leaky Integrate-and-Fire (LIF) implementations. We further analyze spiking entropy and feature usage and introduce synthetic knowledge distillation, reducing the LEE score by more than twofold. Finally, we improve VSN through a ReLU-based spiking loss and graph-neighbor thresholding. On the Heat Exchanger dataset, these approaches reduce L2 error by more than twofold and nearly sevenfold, respectively, while reducing spiking and spatial aggregation. Overall, the work presented is a step towards energy-efficient virtual sensing by providing an alternative framework that can be positioned towards neuromorphic or other edge device integration that can be a gold standard to compare latency, energy, and error performance for future efficient designs that are sparsity or brain-inspired spiking based.
|
| 2131 |
Flow-JEPA: Robust Latent Dynamics for JEPA World Models via Flow Matching
2608.29029
|
cs.LGcs.AI
|
Yanchen Huo, Ziying Song, Yadan Luo |
Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free manner. Although numerous JEPA-based approaches have been proposed to mitigate representation collapse, our experiment...Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free manner. Although numerous JEPA-based approaches have been proposed to mitigate representation collapse, our experiments on localized, out-of-distribution visual noise reveal that performance degradation remains pronounced and unresolved. We propose Flow-JEPA (F-JEPA), a flow-based latent dynamics model that jointly generates a sequence of future latent states conditioned on the current observation and actions. A Gaussian distribution serves as the flow source, exposing the vector field to perturbed latent trajectories as it learns to transport them toward clean future representations. This formulation retains the reconstruction-free JEPA framework while switching from pointwise transition regression to stochastic trajectory-level prediction. F-JEPA raises mean success from $86\%$ to $92\%$ under clean observations and from $67\%$ to $86\%$ under noisy conditions. Further evaluations over varying perturbation severity and inference settings show that the performance advantage persists across a broad range of conditions. These results suggest that conditional flow matching provides a promising alternative to deterministic autoregressive prediction as a dynamics formulation in JEPA world models.
|
| 2132 |
PokaiTrainer: Scaling Equilibrium Search to Competitive Pok\'emon VGC
2608.29197
|
cs.LG
|
Max Yu |
Decision-time equilibrium search carried poker to superhuman play, but it has so far relied on tractable subgames: a handful of actions per decision, chance confined to card deals, one player moving at a time. Competitive Pok\'emon in its official doubles form...Decision-time equilibrium search carried poker to superhuman play, but it has so far relied on tractable subgames: a handful of actions per decision, chance confined to card deals, one player moving at a time. Competitive Pok\'emon in its official doubles format (VGC) breaks all three assumptions at once. Both players act simultaneously from joint menus in the hundreds, each joint action resolves to hundreds of stochastic outcomes, and the opponent's reserves and stat allocations are hidden. No prior Pok\'emon agent performs equilibrium search, and whether it scales to this regime was open; we show that it does, and report what it took. PokaiEngine, our Rust battle engine, enumerates a joint action's full weighted outcome distribution in one pass, at ${\sim}99\%$ parity with Pok\'emon Showdown and a fraction of the cost of sampling it. PokaiTrainer adapts Student of Games to this scale and trains it by self-play over hundreds of human teams. Each decision is solved by counterfactual regret minimization as a Bayesian matrix game over public belief states, subgames grow under an explicit compute budget, and value targets are harvested from the interior of every solve and grounded by realized outcomes. The strength is in the search. The network's policy alone loses even to a shallow heuristic search. PokaiTrainer is, to our knowledge, the first VGC agent rated on the live Showdown ladder. Under open team sheets it wins 59% of 150 best-of-three sets against a human field averaging ${\sim}1320$ Elo, holds a 1350-1400 Elo band, and at its peak reached 1492 Elo, entering the format's top 500.
|
| 2133 |
Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals
2608.29434
|
cs.LG
|
Fabio F. Oberweger, Michael Schwingshackl, Markus Murschitz |
Latent action world models let agents plan new behaviors at test time by predicting how actions change the environment, and joint-embedding predictive architectures (JEPAs) do so by forecasting future latent states rather than pixels. Yet nearly all such model...Latent action world models let agents plan new behaviors at test time by predicting how actions change the environment, and joint-embedding predictive architectures (JEPAs) do so by forecasting future latent states rather than pixels. Yet nearly all such models see the world through a camera, even though robotic manipulation is fundamentally geometric: in robotics goals for manipulation are traditionally specified by target object poses, not by images of the object once placed. We ask whether latent planning survives a shift from appearance to geometry, on the observation side as well as on the goal specifications side. To answer this, we extend the stable-worldmodel evaluation platform with simulated LiDAR-style raycast point clouds as a new sensor modality, and adapt three JEPA designs to point clouds: a frozen-encoder model built on Utonia features, a distribution-prior model based on LeWM, and an action-sensitive model based on Delta-JEPA. We further introduce a goal-encoding mechanism that constructs the goal latent from the current latent and a 3D target pose, removing the need for goal images or goal point clouds. A comparative evaluation of the different anti-collapse mechanisms shows that point-cloud world models can match their image-based counterparts, demonstrating that the modality shift from appearance to geometry is achievable. All models are released as open weights with open-source training and inference code, to make world-model planning accessible for LiDAR-driven and pose-directed robotic tasks.
|
| 2134 |
Seasonality-Aware Hybrid Convolutional Transformer for Antarctic Sea Ice Concentration Forecasting
2608.30654
|
cs.LG
|
Danyang Li, John Taylor, Thang Bui, Quanling Deng |
Antarctic sea ice concentration (SIC) forecasting is an important yet challenging task due to the coexistence of complex spatial structure, long-range temporal dependencies, and strong seasonal variability. Conventional convolution-based models are effective a...Antarctic sea ice concentration (SIC) forecasting is an important yet challenging task due to the coexistence of complex spatial structure, long-range temporal dependencies, and strong seasonal variability. Conventional convolution-based models are effective at capturing local spatial patterns, but often have limited ability to model long-term temporal evolution. To address these challenges, we build on a hybrid convolutional transformer forecasting framework for monthly Antarctic SIC forecasting. This framework combines convolutional encoding for spatial feature extraction with space-time factorised self-attention for SIC modelling. We further introduce two season-based mechanisms: a month-aware positional encoding that injects calendar-month information into the token representation, and a seasonal temporal bias that encourages attention to periodically related historical states. Experimental results show that the proposed framework achieves better performance than convolutional and recurrent baselines as well as ECMWF's physics-based dynamical model SEAS5 under both classification and regression metrics. Ablation studies further indicate that the seasonality-aware components provide consistent additional gains in both short- and long-horizon prediction. These results demonstrate the value of combining convolutional structures, attention mechanisms, and periodic prior information for Antarctic SIC forecasting.
|
| 2135 |
Nonparametric Contextual Pricing and Inventory Learning under Censored Demand
2608.30944
|
cs.LG
|
Zean Han, Jing Liang, Ruihan Lin, Zezhen Ding, Jiheng Zhang |
In online retailing, when a product sells out, a retailer often sees only the units sold, not how many customers would have bought it had inventory been available. However, the inventory level determines how much demand is revealed, and this information can in...In online retailing, when a product sells out, a retailer often sees only the units sold, not how many customers would have bought it had inventory been available. However, the inventory level determines how much demand is revealed, and this information can influence subsequent decisions and future profits. We study an online selling problem in which, in each round, the seller observes a market context and then makes pricing and stocking decisions based on censored sales data from previous rounds. The challenge is to learn a context-dependent pricing and stocking policy without assuming a particular formula for demand or observing realized profit. To overcome this difficulty, we propose a Mean-Calibrated Kernel UCB (MCK-UCB) algorithm that turns each incomplete sales record into a reliable guide for both inventory and price decisions, using data from past rounds with similar market conditions. This design allows us to learn while serving customers, without a separate exploration phase or the need to recover all demand hidden by stockouts. We prove the minimax optimality of the proposed algorithm, with strictly faster rates when expected profit varies more smoothly with price. Comprehensive numerical experiments have been conducted to confirm the effectiveness of the proposed algorithm.
|
| 2136 |
EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction
2609.00653
|
cs.LG
|
Yunzhen Zhang, Ruoxi Piao, Muhammad Ibrahim Ali Shah, Hasan Onur Keles, Mustafa Misir |
Electroencephalography (EEG) is a non-invasive technique for measuring neural activity and has been widely used in neuroscience applications. Recent advances in EEG foundation models have enabled strong performance across diverse neural decoding tasks. However...Electroencephalography (EEG) is a non-invasive technique for measuring neural activity and has been widely used in neuroscience applications. Recent advances in EEG foundation models have enabled strong performance across diverse neural decoding tasks. However, no single foundation model consistently performs best across datasets or individual EEG instances, while instance-level model selection remains largely unexplored. To address this limitation, we formulate EEG foundation model selection as an instance-level Algorithm Selection (AS) problem. We propose \textbf{EEG-AS}, an instance-level algorithm selection framework that characterizes each EEG instance using inference-available latent EEG embeddings, handcrafted neurophysiological features, and an anchor foundation model. During training, EEG-AS learns to reconstruct unavailable foundation-model behaviors from privileged prediction tokens conditioned on an anchor foundation model, while during inference it estimates these behaviors without executing the entire model portfolio, enabling efficient selection from seven EEG foundation models. Experiments on seven public EEG benchmarks demonstrate that EEG-AS substantially narrows the gap between the Single Best Solver (SBS) and the oracle upper bound for each instance. These results highlight the effectiveness of instance-level AS for adaptive deployment of EEG foundation models.
|
| 2137 |
Linear Reusable Neural Bases Architecture for Network Compression
2609.01550
|
cs.LG
|
Binshuai Wang, Peng Wei, Mahyar Ghazanfari |
Memory constraints remain a critical bottleneck in the deployment of large-scale AI models. Parameter sharing across network depth reduces model storage, but repeatedly applying an identical transformation limits flexibility across layers. Inspired by time--me...Memory constraints remain a critical bottleneck in the deployment of large-scale AI models. Parameter sharing across network depth reduces model storage, but repeatedly applying an identical transformation limits flexibility across layers. Inspired by time--memory trade-offs in classical algorithms, we introduce the Linear Reusable Neural Bases (LRNB) architecture, an RNN-based framework that improves parameter efficiency through parameter reuse at the \textit{neuron level}. Each feedforward residual module is represented as a linear combination of shared neural bases, with depth-specific learnable coefficients and optional shifts providing flexibility across layers. This formulation reduces parameter redundancy across depth and enables the construction of wider and deeper networks within a fixed parameter budget. We further provide a geometric interpretation of the neural bases from a vector-field perspective and extend the framework to linear projection modules. Experiments demonstrate that the LRNB architecture achieves comparable or lower final training loss than independently parameterized baselines while using fewer parameters and maintaining stable training dynamics. These findings support neuron-level reuse as a practical approach to parameter-efficient network design.
|
| 2138 |
Percolation Dynamics in Optimization: Variance Cascades and Nested Symmetry
2609.02373
|
cs.LG
|
Sai Niranjan Ramachandran, Suvrit Sra |
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the...We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment. These structural transitions register as variance spikes in a macroscopic order parameter echoing physical phase transitions. We further state sufficient conditions under which the trapping argument carries over to Adam and AdamW under heavy-tailed gradient noise and measure them on a trained Transformer.
|
| 2139 |
TRACE: A spatiotemporal contact memory graph network simulator for granular dynamics
2609.02991
|
cs.LG
|
Changjian Zhou, Negin Yousefpour, Jie Qi, Junfeng Fang, Guillermo A. Narsilio |
Learned graph simulators provide an efficient alternative to high-fidelity solvers for granular dynamics. However, granular motion depends strongly on inter-granular contact history, which is difficult to preserve when particle contacts form, break, and rearra...Learned graph simulators provide an efficient alternative to high-fidelity solvers for granular dynamics. However, granular motion depends strongly on inter-granular contact history, which is difficult to preserve when particle contacts form, break, and rearrange. Existing simulators mainly store temporal information in node features or node-level memory. Here we introduce TRACE, a graph-network simulator that stores interaction history directly on contact edges. Each edge maintains a persistent memory updated by attention-based message passing and a gated recurrent unit, while an edge-identity dictionary preserves this memory as the contact graph changes. A physics-structured decoder predicts inter-granular normal and tangential contact forces, enforces the Coulomb friction limit, and applies equal-and-opposite internal forces. The model is trained with single-step pretraining followed by autoregressive rollout fine-tuning. We evaluate TRACE on 2D and 3D granular column-collapse benchmarks. In both cases, TRACE produces stable, physically consistent long-horizon rollouts, closely reproducing the final deposit geometry and the kinetic energy released during collapse. Compared with graph network simulator (GNS) and node-memory graph neural simulator (NMGNS), TRACE reduces long-rollout position error by 31-62% and final-deposit error by 58-89% across the two benchmarks, while using fewer parameters and maintaining near-zero particle interpenetration. TRACE also achieves 12.2$\times$ and 8.9$\times$ speedups over the material point method (MPM) reference solver in 2D and 3D, respectively. Our code is available at https://github.com/Data-Driven-Computational-Geotechnics/TRACE.
|
| 2140 |
TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables
2609.07441
|
cs.LG
|
Jules Kreuer, Sofiane Ouaari, Julia Hellmig, Julius Braitinger, Nico Pfeifer |
Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. We introduce TabBench-Bio, a living and interactive benchmark of 43 biome...Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. We introduce TabBench-Bio, a living and interactive benchmark of 43 biomedical datasets spanning multiple domains. Under a shared cross-validation protocol, we compare classical estimators, neural networks, and tabular foundation models across 28 feature-by-sample operating points. At the reference cell of 10,000 features and 100 training samples, RealTabPFN 2.5 has the highest point estimate, closely followed by TabPFN 3 and Logistic Regression. A paired bootstrap over the target pool separates RealTabPFN 2.5 from TabPFN 3 by 87 Elo (95\% interval [43, 129]). Tabular foundation models generally occupy the leading ranks, while the strongest configuration depends on the operating point and biomedical modality. The AutoML framework AutoGluon, using its one-hour "extreme" preset, is configured as a separate resource-intensive reference and reported here at the reference and full cell. Fold-level predictions, run status, and deterministic aggregations make every reported result reproducible and reusable. The benchmark is open to contributions of new biomedical datasets. The interactive leaderboard is available at https://tabbench-bio.eu.
|
| 2141 |
No-Regret Mixing of LRU and LFU with Optimal Switching Cost
2609.07566
|
cs.LG
|
Younes Ben Mazziane, Xinying Zou |
Caching systems often rely on simple eviction policies such as Least Recently Used (LRU) and Least Frequently Used (LFU), which perform well in complementary request regimes. Recent policies such as LeCar and Cacheus combine LRU and LFU using ideas from the ex...Caching systems often rely on simple eviction policies such as Least Recently Used (LRU) and Least Frequently Used (LFU), which perform well in complementary request regimes. Recent policies such as LeCar and Cacheus combine LRU and LFU using ideas from the experts problem in online learning. Specifically, upon a miss, they randomize between the two eviction rules using probabilities derived from scores updated by tracking the history of past evictions. While these policies exhibit strong empirical performance, it remains unclear whether they are guaranteed, on every request sequence, to perform asymptotically as well as the better of LRU and LFU, i.e., whether they achieve sublinear regret with respect to this benchmark. We first show that LeCar suffers linear regret against an oblivious adversary, even with unbounded history. We then propose H-MC, a Hedge-based mixture of virtual LRU and LFU caches that preserves Hedge's selection probabilities, and hence its regret guarantees, while minimizing the switching cost among all joint selection rules with these marginals.
|
| 2142 |
Beyond the Matrix Sign: Quadratic Spectral Descent
2609.07597
|
cs.LG
|
Qiaozhe Zhang, Jun Sun, Yingzhuang Liu |
Muon emerges as a strong competitor of the AdamW for LLM pretraining, because the matrix-wise update it employs can potentially incur smaller second-order penalty than the once dominating AdamW, which performs coordinate-wise update. However, the spectral flat...Muon emerges as a strong competitor of the AdamW for LLM pretraining, because the matrix-wise update it employs can potentially incur smaller second-order penalty than the once dominating AdamW, which performs coordinate-wise update. However, the spectral flattening procedure in Muon is quite debatable since it discards the spectral amplitude information totally. To seek for better spectral allocation (and the associated spectral subspace), we propose to solve the quadratic model of loss function under the spectral norm constraint \textit{directly} (i.e., in a genuinely Newtonian way) and thus obtaining the Quadratic Spectral Descent (QSD) algorithm. In contrast, many existing curvature-aware methods either exploit the second-order information in an \textit{implicit} way by changing the weight update geometry (such as Mousse, FISMO) or rely on strong assumptions (such as the weight displacement isotropy assumption in Newton-Muon). QSD's potential advantage over these methods is best illustrated in the isotropic curvature scenario, where Mousse, FISMO and Newton-Muon all reduce to Muon while the spectral allocation in QSD is still \textit{non-flat} (since the spectral allocation in QSD depends on the \textit{gradient to curvature ratio}). Meanwhile, to control the complexity of QSD, we employ inversion-free K-FAC and \textit{online} Frank-Wolfe update which is essentially a matrix sign operator. Overall, the complexity increase can be rather mild. Experiments on GPT pre-training show that QSD consistently improves validation loss over Muon and recent Muon variants, while achieving up to an $8.49\%$ wall-clock speedup at matched validation loss.
|
| 2143 |
KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization
2609.08135
|
cs.LG
|
Lexington Whalen, Yuki Ito, Ryo Sakamoto |
We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise ...We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the participation factor $\kappa$, yielding a closed-form signal-to-noise-ratio law. The resulting functional also admits a closed-form upper bound $\kappa^{*}$ that no function-preserving linear transform can exceed and that is attained by a recent state-of-the-art method. Building on this analysis, we introduce KBBQ (\textbf{K}appa-\textbf{B}raked \textbf{B}lockwise \textbf{Q}uantization), which parameterizes the extent to which a transform approaches this ceiling. At W4A4, across four base models and two FP4 formats, KBBQ outperforms the prior state of the art without additional deployment-time computation.
|
| 2144 |
PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games
2609.09059
|
cs.LG
|
Ryan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie |
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. Here we presen...While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. Here we present PlayTrain, an RL framework that combines the abilities of large language models (LLMs) to robustly generate JavaScript (JS) games from a minimal human prompt, and an efficient pipeline that can run any JS game in a standard 'gym' environment. Not only are recent LLMs particularly good at writing JS code, but the JS format also allows users to easily play generated VGEs, while PlayTrain enables us to train RL agents on the exact same games. We demonstrate multiple use cases of PlayTrain, including cloning well-known Atari and ProcGen games in simple JS, where PlayTrain trains pixel-based agents end-to-end at over 1M agent-decisions per second on a single GPU node; and creating modified versions thereof (e.g., that support novel test sets, procedural generation logics, or game dynamics). Through PlayTrain, we reimagine RL VGE development: all we need is a single JS file, generated and modified through an LLM. We discuss promising future RL research directions that PlayTrain unlocks.
|
| 2145 |
Solving Few-Shot Multiobjective Multitask Optimization via Iterative Sequential Transfer
2609.11228
|
cs.LG
|
Tingyang Wei, Haofeng Wu, Ananda Phan Iman, Zhao Wei, Jiao Liu |
Applying knowledge transfer across multiple optimization tasks, multitask optimization (MTO) emerges as a promising approach to solving synergistic optimization tasks simultaneously. However, the development of effective knowledge transfer mechanisms in MTO fu...Applying knowledge transfer across multiple optimization tasks, multitask optimization (MTO) emerges as a promising approach to solving synergistic optimization tasks simultaneously. However, the development of effective knowledge transfer mechanisms in MTO fundamentally relies on aligning elite solution distributions across tasks. This dependency creates a critical bottleneck in few-shot optimization regimes, as restricted evaluation budgets impede the identification of elite solution distributions required for beneficial transfer. This challenge is exacerbated in multiobjective multitask problems, where each optimizer must approximate a continuous Pareto manifold rather than a single optimal point. This paper introduces Iterative Sequential Transfer (IST) to circumvent this bottleneck. We model MTO as a sequence of sequential transfer optimization problems, concentrating evaluations on a single target per iteration. We propose a likelihood-informed task prioritization mechanism to maximize transfer utility by identifying the task most likely ready for knowledge integration. Empirical results on benchmark and real-world problems verify the effectiveness of the proposed method under tight budgets.
|
| 2146 |
Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning
2609.12014
|
cs.LG
|
Adam Haroon, Cody Fleming |
Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answ...Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free $(\alpha, \delta)$ bound on the unsafe fraction of the selection, and behavior cloning follows. What is certified is the training set, not the policy, whose cost we report rather than bound. Where no threshold attains the target, the procedure refuses. We are not aware of prior work certifying the composition of a training set for offline RL or imitation. Oracle controls justify the design: reweighting individual transitions fails even with an exact value, so the value selects whole trajectories. The policies satisfy the cost budget on twelve of fifteen DSRL tasks, matching a clone of the ground-truth safe subset, which needs a label on every trajectory. Retrained on the certified selection, the strongest full-label method becomes safe where no setting of its own cost target rescues it. Refusal is predictable: the certificate's probability has a closed form in the purity of the pool's top quantile, which the calibration sample estimates and through which the scorer enters.
|
| 2147 |
JumpStart Your Policy Learning with Lessons from 160,000 Training Runs
2609.13730
|
cs.LG
|
Nabil Omi, Eric Bae, Chung Yik Edward Yeung, Siddhartha Sen, Ali Farhadi |
Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has shown that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties, b...Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has shown that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties, but these sources of variability have not been systematically investigated together at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 160,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance among the strongest methods is often close, but the leaders differ substantially across environments. We find that proper hyperparameter tuning frequently reshuffles perceived algorithm rankings and that benchmark composition can produce conflicting conclusions. We also study hyperparameter sensitivity and transfer across environments, identifying a simple strategy for deriving strong default configurations. We use our findings to develop a dataset-conditioned recommender that provides task-specific algorithm recommendations for practitioners. Finally, we release JumpStart: a resource suite containing every trained policy, per-model scores and hyperparameters, strong baselines across all environments, training and evaluation code, and an extensible website for retrieving, analyzing, and contributing results. Together, these resources aim to make offline policy-learning research more reliable and enable future work beyond the scope of this study.
|
| 2148 |
Pathwise Individual Rationality in Federated Learning: A Mechanism-Architecture Co-Design
2609.14591
|
cs.LG
|
Amin Meghrazi, Srinivasan Parthasarathy, Andrew Perrault |
Participation in federated learning (FL) comes at a cost. Clients trade off privacy, communication, and compute costs for potentially greater gains in model efficacy. This paper explores this tradeoff under the aegis of individual rationality (IR) versus autar...Participation in federated learning (FL) comes at a cost. Clients trade off privacy, communication, and compute costs for potentially greater gains in model efficacy. This paper explores this tradeoff under the aegis of individual rationality (IR) versus autarky, the basic game-theoretic requirement that the federation provide utility no worse than local training. Using the above as the design target, we examine pathwise performance of FL, as a per-round bound on cumulative surplus rather than only an asymptotic equilibrium guarantee under different degrees of client data heterogeneity. Along the learning path, clients can remain below their local-training baseline for hundreds of rounds. The natural remedy is to cap each client's per-round contribution so that this shortfall stays bounded. We prove that such a cap drives all contributions to zero, with high probability, when its tolerance is small relative to the warm-up deficit, and we observe learning collapse under it even at low heterogeneity. We then propose a design that combines a short-term participation guarantee, enforced after an announced grace window, with personalized model evaluation, while preserving the incentive properties of the underlying mechanism. We provide a theoretical basis for this approach and empirically demonstrate that, after the grace window, clients meet IR on nearly all rounds without harming overall performance under low to moderate heterogeneity; under severe heterogeneity, the design shows promising outcomes for clients compared to their local baseline at some cost in accuracy.
|
| 2149 |
WaterKron and FlipFlop Hessian: Information-Theoretically Grounded Quantization with Kronecker-factored Hessians
2609.14706
|
cs.LG
|
Johann Birnick, Rayan Saab |
How should a Kronecker-factored Hessian approximation be chosen for post-training quantization? We address this question through WaterKron, which combines two-sided GPTQ with row- and column-dependent waterfilling scales and entropy coding. We derive its high-...How should a Kronecker-factored Hessian approximation be chosen for post-training quantization? We address this question through WaterKron, which combines two-sided GPTQ with row- and column-dependent waterfilling scales and entropy coding. We derive its high-rate distortion with respect to the full Hessian using an explicit Kronecker-Hessian mismatch factor $\Phi$. This factor quantifies the asymptotic distortion penalty due to the Kronecker Hessian approximation and provides a criterion for selecting the factors optimally. Minimizing $\Phi$ leads to a Gaussian covariance-fitting problem with classical ``flip-flop'' updates. We thus give a rate-distortion justification for using the resulting FlipFlop Hessian in quantization. We evaluate it empirically, finding that the FlipFlop Hessian consistently improves KL divergence and perplexity over input-only, marginal, and Frobenius-based Hessian choices.
|
| 2150 |
Refinement-Based Flow Policy Optimization
2609.15123
|
cs.LG
|
Bumgeun Park, Hyukjun Yang, Donghwan Lee |
Flow-based policies offer an expressive representation for online reinforcement learning, but conventional flow matching requires samples drawn from the distribution to be modeled. This poses a challenge when the desired action distribution is defined only imp...Flow-based policies offer an expressive representation for online reinforcement learning, but conventional flow matching requires samples drawn from the distribution to be modeled. This poses a challenge when the desired action distribution is defined only implicitly by a Q-function, since directly sampling actions from the resulting distribution is generally intractable. We propose Refinement-Based Flow Policy Optimization (RFPO), a novel framework for training a flow policy in online reinforcement learning by alternating between Q-guided sample refinement and self-target flow matching. RFPO first generates actions from Gaussian noise using the current flow policy and then uses a finite-step stochastic refinement procedure to move them toward an energy-based distribution induced by the Q-function. Each refined action is then paired with its corresponding initial noise sample and used as a fixed target for flow-matching training. By repeatedly refining its own outputs and learning from the resulting targets, RFPO incorporates Q-guidance into the policy without requiring direct samples from the target distribution, while retaining the capacity to represent multiple action modes. We further provide a theoretical analysis of the distributional dynamics induced by RFPO. Across six continuous-control tasks, RFPO matches or outperforms a standard Gaussian-policy baseline on almost every task. Experiments on six synthetic two-dimensional target distributions with diverse geometries demonstrate that RFPO captures complex multimodal structure without mode collapse.
|
| 2151 |
Sharp Regret Bounds and a Task-Covariance Correction for Spectral Representation Learning
2609.15825
|
cs.LG
|
Dier Tang, Jing Yee Tan, Guangyue Han |
Spectral features can remain optimal under strongly uneven task preferences when they retain the directions most useful to the tasks. In a local-task model, expected probing gain depends on the task prior only through its covariance $\Lambda$: with $B$ recordi...Spectral features can remain optimal under strongly uneven task preferences when they retain the directions most useful to the tasks. In a local-task model, expected probing gain depends on the task prior only through its covariance $\Lambda$: with $B$ recording dependence between views, $k$ spectral features span the leading eigenspace of $BB^\top$, while task-optimal features span that of $B\Lambda B^\top$. We derive alignment-dependent regret bounds, matching worst-case lower bounds for a flat leading spectrum, and bounds using the leading $2k$ directions with a spectral-tail term; a spectral gap makes the bound quadratic in small anisotropy. We also bound the imbalance from random preference patterns and from averaging independent tasks with an isotropic population covariance. With known $\Lambda$, changing one term of the spectral contrastive loss selects task-optimal features; with a labelled task bank, we propose a diagnostic and test correction without retraining. Correction reduces synthetic held-out regret from $0.858$ to $0.003$ with $2000$ tasks when less task-relevant directions dominate, and on CIFAR-100 lowers empirical regret by $0.031$--$0.053$ on held-out fine-label tasks using the same images. With a small task bank, however, correction can hurt.
|
| 2152 |
Scaling Laws for Physics-Aware ACOPF Surrogate Learning
2609.16282
|
cs.LG
|
Yijiang Li, Emon Dey, Stefano Fenu, Massimiliano Lupo Pasini, Teja Kuruganti |
Learning-based surrogates for AC optimal power flow (ACOPF) promise large speedups over classical solvers, but their operational value depends on physical feasibility as much as predictive accuracy. Physics-aware objectives such as the augmented Lagrangian (AL...Learning-based surrogates for AC optimal power flow (ACOPF) promise large speedups over classical solvers, but their operational value depends on physical feasibility as much as predictive accuracy. Physics-aware objectives such as the augmented Lagrangian (AL) improve constraint satisfaction at additional per-step cost, yet how this trade-off behaves with scale is uncharacterized. We sweep model and dataset sizes under MSE and AL training, and measure how violation changes with network size. Both objectives improve as power laws: MSE prediction loss depends more on model size than on data, while the AL composite of prediction loss and violation improves comparably along both axes. Violation grows markedly more slowly with network size under AL. On matched hardware and equal training samples, AL reduces violation about $19\times$ with $2.4\times$ the training time and negligible added memory. The training objective shapes not only where a surrogate lands but how its quality evolves with scale.
|
| 2153 |
Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not
2609.19363
|
cs.LG
|
Rub\'en Dar\'io Guerrero |
The query and key projections $\WQ,\WK$ in attention are almost always trained by Euclidean optimizers with no geometric constraint. We constrain them to the Stiefel manifold and optimize with a Riemannian Adam carrying one scalar second moment per frame---the...The query and key projections $\WQ,\WK$ in attention are almost always trained by Euclidean optimizers with no geometric constraint. We constrain them to the Stiefel manifold and optimize with a Riemannian Adam carrying one scalar second moment per frame---the form of \citet{becigneul2019}, here extended to the compact, non-Hadamard $\St(d,r)$ with a tangent projector, step-norm cap, and polar retraction. Four propositions prove steepest descent in the embedded metric, gradient-scale independence, well-conditioning, and exact $\mathrm{O}(d)$-equivariance. A fifth records that weight decay has \emph{identically zero} Riemannian gradient on $\St(d,r)$ ($W{=}WI_r$ lies in the normal space), so decay cannot act on the constrained frames. On a CIFAR-10 patch benchmark at $n{=}10\mathrm{k}$ this rule gains $\mathbf{+6.79}$\,pp over AdamW across 12 paired starts ($t{=}38.33$, $12/12$); earlier fixed-step Riemannian SGD gains $+1.97$\,pp, of which $+1.69$\,pp comes from frozen orthonormal initialization alone. The corrected Adam's lead grows with data: $+1.9$\,pp at $n{=}1\mathrm{k}$ to $+6.7$\,pp at $n{=}50\mathrm{k}$. A 12-seed ablation credits all gain to the scale-free step ($+4.63$\,pp, $12/12$), nothing to the projector or equivariance; a targeted $\varepsilon$-sweep causally confirms the mechanism ($-2.6$\,pp at $\varepsilon{=}0.1$, $p{<}0.001$). Two five-seed grokking studies confirm the constrained arm does not grok better than the baseline ($p{=}0.019$, A2 wins): the weight-decay exemption has no grokking consequence. A single-seed pilot exploiting this localization achieves the first stable grokking under slingshot conditions---Stiefel + targeted circuit regularization keeps routing-frame isometry error $10^6\times$ lower t
|
| 2154 |
Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks
2609.20269
|
cs.LG
|
Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi, Jaewon Jang |
Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choi...Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model ($\approx$2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a $7\times7$ Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a $4\times4$ Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16\%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59\% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68\% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16$\times$ larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63\% and removing the SSM-family mechanism produces a 3.20\% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.
|
| 2155 |
Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining
2609.21876
|
cs.LG
|
Ang-Kun Wu, Fangdi Wen, Jingtao Zhang |
As an alternative to the additive and extremal biases of average and max pooling, we introduce Geometric Mean Pooling (GMP), a signed pooling operator that combines the product of feature signs with the geometric mean of feature magnitudes. Motivated by local-...As an alternative to the additive and extremal biases of average and max pooling, we introduce Geometric Mean Pooling (GMP), a signed pooling operator that combines the product of feature signs with the geometric mean of feature magnitudes. Motivated by local-to-global composition in quantum many-body physics, GMP retains both joint sign information and a characteristic multiplicative scale without introducing learnable pooling parameters. We show that non-overlapping hierarchical GMP preserves the corresponding global multiplicative statistic and evaluate it on synthetic sequence tasks, iterative coarse-graining, image classification, and molecular lipophilicity regression. On the synthetic tasks, GMP recovers product-based signals more accurately than average and max pooling and maintains predictive performance under the tested levels of multiplicative input noise. On image and molecular data, however, its effectiveness depends on the representation, target parameterization, and placement of local and global pooling. These results position GMP as a complementary, regime-dependent inductive bias for tasks in which equal-weight multiplicative composition is plausible, rather than as a universal replacement for standard pooling operators.
|
| 2156 |
PRQuant: Permutation Residual Quantization for Low-Overhead Inference
2609.22106
|
cs.LG
|
Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming |
Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online op...Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (Permutation Residual Quantization), a training-free framework that combines channel permutation with offline weight residual compensation. PRQuant identifies the scaled-weight columns with the largest quantization errors and permutes them into contiguous tail blocks. This structure allows the corresponding weight residuals to be precomputed entirely offline, while replacing scattered activation gathering with simple contiguous access during inference, yielding a single regular MXFP4 GEMM for compensated computation. Experiments show that PRQuant substantially reduces down-projection reconstruction error, with scaling and residual compensation providing the main numerical gains while permutation enables a hardware-friendly contiguous layout. Comprehensive experiment results on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 illustrate that PRQuant achieves up to averagelly 2.6x and 1.8x operator speedup over BF16 respectively, while preserving near plain MXFP4 end-to-end decoding efficiency. Across five downstream benchmarks, PRQuant achieves the best average accuracy among the quantized methods, improving accuracy over MXFP4 by 1.24 and 0.55, respectively.
|
| 2157 |
A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting
2609.22836
|
cs.LG
|
Li Lin, Zhihao Lin, Qi Zhang, Kaiwen Xia, Shuai Wang |
Time series foundation models (TSFMs) have recently delivered impressive zero-shot performance across diverse forecasting tasks. However, real-world decision-making frequently relies on \emph{irregular multivariate time series} (IMTS), where inconsistent inter...Time series foundation models (TSFMs) have recently delivered impressive zero-shot performance across diverse forecasting tasks. However, real-world decision-making frequently relies on \emph{irregular multivariate time series} (IMTS), where inconsistent inter-observation intervals and asynchronous sampling across variables coexist with informative missingness. Existing TSFMs handle such inputs either through imputation that injects spurious values or through index-based positional encodings that ignore continuous time. There is still a gap in the foundation model that follows the original IMTS patterns. In this paper, we propose a hybrid attention model that learns a unified time-aware patch representation for IMTS forecasting. We first design a \emph{time-aware patch encoding} that maps a variable number of intra-patch timestamps into a fixed-size embedding, producing a uniform format for irregular patches without resorting to imputation. We then introduce a \emph{time bias attention} mechanism that calibrates inter-patch temporal misalignment and asynchronous cross-channel dependencies as auxiliary attention offset. Finally, on top of a decoder-only Transformer backbone, we adopt a \emph{hybrid causal mask} that preserves a bidirectional full view over the historical context while keeping the forecast horizon strictly autoregressive. To support large-scale pretraining under irregular settings, we also curate VersaTSA, an archive of $30$B observations that retains the native sampling sparsity of its sources. Experiments on three IMTS benchmarks and a standard regular-MTS benchmark show that our model achieves state-of-the-art zero-shot performance on IMTS and remains competitive when transferred to regular forecasting.
|
| 2158 |
Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?
2609.22850
|
cs.LG
|
Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng |
Claims about what an internal direction in an LLM represents need evidential constraints beyond an observer's prior beliefs about the system. Decodability and successful activation steering do not, by themselves, establish which construct the direction tracks....Claims about what an internal direction in an LLM represents need evidential constraints beyond an observer's prior beliefs about the system. Decodability and successful activation steering do not, by themselves, establish which construct the direction tracks. This gap is especially consequential for welfare-relevant interpretations, where a proposed functional state must be distinguished from correlated features of the extraction contrast. We treat the question as one of construct validity and study a good-bad outcome direction in a maze task, using controlled interventions that separate the realised outcome from the informational history through which it became known. Across multiple LLM checkpoints, directions fitted on one explicit outcome encoding transfer well to another, indicating that the readout is not tied to surface form. When the same realised outcome is reached through announced and unannounced histories, however, transfer degrades substantially: even after both histories receive the same explicit outcome, the post-event readout remains strongly conditioned on the earlier announcement. In the base model, steering along the direction changes actions, yet removing it leaves the natural cue effect almost intact. In a matched maze-RL run, the post-RL direction becomes substantially more predictive of reference-MDP remaining return and the policy becomes more dependent on it at the tested sites, while the history dependence persists. These dissociations support a functional, value-related interpretation of the direction, but not its identification with a history-invariant scalar valence state.
|
| 2159 |
WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
2609.23033
|
cs.LG
|
Hyeongju Ha, Jae-Joon Kim |
Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue...Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore concurrently batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B. The code is available at https://github.com/summerbro-hhj/wavefront-decoding.
|
| 2160 |
CSC: Calibrated Simplicity for Conflict-Aware Social Bot Detection in the LLM Era
2609.23320
|
cs.LG
|
Yipeng Qian, Pengjie Zhao, Chaoxi Niu |
Social bot detection is essential for protecting online platforms from misinformation amplification, coordinated manipulation, and distorted public discourse. However, large language models have made social bots much harder to detect from text alone because se...Social bot detection is essential for protecting online platforms from misinformation amplification, coordinated manipulation, and distorted public discourse. However, large language models have made social bots much harder to detect from text alone because semantic camouflage is now cheap, fluent, and scalable. The resulting challenge is modality conflict: an account may look human-like in semantics while remaining suspicious in graph structure, profile attributes, or cross-modal consistency. Recent graph-based detectors tackle this limitation by adding graph-side complexity, such as sparse prototype selection, adaptive gating, or architecture-specific control logic, yet our experiments suggest that complexity alone is not the most reliable way to resolve such conflict. We therefore propose CSC, a calibrated-simplicity framework for conflict-aware LLM-era social bot detection. The framework combines three design choices: a simplified prototype-guided graph expert that retains useful structural biases while removing unstable graph-side heuristics, calibrated simplex-constrained fusion that aligns heterogeneous confidence spaces before late fusion, and a lightweight inconsistency expert that models cross-modal disagreement. Experiments on TwiBot-22, TwiBot-20, and MGStBot-large show that \textsc{CSC} improves calibrated operating-point decision quality while remaining competitive across external benchmarks. Further analyses show that calibration improves confidence reliability, the inconsistency expert mainly provides localized corrections in high-conflict or near-threshold regions, and simplified graph-side control yields a better stability-cost trade-off. A targeted semantic-camouflage stress test further shows that replacing selected bot text with matched human text sharply degrades the standalone text expert while leaving graph and fused evidence stable on a balanced challenge set.
|
| 2161 |
Physics-residual machine learning predicts oxygen-evolution catalyst activity beyond the training range from sparse polarization measurements
2609.23549
|
cs.LG
|
Yong-Woon Kim, Jihyeok Lee, Sungtae Park, Sooseok Choi, Yung-Cheol Byun |
Discovery campaigns for oxygen evolution reaction catalysts repeatedly choose, make and measure catalysts. High-throughput platforms stop polarization curves below potentials that damage the catalyst, so the endpoint, the activity at a target potential or curr...Discovery campaigns for oxygen evolution reaction catalysts repeatedly choose, make and measure catalysts. High-throughput platforms stop polarization curves below potentials that damage the catalyst, so the endpoint, the activity at a target potential or current density, often lies beyond the measured window, and the catalysts of most interest are more active than any measured before. Existing methods do not predict these endpoints accurately when few or no endpoints of a new library have been measured. Here we present physics-residual machine learning (PR-ML), which predicts each endpoint as the sum of a Tafel term, computed from the catalyst's own measured curve with an estimated slope, and a residual term learned from labelled catalysts. In twelve Ni-Pd-Pt-Ru thin-film libraries, the current density at 1.70 V$_{RHE}$ was predicted from the currents at 1.40 and 1.55 V$_{RHE}$. Fitted only on earlier libraries, with ridge regression as the residual learner, PR-ML predicted the Ni--Ru library, whose currents mostly exceed theirs, with a mean absolute error of 0.194 mA cm$^{-2}$, against 1.330-1.882 for data-driven models. With five endpoints from the new library and extremely randomized trees as the residual learner, PR-ML gave a similar error, which the same learner used alone reached only with 20, and identified 63-83% of the catalysts more active than the best labelled catalyst, against 2%. In two independent datasets, this fraction rose from at most 1% to 33-95%. Our approach supplies catalyst selection with accurate endpoints beyond the measured part of each curve and above all earlier measurements.
|
| 2162 |
Tail-Weight Control and Localized Generalization in Nearly Low-Rank Adversarial Classification
2609.23688
|
cs.LGcs.AI
|
Kunyu Wang, Dehan Wang, Wenjun Chen |
Empirical ramp fitting can assign weight to pure-noise features even when the population optimum ignores them. We quantify this gap for norm-constrained adversarial classification with Gaussian signal and noise. The variance cost relative to normalized signed ...Empirical ramp fitting can assign weight to pure-noise features even when the population optimum ignores them. We quantify this gap for norm-constrained adversarial classification with Gaussian signal and noise. The variance cost relative to normalized signed mean separates into two factors: selecting observations inside the active margin window and the curvature induced by the norm constraint. Changing the tail variance leaves the activewindow probability unchanged but changes the second factor. With positive attack budget and a signal-only predictor of risk below one half, we prove a uniform quadratic tail-deletion bound, including at zero tail variance. Sufficiently accurate approximate global empirical minimizers admit exact fixeddimensional asymptotic covariances in the low-risk regime with isotropic principal covariance. For positive tail variance at most principal variance, the product exceeds one; an additional moment condition transfers it to expected excess ramp and robust classification risks. A wide window analysis characterizes when this ordering reverses. Controlled experiments test the decomposition, and a separate contamination study examines its scope outside the Gaussian training model.
|
| 2163 |
Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models
2609.23999
|
cs.LG
|
Star S. D. Liu, Xiyu Ding, Robert B. Barrett, Alberto Santamaria-Pang, Nic Dobbins |
How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations rela...How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC $\approx 0.83$), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.
|
| 2164 |
The Undetected Damage of Quantization on Retrieval and How to Fix It
2609.24322
|
cs.LG
|
Luca Zhou, Alessandro Zirilli, Daniele Solombrino, Roberto Dess\`i, Emanuele Rodol\`a |
We show that a quantized model that keeps its classification accuracy still changes $14$ to $46\%$ of its top-1 retrieval results, and that aggregate ranking metrics reveal only part of this damage. We tie this failure to the gap between the two highest model ...We show that a quantized model that keeps its classification accuracy still changes $14$ to $46\%$ of its top-1 retrieval results, and that aggregate ranking metrics reveal only part of this damage. We tie this failure to the gap between the two highest model scores and use that gap to decide when to trust a quantized answer and where additional precision should be spent. We show that the top-1 result is guaranteed to survive quantization only when this gap exceeds twice the largest rounding error. In classification, the scores are logits, and training compares the correct class against every other class, which encourages this gap. In retrieval, the scores are query-document similarities, and training compares each positive only against sampled negatives, so nothing separates the top-1 item from the second. This gap can be measured without labels. Before deployment, it predicts which models will break under quantization, and at deployment time it tells, per input, whether the quantized answer still matches the full-precision answer. Most classification inputs have a gap wide enough to trust the quantized answer, but few retrieval queries do. That gap motivates a different fix in each task. In retrieval, spending extra bit-width on the layers whose quantization moves the gap most recovers up to three-quarters of an extra bit's benefit for half its cost. In classification, routing the few low-gap inputs to full precision recovers most of the lost accuracy at a fraction of the cost.
|
| 2165 |
iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs
2609.24646
|
cs.LG
|
Ahmed Khaled Khamis, Xiaotong Ji, Hassan Jaber, Rasul Tutunov, Matthieu Zimmer |
On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no con...On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state. We introduce Information-Proximal SDFT (iSDFT), which instead treats the teacher as a budgeted source of information. At each token, iSDFT selects the distribution closest to the current student that satisfies a prescribed teacher-information constraint, yielding a closed-form exponential target with a locally determined tilt. To control cumulative drift, we further anchor the student to its frozen base policy. Across four heterogeneous LLM backbones and two specialisation tasks, iSDFT improves vanilla SDFT in 7 of 8 model-task settings and matches it in the remaining one. It also provides tighter retention on the original SDFT benchmark suite, with 73% of evaluations remaining within 0.5 points of the base model versus 52% for the strongest baseline, while achieving the largest mean improvement on all ten additional mathematics, coding, and competition-mathematics benchmarks. These results show that controlling how much and when teacher information is introduced improves specialisation while preserving broader capability.
|
| 2166 |
What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation
2609.25623
|
cs.LG
|
Kanghui Tian, Siyuan Liu, Tianxiang Jiang, Shuai Dong, Yizhuo Li |
More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a self-teacher scores the student's own rollouts under privileged context, conventionally a complete reference solution that b...More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a self-teacher scores the student's own rollouts under privileged context, conventionally a complete reference solution that bundles the final answer with one particular reasoning path. Holding the student view and training fixed within each scale, we compare that default against three abstractions compiled offline, a named strategy, a method-independent framing, and a problem category, and against an answer-only control that keeps the destination but removes the path. In the primary runs on competition mathematics, the best intermediate contexts improve the in-domain peak mean over the full solution by 1.4 points at 4B and 1.6 at 8B, while storing an order of magnitude fewer hint tokens. Comparisons across three seeds also show positive mean gains for the framing and category contexts at both scales. Answer-only conditioning remains competitive in the primary runs, within 0.2 points of the full solution at these scales. The preferred context varies with student scale and task. Initial teacher-student KL does not order downstream performance. What a self-teacher should see is therefore not everything it could, but the level of abstraction its student can still act on.
|
| 2167 |
Graph Domain Adaptation Does Not End with Representation Learning
2609.25692
|
cs.LG
|
Ziqian Liu, Yongxue Xu, Enze Zhang, Jiaqi Zhang, Hao Wang |
Graph domain adaptation (GDA) transfers knowledge from a labeled source graph to an unlabeled target graph under shifts in both node attributes and graph structure. Existing methods primarily adapt graph representations through propagation redesign, distributi...Graph domain adaptation (GDA) transfers knowledge from a labeled source graph to an unlabeled target graph under shifts in both node attributes and graph structure. Existing methods primarily adapt graph representations through propagation redesign, distribution alignment, or source-to-target transition modeling, but still rely on a single graph-propagating path for target prediction. This leaves open whether an adapted graph representation exhausts the predictive evidence available in the target domain, since the graph-aware expert and graph-free local expert may exhibit different failure modes under topological shifts. To address this limitation, we propose EviGDA, an Evidence-Augmented Graph Domain Adaptation framework that complements graph representation adaptation with a graph-free local expert. The graph-aware expert performs message passing and entropy-aware marginal alignment, while the graph-free local expert learns solely from source node features and labels without graph propagation or target alignment. The two experts are optimized independently and combined only at inference through a task-level constant probability mixture, preserving complementary evidence without joint training, learned routing, or target pseudo-labels. Extensive experiments on ten datasets and 16 transfer tasks show that EviGDA outperforms state-of-the-art baselines. Complementary prediction paths add value beyond graph representation alignment.
|
| 2168 |
In-Context Guidance: Learning Inter-Task Synergies via Numerical Foundational Models for Few-Shot Multitask Optimization
2609.25836
|
cs.LG
|
Tingyang Wei, Haofeng Wu, Jiao Liu, Zhao Wei, Puay Siew Tan |
Multi-task optimization (MTO) addresses a set of optimization tasks simultaneously, often suffering from inaccurate inter-task relationship estimation under limited evaluation budgets, leading to negative transfer. This paper introduces In-Context Guidance Mul...Multi-task optimization (MTO) addresses a set of optimization tasks simultaneously, often suffering from inaccurate inter-task relationship estimation under limited evaluation budgets, leading to negative transfer. This paper introduces In-Context Guidance Multitask Optimization (ICG-MTO), a novel framework that leverages numerical foundational models to improve inter-task coupling estimation in few-shot scenarios. Unlike conventional methods that rely solely on scarce observed data, ICG-MTO employs a frozen foundational model to infer auxiliary guidance through in-context learning. The framework operates through three stages: constructing an algorithm-specific in-context query from evaluated solutions, using the foundational model to infer a guidance signal characterizing predictive relationships among tasks, and translating this signal into algorithm-specific guidance for maximum-a-posteriori coupling estimation. This approach provides regularization during the early, data-scarce stages of optimization and gradually relinquishes control as task-specific observations accumulate. We instantiate the framework in multitask Bayesian optimization as ICG-MTBO, using directional fitness-class queries to guide inter-task coupling estimation, and further instantiate it in MFEA-II using decision-space-overlap queries to guide random mating probability estimation. Experiments across synthetic benchmarks and a real-world robot arm control problem, together with evaluations under different acquisition functions and evolutionary multitasking, demonstrate the effectiveness and generality of ICG-MTO for few-shot multitask optimization.
|
| 2169 |
PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation
2609.26890
|
cs.LG
|
Yuta Tarumi |
Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of...Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of the prescribed simulator. We introduce PR-Smoother, a simulator-preserving amortized smoother designed for this prescribed-simulator DA regime. Its key design principle is to keep the prescribed simulator explicit in both the evidence lower bound and the variational family: rather than learning replacement dynamics or a learned trajectory prior, PR-Smoother learns only future-conditioned corrections around the prescribed rollout. This yields an explicit non-Gaussian smoothing distribution over physical trajectories and supports joint state, parameter, and sensor-bias learning from observations alone. The variational family contains the exact smoother in deterministic and linear-Gaussian limits. Empirically, PR-Smoother captures multimodal posteriors in 4-dimensional Lorenz-96, remains accurate under ambiguous nonlinear observations and process noise in 40-dimensional Lorenz-96, and scales to joint state-parameter-bias inference in 16,384-dimensional Kolmogorov flow.
|
| 2170 |
Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated Learning
2609.27760
|
cs.LG
|
Srinivasan Subramanian, Md. Abdullah Al Hafiz Khan, Kazi Aminul Islam |
Federated learning enables distributed training without requiring clients to share their raw data. However, its reliance on the integrity of the client-submitted updates exposes the global model to stealthy backdoor poisoning. Existing defenses often rely on i...Federated learning enables distributed training without requiring clients to share their raw data. However, its reliance on the integrity of the client-submitted updates exposes the global model to stealthy backdoor poisoning. Existing defenses often rely on individual evidence sources, but stealth-constrained attacks can adapt to these signals. Such attacks can suppress anomaly signals they are optimized to evade, yet their poisoned updates still leave residual structural traces. We propose FedMAST, a Federated Multi-Axis Structural Tracing defense for backdoor detection in federated learning. FedMAST scores client updates using complementary structural, spectral, and historical evidence and then applies tiered filtering and round-level containment to limit adversarial influence. To capture traces that isolated signals may miss, FedMAST uses squeeze-pair coherence scoring to expose coupled feature distortions and signed spectral-drift tracking to reveal persistent directional changes over time. Across six backdoor attacks, FedMAST achieves lower attack success rate (ASR) than baseline defenses in all nine evaluated comparisons, averaging 1.51% ASR and 94.84% main-task accuracy (MTA) across the complete 200-round runs. Over the full 200-round method-aware CovertLayers run, FedMAST achieves 1.53% ASR and 92.26% MTA, compared with ASRs of 100.00%, 99.67%, 99.53%, and 32.84% for FedAvg, MultiKrum, AlignIns, and FLAME, respectively.
|
| 2171 |
Physics and Data Driven Transformer-Mamba Framework for Flow Field
2609.29087
|
cs.LG
|
Zhuo Zhang, Shun Zou, Canqun Yang, Xi Yang |
While deep learning accelerates expensive partial differential equation solving in computational fluid dynamics (CFD), existing methods like PINNs and FNOs often struggle with generalization, noise robustness, and physical consistency. We introduce the Transfo...While deep learning accelerates expensive partial differential equation solving in computational fluid dynamics (CFD), existing methods like PINNs and FNOs often struggle with generalization, noise robustness, and physical consistency. We introduce the Transformer-Mamba for Flow Field (TM4FF) framework, a physics-constrained operator learning model with three key innovations: a Residual Wavelet Mamba (RWM) layer for feature denoising, a Transformer-based attention mechanism for enhanced feature fusion, and a physics-informed loss using Fourier derivatives to enforce the Navier-Stokes equations. Experiments on four CFD datasets show TM4FF achieves high accuracy and robust generalization across varying flow conditions.
|
| 2172 |
Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD
2609.29142
|
cs.LG
|
Yibo Zhao, Zixuan Yang, Yunshi Lan, Xiang Li |
Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the student's own r...Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the student's own rollouts. This transfer rewards the policy shift at every state, yet the log-ratio measures only relative change: it can stay fixed even as the probability mass that both checkpoints assign to the student's candidate tokens vanishes. Through an exact construction, we show that the Direct-OPD reward and its update can remain unchanged while the Jensen-Shannon divergence (JSD) and both KL directions between the checkpoints vanish with this mass, and we note that a small JSD bounds how much the teacher's behavior changed. Motivated by this analysis, we propose Selective Supervision for Direct-OPD (S$^2$D-OPD), which ranks student-sampled states by their teacher-reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% of states per response. Across two teacher pairs and four student models ranging from 1.7B to 8B parameters, S$^2$D-OPD improves held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings and matches it in the eighth, without extra forward passes.
|
| 2173 |
Beyond Feature Reliability: Repeat-Informed Multifractal Curve Regression for Brain-Age Prediction
2609.29307
|
cs.LG
|
Yu Chang, Anzhe Cheng, Jiahao Chen, Heng Ping, Peiyu Zhang |
Brain-age prediction from resting-state fMRI provides a quantitative framework for characterizing age-related changes in spontaneous brain dynamics and for identifying functional signatures. Existing studies have linked fractal and multifractal scaling to age ...Brain-age prediction from resting-state fMRI provides a quantitative framework for characterizing age-related changes in spontaneous brain dynamics and for identifying functional signatures. Existing studies have linked fractal and multifractal scaling to age and examined the reliability of individual features. However, prediction repeatability depends on how features fluctuate jointly and how a predictor combines them, which feature-wise reliability assessments do not capture. To address this problem, we propose Repeat-informed Multifractal Curve Regression (RMCR), a structured framework for learning stable age-predictive patterns from multifractal curves. By jointly modeling curve structure and repeat-scan variability, RMCR learns predictive combinations of fluctuation orders that target both accuracy and within-subject consistency. Relative to a matched run-level ridge baseline, RMCR reduces single-run MAE by 6.1% on HCP-A and 7.9% on an external Cam-CAN cohort, and within-visit repeat absolute difference by 18.5% on HCP-A, using a single scan at inference.
|
| 2174 |
FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates
2609.29812
|
cs.LG
|
Wanqi Yang, Shiwei Liu |
Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain un...Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, FlashLoop delivers lossless accuracy while achieving up to 1.64$\times$ end-to-end speedup and up to 6$\times$ KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.
|
| 2175 |
Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation
2609.29931
|
cs.LG
|
Nathan Le, Magdalini Paschali, Arogya Koirala, Andrew Johnston, Zhongnan Fang |
Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities accurately reflect true risk. M...Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities accurately reflect true risk. Many standard techniques for improving calibration, such as MC Dropout and Deep Ensembles, require access to model parameters or retraining. However, proprietary clinical AI systems operate as black boxes, preventing access to the model's internals. To that end, we propose a model-agnostic framework for improving calibration of black-box models using clinically grounded test-time augmentation (TTA). Our framework applies geometric and physics-inspired 3D CT perturbations and learns probability-level aggregation strategies without access to model internals or the original training data. Across pulmonary embolism and intracranial hemorrhage detection tasks, DualTTA achieved the strongest overall calibration among TTA methods, reducing the Expected Calibration Error by 54% (0.239 -> 0.109) and 43% (0.051 -> 0.029), respectively, while requiring only input-output access. Additionally, DualTTA outperformed uncertainty estimation techniques that require access to model internals, such as Temperature Scaling, MC Dropout, and Deep Ensembles, in most calibration metrics. These results demonstrate that learned TTA aggregation can improve the calibration of clinical AI systems, providing a practical approach for improving the reliability of black-box medical AI.
|
| 2176 |
MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
2609.30837
|
cs.LG
|
Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma |
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels res...Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model. Its plug-in interface supports different metrics for selecting and weighting teacher-specific OPD signals. Within this interface, we propose ExpertAlign, which scores each teacher by whether its correction to the student at the current token expresses the specialization that teacher acquired during post-training, and compare it against two reference metrics built on teacher confidence (Entropy) and teacher-student discrepancy (Novelty). Experiments on unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation scenarios show that ExpertAlign achieves the strongest overall performance in all four settings. On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels. These results demonstrate token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment. Code is available at: https://github.com/TURLEing/MOPD-Router.
|
| 2177 |
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
2609.30918
|
cs.LG
|
Marcin Kostrzewa, Maciej Zi\k{e}ba |
Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are dif...Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. The benchmark compares six robust methods and two standard baselines on four tabular datasets. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95\% of test predictions on average, compared with 4.9\% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.
|
| 2178 |
Softmax Reparameterization for Output-Head Quantization
2609.31291
|
cs.LG
|
Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng |
Large vocabularies make output heads a substantial inference cost in small language models. We introduce softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization. The method subtracts a scal...Large vocabularies make output heads a substantial inference cost in small language models. We introduce softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL. For linear-softmax heads, these shifts preserve full-precision predictions exactly and require no decoder retraining; a rank-one correction extends the construction to nonlinear logit paths. Across seven output heads and three quantizers, W4 gains are largest where baseline quantization substantially distorts predictions: test KL falls by 93% on XGLM under RTN and by 73--77% on Phi, BLOOM, and BLOOMZ under activation-weighted MSE. Heads with low baseline error change little; at W2, used as a compression stress test, benefits extend more broadly. On Phi, the gains persist under stronger GPTQ calibration; a separate untouched holdout reproduces the improvements on Phi and BLOOM. Frozen WikiText-selected coefficients also transfer without retuning to C4 and OpenWebMath. Residual analysis on Phi shows how fidelity can improve despite greater total logit error: the selected representative reduces error on likely outputs and lowers its Fisher-weighted cost. For shift-compatible heads, the shift adds no inference operation. With the decoder held in BF16, a packed W4 Phi output head reduces batch-one generation latency by 10.8%, and reparameterization preserves this speedup.
|
| 2179 |
Benchmarking Attention for Tabular Foundation Models
2609.31306
|
cs.LG
|
Maximilian Schambach, Clemens Biehl, Sam Thelin |
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention invol...Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and SageAttention -- measuring forward and backward throughput across realistic tabular shapes on three GPU generations (A100, H100, B200). We find that the optimal backend choice differs between column and row attention and varies across hardware as well as model specifics: While the FlashAttention implementations tailored for each GPU generation perform overall best, they are at times outperformed by CuDNN in the case of column attention at longer sequences with cross-over points depending on the head dimension. Among inference-only backends, SageAttention performs well for row attention and large sequences beyond 16\,k rows. Our reproducible benchmark lays the foundation for future improvements to table-native attention. The self-contained benchmarking and evaluation code is openly available at: https://github.com/SAP-samples/tabular-attention-benchmark
|
| 2180 |
Bridging Generative and Discriminative Noisy-Label Learning via Direction-Agnostic EM Formulation
2308.01184
|
cs.LG
|
Fengbei Liu, Chong Wang, Yuanhong Chen, Yuyuan Liu, Gustavo Carneiro |
Although noisy-label learning is often approached with discriminative methods for simplicity and speed, generative modeling offers a principled alternative by capturing the joint mechanism that produces features, clean labels, and corrupted observations. Howev...Although noisy-label learning is often approached with discriminative methods for simplicity and speed, generative modeling offers a principled alternative by capturing the joint mechanism that produces features, clean labels, and corrupted observations. However, prior work typically (i) introduces extra latent variables and heavy image generators that bias training toward reconstruction, (ii) fixes a single data-generating direction (\(Y\rightarrow\!X\) or \(X\rightarrow\!Y\)), limiting adaptability, and (iii) assumes a uniform prior over clean labels, ignoring instance-level uncertainty. We propose a single-stage, EM-style framework for generative noisy-label learning that is \emph{direction-agnostic} and avoids explicit image synthesis. First, we derive a single Expectation-Maximization (EM) objective whose E-step specializes to either causal orientation without changing the overall optimization. Second, we replace the intractable \(p(X\mid Y)\) with a dataset-normalized discriminative proxy computed using a discriminative classifier on the finite training set, retaining the structural benefits of generative modeling at much lower cost. Third, we introduce \emph{Partial-Label Supervision} (PLS), an instance-specific prior over clean labels that balances coverage and uncertainty, improving data-dependent regularization. Across standard vision and natural language processing (NLP) noisy-label benchmarks, our method achieves state-of-the-art accuracy, lower transition-matrix estimation error, and substantially less training compute than current generative and discriminative baselines. Code: https://github.com/lfb-1/EMNL
|
| 2181 |
Damped Gauss Newton Search for Multi Metric Hyperparameter Optimization
2401.03580
|
cs.LG
|
Qinwu Xu, Zhanyu Ge, Yifan Jiang |
We study hyperparameter optimization (HPO) from a numerical-optimization perspective and propose a black-box, target-seeking damped Gauss--Newton method for multiple validation responses. The method estimates trajectory-based numerical sensitivities from succe...We study hyperparameter optimization (HPO) from a numerical-optimization perspective and propose a black-box, target-seeking damped Gauss--Newton method for multiple validation responses. The method estimates trajectory-based numerical sensitivities from successive changes in hyperparameters and validation performance. These full-vector observations construct an iterative secant approximation of the local sensitivity, and a damped Gauss--Newton system jointly updates all hyperparameters. The method does not differentiate through model training or require coordinate-wise perturbation runs: after two initial evaluations establish the first secant, each subsequent iteration requires one new full-vector evaluation. Tikhonov regularization stabilizes the underdetermined local inverse problem. Unlike multi-objective HPO that approximates a Pareto front, our formulation seeks a specified operating point in multi-response performance space. We evaluate the method on four-dimensional XGBoost HPO across three classification datasets, six-dimensional LoRA/SFT post-training of Qwen3-VL-8B on a binary VQAv2 subset, and eight-dimensional threshold optimization on fixed classifier outputs. XGBoost results are comparable to grid, random, and TPE search with fewer model evaluations. The VLM study reveals non-monotonic response and sensitivity trajectories, while the threshold study demonstrates target seeking in an underdetermined system and sensitivity to initialization. Overall, the results support trajectory-based secant sensitivity as an evaluation-efficient local optimization mechanism complementary to global black-box HPO.
|
| 2182 |
Revisiting Inexact Fixed-Point Iterations for Min-Max Problems: Stochasticity and Structured Nonconvexity
2402.05071
|
cs.LG
|
Ahmet Alacaoglu, Donghwan Kim, Stephen J. Wright |
We focus on constrained, $L$-smooth, potentially stochastic and nonconvex-nonconcave min-max problems either satisfying $\rho$-cohypomonotonicity or admitting a solution to the $\rho$-weakly Minty Variational Inequality (MVI), where larger values of the parame...We focus on constrained, $L$-smooth, potentially stochastic and nonconvex-nonconcave min-max problems either satisfying $\rho$-cohypomonotonicity or admitting a solution to the $\rho$-weakly Minty Variational Inequality (MVI), where larger values of the parameter $\rho>0$ correspond to a greater degree of nonconvexity. These problem classes include examples in two player reinforcement learning, interaction dominant min-max problems, and certain synthetic test problems on which classical min-max algorithms fail. It has been conjectured that first-order methods can tolerate a value of $\rho$ no larger than $\frac{1}{L}$, but existing results in the literature have stagnated at the tighter requirement $\rho < \frac{1}{2L}$. With a simple argument, we obtain optimal or best-known complexity guarantees with cohypomonotonicity or weak MVI conditions for $\rho < \frac{1}{L}$. First main insight for the improvements in the convergence analyses is to harness the recently proposed $\textit{conic nonexpansiveness}$ property of operators. Second, we provide a refined analysis for inexact Halpern iteration that relaxes the required inexactness level to improve some state-of-the-art complexity results even for constrained stochastic convex-concave min-max problems. Third, we analyze a stochastic inexact Krasnosel'ski\u{\i}-Mann iteration with a multilevel Monte Carlo estimator when the assumptions only hold with respect to a solution.
|
| 2183 |
Generalization Analysis of Online Stochastic Gradient Descent for Overparameterized Two-Layer Neural Networks
2407.07670
|
cs.LG
|
Dinghao Cao, Zheng-Chu Guo, Lei Shi |
We investigate the generalization performance of online stochastic gradient descent (SGD) for overparameterized two-layer neural networks under the Neural Tangent Kernel (NTK) regime. Leveraging the NTK approximation together with the convergence theory of sto...We investigate the generalization performance of online stochastic gradient descent (SGD) for overparameterized two-layer neural networks under the Neural Tangent Kernel (NTK) regime. Leveraging the NTK approximation together with the convergence theory of stochastic approximation in reproducing kernel Hilbert spaces (RKHSs), we develop a unified framework that connects the optimization dynamics of neural networks with kernel-based statistical learning. This framework enables us to derive sharp convergence rates for the \emph{generalization error of the last iterate} of online SGD. The obtained rates coincide with the optimal statistical rates for kernel methods under appropriate source assumptions on target function. In contrast to existing NTK analyses, which primarily establish optimization convergence or analyze averaged SGD, our results directly characterize the statistical behavior of the practically implemented last-iterate online SGD algorithm for streaming data. Moreover, our analysis substantially relaxes the required degree of overparameterization by reducing the network-width requirement from exponential to polynomial dependence on the sample size or the number of optimization iterations. These results provide a unified theoretical perspective on stochastic optimization, kernel methods, and statistical learning for overparameterized neural networks, while significantly narrowing the gap between existing theory and practical deep learning.
|
| 2184 |
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
2407.08983
|
cs.LG
|
David N. Palacio, Daniel Rodriguez-Cardenas, Alejandro Velasco, Dipin Khati, Kevin Moran |
Trustworthiness and interpretability are inextricably linked concepts for LLMs. The more interpretable an LLM is, the more trustworthy it becomes. However, current techniques for interpreting LLMs when applied to code-related tasks largely focus on accuracy me...Trustworthiness and interpretability are inextricably linked concepts for LLMs. The more interpretable an LLM is, the more trustworthy it becomes. However, current techniques for interpreting LLMs when applied to code-related tasks largely focus on accuracy measurements, measures of how models react to change, or individual task performance instead of the fine-grained explanations needed at prediction time for greater interpretability, and hence trust. To improve upon this status quo, this paper introduces ASTrust, an interpretability method for LLMs of code that generates explanations grounded in the relationship between model confidence and syntactic structures of programming languages. ASTrust explains generated code in the context of syntax categories based on Abstract Syntax Trees and aids practitioners in understanding model predictions at both local (individual code snippets) and global (larger datasets of code) levels. By distributing and assigning model confidence scores to well-known syntactic structures that exist within ASTs, our approach moves beyond prior techniques that perform token-level confidence mapping by offering a view of model confidence that directly aligns with programming language concepts with which developers are familiar. To put ASTrust into practice, we developed an automated visualization that illustrates the aggregated model confidence scores superimposed on sequence, heat-map, and graph-based visuals of syntactic structures from ASTs. We examine both the practical benefit that ASTrust can provide through a data science study on 12 popular LLMs on a curated set of GitHub repos and the usefulness of ASTrust through a human study.
|
| 2185 |
Detection and Characterization of Coordinated Online Behavior: A Survey
2408.01257
|
cs.LG
|
Lorenzo Mannocci, Michele Mazza, Anna Monreale, Maurizio Tesconi, Stefano Cresci |
Coordination is a fundamental aspect of life. The advent of social media has made it integral also to online human interactions, such as those that characterize thriving online communities and social movements. At the same time, coordination is also core to ef...Coordination is a fundamental aspect of life. The advent of social media has made it integral also to online human interactions, such as those that characterize thriving online communities and social movements. At the same time, coordination is also core to effective disinformation, manipulation, and hate campaigns. This survey collects, categorizes, and critically discusses the body of work produced as a result of the growing interest on coordinated online behavior. We reconcile industry and academic definitions, propose a comprehensive framework to study coordinated online behavior, and review and critically discuss the existing detection and characterization methods. Our analysis identifies open challenges and promising directions of research, serving as a guide for scholars, practitioners, and policymakers in understanding and addressing the complexities inherent to online coordination. We also provide an interactive companion website for exploring the surveyed literature.
|
| 2186 |
A Survey on Reinforcement Learning Applications in SLAM
2408.14518
|
cs.LG
|
Mohammad Dehghani Tezerjani, Mohammad Khoshnazar, Mohammadhamed Tangestanizadeh, Arman Kiani, Qing Yang |
Simultaneous localization and mapping (SLAM) allows a mobile robot or autonomous vehicle to build a map of an unknown environment while estimating its own pose within that map. Reinforcement learning (RL), in which an agent learns a decision policy from intera...Simultaneous localization and mapping (SLAM) allows a mobile robot or autonomous vehicle to build a map of an unknown environment while estimating its own pose within that map. Reinforcement learning (RL), in which an agent learns a decision policy from interaction and reward, has been applied to decide how such systems move, explore, and recognize places they have visited before. This survey reviews the applications of RL in SLAM. We first distinguish passive SLAM, in which the robot's motion is not chosen by the SLAM system, from active SLAM, in which it is, and summarize the sensors that provide the input to SLAM. We then introduce the RL methods used in this literature, from value-based and policy-based methods to actor-critic and deep RL. Next, we classify RL applications in SLAM into three categories: path planning, including environment exploration and obstacle avoidance; loop closure detection; and active SLAM. Thirteen representative studies are compared in terms of their simulation environment, deep learning method, SLAM method, and RL algorithm. Most of these studies are evaluated mainly in simulation, and value-based methods from the deep Q-network family are the most common. Finally, we discuss the challenges of applying RL to SLAM, namely computational demands, safety, generalization, high-dimensional state and action spaces, sample efficiency, and sensor and actuator delays, and we outline directions for future research.
|
| 2187 |
Help Me Help You: The Aggregate Value of Source and Target Data in Transfer Learning
2408.16189
|
cs.LG
|
Steve Hanneke, Samory Kpotufe |
Transfer learning, in regression or classification, concerns problems where one aims to use imperfect data from a source distribution $P$ to improve prediction w.r.t. a target distribution $Q$. While there is now much theoretical understanding of the unsupervi...Transfer learning, in regression or classification, concerns problems where one aims to use imperfect data from a source distribution $P$ to improve prediction w.r.t. a target distribution $Q$. While there is now much theoretical understanding of the unsupervised setting of this problem (where no labeled target data is available), the supervised setting with labeled target data remains under-studied. We identify and characterize an interesting dichotomy in the rates achievable in the supervised setting: there is a weak transfer regime where optimal rates are determined by information in either the source or target dataset but not both, and a strong transfer regime where much faster target rates are achievable, accounting for complementary information in source and target data. Furthermore, while previous theoretical tools for the unsupervised setting can be extended to capture the weak transfer regime, we argue that they cannot capture the strong regime due to limitations inherent in the various notions of relatedness they rely on. Instead, we provide an appropriate refinement of relatedness notions---termed weak and strong moduli of transfer---that reveal the two regimes and lead to a nontrivial characterization of which pairs of distributions $P$ and $Q$ adhere to which regime. Finally, the above dichotomy reveals a new categorization of existing algorithmic approaches in terms of their optimality or lack-thereof across supervised regimes, and we provide a generic approach that remains nearly optimal in all regimes without prior distributional knowledge.
|
| 2188 |
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
2412.01488
|
cs.LG
|
Hugo Malard, Michel Olvera, Stephane Lathuiliere, Slim Essid |
Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the specific task of sound-prompted segmentation, aiming to segment image regions corre...Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the specific task of sound-prompted segmentation, aiming to segment image regions corresponding to objects heard in an audio signal. Most existing approaches tackle this problem by fine-tuning pre-trained models or by training additional modules specifically for the task. We adopt a different strategy: we introduce a training-free approach that leverages Non-negative Matrix Factorization (NMF) to co-factorize audio and visual features from pre-trained models so as to reveal shared interpretable concepts. These concepts are passed on to an open-vocabulary segmentation model for precise segmentation maps. By using frozen pre-trained models, our method achieves high generalization and establishes state-of-the-art performance in unsupervised sound-prompted segmentation, significantly surpassing previous unsupervised methods.
|
| 2189 |
Coreset-Based Task Selection for Sample-Efficient Meta-Reinforcement Learning
2502.02332
|
cs.LG
|
Donglin Zhan, Leonardo F. Toso, James Anderson |
We study task selection to enhance sample efficiency in model-agnostic meta-reinforcement learning (MAML-RL). Traditional meta-RL typically assumes that all available tasks are equally important, which can lead to task redundancy when they share significant si...We study task selection to enhance sample efficiency in model-agnostic meta-reinforcement learning (MAML-RL). Traditional meta-RL typically assumes that all available tasks are equally important, which can lead to task redundancy when they share significant similarities. To address this, we propose a coreset-based task selection approach that selects a weighted subset of tasks based on how diverse they are in gradient space, prioritizing the most informative and diverse tasks. Such task selection reduces the number of samples needed to find an $\epsilon$-close stationary solution by a factor of O(1/$\epsilon$). Consequently, it guarantees a faster adaptation to unseen tasks while focusing training on the most relevant tasks. As a case study, we incorporate task selection to MAML-LQR (Toso et al., 2024b), and prove a sample complexity reduction proportional to O(log(1/$\epsilon$)) when the task specific cost also satisfy gradient dominance. Our theoretical guarantees underscore task selection as a key component for scalable and sample-efficient meta-RL. We numerically validate this trend across multiple RL benchmark problems, illustrating the benefits of task selection beyond the LQR baseline.
|
| 2190 |
Learning Where and What to Restore for Composite Image Restoration
2503.17915
|
cs.LG
|
Jiachen Jiang, Tianyu Ding, Ke Zhang, Jinxin Zhou, Tianyi Chen |
Real-world degraded images often contain multiple co-occurring degradation types, making composite image restoration fundamentally different from the single-degradation setting assumed by most existing all-in-one methods. These methods typically apply uniform ...Real-world degraded images often contain multiple co-occurring degradation types, making composite image restoration fundamentally different from the single-degradation setting assumed by most existing all-in-one methods. These methods typically apply uniform spatial computation and single-label task conditioning, limiting both spatial adaptivity and explicit modeling of degradation mixtures. We propose CART (Composite-Adaptive Routing and Task Conditioning), a unified framework that jointly decides where to spend computation and what degradation cues should guide restoration: a spatial mixer routes patches to graded-capacity experts by local restoration difficulty, and a channel mixer routes each image through degradation-specific experts, conditioned on the global task feature with multi-hot classification supervision. Empirically, CART achieves state-of-the-art performance for composite degradation restoration on CDD-11 and delivers the best results to date on the conventional 3-task and 5-task all-in-one benchmarks.
|
| 2191 |
Optimization on the Oblique Manifold for Sparse Simplex Constraints via Multiplicative Updates
2503.24075
|
cs.LG
|
Flavia Esposito, Andersen Ang |
Low-rank optimization problems with sparse simplex constraints involve variables that must satisfy nonnegativity, sparsity, and sum-to-1 conditions, making their optimization particularly challenging due to the interplay between low-rank structures and constra...Low-rank optimization problems with sparse simplex constraints involve variables that must satisfy nonnegativity, sparsity, and sum-to-1 conditions, making their optimization particularly challenging due to the interplay between low-rank structures and constraints. These problems arise in various applications, including machine learning, signal processing, environmental fields, and computational biology. In this work, we propose a novel manifold optimization approach to efficiently tackle these problems. Our method leverages the geometry of oblique manifolds to reformulate the problem and introduces a new Riemannian optimization method based on Riemannian gradient descent that strictly maintains the simplex constraints. By exploiting the underlying manifold structure, our approach improves optimization efficiency. Experiments on synthetic and real datasets demonstrate the effectiveness of the proposed method compared to standard Euclidean and Riemannian methods, paving the way for broader applications.
|
| 2192 |
Towards Weaker Variance Assumptions for Stochastic Optimization
2504.09951
|
cs.LG
|
Ahmet Alacaoglu, Yura Malitsky, Stephen J. Wright |
We revisit a classical assumption for analyzing stochastic gradient algorithms where the squared norm of the stochastic subgradient (or the variance for smooth problems) is allowed to grow as fast as the squared norm of the optimization variable. We contextual...We revisit a classical assumption for analyzing stochastic gradient algorithms where the squared norm of the stochastic subgradient (or the variance for smooth problems) is allowed to grow as fast as the squared norm of the optimization variable. We contextualize this assumption in view of its inception in the 1960s, its seemingly independent appearance in the recent literature, its relationship to weakest-known variance assumptions for analyzing stochastic gradient algorithms, and its relevance in deterministic problems for non-Lipschitz nonsmooth convex optimization. We build on and extend a connection recently made between this assumption and the Halpern iteration. For convex nonsmooth, and potentially stochastic, optimization, we analyze horizon-free, anytime algorithms with last-iterate rates. For problems beyond simple constrained optimization, such as convex problems with functional constraints or regularized convex-concave min-max problems, we obtain rates for optimality measures that do not require boundedness of the feasible set.
|
| 2193 |
FEAT: Free energy Estimators with Adaptive Transport
2504.11516
|
cs.LG
|
Jiajun He, Yuanqi Du, Francisco Vargas, Yuanqing Wang, Carla P. Gomes |
We present Free energy Estimators with Adaptive Transport (FEAT), a novel framework for free energy estimation -- a critical challenge across scientific domains. FEAT leverages learned transports implemented via stochastic interpolants and provides consistent,...We present Free energy Estimators with Adaptive Transport (FEAT), a novel framework for free energy estimation -- a critical challenge across scientific domains. FEAT leverages learned transports implemented via stochastic interpolants and provides consistent, minimum-variance estimators based on escorted Jarzynski equality and controlled Crooks theorem, alongside variational upper and lower bounds on free energy differences. Unifying equilibrium and non-equilibrium methods under a single theoretical framework, FEAT establishes a principled foundation for neural free energy calculations. Experimental validation on toy examples, molecular simulations, and quantum field theory demonstrates improvements over existing learning-based methods. Our PyTorch implementation is available at https://github.com/jiajunhe98/FEAT.
|
| 2194 |
When do Random Forests work?
2504.12860
|
cs.LG
|
C. Revelas, O. Boldea, B. J. M. Werker |
We study the effectiveness of randomizing split-directions in random forests. Prior literature has shown that, randomization can reduce variance through decorrelation, and, randomization regularizes and works in low signal-to-noise ratio (SNR) environments. Fi...We study the effectiveness of randomizing split-directions in random forests. Prior literature has shown that, randomization can reduce variance through decorrelation, and, randomization regularizes and works in low signal-to-noise ratio (SNR) environments. First, we revisit decorrelation and regularization by presenting a systematic analysis of out-of-sample mean-squared error (MSE) for different SNR scenarios based on commonly-used data-generating processes. We point out that there is no consensus in the literature on how the MSE is computed and the terms "variance" and "bias" are not uniquely defined. We obtain MSE decompositions by averaging over training sets first and then covariates. We find that variance reduction tends to increase with the SNR and forests outperform bagging when the SNR is low because, in low SNR cases, variance dominates bias for both methods. Second, we show that the effectiveness of randomization is a question that goes beyond SNRs. We present a simulation study with fixed and moderate SNR, in which we examine the effectiveness of randomization for other data characteristics. In particular, we find that (i) randomization can increase bias in the presence of fat tails in the distribution of covariates; (ii) in the presence of irrelevant covariates randomization is ineffective because bias dominates variance; and (iii) when covariates are mutually correlated randomization tends to be effective because variance dominates bias. Beyond randomization, we find that, for both bagging and random forests, bias can be significantly reduced in the presence of correlated covariates. This last finding goes beyond the prevailing view that averaging mostly works by variance reduction. Given that in practice covariates are often correlated, our findings on correlated covariates could open the way for a better understanding of why random forests work well in many applications.
|
| 2195 |
Accelerating Natural Gradient Descent for PINNs with Randomized Numerical Linear Algebra
2505.11638
|
cs.LG
|
Ivan Bioli, Carlo Marcati, Giancarlo Sangalli |
Natural Gradient Descent (NGD) has emerged as a promising optimization algorithm for training neural network-based solvers for partial differential equations (PDEs), such as Physics-Informed Neural Networks (PINNs). However, its practical use is often limited ...Natural Gradient Descent (NGD) has emerged as a promising optimization algorithm for training neural network-based solvers for partial differential equations (PDEs), such as Physics-Informed Neural Networks (PINNs). However, its practical use is often limited by the high computational cost of solving linear systems involving the Gramian matrix. While matrix-free NGD methods based on the conjugate gradient (CG) method avoid explicit matrix inversion, the ill-conditioning of the Gramian significantly slows the convergence of CG. In this work, we extend matrix-free NGD to broader classes of problems than previously considered and propose the use of Randomized Numerical Linear Algebra (RandNLA) techniques for efficient preconditioning of the inner CG solver. The resulting algorithms demonstrate substantial performance improvements over existing NGD-based methods on a range of PDE problems discretized using neural networks, and offer competitive results compared to other state-of-the-art optimizers.
|
| 2196 |
Squeeze3D: Extreme Neural Compression with Latent Space Bridging
2506.07932
|
cs.LG
|
Rishit Dagli, Yushi Guan, Sankeerth Durvasula, Mohammadreza Mofayezi, Nandita Vijaykumar |
We propose Squeeze3D, a novel framework that leverages implicit prior knowledge learnt by existing pre-trained encoders and decoders to compress 3D data at extremely high compression ratios. Our approach bridges the latent spaces between a pre-trained encoder ...We propose Squeeze3D, a novel framework that leverages implicit prior knowledge learnt by existing pre-trained encoders and decoders to compress 3D data at extremely high compression ratios. Our approach bridges the latent spaces between a pre-trained encoder and a pretrained decoder model through trainable mapping networks. Any 3D asset represented as a mesh, point cloud, or radiance field is first encoded by the pre-trained encoder and then transformed (i.e. compressed) into a highly compact latent code by a mapping network. This latent code can effectively be used as an extremely compressed representation of the mesh, point cloud, or radiance field. A mapping network transforms the compressed latent code into the latent space of a powerful generative model; the decoder of this generative model then recreates the original 3D asset (i.e. decompression). Squeeze3D is trained entirely on generated synthetic data and does not require any 3D datasets. The Squeeze3D architecture can be flexibly used with existing pre-trained 3D encoders and existing generative models. It can flexibly support different formats, including meshes, point clouds, and radiance fields. Our experiments demonstrate that Squeeze3D achieves compression ratios of up to 2187$\times$ for textured meshes, 58.5$\times$ for point clouds, and more than 650$\times$ for radiance fields while maintaining visual quality comparable to many existing methods. Squeeze3D only incurs a small compression and decompression latency since it does not involve training object-specific networks to compress an object.
|
| 2197 |
A family of graph GOSPA metrics for graphs with different sizes
2506.17316
|
cs.LG
|
Jinhao Gu, \'Angel F. Garc\'ia-Fern\'andez, Robert E. Firth, Lennart Svensson |
This paper proposes a family of graph metrics for measuring distances between graphs of different sizes. The proposed metric family defines a general form of the graph generalised optimal sub-pattern assignment (GOSPA) metric and is also proved to satisfy the ...This paper proposes a family of graph metrics for measuring distances between graphs of different sizes. The proposed metric family defines a general form of the graph generalised optimal sub-pattern assignment (GOSPA) metric and is also proved to satisfy the metric properties. Similarly to the graph GOSPA metric, the proposed graph GOSPA metric family also penalises the node attribute costs for assigned nodes between the two graphs, and the number of unassigned nodes. However, the proposed family of metrics provides more general penalties for edge mismatches than the graph GOSPA metric. This paper also shows that the graph GOSPA metric family can be approximately computed using linear programming. Simulation experiments are performed to illustrate the characteristics of the proposed graph GOSPA metric family with different choices of hyperparameters. The benefits of the proposed graph GOSPA metric family for classification tasks are also shown on real-world datasets.
|
| 2198 |
Identifiable Convex-Concave Regression via Sub-gradient Regularised Least Squares
2506.18078
|
cs.LG
|
William Chung |
We propose a novel nonparametric regression method that models complex input-output relationships as the sum of convex and concave components. The method-Identifiable Convex-Concave Nonparametric Least Squares (ICCNLS)-decomposes the target function into addit...We propose a novel nonparametric regression method that models complex input-output relationships as the sum of convex and concave components. The method-Identifiable Convex-Concave Nonparametric Least Squares (ICCNLS)-decomposes the target function into additive shape-constrained components, each represented via sub-gradient-constrained affine functions. To address the affine ambiguity inherent in convex-concave decompositions, we introduce global statistical orthogonality constraints, ensuring that residuals are uncorrelated with both intercept and input variables. This enforces decomposition identifiability and improves interpretability. We further incorporate L1, L2 and elastic net regularisation on sub-gradients to enhance generalisation and promote structural sparsity. The proposed method is evaluated on synthetic and real-world datasets, including healthcare pricing data, and demonstrates improved predictive accuracy and model simplicity compared to conventional CNLS and difference-of-convex (DC) regression approaches. Our results show that statistical identifiability, when paired with convex-concave structure and sub-gradient regularisation, yields interpretable models suited for forecasting, benchmarking, and policy evaluation.
|
| 2199 |
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
2507.02915
|
cs.LG
|
Ludovic Tuncay (IRIT-SAMoVA), Etienne Labb\'e (IRIT-SAMoVA), Emmanouil Benetos (QMUL), Thomas Pellegrini (IRIT-SAMoVA) |
Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio Joint-Embedding Predictive Ar...Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio Joint-Embedding Predictive Architecture), tailored specifically for audio data. Audio-JEPA uses a simple Vision Transformer backbone to predict latent representations of masked spectrogram patches rather than reconstructing raw audio. We pre-train on unlabeled AudioSet clips (10s, 32kHz) with random patch masking on mel-spectrograms. We evaluate on the X-ARES suite covering speech, music, and environmental sound tasks. Although our implementation is a straightforward translation of the original model to audio, the results still show comparable performance to wav2vec 2.0 and data2vec while using less than one-fifth of their training data and with no hyper-parameter tuning. All code and pretrained checkpoints are available on GitHub.
|
| 2200 |
Adaptively Truncated Signature-based Logistic Regression for Semi-parametric Functional Classification
2507.06637
|
cs.LG
|
Pengcheng Zeng, Siyuan Jiang |
Modern sensing technologies routinely generate multi-dimensional functional data, often accompanied by scalar covariates, motivating statistical methods that can jointly accommodate nonlinear functional effects, complex cross-component dependencies, and irregu...Modern sensing technologies routinely generate multi-dimensional functional data, often accompanied by scalar covariates, motivating statistical methods that can jointly accommodate nonlinear functional effects, complex cross-component dependencies, and irregular sampling patterns. Existing functional logistic regression approaches rely on basis expansions whose performance is critically sensitive to the choice of basis, smoothing parameters, and sampling design; meanwhile, current signature-based methods require heuristic selection of the signature truncation order. We propose Adaptively Truncated Signature-based Logistic Regression (ATSLR), a semi-parametric classification framework that integrates path-signature representations with a data-driven procedure for selecting the signature truncation order via penalized empirical risk minimization. Our method is supported by rigorous theoretical guarantees, including the existence of an optimal truncation order, consistency of its adaptive estimator, convergence rates for the classifier risk, a finite and computable search bound, and a general error-propagation analysis for irregular sampling. Experiments on synthetic and real-world datasets demonstrate that our framework consistently outperforms traditional functional classifiers and fixed-order signature baselines in accuracy, robustness, and interpretability. Our results highlight the practical and theoretical value of integrating rough path theory with adaptive model complexity control.
|
| 2201 |
AI-driven Dispensing of Coral Reseeding Devices for Broad-scale Restoration of the Great Barrier Reef
2509.01019
|
cs.LG
|
Scarlett Raine, Emilio Olivastri, Benjamin Moshirian, Tobias Fischer |
Coral reefs are on the brink of collapse, with climate change, ocean acidification, and pollution leading to a projected 70-90% loss of coral species within the next decade. Reef restoration is crucial, but its success hinges on introducing automation to upsca...Coral reefs are on the brink of collapse, with climate change, ocean acidification, and pollution leading to a projected 70-90% loss of coral species within the next decade. Reef restoration is crucial, but its success hinges on introducing automation to upscale efforts. In this work, we present a highly configurable AI pipeline for the real-time deployment of coral reseeding devices. The pipeline consists of three core components: (i) the image labeling scheme, designed to address data availability and reduce the cost of expert labeling; (ii) the classifier which performs automated analysis of underwater imagery, at the image or patch-level, while also enabling quantitative coral coverage estimation; and (iii) the decision-making module that determines whether deployment should occur based on the classifier's analysis. By reducing reliance on manual experts, our proposed pipeline increases operational range and efficiency of reef restoration. We validate the proposed pipeline at five sites across the Great Barrier Reef, benchmarking its performance against annotations from expert marine scientists. The pipeline achieves 77.8% deployment accuracy, 89.1% accuracy for sub-image patch classification, and real-time model inference at 5.5 frames per second on a Jetson Orin. To address the limited availability of labeled data in this domain and encourage further research, we publicly release a comprehensive, annotated dataset of substrate imagery from the surveyed sites.
|
| 2202 |
Design of Experiment for Discovering Directed Mixed Graph
2509.01887
|
cs.LG
|
Haijie Xu, Chen Zhang |
We study the design of interventions for causal discovery in simple structural causal models whose causal graphs are directed mixed graphs (DMGs) that may contain directed cycles and bidirected edges representing latent confounding. In such case, observational...We study the design of interventions for causal discovery in simple structural causal models whose causal graphs are directed mixed graphs (DMGs) that may contain directed cycles and bidirected edges representing latent confounding. In such case, observational conditional-independence (CI) information may not identify even the graph skeleton, while CI alone cannot generally detect a bidirected edge coexisting with a directed edge. To this end, we propose a stage-wise framework based on tailored separating systems. Separating-system interventions first recover descendant relations and strongly connected components (SCCs). The SCCs are then ordered by ancestry, and an SCC-Anc separating system recovers the directed subgraph. Given this subgraph, further systems use CI tests interpreted through $d$- or $\sigma$-separation to recover non-adjacent bidirected edges, whose endpoints share no directed edge, and do-see comparisons to recover those coexisting with exactly one directed edge. Under our assumptions, the framework recovers the directed subgraph and every bidirected edge except double-adjacent ones, whose endpoints are connected by a directed edge in each direction. We develop algorithms for unrestricted and $M$-bounded settings, with each experiment targeting at most $M$ variables in the latter. For recovering the directed subgraph and non-adjacent bidirected edges, our upper bounds on the number and maximum size of experiments match corresponding worst-case lower bounds up to logarithmic factors.
|
| 2203 |
Synthetic data for ratemaking: imputation-based methods vs adversarial networks and autoencoders
2509.02171
|
cs.LG
|
Yevhen Havrylenko, Meelis K\"a\"arik, Artur Tuttar |
Actuarial ratemaking depends on high-quality data, yet access to such data is often limited by the cost of obtaining new data, privacy concerns, etc. In this paper, we explore synthetic-data generation as a potential solution to these issues. In addition to ge...Actuarial ratemaking depends on high-quality data, yet access to such data is often limited by the cost of obtaining new data, privacy concerns, etc. In this paper, we explore synthetic-data generation as a potential solution to these issues. In addition to generative methods previously studied in the actuarial literature, we explore and benchmark another class of approaches based on Multivariate Imputation by Chained Equations (MICE). In a comparative study using an open-source dataset, MICE-based models are evaluated against other generative models like Variational Autoencoders and Conditional Tabular Generative Adversarial Networks. We assess how well synthetic data preserves the original marginal distributions of variables as well as the multivariate relationships among covariates. The consistency between Generalized Linear Models (GLMs) trained on synthetic data with GLMs trained on the original data is also investigated. Furthermore, we assess the ease of use of each generative approach and study the impact of generically augmenting original data with synthetic data on the estimation of GLMs for predicting claim counts. Our results highlight the potential of MICE-based methods in creating high-fidelity tabular data while offering lower implementation complexity compared to deep generative models.
|
| 2204 |
A Kernel-based Stochastic Approximation Framework for Nonlinear Operator Learning
2509.11070
|
cs.LG
|
Jia-Qi Yang, Lei Shi |
We develop a stochastic approximation framework for learning nonlinear operators between infinite-dimensional spaces utilizing general Mercer operator-valued kernels. Our framework encompasses two key classes: (i) operator-valued kernels whose associated integ...We develop a stochastic approximation framework for learning nonlinear operators between infinite-dimensional spaces utilizing general Mercer operator-valued kernels. Our framework encompasses two key classes: (i) operator-valued kernels whose associated integral operators are compact and hence admit discrete spectral decompositions, and (ii) separable kernels of the form $K(x,x')=k(x,x')T$, where $k$ is a scalar-valued kernel and $T$ is a positive operator on the output space. This broad setting induces expressive vector-valued reproducing kernel Hilbert spaces (RKHSs) that generalize the classical $K=kI$ paradigm, thereby enabling rich structural modeling with rigorous theoretical guarantees. To address target operators lying outside the RKHS, we introduce vector-valued interpolation spaces to precisely quantify misspecification error. Within this framework, we establish non-asymptotic convergence rates for prediction, estimation, and misspecification errors in the online and finite-horizon settings. Importantly, the framework also accommodates a range of operator learning settings, from Fredholm integral operators to encoder--decoder architectures. Numerical experiments on the two-dimensional Navier--Stokes equations illustrate the proposed approach.
|
| 2205 |
The Platonic Universe: Do Foundation Models See the Same Sky?
2509.19453
|
cs.LG
|
UniverseTBD, :, Trinidad Borrell, Steven Dillmann, Kshitij Duraphe |
We investigate when foundation models converge towards shared representations, and how this convergence depends on model capacity, training regime, and model architecture. We take a `science-for-AI' approach, using astronomy as an experimental instrument to te...We investigate when foundation models converge towards shared representations, and how this convergence depends on model capacity, training regime, and model architecture. We take a `science-for-AI' approach, using astronomy as an experimental instrument to test the Platonic Representation Hypothesis and its Aristotelian refinement against an external physical reference. The historical success of astrophysics is evidence that a compact, modality-invariant description of galaxy observables exists, and so representation convergence toward reality should be measurable against the physical parameters astronomers already use. Given this framework, we evaluate eleven foundation model families (spanning classification, self-distillation, joint-embedding prediction, autoencoding, vision-language pre-training, and astro-specific architectures from $\mathcal{O}$(10M)${\to}\mathcal{O}$(10B) parameters) on crossmatched JWST, HSC, and Legacy imagery, and DESI spectroscopy. All models are evaluated frozen, with no astronomy-specific fine-tuning. We probe redshift, stellar mass, and sSFR via linear probes, and local (MKNN) and global (CKA) embedding geometry within families, between modalities, and across architectures. We find that physics performance scales predictably with capacity; probe directions align consistently with expected astrophysical correlations and selection effects; and local (not global) embedding alignment tracks physics performance, including between DESI spectra and HSC imagery---modalities that share essentially no low-level statistics. Our results support the ARH over the strict PRH, demonstrate astronomy's value as an experimental framework for neural representation learning, and suggest that astro-foundation models can build on general-purpose pre-trained architectures, capitalizing on the broader open machine learning community's already-spent computational investment.
|
| 2206 |
Learning to Generate Rigid Body Interactions with Video Diffusion Models
2510.02284
|
cs.LG
|
David Romero, Ariana Bermudez, Viacheslav Iablochnikov, Hao Li, Fabio Pizzati |
Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision makin...Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, current approaches still struggle to generate physically plausible object interactions and lack object-level control mechanisms. To address these limitations, we introduce KineMask, an approach for video generation that enables realistic rigid body control, interactions, and effects. Given a single image and a specified object velocity, our method generates videos with inferred motions and future object interactions. We propose a two-stage training strategy that gradually removes future motion supervision via object masks. Using this strategy we train video diffusion models (VDMs) on synthetic scenes of simple interactions and demonstrate significant improvements and generalization to rigid body and hand-object interactions in real scenes. Furthermore, KineMask integrates low-level motion control with high-level textual conditioning via predicted scene descriptions, leading to support for synthesis of complex dynamical phenomena. Our experiments show that KineMask generalizes to different VDMs and achieves strong improvements over recent models of comparable size. Ablation studies further highlight the complementary roles of low- and high-level conditioning in VDMs.
|
| 2207 |
Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM
2510.05632
|
cs.LG
|
Tianhao Zhu, Dahu Feng, Erhu Feng, Yubin Xia |
With the widespread adoption of Large Language Models (LLMs), the demand for high-performance LLM inference services continues to grow. Multi-core AI accelerators, such as Groq, Graphcore IPU, and Cerebras WSE, provide promising platforms for LLM serving, but ...With the widespread adoption of Large Language Models (LLMs), the demand for high-performance LLM inference services continues to grow. Multi-core AI accelerators, such as Groq, Graphcore IPU, and Cerebras WSE, provide promising platforms for LLM serving, but their distributed memory systems require careful coordination between hardware configuration and serving policies. Otherwise, mismatched tensor partitioning, data placement, and memory management can substantially underutilize compute and communication resources. To address these challenges, we present WaferAI-SIM, a multi-level simulation framework that combines transaction-level simulation with an analytical performance model. WaferAI-SIM enables simulator-driven co-design of LLM serving strategies and multi-core accelerator architectures, targeting the early design stage in which emerging platforms are not yet broadly available for empirical serving studies. It captures how LLM serving policies interact with compute-core count, memory hierarchy, and interconnect topology, enabling architecture-aware exploration beyond GPU-centric assumptions. We evaluate representative LLMs across a range of chip configurations and serving scenarios. Across the evaluation, WaferAI-SIM reports 1.32$\times$--6.03$\times$ latency improvements, where the lower endpoint comes from ring-based placement at TP=16 over the placement baselines, and the upper endpoint comes from K-dimension TP over MN-dimension TP for Qwen3\_4B at TP=4 with sequence length 256. For LLM serving, our findings provide guidance for co-designing hardware architectures and serving strategies for multi-core AI accelerators across diverse LLM workloads.
|
| 2208 |
AlignBeat: A Latent Variable Model for Multi-Class Beat Tracking from Partially Labeled Data
2510.14391
|
cs.LG
|
Jaehoon Ahn, Tae Gum Hwang, Moon-Ryul Jung |
Recent neural beat trackers predict beats and downbeats with two independent frame-wise binary classification heads, a convention adopted because a single multi-class head cannot fully exploit datasets that annotate only beats. The heads can disagree, so Beat ...Recent neural beat trackers predict beats and downbeats with two independent frame-wise binary classification heads, a convention adopted because a single multi-class head cannot fully exploit datasets that annotate only beats. The heads can disagree, so Beat This moves each predicted downbeat to the nearest predicted beat. We instead predict a sparse set of grid points, each with an event time and one distribution over downbeat, beat and no event, so a downbeat is a beat by construction and no reconciliation is needed. The alignment between grid points and annotated events is latent, and we fit the model by expectation-maximization. Where no downbeat labels were recorded the event class is latent too and is marginalized out. Over eight-fold cross-validation on eighteen datasets, retaining rather than discarding the beat-only labeled data raises beat CMLt by 4.6 points on the datasets that provide it; downbeat CMLt rises 3.6 points over Beat This, with both F-measures moving by less than one point and no post-processing at any stage.
|
| 2209 |
Reinforcement Learning and Consumption-Savings Behavior
2510.20748
|
cs.LG
|
Brandon Gary Kaplowitz |
This paper demonstrates how reinforcement learning can explain two puzzling empirical patterns in household consumption behavior during economic downturns. I develop a model where agents use Q-learning with neural network approximation to make consumption-savi...This paper demonstrates how reinforcement learning can explain two puzzling empirical patterns in household consumption behavior during economic downturns. I develop a model where agents use Q-learning with neural network approximation to make consumption-savings decisions under income uncertainty, departing from standard rational expectations assumptions. The model replicates two key findings from recent literature: (1) unemployed households with previously low liquid assets exhibit substantially higher marginal propensities to consume (MPCs) out of stimulus transfers compared to high-asset households (0.50 vs 0.34), even when neither group faces borrowing constraints, consistent with Ganong et al. (2024); and (2) households with more past unemployment experiences maintain persistently lower consumption levels after controlling for current economic conditions, a "scarring" effect documented by Malmendier and Shen (2024). Unlike existing explanations based on belief updating about income risk or ex-ante heterogeneity, the reinforcement learning mechanism generates both higher MPCs and lower consumption levels simultaneously through value function approximation errors that evolve with experience. Simulation results closely match the empirical estimates, suggesting that adaptive learning through reinforcement learning provides a unifying framework for understanding how past experiences shape current consumption behavior beyond what current economic conditions would predict.
|
| 2210 |
A Survey on Efficient Vision-Language-Action Models
2510.24795
|
cs.LG
|
Zhaoshu Yu, Bo Wang, Pengpeng Zeng, Haonan Zhang, Ji Zhang |
Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computat...Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. To this end, recent studies improve VLA efficiency from different views, e.g., real-time inference, training computation, and scalable data collection. However, these efforts are mostly studied separately. A unified view is still missing for understanding how efficiency should be optimized across the full VLA lifecycle. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey provides an organized reference for the community and summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/.
|
| 2211 |
Action-Driven Processes for Continuous-Time Control
2510.26672
|
cs.LG
|
Ruimin He, Shaowei Lin |
At the heart of reinforcement learning are actions -- decisions made in response to observations of the environment. Actions are equally fundamental in the modeling of stochastic processes, as they trigger discontinuous state transitions and enable the flow of...At the heart of reinforcement learning are actions -- decisions made in response to observations of the environment. Actions are equally fundamental in the modeling of stochastic processes, as they trigger discontinuous state transitions and enable the flow of information through large, complex systems. In this paper, we unify the perspectives of stochastic processes and reinforcement learning through action-driven processes, and illustrate their application to spiking neural networks. Leveraging ideas from control-as-inference, we show that minimizing the Kullback-Leibler divergence between a policy-driven true distribution and a reward-driven model distribution for a suitably defined action-driven process is equivalent to maximum entropy reinforcement learning.
|
| 2212 |
Disciplined Biconvex Programming
2511.01813
|
cs.LG
|
Hao Zhu, Joschka Boedecker |
We introduce disciplined biconvex programming (DBCP), a modeling framework for specifying and solving biconvex optimization problems. Biconvex optimization problems arise in various applications, including machine learning, signal processing, computational sci...We introduce disciplined biconvex programming (DBCP), a modeling framework for specifying and solving biconvex optimization problems. Biconvex optimization problems arise in various applications, including machine learning, signal processing, computational science, and control. Solving a biconvex optimization problem in practice usually involves heuristic methods based on alternate convex search (ACS), which iteratively optimizes over one block of variables while keeping the other fixed, so that the resulting subproblems are convex and can be efficiently solved. However, designing and implementing an ACS solver for a specific biconvex optimization problem usually requires significant effort from the user, which can be tedious and error-prone. DBCP extends the principles of disciplined convex programming to biconvex problems, allowing users to specify biconvex optimization problems in a natural way based on a small number of syntax rules. The resulting problem can then be automatically split and transformed into convex subproblems, for which a customized ACS solver is then generated and applied. DBCP allows users to quickly experiment with different biconvex problem formulations, without expertise in convex optimization. We implement DBCP in the open-source Python package dbcp, as an extension to the well known domain specific language CVXPY for convex optimization.
|
| 2213 |
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
2511.10909
|
cs.LG
|
Peichen Xie, Shuotao Xu, Yang Wang, Fan Yang, Mao Yang |
Modern AI accelerators rely on matrix multiply-accumulate units (MMAUs), such as NVIDIA Tensor Cores and AMD Matrix Cores, to accelerate deep neural network workloads. MMAUs expose only instruction-level or API-level interfaces of matrix multiply-accumulate (M...Modern AI accelerators rely on matrix multiply-accumulate units (MMAUs), such as NVIDIA Tensor Cores and AMD Matrix Cores, to accelerate deep neural network workloads. MMAUs expose only instruction-level or API-level interfaces of matrix multiply-accumulate (MMA) operations, while leaving internal floating-point arithmetic behavior undocumented. Consequently, MMAUs across vendors and architectural generations often produce numerical discrepancies for identical inputs, and sometimes exhibit reduced numerical accuracy that can cause training instability. Diagnosing and understanding the root causes of these effects is challenging without white-box models of their arithmetic behavior. This paper proposes closed-loop feature probing (CLFP), a generic and systematic framework for constructing bit-accurate arithmetic behavior models of MMA operations. Based on this framework, we analyze all MMA instructions on ten GPU architectures spanning NVIDIA Volta through RTX Blackwell and AMD CDNA1 through CDNA3, and derive the first bit-accurate arithmetic models for these MMAUs. Our models explain previously observed cross-platform numerical discrepancies and accuracy issues, enable white-box numerical error analysis, reveal four types of precision bottlenecks and one type of numerical asymmetry, and inform software workarounds as well as design suggestions for future MMAUs. This work is open-source at https://github.com/microsoft/MMA-Sim
|
| 2214 |
AI-driven ionic liquid discovery with unified chemical intelligence
2511.11257
|
cs.LG
|
Yuqi Yin, Yibo Fu, Siyuan Wang, Lizhen Huo, Peng Sun |
The rational discovery of functional solvents remains challenging because macroscopic solvent behaviors emerge from complex many-body interactions across vast chemical spaces. Ionic Liquids (ILs) constitute one of the most chemically diverse solvent families, ...The rational discovery of functional solvents remains challenging because macroscopic solvent behaviors emerge from complex many-body interactions across vast chemical spaces. Ionic Liquids (ILs) constitute one of the most chemically diverse solvent families, yet their exceptional structural tunability remains only partially explored owing to the limitations of existing experimental and computational approaches. Here we introduce AIonopedia, a large-language-model-orchestrated agentic framework for IL research. Built around a domain-specific multimodal foundation model, it integrates information retrieval, chemical reasoning and autonomous tool execution within a unified architecture. Pretrained on large-scale unlabeled corpora and fine-tuned using over 80,000 curated measurements, AIonopedia learns expressive joint representations of multicomponent ionic systems, enabling accurate prediction across diverse physicochemical properties. Through a series of literature-grounded evaluations, AIonopedia exhibits robust out-of-distribution generalization, while many-body-expansion-based analysis provides physically grounded explanations of structure-property relationships. By coupling molecular screening and optimization with wet-lab validation in a closed-loop workflow, AIonopedia identifies high-performance ILs and reveals transferable design principles for the capture of various volatile organic compounds. This work establishes an end-to-end AI-driven paradigm for understanding and exploring complex ionic systems, demonstrating the promise of chemically specialized foundation models as a generalizable strategy for scientific discovery.
|
| 2215 |
Server-Enforced Watermarking in U-Shaped Split Federated Learning
2511.14422
|
cs.LG
|
Zhengchunmin Dai, Jiaxiong Tang, Peng Sun, Honglong Chen, Liantao Wu |
U-shaped split federated learning (U-SFL) enables resource-constrained Internet of Things devices to collaboratively train models with an edge server while retaining raw data and labels locally. We consider model-service deployments in which a provider supplie...U-shaped split federated learning (U-SFL) enables resource-constrained Internet of Things devices to collaboratively train models with an edge server while retaining raw data and labels locally. We consider model-service deployments in which a provider supplies proprietary models whose client-side segments reside on participating devices, creating risks of unauthorized copying and redistribution. Protecting these segments is challenging because the server cannot directly access client data, labels, or model parameters, and potentially malicious clients may refuse to perform watermark embedding. To address this challenge, we propose Sigil, a server-enforced watermarking framework for U-SFL. Sigil defines a secret watermark constraint in the server-visible activation space and embeds the watermark into client-side models by injecting a watermark gradient into the gradients returned during training. This mechanism requires neither access to clients' raw data and labels nor client-side watermarking operations. To limit interference with the main task and reduce detectability by gradient anomaly detectors, Sigil adaptively clips the watermark gradient relative to the main-task gradient. Experiments on two datasets and four model architectures demonstrate high watermark detection rates with limited impact on task accuracy, robustness against the evaluated removal attacks, and stealthiness against the evaluated gradient anomaly detector.
|
| 2216 |
Digital Agriculture Sandbox for Collaborative Research
2511.15990
|
cs.LG
|
Osama Zafar, Rosemarie Santa Gonz\'alez, Alfonso Morales, Erman Ayday |
Digital agriculture is transforming the way we grow food by utilizing technology to make farming more efficient, sustainable, and productive. This modern approach to agriculture generates a wealth of valuable data that could help address global food challenges...Digital agriculture is transforming the way we grow food by utilizing technology to make farming more efficient, sustainable, and productive. This modern approach to agriculture generates a wealth of valuable data that could help address global food challenges, but farmers are hesitant to share it due to privacy concerns. This limits the extent to which researchers can learn from this data to inform improvements in farming. This paper presents the Digital Agriculture Sandbox, a secure online platform that solves this problem. The platform enables farmers (with limited technical resources) and researchers to collaborate on analyzing farm data without exposing private information. We employ specialized techniques such as federated learning, differential privacy, and data analysis methods to safeguard the data while maintaining its utility for research purposes. The system enables farmers to identify similar farmers in a simplified manner without needing extensive technical knowledge or access to computational resources. Similarly, it enables researchers to learn from the data and build helpful tools without the sensitive information ever leaving the farmer's system. This creates a safe space where farmers feel comfortable sharing data, allowing researchers to make important discoveries. Our platform helps bridge the gap between maintaining farm data privacy and utilizing that data to address critical food and farming challenges worldwide.
|
| 2217 |
MASTEST: A LLM-Based Multi-Agent System For Testing RESTful APIs
2511.18038
|
cs.LG
|
Xiaoke Han, Hong Zhu |
Testing RESTful API is increasingly complicated but indispensable to quality assurance of cloud-native applications. This paper reports a multi-agent system called MASTEST that combines LLM-based intelligent agents and programmed agents to automate REST API te...Testing RESTful API is increasingly complicated but indispensable to quality assurance of cloud-native applications. This paper reports a multi-agent system called MASTEST that combines LLM-based intelligent agents and programmed agents to automate REST API testing. They form a complete tool chain covering the whole workflow of REST API test with API specification in the OpenAPI Swagger format as the input. It also incorporates human testers in the process to review and correct LLM generated test artefacts to control the quality of testing activities. MASTEST is evaluated on two LLMs, GPT-4o and DeepSeek V3.1 Reasoner with five public APIs. Its performances on various testing activities are measured by a wide range of metrics, including adequacy and coverage metrics, the syntax and data type correctness of generated test scripts, the usability of LLM generated test cases and scripts, as well as the bug detection ability. Experiment results demonstrated that both DeepSeek and GPT-4o achieved a high overall performance but had strengths and weaknesses on different testing activities. MASTEST generated test cases achieved 94% and 98% unit test coverage and 79% and 78% system test coverage for GPT-4o and DeepSeek respectively in comparison with human designed test cases. The generated test scripts maintained 100% syntax correctness and only required minimal manual edits for semantic correctness. The generated test scripts contain assertions on the expected status code as well as contents in the response messages. They are highly capable of detecting bugs in the REST APIs. Experiment data shows that the bug detection rates are between 2.13 to 4.50 per operation. These findings indicate that MASTEST is highly efficient and effective.
|
| 2218 |
Reinforcement Learning to Initialize Newton-Raphson for AC Power Flow with Quantum Annealing-Based Environment Updates
2511.20237
|
cs.LG
|
Zeynab Kaseb, Matthias Moller, Lindsay Spoor, Jerry J. Guo, Yu Xiang |
The Newton-Raphson (NR) method is widely used for solving power flow (PF) equations due to its quadratic convergence. However, its performance deteriorates under poor initialization or extreme operating scenarios, e.g., high levels of renewable energy penetrat...The Newton-Raphson (NR) method is widely used for solving power flow (PF) equations due to its quadratic convergence. However, its performance deteriorates under poor initialization or extreme operating scenarios, e.g., high levels of renewable energy penetration. We propose the use of reinforcement learning (RL) to optimize the initialization of NR, and introduce a quantum-enhanced RL environment update mechanism that addresses the combinatorially large action space at each RL timestep by formulating the voltage adjustment task as a Quadratic Unconstrained Binary Optimization (QUBO) problem, solved with an Ising machine. RL initialization is benchmarked against flat start and start from the DC (linearized) PF solution on a standard 4-bus system, Iwamoto's ill-conditioned 11-bus system, and the IEEE 118-bus system under normal and stressed loading and reactive power limits, with verified operational solutions. On all systems, a supervised initializer refined by RL requires fewer NR iterations than flat and DC starts and than the same initializer without RL, for all seeds. For example, on the 118-bus system under normal and stressed loading, it reached 2.04 and 2.86 NR iterations, compared with 3.02 and 5.13 from DC start and 2.61 and 3.09 without RL. In wall-clock time, this pays off only for an initializer integrated into the solver and reused for many solves on a fixed topology. On the 4-bus system, a quantum-enhanced RL agent with a quantum-inspired annealer moved challenging initial states that required 29 and 44 NR iterations to initializations that required three NR iterations within one RL timestep.
|
| 2219 |
On Memory: A comparison of memory mechanisms in world models
2512.06983
|
cs.LG
|
Eli J. Laird, Corey Clark |
World models enable agents to plan within imagined environments by predicting future states conditioned on past observations and actions. However, their ability to plan over long horizons is limited by the effective memory span of the backbone architecture. Th...World models enable agents to plan within imagined environments by predicting future states conditioned on past observations and actions. However, their ability to plan over long horizons is limited by the effective memory span of the backbone architecture. This limitation leads to perceptual drift in long rollouts, degrading the model's capacity to recall recently observed scenes. In this work, we investigate the effective memory span of transformer-based world models through an analysis of memory augmentation mechanisms. We introduce a taxonomy that distinguishes between memory \emph{encoding} and memory \emph{injection} mechanisms, motivating their roles in extending the world model's memory through the lens of residual stream dynamics. We evaluate twenty combinations of four encoding methods and five injection methods in the MemoryMaze environment. Using a state recall evaluation task across multiple imagination horizons, we measure the memory recall capacity of each mechanism and analyze their respective trade-offs in reconstruction quality, latent prediction error, and computational cost. We further ablate the effect of injection depth and compare the best memory-augmented vision transformer against a pure state-space model backbone. Our central finding is that the mLSTM memory encoder outperforms all alternatives in both reconstruction and latent fidelity metrics. Paired with additive injection, it exhibits the strongest recall capabilities at a moderate computational cost while matching or slightly exceeding a pure Mamba backbone. This evaluation is limited to a single environment and does not explore the effect of each mechanism on downstream task performance. We believe these research directions merit an in-depth study of their own to clearly isolate the effects, and therefore are left for future work.
|
| 2220 |
Manifold limit for the training of shallow graph convolutional neural networks
2601.06025
|
cs.LG
|
Johanna Tengler, Christoph Brune, Jos\'e A. Iglesias |
We study the discrete-to-continuum consistency of the training of shallow graph convolutional neural networks (GCNNs) on proximity graphs of sampled point clouds under a manifold assumption. Graph convolution is defined spectrally via the graph Laplacian, whos...We study the discrete-to-continuum consistency of the training of shallow graph convolutional neural networks (GCNNs) on proximity graphs of sampled point clouds under a manifold assumption. Graph convolution is defined spectrally via the graph Laplacian, whose low-frequency spectrum approximates that of the Laplace-Beltrami operator of the underlying smooth manifold, and shallow GCNNs of possibly infinite width are linear functionals on the space of measures on the parameter space. From this functional-analytic perspective, graph signals are seen as spatial discretizations of functions on the manifold, which leads to a natural notion of training data consistent across graph resolutions. To enable convergence results, the continuum parameter space is chosen as a weakly compact product of unit balls, with Sobolev regularity imposed on the output weight and bias, but not on the convolutional parameter. The corresponding discrete parameter spaces inherit the corresponding spectral decay, and are additionally restricted by a frequency cutoff adapted to the informative spectral window of the graph Laplacians. Under these assumptions, we prove $\Gamma$-convergence of regularized empirical risk minimization functionals and corresponding convergence of their global minimizers, in the sense of weak convergence of the parameter measures and uniform convergence of the functions over compact sets. This provides a formalization of mesh and sample independence, and hence of transferability across discretizations, for the training of such networks.
|
| 2221 |
Contextual Distributionally Robust Optimization with Causal and Continuous Structure
2601.11016
|
cs.LG
|
Fenglin Zhang, Jie Wang |
We propose a framework for contextual distributionally robust optimization (DRO) that considers the causal and continuous structure of the underlying distribution, and we develop an interpretable and tractable decision rule. We first introduce the causal Sinkh...We propose a framework for contextual distributionally robust optimization (DRO) that considers the causal and continuous structure of the underlying distribution, and we develop an interpretable and tractable decision rule. We first introduce the causal Sinkhorn discrepancy (CSD), an entropy-regularized causal Wasserstein distance that encourages continuous transport plans while preserving causal consistency. We then formulate a contextual DRO model with a CSD-based ambiguity set, termed Causal Sinkhorn DRO (Causal-SDRO), and derive its strong dual reformulation, where the worst-case distribution is characterized as a mixture of Gibbs distributions. To obtain an (infinite-dimensional) optimal policy, we propose a soft regression forest (SRF) decision rule: it preserves the interpretability of classical decision trees while being fully parametric, differentiable, and Lipschitz-smooth, enabling intrinsic interpretation from both global and local perspectives. To solve the Causal-SDRO with parametric decision rules, we develop an efficient stochastic compositional gradient algorithm that converges to an $\varepsilon$-stationary point at a rate of $\mathcal{O}(\varepsilon^{-4})$, matching that of standard stochastic gradient descent. Finally, we validate our method through numerical experiments on synthetic and real-world datasets, demonstrating its superior performance and interpretability.
|
| 2222 |
Model-Free Output Feedback Stabilization via Policy Gradient Methods
2601.19284
|
cs.LG
|
Ankang Zhang, Ming Chi, Xiaoling Wang, Lintao Ye |
Stabilizing a dynamical system is a fundamental problem that serves as a cornerstone for many complex tasks in the field of control systems. The problem becomes challenging when the system model is unknown. Among the Reinforcement Learning (RL) algorithms that...Stabilizing a dynamical system is a fundamental problem that serves as a cornerstone for many complex tasks in the field of control systems. The problem becomes challenging when the system model is unknown. Among the Reinforcement Learning (RL) algorithms that have been successfully applied to solve problems pertaining to unknown linear dynamical systems, the policy gradient (PG) method stands out due to its ease of implementation and can solve the problem in a model-free manner. However, most of the existing works on PG methods for unknown linear dynamical systems assume full-state feedback. In this paper, we take a step towards model-free learning for partially observed linear dynamical systems with output feedback and focus on the fundamental stabilization problem of the system. We propose an algorithmic framework that stretches the boundary of PG methods to the problem without global convergence guarantees. We show that by leveraging zeroth-order PG update based on system trajectories and its convergence to stationary points, the proposed algorithms return a stabilizing output feedback policy for discrete-time linear dynamical systems. We also explicitly characterize the sample complexity of our algorithm and verify the effectiveness of the algorithm using numerical examples.
|
| 2223 |
Hypersolid: Emergent Vision Representations via Short-Range Repulsion
2601.21255
|
cs.LG
|
Esteban Rodr\'iguez-Betancourt, Edgar Casasola-Murillo |
A central problem in self-supervised learning is preventing representation collapse. Most methods avoid it through global mechanisms, such as contrastive expansion, variance constraints, decorrelating dimensions, or enforcing certain output distributions. In t...A central problem in self-supervised learning is preventing representation collapse. Most methods avoid it through global mechanisms, such as contrastive expansion, variance constraints, decorrelating dimensions, or enforcing certain output distributions. In this work, we study a different design: short-range repulsion. We introduce Hypersolid, a self-supervised objective that combines view alignment with local collision avoidance. Our method induces a latent geometry of compact, semantically aligned neighborhoods with low anisotropy. This geometry is especially effective for unsupervised clustering and fine-grained separation, although it comes at the cost of weaker transferability.
|
| 2224 |
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
2602.00449
|
cs.LG
|
Jia Liang, Liangming Pan |
Latent Chain-of-Thought aims to enable step-by-step computation without emitting long rationales, yet its internal mechanisms remain unclear. We study CODI, a continuous-thought teacher-student distillation model, on strictly sequential polynomial-iteration ta...Latent Chain-of-Thought aims to enable step-by-step computation without emitting long rationales, yet its internal mechanisms remain unclear. We study CODI, a continuous-thought teacher-student distillation model, on strictly sequential polynomial-iteration tasks with known intermediate states. Using logit-lens decoding, linear probes, attention analysis, and activation patching, we localize intermediate-state representations and trace how they are routed to the final readout. In short-horizon, low-hop tasks, CODI forms faithful bridge states across latent-thought positions, while the final input follows a separate near-direct route; predictions arise through late fusion at the answer readout. As task depth and difficulty increase, however, CODI does not reliably sustain a full latent rollout: it either compresses computation into a partial late-intermediate pathway or, in harder regimes, loses the latent reasoning signature altogether. To explain this transition, we show theoretically that the task's algebraic structure controls its effective memory: compressible regimes support late-bottleneck reasoning, while incompressible regimes preserve full-history dependence and destabilize latent rollouts. Overall, our results characterize when CODI-style latent-CoT yields faithful iterative computation versus compressed or shortcut strategies.
|
| 2225 |
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
2602.01750
|
cs.LG
|
Mohammad Beigi, Ming Jin, Junshan Zhang, Qifan Wang, Lifu Huang |
Reinforcement Learning from Human Feedback (RLHF) remains vulnerable to reward hacking, where models exploit spurious correlations in learned reward models to achieve high scores while violating human intent. Existing mitigations rely on static defenses that c...Reinforcement Learning from Human Feedback (RLHF) remains vulnerable to reward hacking, where models exploit spurious correlations in learned reward models to achieve high scores while violating human intent. Existing mitigations rely on static defenses that cannot adapt to novel exploitation strategies. We propose Adversarial Reward Auditing (ARA), a framework that reconceptualizes reward hacking as a dynamic, competitive game. ARA operates in two stages: first, a Hacker policy discovers reward model vulnerabilities while an Auditor learns to detect exploitation from latent representations; second, Auditor-Guided RLHF (AG-RLHF) gates reward signals to penalize detected hacking, transforming reward hacking from an unobservable failure into a measurable, controllable signal. Experiments across three hacking scenarios demonstrate that ARA achieves the best alignment-utility tradeoff among all baselines: reducing sycophancy to near-SFT levels while improving helpfulness, decreasing verbosity while achieving the highest ROUGE-L, and suppressing code gaming while improving Pass@1. Beyond single-domain evaluation, we show that reward hacking, detection, and mitigation all generalize across domains -- a Hacker trained on code gaming exhibits increased sycophancy despite no reward for this behavior, and an Auditor trained on one domain effectively suppresses exploitation in others, enabling efficient multi-domain defense with a single model.
|
| 2226 |
All-Atom GPCR-Ligand Dynamics Simulation via a Residual Latent Flow Mode
2602.03902
|
cs.LG
|
Jiying Zhang, Shuhao Zhang, Pierre Vandergheynst, Patrick Barth |
G-protein-coupled receptors (GPCRs), which are targeted by over one-third of approved drugs, undergo intricate conformational transitions to transduce signals. While Molecular Dynamics (MD) is essential for elucidating this transduction process, particularly w...G-protein-coupled receptors (GPCRs), which are targeted by over one-third of approved drugs, undergo intricate conformational transitions to transduce signals. While Molecular Dynamics (MD) is essential for elucidating this transduction process, particularly within ligand-bound complexes, conventional all-atom MD simulation is computationally prohibitive. In this paper, we introduce GPCRLMD, a deep generative framework for efficient all-atom GPCR-ligand simulation. GPCRLMD employs an Atom-Anchored Thermal Variational Autoencoder (AAT-VAE) to first map the complex into a regularized dimension-preserving latent space, maintaining geometric topology via physics-informed constraints. Within this latent space, a Residual Latent Flow samples evolution trajectories, which are subsequently decoded back to atomic coordinates. By capturing temporal dynamics via relative displacements anchored to the initial structure, this residual mechanism effectively decouples static topology from dynamic fluctuations. Experimental results demonstrate that GPCRLMD achieves state-of-the-art performance in GPCR-ligand dynamics simulation, faithfully reproducing ensemble observables and critical ligand-receptor interactions.
|
| 2227 |
Instance-optimal high-precision shadow tomography with few-copy measurements: A metrological approach
2602.04952
|
cs.LG
|
Senrui Chen, Weiyuan Gong, Sisi Zhou |
We study the sample complexity of shadow tomography in the high-precision regime under realistic measurement constraints. Given an unknown $d$-dimensional quantum state $\rho$ and a known set of observables $\{O_i\}_{i=1}^m$, the goal is to estimate expectatio...We study the sample complexity of shadow tomography in the high-precision regime under realistic measurement constraints. Given an unknown $d$-dimensional quantum state $\rho$ and a known set of observables $\{O_i\}_{i=1}^m$, the goal is to estimate expectation values $\{\mathrm{tr}(O_i\rho)\}_{i=1}^m$ to accuracy $\epsilon$ in $L_p$-norm, using possibly adaptive measurements that act on $O(\mathrm{polylog}(d))$ number of copies of $\rho$ at a time. We focus on the regime where $\epsilon$ is below an instance-dependent threshold. Our main contribution is an instance-optimal characterization of the sample complexity as $\tilde{\Theta}(\Gamma_p/\epsilon^2)$, where $\Gamma_p$ is a function of $\{O_i\}_{i=1}^m$ defined via an optimization formula involving the inverse Fisher information matrix. Previously, tight bounds were known only in special cases, e.g. Pauli shadow tomography with $L_\infty$-norm error. Concretely, we first analyze a simpler oblivious variant where the goal is to estimate an observable of the form $\sum_{i=1}^m \alpha_i O_i$ with $\|\alpha\|_q = 1$ (where $q$ is dual to $p$) revealed after the measurement. For single-copy measurements, we obtain a sample complexity of $\Theta(\Gamma^{\mathrm{ob}}_p/\epsilon^2)$. We then show $\tilde{\Theta}(\Gamma_p/\epsilon^2)$ is necessary and sufficient for the original problem, with the lower bound applying to unbiased, bounded estimators. Our upper bounds rely on a two-step algorithm combining coarse tomography with local estimation. Notably, $\Gamma^{\mathrm{ob}}_\infty = \Gamma_\infty$. In both cases, allowing $c$-copy measurements improves the sample complexity by at most $\Omega(1/c)$. Our results establish a quantitative correspondence between quantum learning and metrology, unifying asymptotic metrological limits with finite-sample learning guarantees.
|
| 2228 |
Semantic Purification for Conditional Representation Learning
2602.05464
|
cs.LG
|
Jiaquan Wang, Yan Lyu, Chen Li, Yuheng Jia |
Conditional representation learning aims to extract criterion-specific features for customized tasks. Recent methods construct conditional subspaces spanned by criterion-specific text bases in the embedding space of vision-language models (VLMs). Image embeddi...Conditional representation learning aims to extract criterion-specific features for customized tasks. Recent methods construct conditional subspaces spanned by criterion-specific text bases in the embedding space of vision-language models (VLMs). Image embeddings are then projected onto these subspaces to obtain conditional representations. However, since VLMs are not explicitly trained to disentangle semantics associated with different criteria, the corresponding conditional subspaces remain coupled. This coupling induces semantic leakage during projection, thereby degrading the semantic purity of conditional representations. To suppress semantic leakage, we propose Semantic Purification for Conditional Representation Learning (SP-CRL). Specifically, SP-CRL first decomposes the original text basis and performs curvature-based adaptive truncation on the resulting basis vectors to construct a purer conditional subspace. It then identifies an appropriate noise subspace and projects image embeddings onto its null space to remove irrelevant semantic components. Extensive experiments across customized clustering, customized few-shot classification, and customized retrieval tasks demonstrate that SP-CRL achieves state-of-the-art performance with superior generalization.
|
| 2229 |
Efficient reduction of stellar contamination and noise in planetary transmission spectra using neural networks
2602.10330
|
cs.LG
|
David S. Duque-Casta\~no, Lauren Flor-Torres, Jorge I. Zuluaga |
The characterization of exoplanetary atmospheres has been transformed by the James Webb Space Telescope (JWST), whose infrared sensitivity enables transmission spectroscopy at unprecedented precision. However, stellar heterogeneities (e.g., spots and faculae) ...The characterization of exoplanetary atmospheres has been transformed by the James Webb Space Telescope (JWST), whose infrared sensitivity enables transmission spectroscopy at unprecedented precision. However, stellar heterogeneities (e.g., spots and faculae) remain a dominant source of contamination that can bias atmospheric retrievals if not properly corrected. We present a methodology for reducing stellar contamination and instrument-specific noise from exoplanet transmission spectra using neural networks, in particular the so-called Denoising AutoEncoders (DAEs). Our goals are to enable fast, accurate corrections that improve the reliability of atmospheric parameter retrievals and to promote the use of unsupervised algorithms for efficient data processing. We designed and trained DAE architectures using large synthetic datasets of terrestrial (TRAPPIST-1e analogues) and sub-Neptune (K2-18b analogues) planets. Atmospheric retrieval experiments were then performed on contaminated spectra in order to compare our deep-learning approach against standard correction methods in terms of accuracy and computational cost. Our autoencoders successfully reconstruct uncontaminated spectra, preserving essential molecular features even in low-S/N regimes. In retrieval tests, the denoising autoencoder pre-processing yields atmospheric parameter estimates broadly comparable to those obtained with simultaneous stellar-contamination fitting. Notably, our method maintains a much lower computational cost, approximately one order of magnitude smaller. These results demonstrate that DAEs outperform conventional correction methods in computational efficiency while maintaining high accuracy, paving the way for their integration into future atmospheric characterization pipelines for both rocky and sub-Neptune exoplanets.
|
| 2230 |
TADA! Tuning Audio Diffusion Models through Activation Steering
2602.11910
|
cs.LG
|
{\L}ukasz Staniszewski, Katarzyna Zaleska, Mateusz Modrzejewski, Kamil Deja |
Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remains challenging, as their internal mechanisms for representing high-level concepts are poorly understood. In this work,...Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remains challenging, as their internal mechanisms for representing high-level concepts are poorly understood. In this work, we use activation patching to demonstrate that recent audio diffusion architectures exhibit a semantic bottleneck, where a small, shared subset of consecutive attention layers controls distinct musical concepts, such as the presence of specific instruments, vocals, or genres. Building on this, we systematically evaluate a broad spectrum of steering paradigms, comparing activation steering against prompt-level, score-space, and weight-space interventions, analyzing the interaction between the steering mechanism and the intervention site. Our new benchmark, supported by an extensive user study, demonstrates that localized activation steering establishes a new state-of-the-art in audio concept modulation.
|
| 2231 |
Adjoint-based shape optimization of a ship hull using a Conditional Variational Autoencoder (CVAE) assisted propulsion surrogate model
2602.14907
|
cs.LG
|
Moloud Arian Maram, Georgios Bletsos, Thanh Tung Nguyen, Ahmed Hassan, Michael Palm |
Adjoint-based shape optimization of ship hulls is a powerful tool for addressing high-dimensional design problems in naval architecture, particularly in minimizing the ship resistance. However, its application to vessels that employ complex propulsion systems ...Adjoint-based shape optimization of ship hulls is a powerful tool for addressing high-dimensional design problems in naval architecture, particularly in minimizing the ship resistance. However, its application to vessels that employ complex propulsion systems introduces significant challenges. They arise from the need for transient simulations extending over long periods of time with small time steps and from the reverse temporal propagation of the primal and adjoint solutions. These challenges place considerable demands on the required storage and computing power, which significantly hamper the use of adjoint methods in the industry. To address this issue, we propose a machine learning-assisted optimization framework that employs a Conditional Variational Autoencoder-based surrogate model of the propulsion system. The surrogate model replicates the time-averaged flow field induced by a Voith Schneider Propeller and replaces the geometrically and time-resolved propeller with a data-driven approximation. Primal flow verification examples demonstrate that the surrogate model achieves significant computational savings while maintaining the necessary accuracy of the resolved propeller. Optimization studies demonstrate that neglecting the propulsion system can result in hull designs whose performance is inferior to that of the initial shape when subsequently validated using a numerically resolved propulsor. In contrast, the proposed method produces shapes that actually achieve more than an 8% reduction in resistance.
|
| 2232 |
Symmetric Composition of Anisotropic Operators for Global Subseasonal-to-Seasonal Climate Forecasting
2602.15040
|
cs.LG
|
Ziyu Zhou, Yuchen Fang, Weilin Ruan, Tian Zhou, Shiyu Wang |
Accurate global Subseasonal-to-Seasonal (S2S) climate forecasting is critical for disaster preparedness and resource management, yet it remains challenging due to chaotic atmospheric dynamics. Despite advances in geometry-aware representations, existing method...Accurate global Subseasonal-to-Seasonal (S2S) climate forecasting is critical for disaster preparedness and resource management, yet it remains challenging due to chaotic atmospheric dynamics. Despite advances in geometry-aware representations, existing methods do not specify how the zonal and meridional interactions that govern real atmospheric dynamics should be explicitly coordinated and composed, leaving an important architectural gap for S2S forecasting. In this paper, we propose AnisoCast, which explicitly models anisotropic global atmospheric dynamics for accurate S2S climate forecasting. It couples: (1) an Anisotropic Embedding strategy that tokenizes the global grid into latitudinal rings, preserving the integrity of zonal periodic structures; and (2) stacked Aniso Blocks that arrange latent Zonal and Meridional Operators in a weight-shared palindromic composition inspired by symmetric operator splitting. Extensive experiments on the ERA5 reanalysis dataset demonstrate that AnisoCast establishes a new state-of-the-art, significantly outperforming existing methods in both forecasting accuracy and computational efficiency.
|
| 2233 |
ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
2602.15983
|
cs.LG
|
Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin |
Large language models (LLMs) can translate natural-language problem descriptions into optimization code, but the code is prone to silent failures: it executes and returns a solver-feasible solution while encoding a semantically incorrect formulation. On compos...Large language models (LLMs) can translate natural-language problem descriptions into optimization code, but the code is prone to silent failures: it executes and returns a solver-feasible solution while encoding a semantically incorrect formulation. On compositional problems, the resulting feasibility-correctness gap reaches 90 percentage points. We introduce ReLoop, which combines two mechanisms. Structured generation decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify) to reduce formulation errors during generation. Behavioral verification detects the errors that remain by testing whether the formulation responds correctly to solver-based parameter perturbation, a signal that comes from the solver rather than from LLM self-review and requires no ground truth. The two mechanisms address different error structures: structured generation gives the largest gain on compositional problems (+8.5pp accuracy on RetailOpt-190 with Claude Opus 4.6), and behavioral verification gives its largest gain on localized defects (+4.4pp on MAMO-ComplexLP). With diagnostic execution recovery, ReLoop reaches 100% executable code on Claude Opus 4.6, and relative to direct generation it raises or preserves every reported metric of the three chat-tuned foundation models on all three benchmarks. For the narrowly fine-tuned SFT model we test, the chain-of-thought prompt conflicts with its learned output format and lowers its accuracy on MAMO-ComplexLP; we document and analyze this interaction. We release RetailOpt-190, 190 compositional retail optimization scenarios in which several constraints interact.
|
| 2234 |
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
2602.19710
|
cs.LG
|
Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling |
Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones opti...Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA), they excel at semantic identification but often overlook subtle 3D state variations that dictate distinct action patterns. To resolve these misalignments, we propose Pose-VLA, a decoupled paradigm that separates VLA training into a pre-training phase for extracting universal 3D spatial priors in a unified camera-centric space, and a post-training phase for efficient embodiment alignment within robot-specific action space. By introducing discrete pose tokens as a universal representation, Pose-VLA seamlessly integrates spatial grounding from diverse 3D datasets with geometry-level trajectories from robotic demonstrations. Our framework follows a two-stage pre-training pipeline, establishing fundamental spatial grounding via poses followed by motion alignment through trajectory supervision. Extensive evaluations demonstrate that Pose-VLA achieves state-of-the-art results on RoboTwin 2.0 with a 79.5% average success rate and competitive performance on LIBERO at 96.0%. Real-world experiments further showcase robust generalization across diverse objects using only 100 demonstrations per task, validating the efficiency of our pre-training paradigm.
|
| 2235 |
Descent-Guided Policy Gradient for Scalable Cooperative Multi-Agent Learning
2602.20078
|
cs.LG
|
Shan Yang, Yang Liu |
Scaling cooperative multi-agent reinforcement learning (MARL) is fundamentally limited by cross-agent noise. When agents share a common reward, each agent's learning signal is computed from a shared return that depends on all agents, so the stochasticity of th...Scaling cooperative multi-agent reinforcement learning (MARL) is fundamentally limited by cross-agent noise. When agents share a common reward, each agent's learning signal is computed from a shared return that depends on all agents, so the stochasticity of the other agents enters the signal as cross-agent noise that grows with $N$. Many engineering systems, such as cloud computing and power systems, have differentiable analytical models that prescribe efficient system states, providing a new reference beyond noisy shared returns. In this work, we propose Descent-Guided Policy Gradient (DG-PG), a framework that augments policy-gradient updates with a low-noise descent signal toward this reference. We prove that DG-PG reduces policy-gradient estimator variance from $\mathcal{O}(N)$ to $\mathcal{O}(1)$, preserves the stationary points of the cooperative objective, and achieves agent-independent sample complexity $\mathcal{O}(1/\epsilon^3)$. We evaluate DG-PG on cloud resource scheduling and power dispatch tasks, which involve discrete and continuous decisions, respectively. In both tasks, DG-PG converges within 10 episodes on average at every scale up to 1500 heterogeneous agents.
|
| 2236 |
ActionEngine: From Reactive to Programmatic Web Agents via State Machine Memory
2602.20502
|
cs.LG
|
Hongbin Zhong, Fazle Faisal, Luis Fran\c{c}a, Tanakorn Leesatapornwongsa, Adriana Szekeres |
Many web agents operate through a reactive execution loop: they observe the current interface, reason about the next action, execute it, and repeat. This design incurs latency and cost that grow with the number of actions, while requiring agents to repeatedly ...Many web agents operate through a reactive execution loop: they observe the current interface, reason about the next action, execute it, and repeat. This design incurs latency and cost that grow with the number of actions, while requiring agents to repeatedly rediscover how the same web application works. We present ActionEngine, a novel architecture that replaces step-by-step reasoning with programmatic execution using reusable knowledge of the application. A Crawling Agent explores the application offline and constructs an updatable state-machine memory that represents its GUI states, the operations available in each state, and the transitions between states. Unlike trajectory memory, this representation stores how the application works rather than solutions to individual tasks. At runtime, an Execution Agent uses this memory to synthesize a complete executable program in a single planning step, which is then executed deterministically without further planning calls. When the interface changes or the memory is incomplete, a reactive fallback repairs the failed action and updates the memory for future tasks. On 655 tasks across four WebArena domains, ActionEngine achieves a 91.2% success rate, outperforming the strongest reactive baseline, Claude Code, by 8.5 percentage points while reducing average task latency by 3.2x and cost by 8x.
|
| 2237 |
Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation
2603.01375
|
cs.LG
|
Chenxing Wei, Hong Wang, Ying He, Zhongxiang Dai, Bo Jiang |
Test-time policy adaptation for multi-turn interactions (T2PAM) is essential for aligning Large Language Models (LLMs) with dynamic user needs during inference time. However, existing paradigms commonly treat test-time adaptation as a single-axis problem, eith...Test-time policy adaptation for multi-turn interactions (T2PAM) is essential for aligning Large Language Models (LLMs) with dynamic user needs during inference time. However, existing paradigms commonly treat test-time adaptation as a single-axis problem, either purely refining instructions (Prompt Engineering) or only adjusting weights (Test-Time Training), ignoring that interaction failures stem from a coupled mix of ambiguity and incapacity. We argue that these two optimization paths are not merely additive but synergistic: semantic clarity acts as a pre-conditioner for effective parameter updates. To this end, we propose ROSA2, a framework that reformulates interaction as a joint optimization problem over the heterogeneous space of Words and Weights. By mathematically decomposing the error signal, ROSA2 utilizes textual gradients to rectify intent ambiguity and parameter updates to bridge capability gaps. Theoretically, we prove that this co-adaptation strictly reduces the required parameter shift for convergence. Empirically, ROSA2 outperforms state-of-the-art baselines by 30% on MATH while reducing interaction turns by 40%, demonstrating that refining the context unlocks the true potential of parameter updates.
|
| 2238 |
Gravity Falls: A Comparative Analysis of Domain-Generation Algorithm (DGA) Detection Methods for Mobile Device Spearphishing
2603.03270
|
cs.LG
|
Adam Dorian Wong, John D. Hastings |
Mobile devices are frequent targets of eCrime threat actors through SMS spearphishing (smishing) links that leverage Domain Generation Algorithms (DGA) to rotate hostile infrastructure, avoid individual domain blocks, and bypass perimeter enterprise defenses. ...Mobile devices are frequent targets of eCrime threat actors through SMS spearphishing (smishing) links that leverage Domain Generation Algorithms (DGA) to rotate hostile infrastructure, avoid individual domain blocks, and bypass perimeter enterprise defenses. Despite this, DGA research and evaluation largely emphasize malware C2 and email phishing datasets, leaving limited evidence on how well detectors generalize to smishing-driven domain tactics outside enterprise perimeters. This work addresses that gap by evaluating traditional and machine-learning DGA detectors against Gravity Falls, a new dataset derived from smishing links delivered between 2022 and 2025. Gravity Falls captures a single threat actor's evolution across four technique clusters, shifting from short randomized strings to dictionary concatenation and themed combo-squatting variants used for credential theft and fee/fine fraud. Two string-analysis approaches (Shannon entropy and Exp0se) and two ML-based detectors (an LSTM classifier and COSSAS DGAD) are assessed using Top-1M domains as benign baselines. Results are strongly tactic-dependent: performance is highest on randomized-string domains but drops on dictionary concatenation and themed combo-squatting, with generally low recall across multiple tool/cluster pairings. Overall, both traditional heuristics and some common ML detection methods are ill-suited for consistently evolving DGA tactics observed in Gravity Falls, motivating more context-aware approaches and providing a reproducible benchmark for future evaluation.
|
| 2239 |
Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
2603.08309
|
cs.LG
|
Yehonatan Elisha, Oren Barkan, Noam Koenigstein |
Vision Transformers (ViTs) often degrade under distribution shifts because they rely on spurious correlations, such as background cues, rather than semantically meaningful features. Existing regularization methods, typically relying on simple foreground-backgr...Vision Transformers (ViTs) often degrade under distribution shifts because they rely on spurious correlations, such as background cues, rather than semantically meaningful features. Existing regularization methods, typically relying on simple foreground-background masks, which fail to capture the fine-grained semantic concepts that define an object (e.g., ``long beak'' and ``wings'' for a ``bird''). As a result, these methods provide limited robustness to distribution shifts. To address this limitation, we introduce a novel finetuning framework that steers model reasoning toward concept-level semantics. Our approach optimizes the model's internal relevance maps to align with spatially grounded concept masks. These masks are generated automatically, without manual annotation: class-relevant concepts are first proposed using an LLM-based, label-free method, and then segmented using a VLM. The finetuning objective aligns relevance with these concept regions while simultaneously suppressing focus on spurious background areas. Notably, this process requires only a minimal set of images and uses half of the dataset classes. Extensive experiments on five out-of-distribution benchmarks demonstrate that our method improves robustness across multiple ViT-based models. Furthermore, we show that the resulting relevance maps exhibit stronger alignment with semantic object parts, offering a scalable path toward more robust and interpretable vision models. Finally, we confirm that concept-guided masks provide more effective supervision for model robustness than conventional segmentation maps, supporting our central hypothesis.
|
| 2240 |
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
2603.09160
|
cs.LG
|
Tzu-Heng Huang, Sirajul Salekin, Javier Movellan, Frederic Sala, Manjot Bilkhu |
Dense image captioning is critical for vision-language pretraining and for text-to-image generation, but scaling expert-quality annotations is prohibitively expensive. While synthesizing captions from strong vision-language models (VLMs) is a practical alterna...Dense image captioning is critical for vision-language pretraining and for text-to-image generation, but scaling expert-quality annotations is prohibitively expensive. While synthesizing captions from strong vision-language models (VLMs) is a practical alternative, supervised distillation often yields limited output diversity and weak generalization. Reinforcement learning (RL) could overcome these, but its successes have so far been concentrated in verifiable domains that rely on deterministic checkers---a luxury not available in open-ended captioning. We address this verification bottleneck with RubiCap, an RL framework that derives fine-grained, sample-specific rewards from LLM-written rubrics. For each image, an LLM rubric writer compares captions from a diverse committee of VLMs to identify consensus strengths and diagnose the current policy's deficiencies. These findings are converted into explicit evaluation criteria, enabling an LLM judge to decompose quality assessment and replace coarse scalar rewards with structured, multi-faceted assessments. Across extensive benchmarks, RubiCap achieves the highest win rates on CapArena, outperforming supervised distillation, prior RL methods, human-expert annotations, and GPT-4V-augmented outputs. On CaptionQA, it shows superior word efficiency: our 7B model matches Qwen2.5-VL-32B-Instruct, and our 3B model surpasses its 7B counterpart. Remarkably, using a compact RubiCap-3B as a captioner produces stronger pretrained VLMs than those trained on captions from proprietary models.
|
| 2241 |
Learning Transferable Sensor Models via Language-Informed Pretraining
2603.11950
|
cs.LG
|
Yuliang Chen, Arvind Pillai, Yu Yvonne Wu, Sudarshan Regmi, Tess Z. Griffin |
Multimodal language models have demonstrated strong semantic understanding and reasoning over physiological and behavioral signals captured from diverse healthcare sensors. However, existing sensor-language models can only process a fixed number of sensor chan...Multimodal language models have demonstrated strong semantic understanding and reasoning over physiological and behavioral signals captured from diverse healthcare sensors. However, existing sensor-language models can only process a fixed number of sensor channels at a fixed sampling rate, making them difficult to adapt to new or different sensor configurations. This inflexibility prevents models from generalizing across heterogeneous health sensing ecosystems, where data spans high-frequency wearable streams to sparse, day-scale clinical signals. To bridge this gap, we introduce SLIP (Sensor Language-Informed Pretraining), an open-source framework for learning language-aligned representations that generalize across diverse sensor setups. SLIP integrates contrastive alignment with sensor-conditioned captioning, facilitating both discriminative understanding and generative reasoning. By repurposing a pretrained decoder-only language model via cross-attention and introducing a flexible patch embedder, SLIP transfers to new sensor configurations at inference without retraining, regardless of their native sampling rate or input length. Across 11 datasets, SLIP demonstrates superior performance in retrieval, signal captioning, and question answering. It achieves a 77.14% average linear-probing accuracy, a 3.56% improvement over the strongest baseline (SensorLM). Beyond classification, SLIP supports open-vocabulary sensor captioning and question answering without any architectural modifications. Our work offers broad implications for the development of sensor-language models, utilizing cross-modal supervision to ensure more robust generalizability in health sensing applications.
|
| 2242 |
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
2603.12478
|
cs.LG
|
Rujie Wu, Haozhe Zhao, Hai Ci, Yizhou Wang |
Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven. We present Goal-Driven Data Optimization (GDO), a framework that computes six sample descriptors...Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven. We present Goal-Driven Data Optimization (GDO), a framework that computes six sample descriptors for each candidate and constructs optimized 1$\times$ training subsets for different goals. Under a fixed one-epoch Qwen3-VL-8B-Instruct training and evaluation recipe on 8 H20 GPUs, GDO uses far fewer training samples than the Uni-10x baseline while converging faster and achieving higher accuracy. Relative to the fixed 512k-sample Uni-10x baseline, GDO reaches the Uni-10x reference after 35.4k samples on MVBench, 26.6k on VideoMME, 27.3k on MLVU, and 34.7k on LVBench, while improving Accuracy by +1.38, +1.67, +3.08, and +0.84 percentage points, respectively. The gains are largest on MVBench and MLVU, while LVBench improves more modestly, consistent with its ultra-long-video setting and the mismatch between that benchmark and the short-video/image-dominant training pool. Across MinLoss, Diverse, Temp, and Temp+, stronger temporal emphasis yields steadily better long-video understanding behavior. Overall, GDO provides a goal-driven data optimization framework that enables faster convergence with fewer training samples under a fixed training protocol. Code is available at https://github.com/rujiewu/GDO.
|
| 2243 |
When Should Humans Step In? Optimal Human Dispatching in AI-Assisted Decisions
2603.13688
|
cs.LG
|
Lezhi Tan, Naomi Sagan, Lihua Lei, Jose Blanchet |
AI systems increasingly assist decision making by producing cheap factor-level assessments of complex inputs, but these assessments can be biased or incomplete. We study selective human evaluation: given AI signals, which factors should be dispatched to human ...AI systems increasingly assist decision making by producing cheap factor-level assessments of complex inputs, but these assessments can be biased or incomplete. We study selective human evaluation: given AI signals, which factors should be dispatched to human experts? We formulate this human--AI collaboration problem as a factor-level information-acquisition problem. Under squared loss, the optimal dispatching rule maximizes a contextual reward; under a linear model, this reward admits a closed-form decomposition into two interpretable terms: predictive relevance and residual uncertainty in the human evaluation after conditioning on the AI signal. We instantiate the framework on peer review, decomposing papers into ten aspects and evaluating up to 3,408 ICLR submissions across three LLMs and multiple regression heads. In this retrospective proxy, where human aspect signals are extracted from existing reviews, predictions that query three of the ten aspects under the optimal dispatching rules match the full ten-aspect human benchmark. The finding holds when the human signal comes from a single noisy reviewer, and across review years. These results position principled dispatching as a practical building block for scalable human--AI decision systems. We also discuss safeguards under which such a rule can serve as decision support that leaves the final decision with a human.
|
| 2244 |
Patient-level validation of a foundation-model pipeline for pediatric lung sounds: physician-annotated adventitious events are recognized, disease groups are not reliably predicted
2603.15688
|
cs.LG
|
Izzet Turkalp Akbasli, Oguzhan Serin |
Background: Lung-sound classifiers are usually evaluated on recording- or event-level splits, although each child contributes many recordings. We evaluated a foundation-model pipeline (PulmoVec) with the patient as the unit of partitioning and asked whether di...Background: Lung-sound classifiers are usually evaluated on recording- or event-level splits, although each child contributes many recordings. We evaluated a foundation-model pipeline (PulmoVec) with the patient as the unit of partitioning and asked whether disease group can be predicted beyond age and sex. Methods: We analyzed 19693 physician-annotated respiratory events from 736 children in the public SPRSound database. A frozen Health Acoustic Representations (HeAR) encoder with low-rank adapters was trained for screening (normal versus adventitious), sound pattern (normal, crackles, wheeze/rhonchi) and disease group (pneumonia, bronchial disease, normal/other); event probabilities were stacked with age, sex and auscultation site. The primary analysis was nested patient-grouped cross-validation; comparators were the majority class, annotated event duration and demographics. Results: The area under the receiver operating characteristic curve (AUC) was 0.95 (95% CI 0.94 to 0.96) for both screening and sound pattern, against 0.78 and 0.75 for event duration alone; the positive predictive value for adventitious events was 0.72. Disease-group prediction reached an AUC of 0.58 (95% CI 0.54 to 0.62), and its accuracy was below the majority-class rate at event level (0.44 versus 0.62) and at patient level (0.46 versus 0.60). Conclusions: Recognition of physician-annotated adventitious events holds under patient-level validation and is not explained by event duration or demographics, although its positive predictive value would limit unaided use. Prediction of disease group, a clinical diagnosis, from lung-sound events alone was not demonstrated. Lung-sound studies should partition data by patient and report non-acoustic comparators.
|
| 2245 |
Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling
2603.18095
|
cs.LG
|
Sooyoung Ryu, Mathieu Salzmann, Saqib Javed |
Post-training quantization (PTQ) is a practical path to deploy large diffusion models, but quantization noise can accumulate over the denoising trajectory and degrade generation quality. We propose Q-Drift, a sampler-side correction that aims to preserve the i...Post-training quantization (PTQ) is a practical path to deploy large diffusion models, but quantization noise can accumulate over the denoising trajectory and degrade generation quality. We propose Q-Drift, a sampler-side correction that aims to preserve the intended sampling marginals through a deterministic drift adjustment motivated by generalized marginal-preserving SDEs. Q-Drift uses the calibrated conditional residual variance of quantization error to determine a correction factor at each step. Our SDXL study shows that calibration with as few as 10 paired full-precision/quantized runs remains effective. The resulting sampler correction is plug-and-play with common samplers, diffusion models, and PTQ methods, while incurring negligible overhead at inference. Across six diverse text-to-image models (spanning DiT and U-Net), three samplers (Euler, flow-matching, DPM-Solver++), and two PTQ methods (SVDQuant, MixDQ), Q-Drift improves FID over the corresponding quantized baseline in all seven settings of our main evaluation, with up to 4.79 FID reduction on PixArt-Sigma (SVDQuant W3A4), while preserving CLIP scores. Code is available at https://github.com/sooyoung-ryu/Q-Drift.
|
| 2246 |
The Exponentially Weighted Signature
2603.19198
|
cs.LG
|
Alexandre Bloch, Benjamin Walker, Jo\"el Mouterde, Samuel N. Cohen, Terry Lyons |
We introduce the exponentially weighted signature (EWS), a continuous-time model that computes iterated integrals of a path, where each increment is weighted by the matrix exponential of a learnable generator over elapsed clock time. We prove that it solves a ...We introduce the exponentially weighted signature (EWS), a continuous-time model that computes iterated integrals of a path, where each increment is weighted by the matrix exponential of a learnable generator over elapsed clock time. We prove that it solves a linear controlled differential equation, keeps the group-like structure and the universality of the signature, and satisfies a modified Chen identity, enabling a parallel scan. At depth one the EWS is a state-space model (SSM), and we map linear time-invariant SSMs, Mamba channels and Mamba-$2$ heads to it in closed form. The EWS extends SSMs through an arbitrary matrix generator, a clock that generalises the step size to causal functionals of the input, and higher truncation depths that are non-linear in the path within a single layer. Empirically, the EWS achieves the highest average accuracy and rank on six long time-series classification datasets, where depth generally helps. Learned clocks prove necessary for state tracking on formal language tasks, and at depth one, the EWS matches or exceeds competing SSMs on regression and forecasting with far fewer parameters.
|
| 2247 |
Hyperspectral Trajectory Image for Multi-Month Trajectory Anomaly Detection
2603.25255
|
cs.LG
|
Md Awsafur Rahman, Chandrakanth Gudavalli, Hardik Prajapati, B. S. Manjunath |
Trajectory anomaly detection underpins applications from fraud detection to urban mobility analysis. Dense GPS preserves fine-grained evidence such as abnormal speeds and short-duration events, but its quadratic cost makes multi-month analysis intractable; spa...Trajectory anomaly detection underpins applications from fraud detection to urban mobility analysis. Dense GPS preserves fine-grained evidence such as abnormal speeds and short-duration events, but its quadratic cost makes multi-month analysis intractable; sparse stay-point methods scale by discarding that evidence and require a separate modeling regime. We argue that this bottleneck is unnecessary: dense and sparse trajectories share a natural two-dimensional cyclic structure along within-day and across-day axes. TITAnD (Trajectory Image Transformer for Anomaly Detection) is the first framework to cast both in a single representation, a Hyperspectral Trajectory Image (HTI), a day $\times$ time-of-day grid whose channels encode spatial, semantic, temporal, and kinematic information. Under this formulation, agent-level detection reduces to image classification and temporal localization to semantic segmentation. The Cyclic Factorized Transformer (CFT) models the two temporal axes directly, reducing attention cost by up to two orders of magnitude. With the HTI, multi-month dense anomaly detection becomes feasible for the first time, and CFT keeps it accurate and fast. Empirically, TITAnD matches or beats every evaluated baseline in AUC-PR on all four benchmarks. It surpasses vision models such as U-Net while using as few as 6.5M parameters, and it runs 11--75$\times$ faster than a capacity-matched flat Transformer on the same trajectory images.
|
| 2248 |
Stringological sequence prediction I: efficient algorithms for predicting highly repetitive sequences
2603.26852
|
cs.LG
|
Vanessa Kosoy |
We propose novel algorithms for sequence prediction based on ideas from stringology. These algorithms are time and space efficient and satisfy mistake bounds related to particular stringological complexity measures of the sequence. In this work (the first in a...We propose novel algorithms for sequence prediction based on ideas from stringology. These algorithms are time and space efficient and satisfy mistake bounds related to particular stringological complexity measures of the sequence. In this work (the first in a series) we focus on two such measures: (i) the size of the smallest straight-line program that produces the sequence, and (ii) the number of states in the minimal automaton that can compute any symbol in the sequence when given its position in base k as input. These measures are interesting because multiple rich classes of sequences studied in combinatorics of words (automatic sequences, morphic sequences, characteristic words) have low complexity and hence high predictability in this sense.
|
| 2249 |
HRIR-Former: Grid-Free Time-Domain Reconstruction of Head-Related Impulse Responses with a Spatially Encoded Transformer
2603.27998
|
cs.LG
|
Shaoheng Xu, Chunyi Sun, Jihui Zhang, Amy Bastine, Prasanga N. Samarasinghe |
Individualized head-related impulse responses (HRIRs) enable binaural rendering, but dense per-listener measurements are costly. We address HRIR spatial up-sampling from sparse per-listener measurements: given a few measured HRIRs for a listener, predict HRIRs...Individualized head-related impulse responses (HRIRs) enable binaural rendering, but dense per-listener measurements are costly. We address HRIR spatial up-sampling from sparse per-listener measurements: given a few measured HRIRs for a listener, predict HRIRs at unmeasured target directions. Prior learning methods often work in the frequency domain, rely on minimum-phase assumptions or separate timing models, and use a fixed direction grid, which can degrade temporal fidelity and spatial continuity. We propose HRIR-Former, a time-domain, grid-free binaural Transformer for reconstructing HRIRs at arbitrary directions from sparse inputs. It uses sinusoidal spatial features, a Conv1D refinement module, and auxiliary interaural time difference (ITD) and interaural level difference (ILD) heads. On SONICOM, it improves normalized mean squared error (NMSE), cosine distance, and ITD/ILD errors over prior methods; ablations validate modules and show minimum-phase preprocessing is unnecessary.
|
| 2250 |
Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
2604.14370
|
cs.LG
|
Carri W. Chan, Yi Han, Hannah Li, Benjamin L. Ranard |
AI tools increasingly drive targeted interventions in various service settings, including healthcare, education, and public services. Algorithms score individuals, trigger outreach to those above a threshold (e.g., high-risk or high-value), and encourage them ...AI tools increasingly drive targeted interventions in various service settings, including healthcare, education, and public services. Algorithms score individuals, trigger outreach to those above a threshold (e.g., high-risk or high-value), and encourage them to request service; then providers deliver service to those who request. Much of the work in this area has focused on improving predictive accuracy, implicitly assuming that better predictions lead to better outcomes. We show that predictions are only one component of a larger service system: when service capacity is limited and behavioral responses to outreach are probabilistic, system efficacy depends on operational forces that predictive accuracy does not capture. In such settings, the optimal score threshold must balance two effects: ensuring all capacity is filled (utilization) and, when capacity is constrained, preventing low-value requests from crowding out high-value ones (cannibalization). We characterize the optimal threshold and prove that thresholds based solely on predictive accuracy are generally suboptimal. Further, algorithm selection metrics such as AUC can be misaligned with operational performance: they weight all thresholds equally, while optimal deployment uses a subset of thresholds that depends on both capacity and compliance behavior. We introduce a new metric, Operational AUC (OpAUC), and show that it identifies the efficacy-optimal algorithm. Finally, we conduct a case study on sepsis early warning data that illustrates the magnitude of the improvement available from better threshold selection and shows that a predictor with lower AUC can achieve higher system efficacy under optimal deployment.
|
| 2251 |
Conformal Robust Set Estimation
2604.18441
|
cs.LG
|
Alejandro Cholaquidis, Emilien Joly, Leonardo Moreno |
We introduce a robust conformal prediction method based on the half-mass radius $\dP(x):=\inf\{r>0:P(B(x,r))>1/2\}$. Its empirical version is exactly the distance from a point to its $\mathbf{k}$-th nearest neighbour, with $\mathbf{k}=\lfloor n/2\rfloor+...We introduce a robust conformal prediction method based on the half-mass radius $\dP(x):=\inf\{r>0:P(B(x,r))>1/2\}$. Its empirical version is exactly the distance from a point to its $\mathbf{k}$-th nearest neighbour, with $\mathbf{k}=\lfloor n/2\rfloor+1$. The resulting conformal sets retain finite-sample marginal validity and asymptotically target the robust population central set $\Qba:=\{x\in\R^d:\dP(x)\le\beta_\alpha\}$. We prove consistency in symmetric-difference probability, derive non-asymptotic exponential bounds and obtain the rate $O((\log n/n)^{1/2})$ under a local mass condition. Under the same local regularity on the uncontaminated law, we also establish quantitative $O(\eta)$ stability, uniformly over arbitrary contaminations of sufficiently small mass $\eta$, prove Hausdorff convergence of the empirical central sets, and provide computable ball-based representations using only pairwise distances.
|
| 2252 |
Sliced-Regularized Optimal Transport
2604.23944
|
cs.LG
|
Khai Nguyen |
We propose a new regularized optimal transport (OT) formulation, termed sliced-regularized optimal transport (SROT). Unlike entropic OT (EOT), which regularizes the transport plan toward an independent coupling, SROT regularizes it toward a smoothened sliced O...We propose a new regularized optimal transport (OT) formulation, termed sliced-regularized optimal transport (SROT). Unlike entropic OT (EOT), which regularizes the transport plan toward an independent coupling, SROT regularizes it toward a smoothened sliced OT (SOT) plan. To the best of our knowledge, SROT is the first approach to leverage a version of SOT plan as a reference to improve classical OT. We provide a formal definition of SROT, derive its dual formulation, and provide a post-Bayesian interpretation of SROT. We then develop a Sinkhorn-style algorithm for efficient computation, retaining the same scalability advantages as EOT. By incorporating a scalable SOT plan as a prior, SROT yields more accurate approximations of the exact OT plan than EOT under the same level of regularization. Moreover, the resulting transport plan improves upon the reference SOT plan itself. We further introduce the corresponding OT divergence induced by SROT, named SROT divergence, and analyze its topological and computational properties. Finally, we validate our approach through experiments on synthetic datasets and color transfer tasks, demonstrating that SROT is better than both EOT and SOT in approximating exact OT. Additional experiments on gradient flows further highlight the advantages of SROT divergence.
|
| 2253 |
Multiscale Euclidean Network Trajectories: Second-Moment Geometry, Attribution, and Change Points
2605.04589
|
cs.LG
|
Haruka Ezoe, Ryohei Hisano |
Understanding a dynamic network requires more than locating when it changes: we also want to know the magnitude and structural direction of change and which nodes contribute. Unfolded spectral embedding places network snapshots in a joint representation, but t...Understanding a dynamic network requires more than locating when it changes: we also want to know the magnitude and structural direction of change and which nodes contribute. Unfolded spectral embedding places network snapshots in a joint representation, but the underlying rectangular latent factorization is invariant under general linear transformations. Such transformations preserve edge probabilities while potentially distorting Euclidean second moments. We introduce Multiscale Euclidean Network Trajectories (MENT), whose central estimands are pairwise second-moment operators of the dynamic latent positions. Under an isotropic normalization of the shared anchor latent positions, these operators are identifiable up to simultaneous orthogonal conjugation. They yield a global trace variation distance and, through an aggregated operator, shared orthogonal directions for mode-wise distances. The squared global distance decomposes exactly across these modes. Classical multidimensional scaling turns the distances into global and mode-wise trajectories of time. A modified unfolded spectral estimator consistently recovers the operators and induced trajectories, and we establish bounds connecting node attributions to trajectory displacements and trajectory estimation to change point localization. Synthetic experiments validate geometric recovery and show improved detection of mode-specific changes against baselines. Analyses of two real networks illustrate interpretable mode- and node-level descriptions of temporal evolution.
|
| 2254 |
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
2605.06643
|
cs.LG
|
Hao Dong, Hongzhao Li, Shupan Li, Muhammad Haris Khan, Eleni Chatzi |
Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of inconsistent evaluation protocols. Current ...Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of inconsistent evaluation protocols. Current research is fragmented, with studies varying significantly across datasets, modality configurations, and experimental settings. Furthermore, existing benchmarks focus predominantly on action recognition, often neglecting critical real-world challenges such as input corruptions, missing modalities, and model trustworthiness. This lack of standardization obscures a reliable assessment of the field's advancement. To address this issue, we introduce MMDG-Bench, the first unified and comprehensive benchmark for MMDG, which standardizes evaluation across six datasets spanning three diverse tasks: action recognition, mechanical fault diagnosis, and sentiment analysis. MMDG-Bench encompasses six modality combinations, nine representative methods, and multiple evaluation settings. Beyond standard accuracy, it systematically assesses corruption robustness, missing-modality generalization, misclassification detection, and out-of-distribution detection. With 7, 402 neural networks trained in total across 95 unique cross-domain tasks, MMDG-Bench yields five key findings: (1) under fair comparisons, recent specialized MMDG methods offer only marginal improvements over ERM baseline; (2) no single method consistently outperforms others across datasets or modality combinations; (3) a substantial gap to upper-bound performance persists, indicating that MMDG remains far from solved; (4) trimodal fusion does not consistently outperform the strongest bimodal configurations; and (5) all evaluated methods exhibit significant degradation under corruption and missing-modality scenarios, with some methods further compromising model trustworthiness.
|
| 2255 |
Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation
2605.06891
|
cs.LG
|
Aditya Parikh, Stella Frank, Sneha Das, Aasa Feragen |
Biases in annotation pipelines may introduce label bias: when annotation quality or style varies systematically across demographic subgroups, causing disparities in performance and auditability. Label bias in image segmentation remains underexplored, as even d...Biases in annotation pipelines may introduce label bias: when annotation quality or style varies systematically across demographic subgroups, causing disparities in performance and auditability. Label bias in image segmentation remains underexplored, as even detecting it typically requires clean, unbiased annotations, which are not readily available. We present a Confident Learning framework adapted to segmentation for auditing label bias directly in the training data without a clean, unbiased ground truth. By comparing the provided training labels to the model's confident predictions, we measure the prevalence and direction of suspected label errors; where standard overlap metrics like Dice fail. We further show that label bias influences subgroup separability in the encoder's feature space, an artifact we leverage for bias mitigation rather than suppressing it. We evaluate three datasets spanning from synthetic to real-life bias and across experimental conditions, showing how our framework produces proxy audit signals without clean labels and mitigates disparities with a correctly identified cleaner reference group.
|
| 2256 |
In-Context Credit Assignment via the Core
2605.06920
|
cs.LG
|
Keegan Harris, Siddharth Prasad, Asher Trockman |
We propose incentive-aligned mechanisms for in-context credit assignment: the task of assigning credit for AI-generated content among creators whose intellectual property appears in the context window. Our approach is based on the least core solution concept f...We propose incentive-aligned mechanisms for in-context credit assignment: the task of assigning credit for AI-generated content among creators whose intellectual property appears in the context window. Our approach is based on the least core solution concept from cooperative game theory, which distributes value in a way that is as stable as possible by ensuring that no subset of creators is significantly under-compensated relative to the value they could generate on their own. We develop algorithms for approximating the least core, which leverage novel routines for constraint seeding and constraint separation. On a web retrieval credit assignment task, we find that our approaches approximate the least core using an order of magnitude fewer LLM calls compared to alternative methods.
|
| 2257 |
TrajGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations
2605.06990
|
cs.LG
|
Maria Despoina Siampou, Gengchen Mai, Ni Lao, Jinmeng Rao, Neha Arora |
Many urban prediction tasks depend not only on the static attributes of a location, such as the visual appearance of its built environment, but also on its dynamic function: how people navigate and use it over time. Predicting congestion, travel demand, or roa...Many urban prediction tasks depend not only on the static attributes of a location, such as the visual appearance of its built environment, but also on its dynamic function: how people navigate and use it over time. Predicting congestion, travel demand, or road safety risks, for instance, requires mobility information that imagery and geographic coordinates alone cannot provide. Trajectories carry this signal, yet current trajectory representations are difficult to align with static geospatial observations: sparse, irregular GPS samples rarely coincide with street-view image coordinates, and existing encoders produce either a single embedding per trajectory or road segment, or embeddings only at the observed GPS samples. We introduce TrajGANR, a trajectory-centric multimodal self-supervised learning framework that represents each trajectory as a continuous, path-conditioned neural field. Reconstructed from the observed GPS samples, this representation can be queried at arbitrary coordinates (including the off-path locations where street-view images were captured), yielding localized mobility embeddings that align with street-view image and location representations at the same coordinate. Across four fine-grained road-level tasks in San Francisco and Porto, TrajGANR consistently outperforms recent geospatial and trajectory foundation models, with ablations attributing the gains to both the mobility signal and the fine granularity at which it is aligned. TrajGANR's location encoder further generalizes beyond the road network to areas where neither street-view imagery nor trajectories are available.
|
| 2258 |
BGM-IV: AI-Powered Bayesian Generative Modeling for Instrumental Variable Regression with High-Dimensional Covariates
2605.07029
|
cs.LG
|
Guyue Luo, Haidong Lu, Andrew J. Loza, Qiao Liu |
Instrumental-variable (IV) regression enables causal estimation under endogeneity, but modern IV problems often involve nonlinear structural effects and high-dimensional covariates. Existing methods typically operate in observed or generic learned feature spac...Instrumental-variable (IV) regression enables causal estimation under endogeneity, but modern IV problems often involve nonlinear structural effects and high-dimensional covariates. Existing methods typically operate in observed or generic learned feature spaces, and they often yield point estimates without uncertainty quantification. We introduce BGM-IV, a Bayesian generative modeling approach that performs nonlinear IV regression through posterior inference in a causally structured latent space. BGM-IV separates covariate variation by the role in the treatment and outcome mechanism, and accounts for endogeneity through an IV-integrated pseudo-likelihood that averages over instrument-induced treatment variation. The resulting model provides both structural-function estimates and predictive intervals for outcomes under intervention. Across various benchmark datasets, BGM-IV outperforms existing nonlinear IV methods overall, with significant gains in high-dimensional settings, while achieving near-nominal predictive coverage. These results highlight structured latent generative modeling as a flexible approach to uncertainty-aware IV inference with rich covariates. The code of BGM-IV is available at https://github.com/liuq-lab/BGM-IV.
|
| 2259 |
RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
2605.07129
|
cs.LG
|
Shijun Li, Pranav Belligundu, Tianxin Wei, Wooseong Yang, Yu Wang |
Large Language Models (LLMs) have emerged as a promising paradigm for next-generation recommender systems, offering strong semantic understanding and natural-language reasoning abilities. Despite recent progress, current LLM-based recommenders still face key c...Large Language Models (LLMs) have emerged as a promising paradigm for next-generation recommender systems, offering strong semantic understanding and natural-language reasoning abilities. Despite recent progress, current LLM-based recommenders still face key challenges in constructing decision-relevant contexts from heterogeneous evidence. First, existing methods often rely on fixed context construction strategies: collaborative behavioral evidence and item-side metadata are typically incorporated through predefined prompts, static retrieval pipelines, or handcrafted injection mechanisms, making it difficult to determine what information is truly beneficial for each instance. Second, heterogeneous evidence introduces a severe context-efficiency bottleneck. Rich metadata and collaborative interaction records can quickly overwhelm the context window, while aggressive compression or heuristic filtering may discard fine-grained evidence critical for accurate recommendation. To address these challenges, we propose RRCM, a ranking-driven retrieval-and-reasoning framework over collaborative and metadata memories for LLM-based agentic recommendation. RRCM starts from a lightweight user-history context and learns whether to recommend directly, retrieve collaborative evidence, retrieve item metadata, or interleave both through reasoning. Both memories are represented in natural language and accessed through a unified retrieval interface, enabling flexible evidence acquisition without handcrafted CF injection or fixed retrieval rules. We optimize this memory-reading policy with an outcome-only ranking reward, instantiated using group relative policy optimization, so that retrieval decisions are directly driven by final top-k recommendation quality. Extensive experiments show that RRCM significantly outperforms traditional baselines and diverse LLM-based recommendation approaches.
|
| 2260 |
Probabilistic Object Detection with Conformal Prediction
2605.07549
|
cs.LG
|
Christopher Ries, Moussa Kassem Sbeyti, Nicolas Bianco, Nadja Klein |
Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a suitable framework for reliable uncertainty quantification in safety-critical object detection. However, objec...Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a suitable framework for reliable uncertainty quantification in safety-critical object detection. However, object detection introduces structured multi-output predictions, complicating the application of classical CP theory developed for single outputs. In addition, standard, unscaled CP produces fixed-width prediction intervals across inputs, leading to unnecessary width for low-uncertainty predictions. While scaled CP addresses this by adapting the interval width to an input-dependent uncertainty estimate, prior work has neither systematically compared unscaled and scaled CP for multi-class object detection, nor integrated CP with a complementary uncertainty quantification method in this setting. We fill this gap by: (i) applying CP coordinate-wise to bounding box corners with a Bonferroni correction for box-level guarantees; (ii) scaling the resulting intervals using per-prediction aleatoric uncertainty estimates derived from a probabilistic object detector trained with loss attenuation, evaluated in uncalibrated and two calibrated variants; (iii) extending to a two-step pipeline that constructs prediction sets for the class using RAPS and conditions the conformalized bounding boxes on the predicted class set. Across three autonomous driving datasets (KITTI, BDD, CODA), including a cross-domain setting under distribution shift, scaled CP consistently improves interval sharpness over unscaled CP, achieving up to 19% higher IoU and 39% lower interval scores, without sacrificing coverage. Class-wise calibration further improves coverage for both variants with a negligible effect on sharpness. Together, these improvements yield more actionable uncertainty estimates for real-time, real-world object detection.
|
| 2261 |
ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
2605.08774
|
cs.LG
|
Youhe Feng, Hansen Shi, Haoyang Li, Xinlei Guo, Yang Wang |
Long-horizon robotic manipulation requires dense feedback that reflects how a task advances through its procedural stages, not merely whether the final outcome is successful. Existing reward models often rely on trajectory-level success labels or time-based in...Long-horizon robotic manipulation requires dense feedback that reflects how a task advances through its procedural stages, not merely whether the final outcome is successful. Existing reward models often rely on trajectory-level success labels or time-based interpolation, which can conflate elapsed time with true task progress and therefore fail to capture unfinished steps, stagnation, and failure states. We present ProcVLM, a progress-aware vision-language model that learns procedure-grounded progress as a dense reward signal for manipulation. Rather than deriving progress from terminal outcomes or temporal proxies, ProcVLM grounds progress estimation in procedural structure and intra-stage visual change, and further adopts a reasoning-before-estimation paradigm that infers the remaining atomic actions before estimating task progress. Specifically, we construct this supervision by synthesizing frame-level subtask-semantic annotations, assigning progress budgets according to subtask structure, and distributing each budget based on intra-subtask visual change. To train ProcVLM at scale, we build a standardized procedural supervision synthesis pipeline and construct ProcCorpus-60M from 30 embodied datasets with 60M annotated frames, from which we derive ProcVQA for procedure-aware pretraining, with progress estimation as the central task alongside action segmentation and future planning. Experiments on ProcVQA and reward-model benchmarks show that ProcVLM improves embodied procedural reasoning and yields more discriminative trajectory-internal progress estimates than representative baselines, supporting its use as a dense reward model for downstream reward-guided policy optimization. Project page: https://procvlm.github.io/
|
| 2262 |
Broximal Gradient Descent: A Projection-Free Sister of Projected Gradient Descent
2605.08850
|
cs.LG
|
Peter Richt\'arik, Kaja Gruntkowska, Hanmin Li |
We propose Broximal Gradient Descent (BroxGD), a projection-free sister method to projected gradient descent for constrained optimization. Its forward--backward construction replaces the proximal backward operation by the broximal operation of Gruntkowska et a...We propose Broximal Gradient Descent (BroxGD), a projection-free sister method to projected gradient descent for constrained optimization. Its forward--backward construction replaces the proximal backward operation by the broximal operation of Gruntkowska et al. (2025). Instead of projecting, each step minimizes a linear function over the intersection of the constraint set $\mathcal{X}$ and a ball $\mathbb{B}(x_k,t_k)$ centered at the current iterate $x_k$, of suitable radius $t_k>0$: \[ x_{k+1}\in\arg\min_{z\in\mathcal{X}\cap\mathbb{B}(x_k,t_k)}\langle\nabla f(x_k),z\rangle. \] We develop a comprehensive convergence theory spanning a wide range of optimization regimes and radius rules. We expect BroxGD to find many applications and inspire numerous extensions, much like projected gradient descent. Our contribution is theoretical; potential applications and toy experiments illustrate the method and suggest directions for future work.
|
| 2263 |
PoDAR: Power-Decoupled Audio Representation for Generative Modeling
2605.10084
|
cs.LG
|
Alejandro Luebs, Mithilesh Vaidya, Ishaan Kumar, Sumukh Badam, Stephen W. Bailey |
The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of...The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of audio codecs, we demonstrate that latent modelability can be significantly improved through explicit factor decoupling. We present PoDAR (Power-Decoupled Audio Representation), a framework that utilizes a randomized power augmentation and latent consistency objective to decouple signal power from invariant semantic content. This factorization makes the latent space easier to model, which both accelerates the convergence of downstream generative models and improves final overall performance. When applied to a Stable Audio 1.0 VAE with an F5-TTS generator, PoDAR achieves about a 2x acceleration in convergence to match baseline performance, while increasing final speaker similarity by 0.055 and UTMOS by 0.22 on the LibriSpeech-PC dataset. Furthermore, isolating power into dedicated channels enables the application of CFG exclusively to power-invariant content, effectively extending the stable guidance regime to higher scales.
|
| 2264 |
Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration
2605.10330
|
cs.LG
|
Btissame El Mahtout, Florian Ziel |
We propose a novel Mixture-of-Experts framework, referred to as Expert Loss Integration MoE (EliMoE), for time series forecasting. EliMoE addresses the optimization problem arising from small gating weights by incorporating expert-specific losses, which provid...We propose a novel Mixture-of-Experts framework, referred to as Expert Loss Integration MoE (EliMoE), for time series forecasting. EliMoE addresses the optimization problem arising from small gating weights by incorporating expert-specific losses, which provide each expert with a direct learning signal independent of the gate-assigned weight. Specifically, the overall objective comprises the base forecasting loss and expert-specific losses, allowing individual expert prediction errors to directly influence parameter updates alongside the aggregate forecasting error. The framework also encourages different experts to learn from different temporal segments of the data. The proposed framework is further combined with a partial online learning strategy that enables efficient incremental updates of model parameters. By integrating expert-level loss information with partial online learning, EliMoE improves forecasting performance while retaining computational efficiency. Empirical results across five benchmark datasets from diverse application domains with different sampling frequencies show that the proposed approach generally outperforms state-of-the-art supervised neural forecasting models, including Transformer-based architectures such as PatchTST, as well as zero-shot time-series foundation models such as TimeMoE. Furthermore, ablation studies confirm the effectiveness of the expert-specific loss integration strategy, highlighting its contribution to enhancing predictive performance.
|
| 2265 |
FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity
2605.11869
|
cs.LG
|
Jian Tang, Jiawei Fan, Qingbin Liu, Zheng Wei, Jiang Bian |
While the overall inference latency of Video Diffusion Transformers (DiTs) can be substantially reduced through model distillation, per-step inference latency remains a critical bottleneck. Existing acceleration paradigms primarily exploit redundancy across th...While the overall inference latency of Video Diffusion Transformers (DiTs) can be substantially reduced through model distillation, per-step inference latency remains a critical bottleneck. Existing acceleration paradigms primarily exploit redundancy across the denoising trajectory; however, we identify a limitation where these step-wise strategies encounter diminishing returns in few-step regimes. In such scenarios, fewer intermediate states can limit opportunities for feature reuse or predictive modeling, motivating complementary forms of acceleration. To overcome this, we propose Frame Interleaved Sparsity DiT (FIS-DiT), a training-free and operator-agnostic framework that shifts the optimization focus from the temporal trajectory to the latent frame dimension. Our approach is motivated by an intrinsic duality within this dimension: the existence of frame-wise sparsity that permits reduced computation, coupled with a structural consistency that calls for maintaining coverage of frame positions in the global spatiotemporal context. Leveraging this insight, we implement Frame Interleaved Sparsity (FIS) as an execution strategy that manipulates frame subsets across the model hierarchy, refreshing all latent positions without requiring full-scale block computation. Empirical evaluations on Wan 2.2 and HunyuanVideo 1.5 demonstrate that FIS-DiT reaches 2.4$\times$ speedup while maintaining comparable performance on VBench, providing a scalable and robust pathway toward real-time high-definition video generation.
|
| 2266 |
EHR-RAGp: Prototype-Guided Retrieval of Longitudinal Electronic Health Records for Clinical Prediction Models
2605.12335
|
cs.LG
|
Saeed Shurrab, Mariam Al-Omari, Dana El Samad, Farah E. Shamout |
Electronic Health Records (EHR) contain rich longitudinal patient information and are widely used in predictive modeling applications. However, effectively leveraging historical data remains challenging due to long trajectories, heterogeneous events, temporal ...Electronic Health Records (EHR) contain rich longitudinal patient information and are widely used in predictive modeling applications. However, effectively leveraging historical data remains challenging due to long trajectories, heterogeneous events, temporal irregularity, and the varying relevance of past clinical context. Existing approaches often rely on fixed windows or uniform aggregation, which can obscure clinically important signals. In this work, we introduce EHR-RAGp, a retrieval-based framework that dynamically integrates the most relevant patient history consisting of diverse clinical event types. We propose a prototype-guided retrieval module that acts as an alignment mechanism and estimates the relevance of retrieved historical chunks with respect to a given prediction task, guiding the model towards the most informative context. Across multiple clinical prediction tasks and two benchmark datasets, EHR-RAGp consistently outperforms state-of- the-art EHR-based and transformer-based baselines. Furthermore, EHR-RAGp is model-agnostic, as integrating it with various backbone models yields substantial performance gains. Overall, EHR-RAGp establishes a novel direction for modeling long-range clinical context to improve downstream performance via retrieval.
|
| 2267 |
Watermarking Should Be Treated as a Monitoring Primitive
2605.13095
|
cs.LG
|
Toluwani Aremu, Jie Zhang, Nils Lukas |
Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models. We argue that it should be evaluated and governed as a monitoring primitive through two complementary observer models. Internal observers use detector or d...Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models. We argue that it should be evaluated and governed as a monitoring primitive through two complementary observer models. Internal observers use detector or decoder access and entity mappings for attribution; external observers learn entity-specific signals from labeled outputs without keys or detectors. With persistent entity bindings and reliable inference, either pathway can support monitoring. We show that even zero-bit watermarking supports internal attribution under per-entity multi-key deployments without explicitly encoding identity, and demonstrate external identification in selected text and image configurations. External exposure depends on persistent, learnable watermark structure and is not universal, while internal attribution also remains conditional on reliability and access. These findings motivate governance of attribution access and deployment choices alongside evaluation of design-dependent entity linkability, de-anonymization or re-identification, beyond per-sample robustness.
|
| 2268 |
Synthetic American Option Pricing via Jump-HMM-Driven Heston Implied Volatility
2605.13998
|
cs.LG
|
Julia Sun, Zheyu Jin, Jiawei Zhang, Jeffrey D. Varner |
Valuing American options along simulated stock paths requires an implied volatility (IV) for every option on every date. A stock-return model alone does not provide it. We built a simulator that assigned IV to American options on simulated dates for any chosen...Valuing American options along simulated stock paths requires an implied volatility (IV) for every option on every date. A stock-return model alone does not provide it. We built a simulator that assigned IV to American options on simulated dates for any chosen stock model. We fitted parametric and neural IV surfaces to vendor option quotes for 31 tickers. Each surface predicted IV from moneyness and time to expiration. Stock paths came from a jump hidden Markov model with capped daily returns. Each day, every option's implied variance moved partway toward its surface prediction, following the form of the Heston variance equation. Random shocks tended to raise IV when the stock fell. A binomial tree converted IV into option values and price sensitivities. Neural surfaces fitted by sector or ticker matched the quotes more closely than one parametric surface. Their errors still varied by date and carried into dollar prices. Repricing identical stock paths under five ways of updating IV changed option values before expiration, and the worst simulated short-position losses at a fixed horizon. Payoffs at expiration did not change. In forecasts of later Goldman Sachs and Eli Lilly option prices, stochastic IV did little better than fixed IV. Rerunning the forecasts with the realized stock paths pointed to stock prediction as a major source of option-price error. The hidden Markov model and an adaptive-volatility model predicted stock prices about as accurately as assuming no change, and tuning found no gain from predicting direction. The adaptive-volatility model improved predicted price ranges for Eli Lilly but not Goldman Sachs. The simulator supported reproducible comparisons of IV and stock assumptions but did not forecast better than simpler alternatives. Complete option histories and prices consistent across strikes are needed before simulated prices can replace market data.
|
| 2269 |
Regret Equals Covariance: A Closed-Form Characterization for Stochastic Optimization
2605.14019
|
cs.LG
|
Irene Aldridge |
Quantifying regret typically involves repeatedly re-solving a stochastic optimization problem across simulated scenarios. Such an approach is costly: the cost scales with the number of scenarios and the cost of each solve. We propose that regret can instead be...Quantifying regret typically involves repeatedly re-solving a stochastic optimization problem across simulated scenarios. Such an approach is costly: the cost scales with the number of scenarios and the cost of each solve. We propose that regret can instead be measured using the joint moments of the cost vector $c$ and the optimal decision $\pi^*(c)$. Working with the optimal concave and Lipschitz under only a feasible region value function $v(c) = \min_{z\in Z} c^\top z$, we show that regret always decomposes as $\mathrm{Regret} = -\mathrm{Cov}(c,\pi^*(c)) - R(c)$ for a residual $R(c)$. We further show that $R(c)=0$ when $\pi^*$ is affine in $c$, which holds for unconstrained and equality-constrained (e.g., budget-constrained Markowitz) quadratic programs and many, but not all, linear programs. For all cases, we give a distribution-free bound $\mathrm{Regret} \le D\sqrt{\mathrm{tr}(\Sigma_c)}$ that depends only on the radius of the feasible region and the cost covariance. We also derive concentration and asymptotic-normality results for the empirical regret, the latter under an explicit and checkable non-degeneracy condition.
|
| 2270 |
CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage
2605.15597
|
cs.LG
|
Jiale Liu, Jungang Li, Jieming Yu, Xinglin Yu, Zihao Dongfang |
Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructions do not directly provide a sparse, comparable, and geometry-consistent panoramic training interface. Dense...Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructions do not directly provide a sparse, comparable, and geometry-consistent panoramic training interface. Dense trajectories duplicate nearby views, source-specific rendering policies yield heterogeneous annotations, and sparse heuristics may miss important regions or introduce depth-inconsistent observations. We study how to convert 3D assets into sparse panoramic RGB-D-pose data that preserves complete scene coverage with low redundancy and auditable provenance. We propose COVER (Coverage-Oriented Viewpoint curation with ERP Range-depth warping), a training-free ERP viewpoint curator that projects geometry observed from selected views into candidate ERP probes, scores incremental coverage, and penalizes depth conflicts. Under bounded proxy error, its greedy coverage proxy preserves the standard coverage-style approximation behavior up to an additive error term. Using COVER, we build CM-EVS (Coverage-curated Metric ERP View Set), a panoramic RGB-D-pose dataset with 36,373 curated ERP frames from 1,275 indoor scenes across Blender indoor, HM3D, and ScanNet++, complemented by outdoor panoramas from TartanGround and OB3D re-encoded into the same schema. Each frame provides full-sphere RGB, metric range depth, calibrated pose; COVER-produced indoor frames include per-step provenance logs. With a median of only 25 frames per indoor scene, CM-EVS covers all 13 unified room types while maintaining compact scene-level coverage. Experiments show that COVER improves the coverage-conflict trade-off, making CM-EVS a sparse, compact, and auditable RGB-D-pose resource for geometry-consistent panoramic 3D learning.
|
| 2271 |
Generalized Functional ANOVA: A Complete Theoretical Framework
2605.18422
|
cs.LG
|
Baptiste Ferrere, Nicolas Bousquet, Fabrice Gamboa, Jean-Michel Loubes |
The functional ANOVA provides a fundamental representation of square-integrable multivariate functions into main effects and higher-order interactions. For independent inputs, the components belong to mutually orthogonal Hilbert subspaces and admit an explicit...The functional ANOVA provides a fundamental representation of square-integrable multivariate functions into main effects and higher-order interactions. For independent inputs, the components belong to mutually orthogonal Hilbert subspaces and admit an explicit representation. For dependent inputs, however, the components are only hierarchically orthogonal: although existence and uniqueness results are available, the Hilbert subspaces underlying the generalized decomposition have remained implicit. We resolve this representation problem for continuous inputs supported on a bounded hyperrectangle whose joint density is bounded above and away from zero. We introduce a distribution-adapted family of functions and prove that it forms a Riesz basis of the $L^2$ space, thereby guaranteeing a unique, stable, and unconditionally convergent representation. We then show that, for every coalition of variables, the corresponding block of this basis exactly characterizes the Hilbert subspace containing the functional ANOVA component. Our construction recovers the classical orthogonal decomposition under input independence. As a direct consequence, computing the generalized functional ANOVA reduces to estimating coefficients in an explicit, distribution-adapted basis. Finally, as a \emph{proof of concept}, we introduce an elementary, fast and model-agnostic estimator based on our theoretical results. Experiments on synthetic and real-world datasets illustrate its connections with established tabular explanation methods and show that low order components often capture most of the signal in the model output.
|
| 2272 |
Uncertainty-aware neural emulation reveals climate-resilient maize traits at scale
2605.22848
|
cs.LG
|
Mojdeh Saadati, Juan Panelo, Gustavo Visentini, Soumik Sarkar, Carlos Messina |
Global food security depends on predicting crop responses to climate variability, yet process based crop models remain too computationally expensive for large scale exploration of genotype and environment interactions. Here we develop a probabilistic neural em...Global food security depends on predicting crop responses to climate variability, yet process based crop models remain too computationally expensive for large scale exploration of genotype and environment interactions. Here we develop a probabilistic neural emulator of APSIM that reproduces key maize growth processes across 13 outputs with high fidelity (with R^2 of 0.93) while reducing simulation time by several orders of magnitude. Trained on two million simulations spanning diverse genetic, soil, and management conditions, and augmented with a convolutional synthetic weather generator that produces physically consistent climate sequences, the framework enables scalable exploration of crop responses under realistic and diverse environmental inputs while providing calibrated predictive uncertainty without costly Bayesian inference. Applying this framework across 100,000 trait configurations, six soil environments in Iowa and Illinois, and climate projections through the year 2100 under two emissions scenarios, we identify 181 maize trait combinations that consistently maintain high yield across all tested conditionsan analysis infeasible with the mechanistic model alone. We further show that radiation use efficiency and temperature driven root dynamics are dominant drivers of yield resilience. Notably, projected yield distributions vary substantially across locations, with some lower productivity sites exhibiting yield increases under future climate scenarios, indicating that climate change may reshape regional yield potential in nonintuitive ways. These results demonstrate how uncertainty aware emulation transforms mechanistic crop simulation from a computational bottleneck into an on demand discovery engine, one capable of interrogating the full genotype, environment and management space at a scale no process-based model can match.
|
| 2273 |
Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
2605.25590
|
cs.LG
|
Joongkyu Lee, Min-hwan Oh |
We study nonstationary generalized linear bandits (GLBs), where the expected reward is modeled through a nonlinear link function with an unknown time-varying parameter. This framework encompasses a broad class of reward models, including linear, Bernoulli, and...We study nonstationary generalized linear bandits (GLBs), where the expected reward is modeled through a nonlinear link function with an unknown time-varying parameter. This framework encompasses a broad class of reward models, including linear, Bernoulli, and binomial rewards. Existing approaches are predominantly based on maximum-likelihood estimation (MLE), using sliding-window, restart, or discounting mechanisms to handle nonstationarity. Although these methods achieve statistically efficient regret guarantees, they generally require revisiting past observations at every round, which leads to computation and memory costs that grow with time; moreover, several of them rely on a non-convex projection step. In this paper, we propose DOMD-GLB, a new algorithm for nonstationary GLBs that utilizes discounted online mirror descent (DOMD) for parameter estimation, thereby incurring only $O(1)$ computation and memory costs per round. We prove dynamic regret bounds of order $\tilde{O} \big(c_\mu^{-1/2} d^{3/4} P_T^{1/4} T^{3/4}\big)$ in drifting environments and $\tilde{O}\big(c_\mu^{-1/3} d^{2/3} \Gamma_T^{1/3} T^{2/3}\big) $in piecewise-stationary environments, where $d$ denotes the feature dimension, $T$ the time horizon, $P_T$ the path length, $\Gamma_T$ the number of change points, and $c_\mu$ a curvature parameter associated with the link function, while substantially improving computational efficiency over prior work. To the best of our knowledge, this is the first algorithm for nonstationary GLBs with per-round computation and memory costs independent of time.
|
| 2274 |
Equivalent Flows, Unequal Learning: Clean-Latent Prediction in Transformers
2605.27102
|
cs.LG
|
Funing Fu, Tenghui Wang, Guanyu Zhou, Junyong Cen, Qichao Zhu |
Flow samplers consume velocity, but the neural network can predict the clean endpoint and convert it to velocity through a fixed affine readout. We study this choice with JLT, a latent Transformer in a frozen variational autoencoder (VAE) representation. For s...Flow samplers consume velocity, but the neural network can predict the clean endpoint and convert it to velocity through a fixed affine readout. We study this choice with JLT, a latent Transformer in a frozen variational autoencoder (VAE) representation. For squared error, the optimal clean and velocity predictors are algebraically equivalent; a finite Transformer assigns different computation to its learned output under the two interfaces. A local Gaussian analysis identifies a known residual response supplied by the readout and isotropic target variance added by velocity prediction. Measured FLUX.2 channel spectra support this geometric distinction: 90% of target variance occupies 83 of 128 clean directions versus 109 velocity directions. Under a matched velocity objective, clean prediction improves ImageNet FID-50K from 6.56 to 2.70 at Base scale and from 2.12 to 1.47 at Large scale, with lower FID at every measured Large checkpoint. Scaling clean prediction to 951M parameters reaches FID-50K 1.19 and IS 271.96. In addition, an objective ablation at Base scale shows that direct clean regression reaches FID-50K 2.38 without time-dependent error weighting. These results show how moving known computation outside the network changes learning under algebraically equivalent flow interfaces. Code: https://github.com/akatsuki-neo/JLT/blob/main/README.md
|
| 2275 |
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
2605.30120
|
cs.LG
|
Lixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng, Stefanie Jegelka |
Multi-vector retrieval (MVR) models, exemplified by ColBERT, have established new benchmarks in retrieval accuracy by preserving fine-grained token-level interactions. However, this granularity imposes prohibitive storage and retrieval efficiency bottlenecks: ...Multi-vector retrieval (MVR) models, exemplified by ColBERT, have established new benchmarks in retrieval accuracy by preserving fine-grained token-level interactions. However, this granularity imposes prohibitive storage and retrieval efficiency bottlenecks: to manage the immense memory footprint and computational overhead of billion-scale token vectors, state-of-the-art systems are forced to rely on aggressive dimension reduction and complex clustering (e.g., K-means). This compromise introduces two critical limitations: excessive indexing latency of clustering large-scale corpora and semantic information loss inherent to compression. In this paper, we propose Single-stage Sparse Retrieval (SSR}, a paradigm shift that replaces expensive clustering with efficient sparse coding. Instead of compressing features into low-dimensional dense vectors, we utilize Sparse Autoencoder (SAE) to project token embeddings into a high-dimensional but highly sparse representation. This transformation enables us to bypass vector clustering entirely and leverage inverted indexing for precise, high-throughput retrieval. Extensive experiments on the BEIR benchmark demonstrate that SSR achieves a "trifecta" of improvements: it reduces indexing time by 15x compared to ColBERTv2, halves retrieval latency, and simultaneously improves retrieval performance over leading baselines.
|
| 2276 |
Diagnosing Visual Ignorance in Vision-Language Models
2606.06890
|
cs.LG
|
Runyu Zhou, Qi Zhang, Yisen Wang |
Vision-language models (VLMs) achieve high accuracy on many visual question-answering benchmarks, yet it remains unclear whether this accuracy reflects reliable use of the visual details a question depends on. We examine this tension through the lens of visual...Vision-language models (VLMs) achieve high accuracy on many visual question-answering benchmarks, yet it remains unclear whether this accuracy reflects reliable use of the visual details a question depends on. We examine this tension through the lens of visual ignorance: cases where a model can represent task-relevant visual evidence without using it to determine the answer. Across three VLMs and twelve benchmarks, we track the answer-relevant information decodable at each decoder layer and measure when it becomes coupled to the final output. The visual evidence needed for the correct answer can already be read out from intermediate layers, yet it has little effect on the output until the final layers of the decoder. Causal residual interchange confirms that shifting the answer requires the late states, not the earlier ones. We then evaluate all twelve benchmarks under progressive Gaussian blur and find that substantial fractions of predictions survive the entire degradation sequence while continuing to receive benchmark credit. These results distinguish the availability of visual information from its influence on prediction, and show that standard accuracy can coexist with limited sensitivity to the tested visual detail. They motivate training and evaluation that make answer-relevant visual distinctions necessary to answer correctly.
|
| 2277 |
Driving Video Retrieval for Complex Queries with Structured Grounding
2606.09109
|
cs.LG
|
Manyi Yao, Sparsh Garg, Christian Shelton, Amit Roy-Chowdhury, Abhishek Aich |
Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods of...Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant motion may not be explicitly described in text or captured by lexical overlap. Rule-based retrieval can encode such events more directly, but it is brittle: generated or hand-written rules often fail when their assumptions do not match real driving data. We propose STRIVE-D, a data-calibrated retrieval framework for driving videos. It uses weakly labeled in-domain videos to estimate when a query rule is reliable, adapt rules that mismatch observed data, and fuse calibrated rule scores with vision-language and keyword-based retrieval signals. Across three driving benchmarks, including newly released human-annotated event data on DrivingDojo, STRIVE-D delivers up to 84% relative improvement in top-1 accuracy over state-of-the-art methods.
|
| 2278 |
PRISM: Recovering Instruction Sets from Language Model Activations
2606.09563
|
cs.LG
|
Gilad Gressel, Rahul Pankajakshan, Julia Diament, Efim Hudis, Krishnashree Achuthan |
As LLMs are deployed as agents, reliable monitoring requires knowing not only what they output, but which instructions are steering their behavior. This is difficult when models infer unintended subgoals, follow contextual cues, or are influenced by prompt inj...As LLMs are deployed as agents, reliable monitoring requires knowing not only what they output, but which instructions are steering their behavior. This is difficult when models infer unintended subgoals, follow contextual cues, or are influenced by prompt injections and hidden objectives. While activation-to-language methods suggest that hidden states can reveal natural-language information, existing approaches are not designed to recover the full set of simultaneous instructions, constraints, prohibitions, and subgoals present in agentic settings. As a first step toward this monitoring goal, we formalize instruction set retrieval and introduce PRISM, an activation-conditioned interpreter that decodes hidden states from a frozen target model into a faithful bullet list of instructions. Unlike prior activation-to-language methods, PRISM is trained to recover instruction sets directly, using judge-guided GRPO to reward covered instructions and penalize unsupported ones. Across benign, constrained, prompt-injection, and hidden-objective settings, PRISM outperforms activation-to-language baselines, especially on security-relevant objectives.
|
| 2279 |
Robust Active Learning for Few-Shot Example Selection in Text-to-SQL
2606.10125
|
cs.LG
|
Arash Pourhabib |
Domain-specific text-to-SQL systems ground a large language model by retrieving annotated few-shot examples, and each example needs expert-written SQL. We treat the choice of which queries to annotate as constrained experimental design on the low-dimensional m...Domain-specific text-to-SQL systems ground a large language model by retrieving annotated few-shot examples, and each example needs expert-written SQL. We treat the choice of which queries to annotate as constrained experimental design on the low-dimensional manifold of query embeddings, with query-dependent annotation noise, a partition matroid constraint that spreads selections across semantic domains, and an unknown covariance structure. We propose a stratified greedy algorithm that maximizes a heteroscedastic information-gain objective. We prove that the objective is monotone and submodular under query-dependent noise, so stratified greedy selection carries a 1/2-approximation guarantee under the partition constraint. Under kernel misspecification the guarantee degrades by an additive spectral term; we compute it on both experimental pools and find it too large for the bound to be quantitatively informative. To connect the design objective to the downstream task, we give a retrieval model that bounds few-shot accuracy from below by per-domain fill distance, demonstration noise, and domain coverage, and we calibrate its locality assumption on both pools. On an enterprise supply-chain corpus and on the BIRD benchmark, the selected banks improve cross-domain retrieval and end-to-end LLM SQL over random and distance-based selection at the same annotation budget. Stratified controls and pre-specified tests show that the gain comes from the partition constraint: uniform sampling within each stratum matches the full method in the oracle-label evaluations, farthest-point selection within strata adds a little at small budgets, and the noise weighting has no measurable effect. The practical advice is to annotate one example per domain per batch from the first batch on.
|
| 2280 |
Trainability of IQP Quantum Circuit Born Machines Under Gaussian Initialization
2606.10179
|
cs.LG
|
Gennaro De Luca, Vinayak Sharma, Aviral Shrivastava |
Quantum Circuit Born Machines (QCBMs) offer a natural approach to generative machine learning by leveraging the Born rule. Recent work has provided a method to classically train QCBMs with Instantaneous Quantum Polynomial (IQP) circuits via the Maximum Mean Di...Quantum Circuit Born Machines (QCBMs) offer a natural approach to generative machine learning by leveraging the Born rule. Recent work has provided a method to classically train QCBMs with Instantaneous Quantum Polynomial (IQP) circuits via the Maximum Mean Discrepancy (MMD) loss. Despite the assumed intractability of sampling from IQP circuits classically, their expectation values can be computed classically, enabling training of these IQP QCBMs. However, quantum machine learning (QML) models have various other challenges, including trainability issues caused by exponential concentration or barren plateaus. While these issues have been explored for parameters sampled from a uniform distribution, little work has been done to rigorously treat the use of arbitrary Gaussian initialization schemes. This work leverages Stein's lemma and Lipschitz concentration bounds for Gaussian random variables to provide an analytical lower bound of the variance of the gradient and a probabilistic concentration bound of the deviation of the gradient from its mean. It discusses strategies to either avoid or encourage exponential concentration, as well as the conditions under which barren plateaus are more likely to occur.
|
| 2281 |
Toward Proactive RF Charging Scheduling: Generative AI for Decision Support
2606.10600
|
cs.LG
|
Amirhossein Azarbahram, Osmel M. Rosabal, David Ernesto Ruiz-Guirola, Melike Erol-Kantarci, Kaibin Huang |
Radio frequency wireless power transfer (RF-WPT) is an enabling technology for supporting uninterrupted communications in future Internet of Things systems by reducing the need for battery replacement and mitigating battery-waste-related issues. For large-scal...Radio frequency wireless power transfer (RF-WPT) is an enabling technology for supporting uninterrupted communications in future Internet of Things systems by reducing the need for battery replacement and mitigating battery-waste-related issues. For large-scale RF-WPT deployment, one of the main challenges is the scheduler-level resource allocation. Specifically, the RF charger must decide how much energy to deliver, when, and to whom, under limited charging resources, incomplete receiver-side information, and uncertain near-future charging conditions. This article positions generative artificial intelligence (GenAI) as a promising tool for this setting because it can foresee multiple plausible charging scenarios conditioned on coarse operational context and receiver-side information. We propose GenAI to act as an uncertainty-aware support layer for the RF-WPT scheduler rather than as a standalone forecasting or decision-making tool. To this end, we first revisit the main challenges of RF-WPT scheduling, and discuss how major GenAI families can support uncertainty-aware charging decisions by generating scenario-based inputs for downstream tasks. We then present a case study showing that distribution-aware prediction can improve robust charging decisions over deterministic, ensemble, and non-learning baselines, particularly under risk-sensitive objectives. Finally, we outline key open challenges and future research directions.
|
| 2282 |
DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace
2606.11687
|
cs.LG
|
Marius Bayizere |
Unmanned Aerial Vehicle (UAV) threats have emerged as a defining security challenge of the 21st century. This paper presents DroneShield-AI, a unified open framework integrating six processing layers: RF signal classification, acoustic motor-signature detectio...Unmanned Aerial Vehicle (UAV) threats have emerged as a defining security challenge of the 21st century. This paper presents DroneShield-AI, a unified open framework integrating six processing layers: RF signal classification, acoustic motor-signature detection, YOLOv8-based visual detection, evidence-weighted sensor fusion, a Behavioral Intent Classification Engine (BICE), and a Graph Neural Network Swarm Intelligence Module (GNN-SIM). This v2 revision reports measured results on the completed implementation (495 automated tests), superseding the v1 simulation-only preprint. On real public data: RF presence detection reaches F1 0.9924; acoustic detection reaches 98.12% accuracy, though a 4-feature statistical baseline reaches 93.25% on the same split, so the model's drone-specific contribution is approximately 4.9 points, a property of dataset provenance rather than a leakage bug; visual detection (YOLOv8) reaches AP50 0.8782 on a held-out test split, from a model still improving when training stopped. No fused-accuracy figure is reportable, since the three real-data layers use independently collected, unsynchronized datasets with no shared ground truth. BICE and GNN-SIM remain evaluated on physics-calibrated simulation data pending real adversarial and swarm field data: BICE shows a consistent robustness advantage over a trivial baseline under distribution shift; GNN-SIM underperforms its own trivial baseline in-distribution but independently replicates BICE's noise-robustness advantage, regarded as this work's most durable finding. All code, trained checkpoints, and the full test suite are publicly released.
|
| 2283 |
OneCanvas: 3D Scene Understanding via Panoramic Reprojection
2606.19253
|
cs.LG
|
Bart{\l}omiej Baranowski, Dave Zhenyu Chen, Matthias Nie{\ss}ner |
Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or large training budgets in pursuit of spatial reasoning. Instead, OneCanvas aggregates patch features from all views onto ...Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or large training budgets in pursuit of spatial reasoning. Instead, OneCanvas aggregates patch features from all views onto a single equirectangular panoramic canvas. Namely, each patch is unprojected to a 3D world coordinate using its depth and camera pose, then placed on the canvas at the continuous longitude and latitude of that point as seen from the canvas origin, with no rasterization or aggregation across overlapping views. A 3D position embedding of the patch's metric coordinates is added to its feature, restoring the depth lost when collapsing the world position to an angular canvas coordinate. Patches from all frames thus share one spatial coordinate system with no fusion or major architectural modifications of the backbone. The pretrained VLM consumes this representation as if it were an ordinary image. Because the canvas can be centered on any pose of interest, the same representation directly supports situated reasoning from a specific viewpoint, a common requirement in robotics and embodied AI. Thanks to this representation, we can also introduce a spatial pretraining curriculum: by procedurally placing patch features of objects, drawn from real images, at chosen 3D world positions on an otherwise empty canvas, we generate on-the-fly supervision spanning a broad range of spatial reasoning tasks, with answer distributions controlled to reduce spatial reasoning shortcuts. OneCanvas achieves state-of-the-art results on SQA3D and VSI-Bench and generalizes to out-of-distribution data on SPBench, using an order of magnitude less training compute than its competitors closest in benchmark performance.
|
| 2284 |
Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence
2606.20592
|
cs.LG
|
Junaid Sajid, Sheikh Salman Hassan, Wenshuai Liu, Yan Kyaw Tun, Yaru Fu |
Embodied artificial intelligence (AI) couples perception and learned decision making to actions that change the physical world. This coupling distinguishes an embodied agent from a conventional connected controller: the agent maintains task state and uncertain...Embodied artificial intelligence (AI) couples perception and learned decision making to actions that change the physical world. This coupling distinguishes an embodied agent from a conventional connected controller: the agent maintains task state and uncertainty, reasons about the consequences of actions, and adapts from subsequent observations. Wireless networking becomes relevant when perception, inference, or coordination is distributed, but it should not replace local safety control. This article develops a tutorial perception--communication--action (PCA) architecture that exposes task state, action deadlines, uncertainty, agent intent, and safety envelopes to a 6G orchestration plane. It separates capabilities already addressed by 5G and 5G-Advanced from functions that motivate 6G, including task-state interfaces, semantic freshness, predictive digital twins, and safety-aware coordination across agents. A multi-robot simulation study is retained to illustrate joint sensing, communication, and computation control. The results show where network orchestration improves task utility and where local autonomy remains essential.
|
| 2285 |
Finite-Sample Performance of Gradient Descent in Logistic Regression with Gaussian Design
2606.21683
|
cs.LG
|
Junren Chen, Arya Mazumdar |
We consider the parameter estimation problem in logistic regression with Gaussian design: the estimation of a fixed unknown parameter $\theta^*\in \mathbb{R}^d$ ($\|\theta^*\|_2\ge 1$) from $n$ i.i.d. samples $\{(x_i,y_i)\}_{i=1}^n$, where $x_i\sim N(0,I_d)$ a...We consider the parameter estimation problem in logistic regression with Gaussian design: the estimation of a fixed unknown parameter $\theta^*\in \mathbb{R}^d$ ($\|\theta^*\|_2\ge 1$) from $n$ i.i.d. samples $\{(x_i,y_i)\}_{i=1}^n$, where $x_i\sim N(0,I_d)$ and $y_i|x_i \sim {\rm Bernoulli}(1/(1+\exp(-x_i^\top \theta^*)))$. Our main aim is to characterize the finite-sample estimation performance and convergence behavior of gradient descent (GD) on the maximum likelihood objective (i.e., the logistic loss). Under small $O(1)$ stepsize and $0$ initialization, we show that GD linearly converges to a small neighborhood of $\theta^*$ achieving an $\ell_2$ error of order $O(\sqrt{\|\theta^*\|_2^5d/n})$. This substantially goes beyond existing theoretical results that lack non-asymptotic estimation error rate and exhibit much slower parameter convergence. We also establish a faster local linear convergence to the same statistical error under a large $\Theta(\|\theta^*\|_2)$ stepsize. The main technical component is to show that the gradient of the logistic loss satisfies a certain approximate invertibility condition (AIC). To that end, we uniformly control the deviation of the gradient from its population counterpart by covering and peeling arguments, and then show that the population GD is a contraction by a delicate analysis based on the eigenvalues of population Hessian matrices. Finally, we build upon the recent work Matsumoto and Mazumdar (2025) and devise a novel efficient estimator that attains a sharper rate in high dimensions. This indicates that the existing non-asymptotic guarantees exhibit sub-optimal dependence on $\|\theta^*\|_2$, and that in many regimes $\Theta(\sqrt{\|\theta^*\|_2d/n})$ is the tight estimation error rate. Numerical examples are provided to corroborate our theoretical results.
|
| 2286 |
Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
2606.21838
|
cs.LG
|
Zhiyuan Tao, Srikumar Sastry, Matthew J Thompson, Elizabeth G Campolongo, Net Zhang |
Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing methods often produce predictions that are inconsistent across taxonomic levels...Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing methods often produce predictions that are inconsistent across taxonomic levels. For example, a model may predict a fine-grained category whose parent category contradicts its simultaneously predicted higher-level label. By analysis, the issue originates from false negative labels when contrastive comparison involves multiple taxonomic levels. To this end, we propose to restrict contrastive comparisons to categories within the same taxonomic level. In addition, we adopt a group-balanced design, ensuring each taxonomic level receives adequate optimization. As a result, the proposed framework improves both hierarchical consistency and classification accuracy from coarse to fine granularity. We train our model with TreeOfLife-10M based on BioCLIP and evaluate it across multiple hierarchical classification benchmarks, where the model demonstrates significantly improved hierarchical consistency in both Euclidean and hyperbolic spaces. Notably, on iNaturalist 2021 (iNat21), our method improves average accuracy across levels by 30.47% over the baseline, highlighting its effectiveness for hierarchical zero-shot classification.
|
| 2287 |
Rethinking Object-Centric Representations for Video Dynamics Modeling
2606.23436
|
cs.LG
|
Amaury Wei, Ismail Nejjar, Olga Fink |
Learning to decompose videos into persistent objects is a fundamental challenge in unsupervised object-centric representation learning. Despite recent progress, existing methods struggle to simultaneously achieve accurate object segmentation, consistent identi...Learning to decompose videos into persistent objects is a fundamental challenge in unsupervised object-centric representation learning. Despite recent progress, existing methods struggle to simultaneously achieve accurate object segmentation, consistent identities over time, and reliable foreground-background separation. To address these challenges, we introduce UniSlot (Unified Slots), an unsupervised framework for learning robust and disentangled object-centric representations from videos. UniSlot explicitly separates object appearance from its 3D-aware geometric pose in the scene, linking object identity to appearance while leveraging depth to better distinguish objects from their surroundings. UniSlot achieves state-of-the-art performance in unsupervised object-centric video decomposition and tracking across synthetic and real-world benchmarks, yielding substantially tighter object masks and reducing background leakage while preserving object identities. Beyond decomposition and tracking, these improved representations translate directly to downstream tasks such as unsupervised object dynamics prediction, enabling more accurate forecasting of future object trajectories.
|
| 2288 |
FacePlex: Toward Natural Full-Duplex Conversational Avatars
2606.30145
|
cs.LG
|
Habin Lim, Hah Min Lew, Jae-Ho Lee, Min-Jae Kim, Seungeun Lee |
Natural human conversation is inherently a real-time interaction in which speech and facial behavior continuously evolve. Enabling such interaction requires a conversational avatar to jointly generate speech and facial motion in real time, prepare facial motio...Natural human conversation is inherently a real-time interaction in which speech and facial behavior continuously evolve. Enabling such interaction requires a conversational avatar to jointly generate speech and facial motion in real time, prepare facial motion for upcoming speech before the corresponding audio is emitted, and produce non-verbal reactions that reflect the ongoing dialogue. However, existing conversational avatar systems cannot address such requirements: audio-driven methods rely on pre-given speech, while joint streaming generation alone does not ensure anticipatory articulation or semantically appropriate reactions. We propose $\textbf{FacePlex}$, a unified framework for full-duplex speech-facial motion generation and real-time avatar rendering. FacePlex jointly coordinates speech, facial motion, and Gaussian splatting rendering on a shared streaming timeline. To prepare facial motion for upcoming speech, FacePlex predicts a short speech continuation and uses its future hidden states through asymmetric conditioning and denoising, while continuously updating them as new user audio arrives without observing future user input. For dialogue-grounded non-verbal behavior, we construct $\textit{SemReact}$, a dataset aligning dialogue context, reaction semantics, and facial motion, and introduce a semantic behavior router that guides continuous motion during both speaking and listening. Extensive experiments show improved audio-visual synchronization and facial articulation, natural dialogue-grounded non-verbal responses, and low-latency of 122 ms end-to-end avatar interaction.
|
| 2289 |
Retrieval Observability Bounds on Provenance Detection for Agent Memory Poisoning: Measured Coverage and a Falsified Standalone Detector
2606.30566
|
cs.LG
|
Jun Wen Leong |
Whether a trajectory-based detector can observe the retrieval provenance of an agent memory-poisoning attack is governed by one architectural variable: retrieval observability, whether the agent's memory access produces a logged tool call. We formalise this as...Whether a trajectory-based detector can observe the retrieval provenance of an agent memory-poisoning attack is governed by one architectural variable: retrieval observability, whether the agent's memory access produces a logged tool call. We formalise this as a framework-agnostic retrieval-to-action provenance graph, classify published attack families by their position relative to it, measure the resulting coverage property against collected agent traces, and then falsify the detector's standalone deployment, reversing our own earlier recommendation to use the signature as a pre-filter ahead of recipient-metadata gating. Under exclusive tool-mediated payload access a retrieval-before-exfiltration event is structurally forced. Across eight non-overlapping zero-event control cells we observe 0 violations in 289 eligible successes over 87 scenario configurations. The invariant is empirically unbroken, but on the clustering unit no registered cell certifies the preregistered <=0.10 upper bound (per-cell bounds 0.10-0.35); a post-hoc pool of two control arms nominally clears it but is not credited. When the scaffold delivers the payload implicitly the signature disappears; we had read this as agents declining an available retrieval, but a follow-up shows it tracks store content instead, and that behavioural reading is withdrawn. The detector is then falsified for standalone use. Benign memory-grounded sends are trajectory-isomorphic: false-positive rate 24.7-57.6% on a 13-model factorial (N=4,160), positive predictive value <=3.85% at 1% attack prevalence, and a recipient-metadata gate flags 129/129 legitimate external emails. Retrieval-before-send is an attack precondition, not a maliciousness predicate; its cross-validated AUC also reflects membership of a 22-vector feature codebook, not generalisation. Companion to arXiv:2605.08442; we release the corpus and scoring code.
|
| 2290 |
OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration for ARC-AGI-3
2607.01531
|
cs.LG
|
David Courtis, Wenhao Li, Scott Sanner |
Learning a useful world model from minimal interaction is central to building agents that adapt to unfamiliar tasks. Programmatic world modeling approaches such as WorldCoder quickly learn transition models when given pre-supplied symbolic representations; how...Learning a useful world model from minimal interaction is central to building agents that adapt to unfamiliar tasks. Programmatic world modeling approaches such as WorldCoder quickly learn transition models when given pre-supplied symbolic representations; however, such approaches are insufficient when the symbolic representation and goal are unknown and continually changing, and are intractable under large action spaces due to rigid planning heuristics. Extending the programmatic world modeling approach, OPINE-World additionally maintains an abstracted representation in which provisional objects, incomplete causal explanations, and unresolved interactions can exist in a Bayesian, exploration-centric hypothesis space. This enables learning an unknown, dynamically evolving state, transition, and goal. Exploration and program revision share this evolving abstraction, allowing new evidence to revise both the hypothesized mechanics and their implementation in tandem at test time. ARC-AGI-3 presents this open-world problem as a general-intelligence learning benchmark. Through our OPINE-world model learning paradigm, we improve the Relative Human Action Efficiency score of Claude Opus 4.8 high from 1.5% to 78.40% on the ARC-AGI-3 benchmark.
|
| 2291 |
Dithered Gaussian Mechanism for Randomness-Efficient Differential Privacy
2607.06320
|
cs.LG
|
Nikita P. Kalinin, Rasmus Pagh |
We present the dithered Gaussian mechanism, an alternative to the discrete Gaussian mechanism for differential privacy that discretizes the private output rather than the noise distribution itself.By interpreting this discretization as post-processing of the G...We present the dithered Gaussian mechanism, an alternative to the discrete Gaussian mechanism for differential privacy that discretizes the private output rather than the noise distribution itself.By interpreting this discretization as post-processing of the Gaussian mechanism, our construction directly inherits the privacy guarantees of the standard Gaussian mechanism while avoiding vulnerabilities caused by finite-precision floating-point outputs. In addition, the mechanism is provably randomness-efficient: by sampling the discretized output values directly, the number of high-quality random bits required for privacy can be reduced significantly and made independent of the noise level. This is achieved by separating the randomness into two sources: a high-quality source used for the privacy-critical sampling step, and a high-performance public source, possibly known to the adversary, that supplies the additional randomness needed for randomized discretization. This separation enables the use of cryptographically secure randomness without substantial performance loss. As an application, we study model training with DP-SGD and show that cryptographically secure noise generation with reduced exposure to floating-point vulnerabilities can be achieved with modest practical overhead.
|
| 2292 |
Artificial Intelligence Across the Cardiac Amyloidosis Diagnostic and Management Pathway: From Single-Modality Detection to Multimodal Clinical Integration
2607.09948
|
cs.LG
|
Diana Shadibaeva, Rochak Dhakal, Kui Zhang, Diana Paez, Xiaofeng Yang |
Cardiac amyloidosis (CA) is increasingly recognized but remains substantially underdiagnosed because its clinical and imaging phenotype overlaps with more common cardiomyopathies. Definitive subtype assignment and management further require integration of mult...Cardiac amyloidosis (CA) is increasingly recognized but remains substantially underdiagnosed because its clinical and imaging phenotype overlaps with more common cardiomyopathies. Definitive subtype assignment and management further require integration of multimodal evidence to distinguish transthyretin from light-chain disease. Machine learning and deep learning have been applied across the diagnostic and management pathway. These applications span electrocardiography (ECG), echocardiography, and health record-based case finding, as well as cardiac magnetic resonance (CMR) and nuclear interpretation, including single-photon emission computed tomography/computed tomography (SPECT/CT) biomarker quantification, prognostic modeling, and treatment response assessment. This narrative review synthesizes these studies by clinical tasks, namely screening, detection, quantification, prognosis, and longitudinal assessment after treatment initiation, rather than by input modality. This task-based organization clarifies why apparently similar AI models require different cohorts, reference standards, evaluation metrics, and implementation thresholds. The evidence reveals a maturity gradient. Binary detection and AI-assisted interpretation of cardiac scintigraphy with bone-avid tracers and SPECT/CT currently represent one of the more mature AI applications in cardiac amyloidosis, supporting standardized image interpretation and quantitative biomarker extraction. However, patient-level diagnosis still requires integration with monoclonal protein testing, SPECT/CT localization, clinical context, and biopsy or tissue typing when indicated.
|
| 2293 |
NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations
2607.10490
|
cs.LG
|
Ruksat Khan Shayoni, Muhammad Faraz Shoaib, S M Asif Hossain, M. F. Mridha |
Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections. We present NetInjectBench, a 130-scenario benchmark that separates untrusted art...Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections. We present NetInjectBench, a 130-scenario benchmark that separates untrusted artifact text, trusted policy metadata, and evaluation labels for network-operation tool use. The sample contains 40 benign, 40 weak-attack, 40 strong-attack, and 10 approved high-impact change scenarios; each is evaluated with Qwen2.5-7B, Llama3.1-8B, and Mistral-7B. Across 240 attack instances, naive execution reached an 82.50% unsafe tool-action rate. Prompt-only safety, Self-Reminder, Spotlighting, and a Two-Pass LLM Judge reduced this rate to 25.63%, 21.67%, 18.33%, and 10.00%, respectively. Static allowlisting reached 5.00% but blocked all approved changes, yielding 0.00% usefulness and 100.00% overblocking on approved cases. Under the stated metadata-integrity assumption, the metadata-aware policy gate produced 0/240 unsafe attack actions, with a 95% Wilson upper bound of 1.58%, while preserving 99.17% attack-scenario usefulness and 100.00% approved-change usefulness. The findings show that network-operation agents need execution-time authorization boundaries alongside prompt-level instruction hygiene.
|
| 2294 |
Diversified Multinomial Logit Contextual Bandits
2607.11684
|
cs.LG
|
Heesang Ann, Taehyun Hwang, Min-hwan Oh |
Existing contextual multinomial logit (MNL) bandits model relevance-driven choice but ignore the potential benefits of within-assortment diversity, while submodular/combinatorial bandits encode diversity in rewards but lack structured choice probabilities. We ...Existing contextual multinomial logit (MNL) bandits model relevance-driven choice but ignore the potential benefits of within-assortment diversity, while submodular/combinatorial bandits encode diversity in rewards but lack structured choice probabilities. We bridge this gap with the $\textit{diversified multinomial logit}$ (DMNL) contextual bandit, which augments MNL choice probabilities with a generally submodular diversity function, thereby formalizing the relevance--diversity trade-off within a single model. Incorporating diversity renders exact MNL assortment optimization intractable. We propose a $\textit{white-box}$ UCB-based algorithm, $\texttt{OFU-DMNL}$, that constructs assortments item-wise by maximizing optimistic marginal gains, avoids black-box optimization oracles. We show that $\texttt{OFU-DMNL}$ achieves at least a $(1-\frac{1}{e+1})$-$\textit{approximate}$ regret bound $\tilde{O}\left(d \sqrt{T/K}\right)$, where $d$ is the context dimension, $K$ the maximum assortment size, and $T$ the horizon, and attains an improved approximation factor over standard submodular baselines. Experiments demonstrate consistent gains and, relative to exhaustive enumeration, comparable regret with substantially lower runtime. Overall, DMNL bandits provide a practical foundation for diversity-aware assortment optimization under uncertainty, and $\texttt{OFU-DMNL}$ offers a statistically and computationally efficient solution.
|
| 2295 |
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
2607.13075
|
cs.LG
|
Dominik Schwarz |
An activation-based safety gate must distinguish harmful assistance from legitimate work on the same topic. We test whether probes trained on broad harmful-versus-harmless data provide this decision boundary. Five checkpoint-specific readouts and their source-...An activation-based safety gate must distinguish harmful assistance from legitimate work on the same topic. We test whether probes trained on broad harmful-versus-harmless data provide this decision boundary. Five checkpoint-specific readouts and their source-derived thresholds are transferred unchanged to four matched-pair cohorts containing 450 pairs, manually reviewed for topic and request-frame similarity. On Twin-70, the primary cohort of 70 pairs, central harmful and harmless score intervals overlap. Area under the receiver operating characteristic curve (AUROC) falls from 0.996-0.999 on the broad source task to 0.701-0.860. Pair ranking remains 0.829-0.929, but the transferred thresholds detect 0.943-1.000 of harmful requests while falsely blocking 0.686-0.986 of matched harmless requests. False blocking remains low on XSTest. Even with thresholds refitted within pair-grouped cross-validation, balanced threshold error is 0.229-0.329. Alternative source-trained logistic, radial-basis, and shallow multilayer probes reduce false blocking but detect only 0.200-0.671 of harmful requests at their transferred thresholds. Three primary readouts meet all prespecified Wall criteria, while two yield mixed evidence. Matched-label training recovers the distinction for the three checkpoints tested, and lexical controls also exhibit transfer loss. An automated post-freeze audit finds little numerical sensitivity to arm-label disagreement. The evaluated source operating points fail to combine strong harmful-request detection with access for legitimate same-topic work. This gap between risk ranking and a usable decision boundary is the security concern captured by the Entanglement Wall.
|
| 2296 |
BARS: Benign-Anchored Ranking and Selection for False Alarm Reduction in Network Intrusion Detection
2607.13203
|
cs.LG
|
Abu Fuad Ahmad, Istiaque Ahmed |
False alarms remain a major barrier to deploying network intrusion detection systems (NIDS). In high-volume environments, even a sub-1% false positive rate can generate tens of thousands of daily alerts. Filter-based feature selection is attractive because it ...False alarms remain a major barrier to deploying network intrusion detection systems (NIDS). In high-volume environments, even a sub-1% false positive rate can generate tens of thousands of daily alerts. Filter-based feature selection is attractive because it operates upstream of the classifier and adds no inference-time cost. However, classical filters use class-symmetric criteria that ignore the asymmetry of intrusion detection, where benign traffic defines the baseline and attacks are deviations from it. A recent class-asymmetric filter, Classwise Mean Deviation (CMD), addresses this issue but anchors its score to a global mean that shifts toward attack distributions under class imbalance, weakening the deviations it aims to capture. We propose Benign-Anchored Ranking and Selection (BARS), a two-stage filter that replaces CMD's global anchor with the benign-class mean and applies an order-preserving decorrelation step. We evaluate BARS on CICIDS2017, CICDDoS2019, and UNSW-NB15 using feature budgets k = {5, 10, 20, 30, 40}. On attack-majority datasets, where global-anchor bias is strongest, BARS reduces false positive rate relative to CMD by 15.4% on UNSW-NB15 at k = 20 and by 21% to 23% on CICDDoS2019 at small feature budgets while preserving true positive rate and macro-F1. On benign-majority data, BARS and CMD converge, consistent with the theoretical limit where global- and benign-anchored scores coincide. BARS reduces false alarms by up to 32% while maintaining similar detection rates, with larger gains under stronger imbalance. Ablation results show that the two stages provide complementary benefits and select a consistent set of robust features. BARS retains linear-time scoring and a low memory footprint, making it suitable for resource-constrained deployments.
|
| 2297 |
The Steering Budget: Examples beat Knobs
2607.14246
|
cs.LG
|
Raj Kumar Rajendran |
Generative models are steered with knobs -- prompts, guidance scales, property tags. Turn one as hard as you like and, past a point, it stops moving the property you care about. We find that ceiling is not a shortcoming of the model but a budget, set by the tr...Generative models are steered with knobs -- prompts, guidance scales, property tags. Turn one as hard as you like and, past a point, it stops moving the property you care about. We find that ceiling is not a shortcoming of the model but a budget, set by the training data before the model is trained: a property's movable range splits in two -- the part a knob can reach, and a second, significant part that only examples -- concrete instances of what you want more of -- can reach. That second part is usually much larger, but not always, and the same budget says so in advance. Reaching that second part takes a different move: instead of turning a knob, you show the model examples, composed from what it already learned rather than added to its training. A cheap audit of the training data measures the budget; we give a recipe for building the example set that reaches all of it. This buys two things a knob can't. Reach: it moves a property across the whole budget, not just the part a knob reaches. Expressiveness: it steers toward targets you can only specify by example -- including ones you can't put into words. We turn these into a handful of falsifiable claims and verify them in two unrelated domains, image and crystal-structure generation -- marking where a knob is enough, and where only examples will do.
|
| 2298 |
Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dynamics
2607.19389
|
cs.LG
|
Vedant Palit, Udvas Das, Brahim Driss, Debabrota Basu |
As AI-driven Decision Makers (ADMs) influence our socioeconomic reality, their roles in both enhancing efficiency and amplifying the social biases have drawn significant attention. In this paper, we revisit the nuances of long-term fairness achievable by an AD...As AI-driven Decision Makers (ADMs) influence our socioeconomic reality, their roles in both enhancing efficiency and amplifying the social biases have drawn significant attention. In this paper, we revisit the nuances of long-term fairness achievable by an ADM-- specifically, through the lens of a loan approver inducing a population-level wealth dynamics. The literature generally (a) considers passive environments, i.e. the decisions of an ADM does not change the population's behaviour, and (b) measures bias in terms of disparity in the instantaneous decisions rather than downstream equity. Modern ADMs challenge both the notions. To address these caveats, we first formalise the wealth dynamics induced by a loan approving ADM interacting with a multi-demographic population as a Performative Markov Decision Process. Then, we mitigate the absence of such a performative test-bed by developing Eutopia: a lending-process simulator enabled with a novel performative data generator to learn long-term fair strategies. With Eutopia, we test reinforcement learning algorithms with different fairness-aware utilities dependent on approval decisions and downstream wealth. Results show that performative dynamics-aware learning with fairness-aware utility that incorporates the downstream outcomes induce better long-term equity and inclusivity.
|
| 2299 |
Priors learned from legacy reconstructions inherit undetectable overconfidence
2607.21721
|
cs.LG
|
Ali Siahkoohi, Sina Alemohammad |
Where truths are scarce (e.g., seismic and medical imaging), learned priors in ill-posed inverse problems are trained on archives of legacy reconstructions---i.e., an older method's outputs---and their reported uncertainty is taken as data-driven. We show that...Where truths are scarce (e.g., seismic and medical imaging), learned priors in ill-posed inverse problems are trained on archives of legacy reconstructions---i.e., an older method's outputs---and their reported uncertainty is taken as data-driven. We show that this prior is, in the population limit, exactly the regularizer that produced its archive of posterior samples, advanced one expectation--maximization step toward the truth. While the step improves the regularizer on the directions the measurements resolve, it leaves the regularizer's assumption on the operator's blind subspace unchanged. An archive of single-best reconstructions collapses the blind interval to zero width. Neither error is detectable in practice, as truths differing only on the blind subspace share the data law, and simulation-based calibration is neutral by construction. We identify from the operator alone which directions the measurements do not inform, and, given a handful of ground-truth models, build intervals there that contain the truth as often as they claim to. We validate these findings on a two-dimensional example with closed-form predictions and in controlled experiments on seismic-imaging and groundwater-flow operators, against priors trained on the truth.
|
| 2300 |
Codifying the Judge: Scalable Evaluation via Program Distillation
2607.22561
|
cs.LG
|
Tzu-Heng Huang, Shengqi Qiu, Frederic Sala |
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, inference latency, and opaque decisions---limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program...LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, inference latency, and opaque decisions---limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at evaluation time, we distill its decision logic into a committee of programs that can score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and eight model families, we show that programmatic judges match the performance of a 13B-size LLM judge at 47x higher throughput. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.
|
| 2301 |
DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
2607.23614
|
cs.LG
|
Xingyang Yu |
We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral...We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency certificate, which states that no tested inconsistency was found, not that the duality is proven. We use the verifier as a repair environment for language-model agents, which receive a deliberately broken claim and must edit it until it certifies. On a preregistered benchmark of 145 broken claims, with the analysis fixed before the first confirmatory model call, verifier-gated retry improves final repair success over a single attempt by +8.3 percentage points (pp) on deepseek-chat and +7.1 pp on qwen-plus (Holm-adjusted p<0.002). Under an equal budget of eleven attempts, the stop-first strategy portfolio underperforms independent verifier-filtered resampling by 10.3 percentage points on deepseek-chat but outperforms it by 14.7 points on qwen-plus, reversing the ordering of the two tested verifier-exploitation policies across the two confirmatory models. On qwen-plus, category-level verifier feedback is worth +8.7 pp over content-free retry, and interpretable obligation identities alone are worth +6.4 pp over structurally identical masked feedback. Neither effect is detected on deepseek-chat. Separately, a preregistered MiniMax-M2.5 extension again finds an iteration gain and independent verifier-filtered resampling outperforming the strategy portfolio. Which policy is better thus differs between the two models, while every winning policy uses the same cheap certificate. The verifier, benchmark, protocol, and all per-attempt records are released.
|
| 2302 |
InferScale: GPU-Native KV Injection for Personalized LLM Serving
2607.27090
|
cs.LG
|
Peter Li, Prashant Pandey |
Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.g., Mem0, MemGPT, and Zep) retr...Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant subset of this memory and inject it into the prompt, forcing the serving engine to repeatedly prefill the same content. As the retrieval budget grows, time-to-first-token (TTFT) increases even though the underlying memory is reused across requests. We present InferScale, a GPU-native LLM memory system that replaces repeated prompt prefilling with reusable KV state. InferScale precomputes each memory fact's KV representation, stores it alongside a semantic embedding on the GPU, retrieves relevant facts at serving time, and injects their KV directly into vLLM's paged cache. To support dynamically assembled memories under rotary position embeddings, we introduce Chunked RoPE, which stores keys before rotation and applies their serving-time positions during injection. However, encoding memory facts independently omits the cross-fact context available during joint prefilling. We mitigate this with Context-Window Encoding, which encodes each memory fact together with a small window of preceding conversation context while caching only the target fact's KV. InferScale is implemented through vLLM's KV-connector interface, requiring neither engine modifications nor model fine-tuning. Across three open-weight models on LoCoMo, InferScale keeps TTFT nearly constant as the retrieval budget increases: at k=50 it reduces TTFT by 72-79% (3.6-4.8x), achieves 60.3% accuracy versus 63.3% for Mem0 without serving-time recomputation, and delivers 3.7-4.5x the throughput under concurrent load. Reusable KV state thus decouples memory-conditioned serving latency from retrieved-context size while preserving application quality.
|
| 2303 |
Fractional Parabolic Partial Differential Equations in Anisotropic Spectral Barron Spaces: Regularity and Neural Approximation
2607.27781
|
cs.LG
|
Jae-Hwan Choi, Hyojae Lim, Jinsol Seo, Young-Jin Sim, Changhoon Song |
We study fractional parabolic initial-value problems with lower-order drift and potential terms in anisotropic spectral Barron spaces, defined by weighted space--time Fourier $L^1$ norms adapted to parabolic scaling. We prove existence, uniqueness, and maximal...We study fractional parabolic initial-value problems with lower-order drift and potential terms in anisotropic spectral Barron spaces, defined by weighted space--time Fourier $L^1$ norms adapted to parabolic scaling. We prove existence, uniqueness, and maximal regularity with a gain of one derivative in time and $\gamma$ derivatives in space, where $\gamma>0$ is the order of the fractional Laplacian. The evolution is defined only for $t\geq0$, whereas the finite-time norm requires a global extension with sufficient temporal Fourier decay. We construct a finite reflected semigroup extension using a Vandermonde system to match derivatives at $t=0$, obtaining temporal Fourier estimates uniform in the semigroup parameter. Combined with Fourier multiplier estimates for the damped principal operator, it yields maximal regularity. Dimension-independent multiplication estimates support a finite regularity bootstrap, while interpolation and sufficient damping absorb the lower-order terms in the base estimate. The a priori estimate and the method of continuity yield maximal regularity without smallness assumptions on the lower-order coefficients. A frequency-localized counterexample shows that a uniform-in-time spatial Barron bound on the forcing does not imply the corresponding two-derivative solution bound, even for the one-dimensional heat equation. Using this regularity, Fourier sampling yields $n^{-1/2}$ approximation rates for the solution in mixed space--time Sobolev norms using shallow networks with suitable activations. Sampling in a product Hilbert space yields a population-level PINN consistency estimate for shallow cosine networks on a bounded cylinder. There exists a single width-$n$ network for which the sum of the squared mixed-Sobolev solution error, the squared $L^2$-norm of the residual for the whole-space fractional equation, and the squared initial-data error is $O(n^{-1})$.
|
| 2304 |
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
2608.00155
|
cs.LG
|
Dong Yan, Jian Liang, Dapeng Hu, Ran He, Nicholas Jing Yuan |
Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming se...Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the \texttt{Isolated}, \texttt{Sequential}, and \texttt{Interleaved} streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and is non-monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.
|
| 2305 |
Tokenizer-Generator Coupling in Medical Image Generation
2608.07713
|
cs.LG
|
Liam Chalcroft |
Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation holds in a controlled ChestMNIST study at $64\times64$ that crosses discrete tokenizers, generator families, and sampler settings under a shared...Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation holds in a controlled ChestMNIST study at $64\times64$ that crosses discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. The rankings depend jointly on the tokenizer, generator, and sampler. The best quantizer changes with the generator, and validation-based sampler selection changes the apparent generator ranking. We interpret this through a rate-distortion-modelability framing in which modelability is conditional on the generator, sampler, and inference budget. The interaction persists when the vocabulary-1024 block is retrained at three seeds, with 4 of 9 pairwise quantizer comparisons exceeding three seed standard deviations, including a reversal between LFQ and FSQ from MaskGIT to D3PM; the rest of the grid was trained at a single seed and is correspondingly less certain. Reconstruction PSNR alone is not a reliable selection criterion. On LFQ-1024, retuning the D3PM and SEDD samplers on a held-out validation split reduces FID-192 from 0.44/0.41 at the default budget to 0.09/0.10 at lower NFE, replicated across seeds, although the continuous references were not given an equivalent sampler sweep. FID-192, our internal ranking metric, ranks consistently with standard FID-2048 (Spearman 0.89) and with a label-free classifier two-sample test (0.86). All experiments are unconditional, use low-resolution $64\times64$ medical-style images and are evaluated with non-clinical FID-based metrics, and our claims are limited to that setting. Implementations are released at https://github.com/liamchalcroft/medtokenizers and https://github.com/liamchalcroft/medlatents.
|
| 2306 |
Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
2608.18484
|
cs.LG
|
Pardis Taghavi, Reza Langari, Gaurav Pandey |
Block-sparse attention accelerates video transformers by grouping queries and keys and computing attention only between selected groups. However, it faces two challenges: (1) queries sharing a block route may attend to different keys; and (2) retaining most at...Block-sparse attention accelerates video transformers by grouping queries and keys and computing attention only between selected groups. However, it faces two challenges: (1) queries sharing a block route may attend to different keys; and (2) retaining most attention mass does not necessarily preserve the attention output. We find that partition choice affects both the overlap among grouped queries' preferred supports and how well an affine function of the sparse output can represent the difference between dense and sparse outputs. We introduce SparsePR, a training-free method combining Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. We first group paired key/value tokens by how their keys respond to sampled queries. We then group queries by their responses to the resulting key-group centroids for shared sparse routing. We compute exact attention on a small set of query rows and fit an affine correction to the differences between their dense and sparse outputs using weighted ridge regression. The fitted map predicts a residual for each unprobed query, which we add to its sparse output. We measure attention-output error as the relative L2 error with respect to each query's dense output. Across four video generation and world models, SparsePR reduces this error, averaged over unprobed queries, by 61.6--90.5% compared with semantic block-sparse attention without residual correction, at 22% executed-pair density including exact probes. SparsePR achieves 1.48x--2.61x end-to-end speedups over dense attention, with dense-reference PSNR of 24.42--31.84 dB at 21.9--26.0% realized executed-pair density. Project page: https://pardistaghavi.github.io/SparsePR-website/
|
| 2307 |
Neural Boltzmann Equations
2608.23022
|
cs.LG
|
Jonas Spinner, Jack D. Shergold |
The dynamics of particles in the early universe are described by Boltzmann equations, which involve high-dimensional phase-space integrals. Classical approaches use quadrature integration and evolve the system on a fixed momentum grid, which scales poorly to c...The dynamics of particles in the early universe are described by Boltzmann equations, which involve high-dimensional phase-space integrals. Classical approaches use quadrature integration and evolve the system on a fixed momentum grid, which scales poorly to complicated systems and parameter scans, severely limiting the complexity of processes that can be studied. We introduce Neural Boltzmann Equations (NBEs), which combine three coupled concepts to overcome these limitations. First, particle properties are encoded in physics-inspired neural distribution functions, with parameters that can be predicted using neural networks, enabling efficient parameter scans. Second, phase-space integrals are evaluated with Monte Carlo, using importance sampling tools from collider physics. Third, we use the natural gradient method to evolve the system. After demonstrating the individual benefits of NBEs, we use the framework to perform the most numerically precise calculation to date of the effective number of relativistic neutrino degrees of freedom, $N_\mathrm{eff}$, in the Standard Model.
|
| 2308 |
Which Histories Matter for Time Series Forecasting? Learning Predictive Relevance with Future Supervision
2608.23221
|
cs.LG
|
Yong-Hoon Choi, Kwang-Hyun Park, Youngjin Cho |
Retrieval-augmented time-series forecasting typically selects historical examples by similarity between observed pasts, although similar pasts can evolve differently. We propose Predictive Relevance Retrieval (PRR), which uses realized future compatibility as ...Retrieval-augmented time-series forecasting typically selects historical examples by similarity between observed pasts, although similar pasts can evolve differently. We propose Predictive Relevance Retrieval (PRR), which uses realized future compatibility as privileged supervision to learn a retrieval function that remains strictly past-only at inference. PRR combines Pearson retrieval with a futuresupervised predictive representation to expand candidate support, then reranks the union using statistical and learned pair relations. Across six datasets and four long horizons, PRR improves Pearson retrieval in 23 of 24 conditions, reducing AnalogFutureMSE by 26.7% on average. A candidate-budget-matched variant, PRR-B100, retains nearly the same retrieval improvement while using at most 100 candidates at inference, showing that the gain is not explained simply by a larger candidate pool. We then connect the retriever to five frozen forecasting backbones using validation-calibrated trust. Downstream effects are heterogeneous, and calibration primarily reduces harmful retrieval use rather than making improved retrieval universally beneficial. These results show that predictive relevance and forecast utility are empirically distinct objectives.
|
| 2309 |
Scaling Reinforcement Learning for Diffusion Models via Velocity Matching
2608.23664
|
cs.LG
|
Jaemoo Choi, Wei Guo, Yuchen Zhu, Arash Vahdat, Molei Tao |
Reward-based fine-tuning of diffusion models has largely inherited likelihood-based policy optimization developed for autoregressive large language models (LLMs). Diffusion models, however, are natively trained through velocity regression and do not directly p...Reward-based fine-tuning of diffusion models has largely inherited likelihood-based policy optimization developed for autoregressive large language models (LLMs). Diffusion models, however, are natively trained through velocity regression and do not directly provide the likelihood of a generated sample. Existing methods address this using transition likelihoods along stochastic denoising trajectories or evidence lower bounds (ELBOs) to approximate generated-sample likelihoods. These approximations arise from applying likelihood-based updates to models whose native training and generation operate through velocity fields. We instead take velocity matching as the starting point for reward fine-tuning. We propose \textbf{reward-based velocity matching (RVM)}, which weights the velocity-matching loss by reward with an additional anchor regression term, requiring neither likelihood estimation nor likelihood ratios. RVM recovers Reinforce Adjoint Matching at the update level and contains DiffusionNFT as a special case, while ELBO-based methods are closely related through the same velocity-regression structure. This unified view isolates reward design and anchor velocity as principal design choices in velocity-based fine-tuning. Across text-to-image, text-to-video, and image-to-video generation, RVM matches or outperforms the evaluated trajectory-based methods. On Wan2.1-T2V-1.3B, it achieves the highest VBench Overall among the evaluated methods, with substantially lower estimated training costs than the trajectory-based approaches. For video fine-tuning, we introduce a dynamic-tracking reward that provides explicit motion feedback and improves both Dynamic Degree and overall VBench performance.
|
| 2310 |
FRAME: Separating sampling variation from performance disparities in medical image analysis
2608.25981
|
cs.LG
|
Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh |
Fairness audits of medical imaging models commonly report the largest performance difference between demographic subgroups and treat any positive value as evidence of bias. Yet a perfectly fair model also produces a positive difference, which grows as subgroup...Fairness audits of medical imaging models commonly report the largest performance difference between demographic subgroups and treat any positive value as evidence of bias. Yet a perfectly fair model also produces a positive difference, which grows as subgroups shrink. Some mitigation methods remove demographic information from models because they assume that it produces the difference. Here we introduce Fair-model Reference And Mechanism Evaluation (FRAME). Its first step compares each reported difference with the difference expected from a perfectly fair model at the same subgroup counts. Its second step tests whether candidate causes of any excess over this fair-model reference change the difference when they are injected into model features. We evaluated 36 encoders on nine datasets in three modalities and audited 89 published differences across six modalities. For 10 frozen encoders and 13 chest radiograph findings, the reference is a median 41% of the race difference and 22% of the age difference. It is exceeded by 40 of 53 published differences in sensitivity or false positive rate and by one of 36 in the area under the receiver operating characteristic curve (AUROC). Adding race information to the features of two encoders does not change the race difference significantly. At matched disease performance, mitigation reduces the race difference by less than its spread across three pretraining seeds (medians 0.005 and 0.012) and the age difference by more (0.008 and 0.006). Image-text pretraining raises the worst-group AUROC for race by a median 0.049 over self-supervised pretraining. Reporting the fair-model reference beside every new and published subgroup difference could direct mitigation to the differences above it.
|
| 2311 |
A Systematic Survey of Agentic Skills: Architecture, Lifecycle, and Security
2608.29596
|
cs.LG
|
Sanket Badhe, Deep Shah, Priyanka Tiwari, Nehal Kathrotia |
Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms struggle...Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms struggle to scale, the field is rapidly converging toward \emph{agentic skills}: modular procedural abstractions that externalize execution knowledge into reusable, executable, and portable artifacts. This paper establishes a unified systems foundation and reference architecture for the agentic skills ecosystem. We formalize skills as externalized procedural knowledge bridging high-level cognitive planning with deterministic execution environments, and systematically delineate the architecture across a nine-stage lifecycle: autonomous discovery, authoring and representation formats, memory storage, dynamic retrieval and routing, composition and orchestration, execution and repair, lifelong adaptation, empirical evaluation, and security governance. We further examine marketplace dynamics, public registries, and emerging adversarial threat vectors, alongside runtime verification and defense mechanisms. Finally, we categorize system implementations across software engineering, operating system navigation, embodied robotics, and scientific discovery, while highlighting critical open challenges in continual learning and benchmark realism. This work establishes agentic skills as a foundational paradigm for building scalable, robust, and verifiable autonomous language agents.
|
| 2312 |
WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading
2609.02177
|
cs.LG
|
Gang-Hyun Park, Ju-Hyeong Lee, Hee-Youl Kwak, Dae-Young Yun |
Multi-bit watermarking for large language models enables content source tracing by embedding user-identifiable messages into generated text. Existing methods face a fundamental trade-off among extraction accuracy, text quality, and payload capacity. We propose...Multi-bit watermarking for large language models enables content source tracing by embedding user-identifiable messages into generated text. Existing methods face a fundamental trade-off among extraction accuracy, text quality, and payload capacity. We propose WeaveMark, a robust and scalable multi-bit LLM watermarking scheme based on coded payload spreading. WeaveMark shifts this trade-off frontier by improving payload capacity through multi-bit-per-token spreading (weaving), improving extraction accuracy through soft-decision error-correcting codes, and preserving text quality through unbiased multilayer reweighting. It further introduces dedicated zero-bit layers for reliable watermark presence detection. Extensive experiments demonstrate substantial gains in extraction performance, especially for long messages and edited text, without degrading text quality. WeaveMark achieves an 89.8% match rate for 32-bit messages at 200 tokens, compared with 20.8% for BiMark. Under 10% substitution attacks on 16-bit messages at 200 tokens, it maintains 86.0% versus 30.7%. Code is available at https://anonymous.4open.science/r/WeaveMark-ED6F.
|
| 2313 |
PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion
2609.08101
|
cs.LG
|
Peining Zhang, Jinbo Bi |
Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a balance between pocket compatibility, molecular properties, and physical geometry. We propose \textbf{PocketVE}, a protein-pocket-conditioned variance...Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a balance between pocket compatibility, molecular properties, and physical geometry. We propose \textbf{PocketVE}, a protein-pocket-conditioned variance-exploding (VE) diffusion framework that couples stable coordinate denoising with inference-time property guidance. Specifically, PocketVE combines an EDM-style training and sampling setup for 3D denoising, classifier-free guidance for multi-property steering, and adaptive protein perturbation as a training-time pocket regularizer.On the CrossDocked2020 benchmark under GenBench3D, PocketVE improves Valid$_{3\text{D}}$ from 58.6 to 80.6 and reduces strain energy from 457.4 to 127.9 relative to its guided TAGMol architectural parent; relative to TargetDiff, it attains comparable Valid$_{3\text{D}}$ with lower strain energy (127.9 vs.\ 306.0), while retaining competitive docking and molecular-property scores under moderate guidance. A guidance-scale study shows that moderate guidance gives a favorable balance between target-related objectives and geometric quality, whereas stronger guidance can degrade geometry and distributional fidelity. Pocket-permutation and PoseCheck diagnostics further support pocket-specific spatial compatibility with reduced steric conflicts. Overall, the results suggest that geometric stability and inference-time property guidance should be considered as coupled design objectives.
|
| 2314 |
ActionSplice: In-Flight Action Editing for Interactive World Models
2609.08230
|
cs.LG
|
Pardis Taghavi, Tingyu Guo, Jonas Lossner, Gaurav Pandey, Reza Langari |
Chunk-autoregressive video world models typically generate each chunk under one action. When an action changes during sampling, waiting until the next chunk delays the response, while directly switching the conditioning leaves the intermediate solver state sha...Chunk-autoregressive video world models typically generate each chunk under one action. When an action changes during sampling, waiting until the next chunk delays the response, while directly switching the conditioning leaves the intermediate solver state shaped by the previous action. Restarting sampling under the revised action avoids this mismatch but repeats completed computation. We introduce \emph{ActionSplice}, which edits the interrupted state through Counterfactual State Transport (CST). A lightweight corrector moves the interrupted solver state toward the matched counterfactual solver state induced by the revised action at the same solver step. The world model and sampler remain frozen, and sampling resumes without replaying completed evaluations. $\mathrm{CST}_{R}$ retargets the entire active chunk, while $\mathrm{CST}_{T}$ preserves the temporal prefix at the intervention step and corrects only the suffix. On minWM--Wan Action2V and HY-WM1.5, we measure fidelity against matched rollback references. Relative to condition swapping, $\mathrm{CST}_{R}$ reduces LPIPS by 61.5\% and 75.9\%, and $\mathrm{CST}_{T}$ reduces suffix LPIPS by 56.1\% and 77.5\%, respectively.
|
| 2315 |
HuRo: Robotizing Human Videos for Scalable VLA Pretraining
2609.10706
|
cs.LG
|
Jinho Jeong, Se June Joo, Jaehyun Kang, Dongyun Kim, Yena Kim |
Human video datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot data. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation...Human video datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot data. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately. In this work, we systematically examine whether robotized human videos can serve as an effective and scalable source of robot-aligned supervision. To this end, we develop a robotization pipeline that converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels. Using this pipeline, we construct the HuRo dataset, comprising about 630K robotized episodes and 142M processed frames from five human-video sources. Across four real-world manipulation tasks, pretraining a VLA policy on increasing amounts of robotized human-video data improves overall completion from 51.5% to 80.3% and OOD completion under spatial and visual shifts from 34.9% to 72.2%. Ablations further show that visual robotization improves OOD robustness and that end-to-end pretraining with retargeted actions outperforms visual-only transfer. Project website: https://3587jjh.github.io/HuRo.
|
| 2316 |
Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead
2609.11807
|
cs.LG
|
Corentin Pla, Hugo Richard, Marc Abeille, Vianney Perchet |
We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of actions before deciding its course of action. Although look-ahead can substantially improve achievable perfo...We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of actions before deciding its course of action. Although look-ahead can substantially improve achievable performance, [1] showed that optimal planning with multi-step transition look-ahead is NP-hard. However, this hardness was established using a discount factor close to one. It was therefore unknown whether the problem remains hard for every discount factor, and whether near-optimal planning can nevertheless be performed efficiently. We resolve both questions. First, we show that for every fixed discount factor, exact planning remains NP-hard. Second, we introduce a randomized polynomial-time approximation scheme for every fixed look-ahead depth. Third, we extend our approach to account for unknown transitions. We empirically validate the soundness of our results on the wind-farm storage-control benchmark of [2], showing that our approach, optimally accounting for -step look-ahead information, offers substantially better performance than existing algorithms. [1] Corentin Pla, Hugo Richard, Marc Abeille, Nadav Merlis, Vianney Perchet : On the Hardness of Reinforcement Learning with Transition Look-Ahead [2] Chenbei Lu, Zaiwei Chen, Tongxin Li, Chenye Wu, Adam Wierman : Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach
|
| 2317 |
Tight Sampling Complexity with Stochastic Gradient Oracles in Fixed Dimensions
2609.12590
|
cs.LG
|
Weiming Ou, Xiao Wang |
We establish the stochastic-gradient query complexity of sampling smooth strongly log-concave distributions in every fixed dimension $d\geq1$. For $\mu$-strongly convex, $L$-smooth potentials with unknown minimizers $x_f^\star$ in the ball $\mathbb{B}(0,\mu^{-...We establish the stochastic-gradient query complexity of sampling smooth strongly log-concave distributions in every fixed dimension $d\geq1$. For $\mu$-strongly convex, $L$-smooth potentials with unknown minimizers $x_f^\star$ in the ball $\mathbb{B}(0,\mu^{-1/2})$, and unbiased gradient oracles with variance at most $\sigma^2$, the minimax worst-case expected query complexity is $\Theta\left(\log(1+\kappa)+\frac{\sigma^2}{\mu\varepsilon}\right)$, jointly optimal for the condition number $\kappa:=L/\mu$, variance $\sigma^2\ge 0$, and TV accuracy $0<\varepsilon\leq 1/10$. In the noiseless setting $\sigma=0$, this tight complexity $\Theta\left(\log(1+\kappa)\right)$ is independent of $\varepsilon$. Moreover, a gradient-only sampler generates an exact sample with ${O}(\log(1+\kappa))$ worst-case expected queries. Without previous ball $\mathbb{B}(0,\mu^{-1/2})$, exact sampling from any initial point $x_0$ can be implemented with expected cost ${O}\left(\log(1+\kappa)+\log(1+\sqrt{\mu}\|x_0-x_f^\star\|)\right)$, without knowing the initial distance. We also show that dependence on initial distance is generally unavoidable.
|
| 2318 |
Learned Bow Control on a Measured Bowed-String Model: a Revised Minimum-Bow-Force Law, a Recurrent Controller, and the Domain of a Supervision Ceiling
2609.14990
|
cs.LG
|
Homayoon Beigi, Grace Conneely |
A finite-difference bowed-string model with implicitly resolved Stribeck friction is presented, with a regime diagnostic. Without implicit resolution no stick phase forms at any bow force. With friction, impedance and quality factor from published measurement ...A finite-difference bowed-string model with implicitly resolved Stribeck friction is presented, with a regime diagnostic. Without implicit resolution no stick phase forms at any bow force. With friction, impedance and quality factor from published measurement rather than fitted, all four strings return a stick fraction of 89.1% against an ideal 90%. Schelleng's maximum bow force is recovered on every string. His minimum, $F_{\min} \propto Z^2 v_b \beta^{-2}$, is replaced by a law, $F_{\min} = C Z v_b / \beta$ with a dimensionless $C = 1.112 \pm 0.017$ that reproduces on sixty-four held-out operating points. Six controllers at matched capacity, on four strings at twenty seeds, place a gated recurrent network ahead of a feedforward one, its margin over the mean of the other five largest when the commanded bow speed is overridden at mid-stroke. The feedforward network completes more strokes, mostly from a cold start no player would use. A minimal gated variant fails because gates computed from the input alone cannot clear a latched state. Training loss selects neither the capacity nor the context length. No learned controller improves on the lookup rule that generated its labels where that rule is correct. Of the rule's 1089 cells, 57 are playable on a properly settled plant and unplayable by its labels, and the controller commands them where the lookup will not. The controller's score regresses on the rule's with a slope of 0.32, more than ten standard errors below unity, so it overtakes the rule where the rule fails and is bounded by it where it holds. Under a rigid finger stop the plant is provably invariant, so transfer loss between pitches is the controller's alone, traced to one feature. A regime classifier without a stick test mistakes small-amplitude periodic slipping for Helmholtz motion, and a harmonicity measure rates a string the bow never grips above it.
|
| 2319 |
SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy
2609.15039
|
cs.LG
|
Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar |
User prompts provided to large language models (LLMs) may contain sensitive or private information that can be misused by remotely deployed models, such as through inadvertent memorization during retraining. One way to protect user prompts is to execute the LL...User prompts provided to large language models (LLMs) may contain sensitive or private information that can be misused by remotely deployed models, such as through inadvertent memorization during retraining. One way to protect user prompts is to execute the LLM inside a trusted execution environment (TEE), with the guarantee that the service provider has no access to computations performed within or information exchanged with the TEE. However, current TEEs are primarily CPU-based and significantly slower than GPUs optimized for LLM inference. To circumvent this, Tramer and Boneh (2019) proposed Slalom, which splits neural network inference between a TEE and an untrusted GPU and encrypts intermediate inputs sent to the GPU. We extend this split-inference architecture to LLM inference and instead protect intermediate inputs using differential privacy. We show that masking intermediate representations is necessary by showing that a prompt-reconstruction attack can recover prompts from these representations with nearly 80% accuracy. Our main contribution is a global sensitivity analysis of key LLM functions, which bounds the required scale of differentially private noise. Unlike encryption, differential privacy avoids quantization, allowing the LLM to remain in the floating-point domain. We also derive an upper bound on floating-point error from masking and noise cancellation in the TEE as a function of the privacy parameter epsilon. We implement our architecture using Intel TDX and evaluate it with two LLMs: Llama-3.2-3B and Qwen3-4B. Our split execution is nearly twice as fast as fully CPU-based inference inside TDX and 5-15 seconds faster than encryption-based Slalom while achieving higher accuracy. Finally, we demonstrate that prompt reconstruction, even with knowledge of the differential privacy mechanism, cannot recover more information than is contained in an unrelated prompt.
|
| 2320 |
Decentralized Gossip Learning and Federated Averaging for Histopathology Image Classification
2609.16448
|
cs.LG
|
Yusuf Ozturk, Bengisu Atli, Enes Goktekin, Akin Ozturk, Ulas Bagci |
Breast histopathology analysis increasingly relies on distributed learning because direct data pooling across institutions is often restricted by privacy, governance, and communication constraints. This study compares server-based Federated Averaging (FedAvg),...Breast histopathology analysis increasingly relies on distributed learning because direct data pooling across institutions is often restricted by privacy, governance, and communication constraints. This study compares server-based Federated Averaging (FedAvg), fully decentralized gossip learning, and Hybrid Gossip-FedAvg for invasive ductal carcinoma (IDC) patch classification. Experiments used 277,524 color image patches with patient-disjoint training, validation, and test partitions and a workload-balanced, Dirichlet-guided allocation across six nodes. Ring, random degree-3, and fully connected gossip topologies were evaluated together with sensitivity analyses for statistical heterogeneity, mixing coefficient, learning rate, model drift, prediction disagreement, calibration, clinically motivated operating points, communication payload, and patient-level IDC burden, together with auxiliary backbone robustness analyses. In the principal alpha=0.3 experiment, Hybrid Gossip-FedAvg achieved a test area under the receiver operating characteristic curve (ROC-AUC) of 0.8811, closely followed by FedAvg at 0.8801 and fully connected gossip at 0.8751. Across three independent patient-level repetitions, FedAvg and Hybrid Gossip-FedAvg obtained the same mean ROC-AUC of 0.9082, with standard deviations of 0.0037 and 0.0043, respectively. Hybrid achieved the highest mean area under the precision-recall curve of 0.8240, whereas FedAvg produced the lowest mean Brier score of 0.1335. Denser gossip graphs improved discrimination but increased theoretical model payload, while ring gossip remained sensitive to learning rate and mixing strength. Overall, FedAvg provided the most consistently reliable server-based baseline, topology-aware gossip offered a viable decentralized alternative, and Hybrid Gossip-FedAvg provided a balanced compromise between peer-to-peer diffusion and periodic global coordination.
|
| 2321 |
VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
2609.18663
|
cs.LG
|
Deyu Cao, Ryuji Oi, Kosuke Matsushima, Yuxuan Pan, Ziheng Wang |
Billion-parameter vision-language-action (VLA) policies run either onboard, consuming substantial power, or on remote servers, adding communication latency. To address these drawbacks and better balance latency and onboard energy consumption, we propose VLA-UL...Billion-parameter vision-language-action (VLA) policies run either onboard, consuming substantial power, or on remote servers, adding communication latency. To address these drawbacks and better balance latency and onboard energy consumption, we propose VLA-ULAP. It partitions inference across decision times, interleaving remote VLA calls with predictions from an Ultra-Lightweight Local Action Predictor (ULAP). A single ULAP has $\sim$7.4M parameters including the frozen vision encoder. It combines current views, proprioception, and executed action history to predict chunks in one pass. Trained independently, it requires no VLA hidden states, online verification, or server round trips. On Jetson Orin Nano, ULAP takes 20.7 ms and 0.122 J of idle-subtracted energy per inference, versus 289.3 ms and 40.46 J for GR00T on RTX A6000. Across four simulated base-policy/benchmark pairs, VLA-ULAP removes 45.3-77.3% of VLA calls while retaining 95.0-98.5% of baseline success rates at selected operating points. On VLA-JEPA, VLA-ULAP also surpasses local acceleration alternatives, using an estimated 49.4% less inference time and 51.5% less GPU energy per successful episode than ACT, and 77.5% less time and 80.7% less energy than SP-VLA at higher success rates. In physical SO-101 trials, it similarly removes 70.3-71.9% of VLA calls without observed success-rate loss at seen or held-out placements. Measured device costs imply 64.6-66.4% less inference time and 70.1-71.7% less idle-subtracted energy per successful episode at these call counts. Beyond these savings, faster responses help VLA-ULAP exceed $\pi_{0.5}$'s success rate by 11.0 and 15.5 percentage points (pp) on two tasks in latency-aware LIBERO-Safety simulation while approximately halving VLA calls.
|
| 2322 |
Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
2609.19122
|
cs.LG
|
Meng'en Qin, Yinchen Liu, Mingxuan Cui, Youlu Xing |
Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components while preserving signal content, but its sparsity coefficient ...Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components while preserving signal content, but its sparsity coefficient is often manually selected during training. We propose a training-adaptive convolutional sparse coding framework for robust visual signal representation. Specifically, we unfold the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) and treat the sparsity coefficient as a differentiable variable jointly learned with the network parameters. From the information bottleneck perspective, this coefficient controls the trade-off between information retention and compression: the sparsity term promotes compact representations, while the reconstruction term together with task loss preserves task-relevant signal content. We further introduce a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed. Experiments on CIFAR and ImageNet demonstrate competitive clean-data recognition and greatly improved robustness under different input perturbations.
|
| 2323 |
When2Think: Learning When and How Much to Reason
2609.19671
|
cs.LG
|
Jaejun Shim, HyunJin Kim, Young Jin Kim, JinYeong Bak |
Large Reasoning Models (LRMs) often overthink easy problems and underthink hard ones, leading to inefficient computation allocation. Existing methods regulate generated computation or select between direct answering and explicit reasoning, but do not jointly c...Large Reasoning Models (LRMs) often overthink easy problems and underthink hard ones, leading to inefficient computation allocation. Existing methods regulate generated computation or select between direct answering and explicit reasoning, but do not jointly control whether}to reason and how much computation to allocate within reasoning. We call the resulting difficulty-dependent loss in accuracy under computation reduction the efficiency tax. We propose When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation. Its core mechanism, Instance-level Difficulty-Aware Control (IDAC), uses cached reference statistics of success and token cost to modulate a correctness-gated efficiency bonus based on generated token count. Importance sampling supports exploration of Think and NoThink, while Batch-Wise Standardization constructs standardized advantages for critic-free optimization. The framework requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates. On AIME24, When2Think improves Pass@3 by 10.0 percentage points while reducing token usage by 27.9% relative to the backbone.
|
| 2324 |
Near-Optimal Single-Loop Predictor--Corrector Extragradient Method for Strongly Convex--Strongly Concave Minimax Optimization
2609.20327
|
cs.LG
|
Minhao Zhang, Zi Xu |
We study smooth strongly convex--strongly concave minimax optimization in the deterministic unconstrained setting, without assuming a bilinear or separable structure. Although existing multi-loop methods attain near-optimal condition-number dependence, standar...We study smooth strongly convex--strongly concave minimax optimization in the deterministic unconstrained setting, without assuming a bilinear or separable structure. Although existing multi-loop methods attain near-optimal condition-number dependence, standard single-loop methods generally exhibit a substantial complexity gap. To close this gap, we propose the Single-Loop Predictor--Corrector Extragradient Method with Damped Momentum (PCE-DM), which combines an extragradient prediction--correction scheme with a novel auxiliary feedback recursion for the weaker-curvature variable. PCE-DM uses fixed parameters and two new full-gradient evaluations per iteration after one initialization query, while requiring no inner solves, accuracy schedules, or staged restarts. We develop a Lyapunov analysis that controls the predictor--corrector mismatch through corrected-gradient increments and establish last-iterate linear convergence. Specifically, PCE-DM computes an $\varepsilon$-accurate relative solution, measured by the squared Euclidean distance to the saddle point, within $\mathcal{O}\!\left(\sqrt{\kappa_x\kappa_y} \log(2\kappa_x\kappa_y/\varepsilon)\right)$ full-gradient queries. This result closes the condition-number complexity gap between standard single-loop methods and near-optimal multi-loop methods, matching the known lower-bound order up to logarithmic factors while retaining fixed, explicit single-loop updates. Numerical experiments on regularized linear regression and AUC maximization demonstrate the computational efficiency of PCE-DM.
|
| 2325 |
Complete Neural Electronic Initialization Accelerates Materials DFT
2609.21759
|
cs.LG
|
Felix {\AE}rtebjerg, Jonas Elsborg, Arghya Bhowmik |
We present the first complete machine learning method for accelerating plane-wave density functional theory (DFT) in materials under the projector augmented wave (PAW) formalism. We formalize seven criteria that a Complete Neural Electronic Initializer must sa...We present the first complete machine learning method for accelerating plane-wave density functional theory (DFT) in materials under the projector augmented wave (PAW) formalism. We formalize seven criteria that a Complete Neural Electronic Initializer must satisfy for practical end-to-end PAW DFT acceleration. Applying these to prior work reveals two structure-dependent components, augmentation occupancies and spin initialization, whose absence prevents existing acceleration methods from providing complete reference-free initialization. We show that omitting these components can eliminate or reverse the acceleration obtained via models that only predict the smooth valence density. We satisfy the missing requirements by introducing AugNet, a general equivariant model for PAW augmentation occupancies, and the first general spin density model for materials, which predicts the smooth spin-difference density and spin-difference PAW augmentation occupancies using predicted magnetic moments to constrain the global magnetic state. Combined with existing valence density models, our full method satisfies all seven criteria and forms a fully reference-free electronic initializer for materials DFT, requiring no electronic quantities from a converged target calculation. We show that perfect initialization could cut PAW DFT wall time by 40-52%, and our method recovers up to 62% of this saving, reducing end-to-end DFT wall time by up to ~25% on unseen structures while preserving converged energies.
|
| 2326 |
Learning and Control Beyond Linearity: Towards a Non-asymptotic Theory for Bilinear Systems
2609.22338
|
cs.LG
|
Yahya Sattar, Yassir Jedra, Robin Str\"asser, Frank Allg\"ower, Maryam Fazel |
This tutorial provides a unified view of the emerging area of bilinear learning and control. Using linear systems as a benchmark, it explains what fundamentally changes in the bilinear settings, how recent theory addresses finite-sample learning and control, a...This tutorial provides a unified view of the emerging area of bilinear learning and control. Using linear systems as a benchmark, it explains what fundamentally changes in the bilinear settings, how recent theory addresses finite-sample learning and control, and how these ideas connect to broader themes in nonlinear control, representation learning, and data-driven decision making. For learning, we emphasize tools that are particularly useful in the bilinear settings, such as one-sided Bernstein's inequality for dependent and heavy-tailed covariates, blocking arguments, and martingale concentration for input-dependent noise. We then apply these tools to obtain finite-sample learning guarantees for fully observed bilinear systems, partially observed bilinear systems, and linear systems with bilinear observations. For control, we discuss quadratic control from bilinear observations, where the classical separation principle fails, and review tractable approaches based on belief-space receding horizon control. We also cover stabilization of bilinear dynamics under state feedback using semi-definite programming, LMI relaxations, sum-of-squares methods, and Koopman-based lifting. We conclude by discussing connections to reinforcement learning and machine learning, and some open problems in combined learning and control of bilinear systems.
|
| 2327 |
RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers
2609.22547
|
cs.LG
|
Chen Yang, Jun Chen |
Reinforcement learning with verifiable rewards (RLVR) often improves pass@1 while falling behind its base model at larger sampling budgets $k$, a crossover read as evidence that RLVR only sharpens existing capability. We identify two limits to this reading. Fi...Reinforcement learning with verifiable rewards (RLVR) often improves pass@1 while falling behind its base model at larger sampling budgets $k$, a crossover read as evidence that RLVR only sharpens existing capability. We identify two limits to this reading. First, a visible crossing need not be statistically established: comparing models on the same prompts, we build confidence bands across sampling budgets $k$ that require evidence of both an early gain and a later loss. Across five public RLVR pairs no crossing is statistically established in the initial evaluations, while a 32k-token evaluation on fresh prompts locates a reversal with first loss between 11 and 61 samples; power analysis shows why failure to detect a crossing need not mean no crossing, and why more prompts can help more than more answers per prompt. Second, base success alone does not determine what RLVR does to a prompt: prompts with the same base success rate have different post-RL success rates, and these differences repeat across independent generation halves. The relationship is a conditional distribution---a Markov kernel---rather than a single curve, and fitting it predicts crossings in independent generations for the same prompts and corrects the simple model's power estimates. Theory further shows how losses on a minority of the hardest prompts can overturn an early lead even when training improves other prompts, separating evidence that a crossover exists from claims about what it means for capability.
|
| 2328 |
A Bayesian Vertical Federated Learning Framework for Multivariate Reduced-Rank High-Dimensional Regression
2609.22654
|
cs.LG
|
Brigham Halverson, Sharmistha Guha, Jessica Bernard, Rajarshi Guhaniyogi |
Federated learning (FL) has emerged as a leading privacy-preserving framework for collaborative machine learning across decentralized environments. While considerable progress has been made in horizontal federated learning (HFL), where data with common feature...Federated learning (FL) has emerged as a leading privacy-preserving framework for collaborative machine learning across decentralized environments. While considerable progress has been made in horizontal federated learning (HFL), where data with common features is distributed across sites, vertical federated learning (VFL), where sites share observations across distinct feature sets, remains less explored. Advancing Bayesian high-dimensional multivariate reduced-rank regression methods for VFL poses unique challenges: (a) stringent privacy regulations preventing local site data sharing, and (b) fitting local regressions overlooks essential modeling aspects like inter-variable correlations. In contrast HFL allows each site to fit a comparable model independently. We present a novel Bayesian VFL framework for multivariate high-dimensional reduced-rank regression, termed BayesVFLReg, which enables precise coefficient estimation while safeguarding both feature and response privacy. Participating sites use a shared random sketching matrix to compress local variables into privacy-preserving sketches. A central server collects these sketches where Bayesian multivariate reduced-rank regression uses Gaussian scale mixture priors. For feature selection, we introduce a single-step post-processing strategy based on mixture-model clustering of the absolute posterior coefficient means to distinguish signal from noise per response variable. BayesVFLReg is computationally scalable for large, high-dimensional datasets and facilitates efficient variable selection. Theoretically, we establish sharp non-asymptotic bounds on the posterior probability that the fitted density falls within a Hellinger ball centered at the true data-generating density. Comparative simulation studies and real-world data analyses show that BayesVFLReg reliably identifies sparse feature effects, even under feature correlation.
|
| 2329 |
PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control
2609.24840
|
cs.LG
|
Lei Ye, Haibo Gao, Yitang Li, Peng Xu, Zetong Jing |
Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving r...Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving recovery and physical execution largely to the tracker. Action-only diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state-action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text-conditioned behavior, while classifier guidance steers predicted states toward test-time objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches 44 of 45 destination targets and achieves a text retrieval score of 0.539, compared with 0.424 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation.
|
| 2330 |
FREESIA: Covariance-Aware Posterior Transport for Expressive and Scalable Data Assimilation
2609.25085
|
cs.LG
|
Shiwei Ni, Yangwen Zhang, Hang Qi, Xiaofei Guan, Lili Ju |
Data assimilation aims to infer the state of complex dynamical systems based on observational data. However, accurate inference of the multimodal posteriors induced by nonlinear or non-injective observation operators remains a key challenge under high-dimensio...Data assimilation aims to infer the state of complex dynamical systems based on observational data. However, accurate inference of the multimodal posteriors induced by nonlinear or non-injective observation operators remains a key challenge under high-dimensional and sparse observation conditions. Ensemble filters scale to high dimensions but are confined by restrictive distributional assumptions, while training-free generative filters (e.g., EnSF, EnFF) alleviate this limitation but may introduce structural errors and hinder information propagation under sparse observations. To address these issues, we propose a training-free, asymptotically exact posterior transport method. Firstly, a covariance-aware posterior transport scheme is designed, which embeds the forecast cross-covariance into flow-based transport and accurately recovers unobserved states while preserving the non-Gaussian posterior structure. Furthermore, the method combines a tractable observation-adaptive proposal with posterior correction, ensuring accurate approximation of the nonlinear posterior distribution. Finally, we establish the corresponding posterior flow theory, from which the asymptotic exactness of the proposed method relative to finite-ensemble surrogates and the Wasserstein error bound are derived. Experiments on Double-Well, Lorenz-96, and Kolmogorov flow show that the proposed method captures complex posterior structure and remains accurate under sparse, nonlinear, and non-injective observations. In the sparse non-injective setting, it reduces RMSE by 56% relative to the best baseline.
|
| 2331 |
GINIO: A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry
2609.25338
|
cs.LG
|
Chankyo Kim, Minghan Zhu, Tzu-Yuan Lin, Avantika Rattan, Maani Ghaffari |
Neural inertial odometry increasingly uses networks as learned measurements inside filtering pipelines. Such measurements should transform consistently under arbitrary IMU mounting conventions: their mean must transform as a vector, and their covariance must t...Neural inertial odometry increasingly uses networks as learned measurements inside filtering pipelines. Such measurements should transform consistently under arbitrary IMU mounting conventions: their mean must transform as a vector, and their covariance must transform congruently as a second-order tensor. We present GINIO, a geometric SO(3)-equivariant interface for neural inertial odometry under arbitrary rotations of the IMU measurement frame. Given calibrated IMU windows, our framework predicts a motion measurement and uncertainty obeying these tensorial laws. To support efficient sensor-frame learning, we introduce Last-Frame Alignment (LFA), a deterministic preprocessing step that is provably equivalent to world-frame training for SO(3)-equivariant predictors. The connected estimator tracks sensor-local states such as IMU bias, separating nuisance estimation from the geometric law enforced by the learned measurement. We instantiate the same interface in filter-connected NIO, AirIO-style recurrent aerial prediction, EqNIO-style full-SO(3) canonicalization, and ResNet-style temporal backbones. On TLIO, GINIO achieves 2.018 m ID/SO(3) ATE while EqNIO degrades to 76.389 m, using 11.6x fewer FLOPs. On NanoBench, our AirIO-style instantiation improves ATE from 5.579 m to 1.430 m without external attitude input, and our ResNet-style instantiation reaches 0.581 m ATE versus 0.645 m for ResNet1D. On Fetch, GINIO empirically reduces unseen physical-remount ATE from 8.15 m to 0.50 m without retraining, demonstrating robustness beyond the exact coordinate-frame guarantee. For uncertainty, spectral covariance reduces covariance-equivariance error by over three orders of magnitude compared with a diagonal head.
|
| 2332 |
SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation
2609.25352
|
cs.LG
|
Fatih Deniz, Yazan Boshmaf, Issa Khalil |
Evaluation of large language models (LLMs) for safety, security, and privacy (SSP) relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. ...Evaluation of large language models (LLMs) for safety, security, and privacy (SSP) relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. As a result, models that perform well on fixed test sets often fail under semantically equivalent rephrasings. We introduce SSP-Bench, a dynamic benchmarking framework that generates evaluation instances on demand while preserving domain consistency. The framework ensures label validity through externally grounded sources, enforces scope via service-specific validation, and calibrates difficulty using a multi-model steering panel. Benchmark construction is formulated as a multi-objective optimization problem over difficulty, separability, novelty, and diversity. Across 24 models and four SSP services, SSP-Bench reveals systematic failures of static evaluation, including near-zero correlation in safety rankings due to construct mixing, strong safety--over-refusal coupling, and hidden within-family regressions. These results show that static benchmarks can misrepresent model behavior, motivating dynamic, deployment-relevant evaluation.
|
| 2333 |
Statistical Gains from Looped Estimation under Parameter Budgets
2609.25778
|
cs.LG
|
Xinyu Tian, Xiaotong Shen |
Memory constraints in artificial intelligence motivate accurate function approximation with fewer parameters. We study looping, which repeatedly composes one update function with shared parameters; each output becomes the next input. A looped Transformer, for ...Memory constraints in artificial intelligence motivate accurate function approximation with fewer parameters. We study looping, which repeatedly composes one update function with shared parameters; each output becomes the next input. A looped Transformer, for example, reuses one block, whereas its conventional untied counterpart uses separately parameterized blocks. We compare their parameter requirements for a given worst-case approximation accuracy, or equivalently, their approximation accuracy under a common budget limiting distinct trainable coefficients. We then ask whether this representational parsimony improves statistical accuracy. For general likelihood models, we establish an upper squared Hellinger risk bound for looped sieve maximum likelihood and a minimax lower bound for the jointly tuned untied family. Further loop iterations improve the approximation bound without adding parameters, while increasing computation and the fitted-class complexity bound. For targets of known H\"older smoothness, looped residual feedforward networks and post-layer-normalized Transformers attain the minimax polynomial rate up to logarithmic factors with a fixed number of bounded real parameters. At sufficiently large fixed budgets, looped worst-case risk vanishes while optimal worst-case untied risk remains bounded away from zero. The loop-to-untied risk ratio also tends to zero under specified growing-budget conditions. Regression, binary response, and energy-based generative models illustrate the theory.
|
| 2334 |
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
2609.27490
|
cs.LG
|
Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li |
AI research agents must predict the effects of computational changes after budgeted experiments. WhatWorkedBench evaluates this experimental understanding through a delivered response surface of configuration scores. Agents inspect workflow code and buy measur...AI research agents must predict the effects of computational changes after budgeted experiments. WhatWorkedBench evaluates this experimental understanding through a delivered response surface of configuration scores. Agents inspect workflow code and buy measurements; exhaustive CPU references score conditional component effects, mean pair interactions, configuration choice, and delivery. The catalog spans 36 task conditions, 30 sources, 8 workflow families, and 1248 indexed configuration records, with 4,206 numerical controls. At eight purchased measurements plus two free anchors, pair-effect ridge selects an exact optimum on 15 of 22 sources; 13 of these cases have at least one conditional-effect error exceeding 10% of the task utility range. Shared-estimator comparisons measure acquisition and reconstruction on common observations. In a prospective typed study on 12 four-factor sources, DeepSeek Flash submits 12/12 direct tables and gains 0.147 recovery over a Gaussian process (GP) fitted to the same observations. Pro delivers 11/12 artifacts, with an all-attempt GP difference of -0.001 and a delivered-only difference of +0.063. On six prespecified new agent-evaluation sources, Flash and Pro gains are 0.149 and 0.074. In eight typed six-factor episodes, seven final tables satisfy verified code equivalences. Four fresh agents pass all six registered rules through named estimators and deliver consistent tables at 0.682 recovery versus 0.710 for separate direct-table runs. WhatWorkedBench links acquisition, inference, program structure, and delivery.
|
| 2335 |
How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
2609.28177
|
cs.LG
|
Chen Yang, Jun Chen |
LLM leaderboard gains can reflect selection among privately evaluated model variants, yet neither the number of variants nor their dependence is public. We ask how many hidden variants a published margin can support while retaining statistical evidence of a pr...LLM leaderboard gains can reflect selection among privately evaluated model variants, yet neither the number of variants nor their dependence is public. We ask how many hidden variants a published margin can support while retaining statistical evidence of a provider's advantage over a fixed comparator. For a fixed candidate family under a Gaussian margin model, we derive a sensitivity curve that reports this maximum count as a function of a lower bound on within-family correlation. The relevant correlation must match the score used for ranking and the sampling model: in a controlled family, pooled item correlation is 0.90, whereas composite-score correlation is 0.46 under item resampling and 0.92 when MMLU subjects are resampled. An item-based audit of 394 adjacent-rank claims on the Open LLM Leaderboard finds that 391 lack statistical support even before accounting for selection. Among claims that pass the uncorrected test, certification can depend on assumptions about the hidden family's correlation. The resulting curves make these assumptions explicit without estimating the unobserved search size.
|
| 2336 |
Non-Commutative State Tracking with Input-Dependent Low-Rank Updates in Mamba-3
2609.28273
|
cs.LG
|
Hiroki Fujii, Masaki Yamakita |
State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracki...State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracking, in which the order of operations matters. The rank-one update couples state coordinates along an input-dependent direction, enabling non-diagonal state transitions within a single Mamba-3 block. The extension preserves Mamba-3's exponential-trapezoidal discretization, rotary embeddings (RoPE), and readout. For training, we adapt chunkwise computation to parallelize the proposed recurrence within each chunk. Experiments cover group word problems with discrete inputs and a shell game with continuous observations, in which a policy is trained by behavioral cloning. Among the models selected for their strong performance under fixed timing, the proposed model maintains higher tracking success on longer swap sequences in the shell game with continuous observations and timing jitter. These experiments show that the proposed method achieves high accuracy on the evaluated non-commutative tracking tasks, improving on standard Mamba-3. The extension thus offers a Mamba-3-based approach to non-commutative state tracking.
|
| 2337 |
HClimRep-Ocean: A Global Ocean Emulator on an Unstructured Mesh
2609.28601
|
cs.LG
|
Kacper Nowak, Aleksei Koldunov, Nikolay Koldunov, Savvas Melidonis, Ankit Patnala |
Machine-learning (ML) emulators for atmospheric processes have advanced rapidly in recent years, transforming weather forecasting. Although early ML ocean forecasting models now exist, they remain less developed than their atmospheric counterparts. Unlike the ...Machine-learning (ML) emulators for atmospheric processes have advanced rapidly in recent years, transforming weather forecasting. Although early ML ocean forecasting models now exist, they remain less developed than their atmospheric counterparts. Unlike the atmosphere, much of the ocean's kinetic energy resides in mesoscale eddies whose characteristic spatial scales are approximately an order of magnitude smaller than those of comparable atmospheric features. Moreover, complex coastlines, narrow straits, and ice-covered seas make boundary representation a central challenge that atmospheric models do not face. Consequently, numerical ocean simulations commonly use locally refined or even completely unstructured meshes. However, their data-driven counterparts have so far been built around latitude-longitude grids. We present HClimRep-Ocean, an ocean emulator that operates directly on the native unstructured mesh of FESOM2. The emulator is trained on a 209-year AWI-CM3 control integration and is run without atmospheric forcing, receiving the atmospheric state only at initialisation time, which isolates the predictability carried by the ocean state itself. Skill is strongly field-dependent: for currents, HClimRep-Ocean outperforms every reference at 30 day forecast, whereas for temperature and salinity a damped-anomaly persistence forecast remains the more accurate estimator. This behaviour is physically interpretable: current variability is largely geostrophic and internally generated, whereas sea-surface temperature and salinity fluctuations are driven by atmospheric forcing through weather state. Evaluated independently on the OceanBench benchmark, a reanalysis-trained variant of HClimRep-Ocean achieves the lowest RMSE against GLORYS reanalysis among all assessed systems, confirming the competitiveness of the native-mesh approach.
|
| 2338 |
Learning a Flow to Self-Supervised Representations
2609.29350
|
cs.LG
|
Yuling Jiao, Wensen Ma, Houduo Qi, Defeng Sun |
Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a...Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to exceed the auxiliary flow dimension d* while retaining structured geometric separation. We assign both augmented views of each image to the same target, while limiting how many images each reference center can receive. An explicit alignment loss further pulls the two views' representations closer together. Experiments across benchmarks ranging from CIFAR to ImageNet show that FBDM achieves performance nearly on par with DM and remains competitive with existing SSL methods. Matched training-cost comparisons show a 1.48- to 1.83-fold speedup over DM with a negligible increase in GPU memory usage. We also provide a theoretical explanation for the usefulness of the learned representations: under stated conditions, we bound the downstream misclassification rate in terms of the FBDM pretraining loss.
|
| 2339 |
Cost-Sensitive Online Window Size Selection for Portfolio Management
2609.29887
|
cs.LG
|
Yi-Chen Liu, Chung-Han Hsieh |
This paper investigates cost-sensitive online window size selection for portfolio management under changing market conditions. Specifically, we propose a two-level framework that constructs portfolios using candidate window sizes and dynamically aggregates the...This paper investigates cost-sensitive online window size selection for portfolio management under changing market conditions. Specifically, we propose a two-level framework that constructs portfolios using candidate window sizes and dynamically aggregates them through online learning. By treating candidate window sizes as ``experts,'' we dynamically update their aggregation weights using turnover-inclusive losses. Moreover, we derive finite-horizon cost-sensitive tracking-regret bounds that account for turnover of the aggregated portfolio, with static regret as a special case. Under bounded losses and cost rates, suitably tuned Fixed Share achieves asymptotically no tracking regret for sublinear switching budgets, with Hedge covering the static case.
|
| 2340 |
Low-Rank Friction for Memory-Efficient Transformer Pretraining
2609.30342
|
cs.LG
|
Rajit Rajpal, Benedict Leimkuhler |
iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathca...iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as Adam's second-moment buffer. Here we replace iKFAD's friction tensor $\xi$ with a rank-1 outer-product factorisation built from row and column momentum statistics, resulting in Rank-1 iKFAD (R-iKFAD). This reduces the friction memory footprint from $\mathcal{O}(mn)$ to $\mathcal{O}(m+n)$ per layer, which approximately halves iKFAD's total optimiser state. Despite this reduction, R-iKFAD maintains parity in performance with iKFAD: experiments on GPT2-Nano, TinyViT, DistilBERT and GPT2-S confirm that it matches or exceeds iKFAD while nearly halving the memory footprint and remaining comparably robust to hyperparameters. We analyse the continuous-time dynamics in two damping regimes. For linear damping ($\gamma>0$) we prove exponential convergence under strong convexity. For $\gamma=0$, the preferred option in our experiments, the friction is generated entirely from past momentum and switches off as the momentum vanishes, so geometric convergence cannot be shown. We nonetheless prove convergence to the minimiser, together with matching upper and lower bounds on the energy: of order $t^{-1}$ when the regularisation scale $\epsilon_{\mathrm{stab}}$ is zero, and of order $t^{-1/2}$ when it is positive. To our knowledge this is the first convergence rate for a rank-1 factored optimiser in continuous time, and the first such result that does not require positive damping.
|
| cs.MM 3 papers | ||||
| 2979 |
Toward Generative Video Communication: A Dual-Stream Digital Transmission Framework
2609.36565
|
cs.MM
|
Bingyan Xie, Longyu Zhou, Tianhao Liang, Yongpeng Wu, Zehui Xiong |
Generative video communication has shown promise for bandwidth-constrained wireless transmission and has the potential to support personalized content delivery. In this article, we propose a dual-stream digital generative video communication (DGVC) framework t...Generative video communication has shown promise for bandwidth-constrained wireless transmission and has the potential to support personalized content delivery. In this article, we propose a dual-stream digital generative video communication (DGVC) framework that integrates a traditional digital link with a generative link. The traditional link provides source-grounded visual references, while the generative link conveys compact semantic and perceptual information for receiver-side generation. We further discuss three bandwidth-dependent operating regimes and key technologies for dual-stream coordination, synchronization, reliability, and latency control. A practical case study demonstrates the perceptual and temporal-quality benefits of DGVC under wireless fading channels. Finally, we discuss open challenges and future research directions for generative video communication.
|
| 2980 |
SPECTRA: On-Device Cognitive Perturbation and Trajectory Analysis for Autonomous Edge-Cloud GUI Grounding
2609.35775
|
cs.MM
|
Zhan Qu, Hui Zang, Ran Chen, Tao Wang, Shengyu Zhang |
The effectiveness of edge-cloud collaboration for GUI grounding depends on autonomous requesting, where the edge agent selectively offloads complex tasks to the powerful cloud. However, in visually dense scenarios, lightweight edge agents often exhibit overcon...The effectiveness of edge-cloud collaboration for GUI grounding depends on autonomous requesting, where the edge agent selectively offloads complex tasks to the powerful cloud. However, in visually dense scenarios, lightweight edge agents often exhibit overconfident hallucinations, leading to a misalignment between confidence and accuracy that hinders reliable autonomous requesting. To address this, we leverage the observation that an agent's cognitive instability leads to significant latent drift under minute perturbations due to steep decision boundaries. We propose SPECTRA, a lightweight autonomous request framework for edge-cloud GUI grounding, comprising (1) Saliency-Guided Targeted Perturbation and (2) Efficient Cognitive Trajectory Analysis. SPECTRA conducts a visual cognitive stress test by injecting masks into critical visual anchors and quantifies the topological divergence of the agent's high-dimensional cognitive trajectories during the prefill phase, avoiding inefficient output decoding. Experiments demonstrate that SPECTRA performs cloud request assessment without autoregressive decoding. Our GTA1-32B+InfiGUI-G1-3B and GTA1-32B+Holo1.5-3B maintain 93.44% and 95.60% of cloud-only performance with average request rates of 37.58% and 39.24%, respectively.
|
| 2981 |
Structured Interaction, Visual Localization, and Robust Execution for Complex Web Tasks: A Technical Report on the WebRetriever Challenge
2609.35904
|
cs.MM
|
Ziqi Zhang, Shaohui Li, Bing Li |
This report presents the web agent system developed for the WebRetriever Challenge. The system follows a structuredinteraction- first strategy, using semantic webpage information for routine browser operations and invoking visual perception only when structure...This report presents the web agent system developed for the WebRetriever Challenge. The system follows a structuredinteraction- first strategy, using semantic webpage information for routine browser operations and invoking visual perception only when structured representations are insufficient. Three key designs are introduced: grid-assisted visual localization for difficult-to-access controls, hierarchical context management for reducing redundant page and interaction history, and fault-aware execution mechanisms for stable multi-browser task processing. The system achieved a pass rate of up to 79% in local evaluation on Protocol 1. In the official Protocol 3 competition, it achieved a 59% pass rate with eight concurrent browser workers and ranked first overall, winning the WebRetriever Challenge.
|
| cs.SD 25 papers | ||||
| 2940 |
Enabling Immersive Audio-Visual Experience from Any Video
2609.36295
|
cs.SDeess.AScs.MM
|
Zitong Lan, Mutian Tong, Jiatao Gu, Mingmin Zhao |
Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial sound...Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial soundscape. Without spatially consistent audio, these expanded visual worlds remain incomplete. This paper presents OmniDream, a training-free framework that transforms a silent monocular video into an immersive audiovisual experience, where viewers can freely look around while sounds remain spatially aligned with the visual scene. At the core of OmniDream is an object-centric audio representation that disentangles each sound source's intrinsic audio content from its scene-dependent acoustic effects, enabling independent audio generation, physics-based simulation of propagation effects, and flexible spatial audio rendering. Experiments show improved audio-visual alignment, spatial correctness, and perceptual immersiveness over baselines. Examples are available on https://huggingface.co/spaces/CuriousAlien000/spatial-audio-360-demo
|
| 2941 |
Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech
2609.36324
|
cs.SD
|
Yentl Collin, Evan Dufraisse, Amr Mohamed, Amine Khelif Khelif, Dani Bouch |
Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically...Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.
|
| 2942 |
Trigger Sound Suppression for Misophonia
2609.36351
|
cs.SD
|
Vaishnavi Vidyasagar, Jasmine Zhang, Mahima Uliyar, Seunghyun Oh, Emily Catherine Gates |
Misophonia, a disorder of decreased tolerance to specific sounds, affects 5-20% of the population, yet sufferers have no good options: therapy helps a minority, and earplugs or noise cancellation silence everything. We present a study for neural trigger sound ...Misophonia, a disorder of decreased tolerance to specific sounds, affects 5-20% of the population, yet sufferers have no good options: therapy helps a minority, and earplugs or noise cancellation silence everything. We present a study for neural trigger sound suppression for misophonia, selectively removing trigger sounds. We curate a dataset covering the 10 most common trigger classes. Using streaming dual-path networks operating on 6 ms audio chunks, we explore both one-hot and multi-hot-conditioned models that suppress 1-3 triggers from the acoustic scene. We validate our model outputs in a listening study with 30 adults with clinically elevated misophonia impairment. Participants reported significantly lower distress and arousal, and improved valence, for suppressed audio.
|
| 2943 |
UDSS-BWE: Uncertainty- and Decision-Science Inspired Swin BandWidth Extension
2609.36379
|
cs.SD
|
Tarikul Islam Tamiti, Sajid Fardin Dipto, David Vergano, Luke Baja-Ricketts, Anomadarshi Barua |
Bandwidth extension (BWE) is fundamentally localized: the most perceptual distortions are not average-case distortions, but rare high-frequency (HF) transients that standard, risk-neutral objectives tend to smooth away. To close this gap, we seek solutions in ...Bandwidth extension (BWE) is fundamentally localized: the most perceptual distortions are not average-case distortions, but rare high-frequency (HF) transients that standard, risk-neutral objectives tend to smooth away. To close this gap, we seek solutions in the risk-sensitive and uncertainty-aware decision science rules and present UDSS-BWE, which introduces five decision-science and uncertainty-aware discriminators: CVaRD (does tail pooling to amplify HF artifacts), CCD (a primal-dual augmented Lagrangian to prevent HF overboost), MCUD (a learnable utility over spectral flatness/ centroid/ rolloff), EDD (captures epistemic uncertainty), and DROD (captures entropic KL-DRO aggregation). UDSS-BWE is also designed as a complex valued adversarial BWE framework that uses Swin-based generators, a lightweight dual-stream shifted-window backbone, to capture local and long-range structure efficiently, while learnable lattice coupling provides controlled cross-stream exchange. UDSS-BWE is optimized extensively and achieves better perceptual quality with 3.89x fewer parameters (72M vs.18.5M) over two English and French datasets under clean and noisy conditions. To the best of our knowledge, this work shows how multi disciplinary decision-science-inspired and uncertainty theories can be successfully used to design efficient discriminators for producing more nuanced audios, establishing a new baseline in the BWE task.
|
| 2944 |
Reconstructing the Vocal Tract with Differentiable Acoustic Simulation
2609.36737
|
cs.SDeess.AS
|
Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann |
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by pr...The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
|
| 2945 |
When Capabilities Fail to Compose: Diagnosing the Compositionality Gap in Large Audio-Language Models
2609.36921
|
cs.SD
|
Chien-Feng Liu, Chih-Kai Yang, Bo-Han Feng, Yu-Hsuan Li Liang, Hung-yi Lee |
Large audio-language models (LALMs) perform strongly on individual audio tasks, but whether these capabilities can be reliably composed remains underexplored. We conduct a controlled diagnostic study of capability composition in LALMs, requiring models to inte...Large audio-language models (LALMs) perform strongly on individual audio tasks, but whether these capabilities can be reliably composed remains underexplored. We conduct a controlled diagnostic study of capability composition in LALMs, requiring models to integrate audio-attribute recognition, cue-conditioned segment selection, and downstream ASR or question answering. We construct two-utterance inputs with distinct acoustic cues to evaluate composition across environmental sound, gender, and emotion cues, with ASR, Math QA, and Factual QA as downstream tasks. Across four open-source LALMs, compositional QA accuracy decreases in 39 of 40 model-task-cue settings, by an average of 26.7 percentage points. ASR exhibits a similarly consistent degradation, with WER increasing in 39 of 40 settings by an average of 28.5 percentage points, while the magnitude of degradation varies across models, cue types, and cue salience. We further probe these failures through output format, positional preference, and chain-of-thought (CoT) analyses. Our study reveals a systematic gap between possessing individual audio capabilities and reliably composing them.
|
| 2946 |
Interpreting and Evaluating Dynamic-Rate Speech Codec Boundaries
2609.36951
|
cs.SD
|
Han Wang, Jiaqi Li, Yingda Shen, Yuxiang Wang, Zhizheng Wu |
Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful...Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful for neural speech reconstruction. This work combines boundary interpretation and controlled reconstruction analysis by comparing predicted boundaries with linguistic and acoustic references. We find that boundary meaning depends on the underlying speech representation. For the ASR-oriented SenseVoice and Whisper encoders, shallow layers emphasize phonetic, voicing, and acoustic transitions, whereas deeper layers shift toward syllable and subword structure. In a comparison of six dynamic-frame-rate algorithms and a uniform (fixed-frame-rate) baseline, higher-level linguistic alignment is associated with lower pooling distortion and better reconstruction from semantic tokens. Frame-rate-matched boundary tests make the distinction concrete: a syllable-derived partition improves over both Uniform and Similarity, while a phoneme-derived partition does not improve over Uniform. We infer that for semantically rich speech representations, useful codec boundaries are best understood as allocation decisions organized around the syllable scale.
|
| 2947 |
RVQ Position Aware Speculative Decoding for On Device Text to Speech
2609.37007
|
cs.SD
|
Berkin Durmus, Eduardo Pacheco, Zach Nagengast, Atila Orhon |
Autoregressive decoding (AR) with Transformer models is memory bandwidth bound at single stream inference, the typical deployment regime for on device text to speech (TTS). Real time streaming with Qwen3-TTS requires more than 200 sequential model calls per se...Autoregressive decoding (AR) with Transformer models is memory bandwidth bound at single stream inference, the typical deployment regime for on device text to speech (TTS). Real time streaming with Qwen3-TTS requires more than 200 sequential model calls per second, dominated by the inner loop MultiCodeDecoder that emits the 15 residual vector quantization (RVQ) codes per 80 ms audio frame. We propose RVQ position aware speculative decoding for the MultiCodeDecoder, attaining 2.47 accepted tokens per model call at $5\times10^{-4}$ percent added parameters and 10 to 20 percent per round speculation/verification overhead, reducing real time synthesis from 200 to 88 sequential model calls per second. The scheme is distributionally lossless under the deployed top-k sampling, and WER parity with the original system is consistent with this guarantee. We deliver 2 to 2.2x speedup for RVQ token generation with Qwen3-TTS 0.6B on recent iPhone and Apple Silicon Mac devices.
|
| 2948 |
ReDimNet2+: Multi-Corpus Data Scaling for Robust Speaker Verification
2609.37014
|
cs.SD
|
Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian |
Automatic speaker verification must remain reliable across devices, rooms, and compression pipelines. We present ReDimNet2+, which scales training of the compact ReDimNet2 backbone across seven public corpora (63,934 speakers, about 8,675 hours). Analysis of a...Automatic speaker verification must remain reliable across devices, rooms, and compression pipelines. We present ReDimNet2+, which scales training of the compact ReDimNet2 backbone across seven public corpora (63,934 speakers, about 8,675 hours). Analysis of a VoxBlink2 subset reveals a shift in predicted spectral coloration, motivating codec and waveform augmentation alongside this multi-corpus training, large-margin fine-tuning (LMFT), and graph-based retrieval reranking. With random 4-second evaluation windows for all models, ReDimNet2+ LMFT reduces pooled VoxCeleb1 EER from 2.42% to 0.82% and a 26-condition robustness stress-test EER from 7.21% to 1.99%. Under this shared local protocol, it reaches 0.35% EER on VoxCeleb1-O versus 0.787% for the best evaluated WeSpeaker checkpoint. On a VoxBlink2 retrieval subset, reranking improves the final model's Pr@k from 0.7413 to 0.7687.
|
| 2949 |
RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning
2609.37028
|
cs.SD
|
Maxim Maslov, Kirill Borodin, Vasilii Kudryavtsev, Nikita Vasiliev, Grach Mkrtchian |
Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these wavefor...Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.
|
| 2950 |
Prediction-Layer Branch Calibration for Multimodal Sentiment Analysis
2609.37100
|
cs.SDcs.MM
|
Yulin Sun, Kele Xu, Yong Dou |
Multimodal sentiment analysis integrates textual, acoustic and visual cues, yet current language-model-based fusion methods typically leave prediction-layer branch allocation implicit. We introduce Branch-Calibrated Multimodal Language Fusion (BC-MLF), which e...Multimodal sentiment analysis integrates textual, acoustic and visual cues, yet current language-model-based fusion methods typically leave prediction-layer branch allocation implicit. We introduce Branch-Calibrated Multimodal Language Fusion (BC-MLF), which explicitly models prediction-layer branch allocation through a Branch-Calibrated Task Head (BCHead), complemented by Fusion Token Contrastive Learning (FTCL) for sentiment-aware fusion-token regularization. FTCL organizes mean-pooled fusion-token representations according to continuous sentiment affinity, while BCHead combines fusion, text and audiovisual predictions through a lightweight sample-adaptive constrained mixture. Without modifying the fusion backbone, BC-MLF consistently improves the reproduced DeepMLF baseline and achieves the strongest results among the compared methods on CMU-MOSEI and CH-SIMS across classification and regression metrics. The controlled ablations show that sample-adaptive prediction-layer branch allocation consistently outperforms static branch aggregation. Code is available at https://github.com/sunyulin0421/BC-MLF.
|
| 2951 |
Bad: Taming the Bioacoustic Data Deluge with a Bat Acticity Detector
2609.37518
|
cs.SD
|
Stefano Ciapponi, Santiago Martinez Balvanera, Andrea Cesaretti, Elisabetta Farella, Kate E. Jones |
Passive Acoustic Monitoring of bats generates massive ultrasonic datasets (>27 GB/night per node), straining edge storage and battery life. Legacy triggers fail against acoustic confusers, while deep models exceed microcontroller limits. We present a hardwa...Passive Acoustic Monitoring of bats generates massive ultrasonic datasets (>27 GB/night per node), straining edge storage and battery life. Legacy triggers fail against acoustic confusers, while deep models exceed microcontroller limits. We present a hardware-aware Bat Activity Detector (BAD) specifically designed to discriminate bat calls from hard biological and environmental confusers across variable sampling rates (192-384 kHz). Tailored for the Silicon Labs EFM32PG26 (MVP) in 8-bit integer precision, our model achieves 100 percent hardware offload across all 14 layers (17.2 KB Flash, 73.1 KB RAM). End-to-end preprocessing (74.00 ms for 76 frames) and inference (30.00 ms) of 100 ms clips at 192 kHz require 104.00 ms per clip. On spatially out-of-domain recordings under a realistic low-prevalence regime (r_pos = 0.05), BAD achieves an AUC-ROC of 0.9748 and suppresses 99.4% of non-target noise frames while retaining 65.3% of bat calls - delivering a >33x precision gain over classical Goertzel baselines.
|
| 2952 |
Rate-Agnostic Bioacoustics: Heterogeneous Multi-Taxa Classification with Continuous Filterbanks and Fourier Neural Operators
2609.37540
|
cs.SD
|
Stefano Ciapponi, Francesco Ardan Dal R{\i}, Nicola Conci, Elisabetta Farella |
Conventional bioacoustic classification models rely on fixed-rate spectral representations, requiring recordings acquired at heterogeneous sampling rates to be resampled before analysis. We propose a Sampling-Frequency-Independent (SFI) frontend that processes...Conventional bioacoustic classification models rely on fixed-rate spectral representations, requiring recordings acquired at heterogeneous sampling rates to be resampled before analysis. We propose a Sampling-Frequency-Independent (SFI) frontend that processes each recording directly at its native sampling rate, coupled with a Fourier Neural Operator (FNO) backbone featuring progressive temporal-scale fusion. This framework avoids fixed-rate resampling and high-frequency information loss while producing fixed-size representations across sampling rates. Mild training-time sampling-rate (\textit{sr}) augmentation further improves robustness to unseen rate variations. Evaluated on a multi-taxa corpus comprising 84 classes and 60 sampling rates, the proposed SFI-FNO configuration outperforms fixed-rate and corpus-maximum-rate baselines, achieving .906 accuracy, .921 balanced accuracy, and a Macro-F1 score of .899.
|
| 2953 |
Do Music Generative Models Understand Musical Qualities? Automatic Music Evaluation with Model-Intrinsic Signals
2609.37710
|
cs.SD
|
Xiaosha Li, Chun Liu, Ziyu Wang |
Current music generative models can produce high-quality music, but does this ability imply that they ``understand'' the musical qualities of their outputs, and is that understanding aligned with human evaluation? Previous attempts to use the likelihood of a g...Current music generative models can produce high-quality music, but does this ability imply that they ``understand'' the musical qualities of their outputs, and is that understanding aligned with human evaluation? Previous attempts to use the likelihood of a generative model to evaluate music, an approach commonly used in text, have proven unsuccessful, leading researchers to rely on standalone supervised music evaluation models. In this paper, we answer this question affirmatively: we show that a model's intrinsic signals---derived from its hidden representations and predictions---are strongly correlated with human ratings. In particular, we study MusicGen and consider three types of features: (1) prediction loss, (2) prediction entropy, and (3) concepts extracted from the model using a sparse autoencoder (SAE). Using these features, we train a lightweight prediction model to estimate subjective ratings. We evaluate these features both individually and in combination. We hypothesize that these signals parallel the listening process: the temporal and frequency-domain structure of loss and entropy reflects listeners' expectation and surprise, while gradient directions in SAE latent space predict perceived quality. Experiments on five human-evaluation benchmarks spanning continuous ratings and pairwise preferences confirm this hypothesis, with SAE latents carrying most of the predictive signal.
|
| 2954 |
2-Dimensional spectral gating for denoising bioacoustics recordings
2609.37910
|
cs.SD
|
Julien Boussard, M\'elisande Teng, Sulagna Saha, Mario Gallego-Abenza |
Isolating vocalizations from noise in bioacoustics recordings is a prerequisite to many ecological analyses, including species identification, animal communication understanding, and individual or population-level variability studies. However, when recordings ...Isolating vocalizations from noise in bioacoustics recordings is a prerequisite to many ecological analyses, including species identification, animal communication understanding, and individual or population-level variability studies. However, when recordings are acquired in open environments, vocalizations, noise, or signal-to-noise ratio can vary widely across individuals, species, environment, and recording conditions, making it hard to develop robust and generalizable methods for ecological analyses. To account for these challenges, noise reduction techniques are used to remove noise before downstream analyses. Popular methods such as Noisereduce rely on spectral gating, which estimates a noise threshold for each frequency channel. We propose a further improvement to Noisereduce, leveraging the fact that most animal sounds have structure across multiple frequencies. We apply our method on bird and marine mammals recordings and show that our extension of Noisereduce leads to improved denoising in both above and under water acoustic recordings, without impacting speed of preprocessing.
|
| 2955 |
Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs
2609.38106
|
cs.SD
|
Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar |
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we s...Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on SLAM-ASR for different demographic groups. Using the Fair-Speech and Common Voice datasets, we found that the pruning does not affect all demographic groups equally; the gap between best- and worst-performing groups increases in fold. These disparities appear across all three encoder scales, but only the largest model initially hides them behind aggregate WER. LoRA adaptation improves WER for every group, but benefits groups already performing well more strongly and widens for certain groups. On Common Voice English, Danish, and Dutch, accent gaps persist but do not clearly widen, showing that the fairness effects of pruning vary across datasets and must be measured directly. Our findings suggest that for pruned models, deployment decisions should include per-group WER, with the worst-performing group's error rate as an explicit criterion.
|
| 2956 |
EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
2609.38157
|
cs.SDeess.AS
|
Kuan-Po Huang, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni |
Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering,...Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.
|
| 2957 |
Beyond Discrimination: Calibrated Geoprior Fusion for Bioacoustic Monitoring
2609.35863
|
cs.SDeess.AS
|
Neha Sajja, Bart van Merri\"{e}nboer, Burcu Karagol Ayan, Tom Denton |
Modern bioacoustic foundation models like Perch and BirdNET can identify species with high discriminative accuracy, yet their confidence scores are often uncalibrated and difficult to interpret as probabilities of real-world occurrence. This limits their use f...Modern bioacoustic foundation models like Perch and BirdNET can identify species with high discriminative accuracy, yet their confidence scores are often uncalibrated and difficult to interpret as probabilities of real-world occurrence. This limits their use for ecological inference beyond threshold-based detection. We leverage a global annotated acoustic dataset (WABAD) to produce calibration priors for an acoustic model, optionally incorporating species-level information. We introduce new methods of fusing the acoustic predictions with geopriors, which empirically improves calibration while preserving discrimination. Together, these results suggest a path to simpler and more reliable acoustic monitoring for broad biodiversity.
|
| 2958 |
InstCharVoice: Grounding Natural-Language Instructions for Character-Level Control in Text-to-Speech
2609.36287
|
cs.SDeess.AS
|
Sihang Nie, Xueru Li, Xiaofen Xing, Deyi Tuo, Cheng-Bin Jin |
Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit...Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.
|
| 2959 |
Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech
2609.36974
|
cs.SD
|
Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian |
Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sen...Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ever repeats back-to-back. Six models from three architectures render the controls almost perfectly and fail the repeated twins: 94.3% against 18.2% exactly right at k >= 6. The gap survives greedy decoding, repetition-penalty sweeps, four independent speech recognisers and 420 analysis specifications without once reversing sign; a held-out fourth architecture lands within a point of its predicted gap, and one of two non-autoregressive baselines shows the same failure. Varying the period of the text shows the failure grows smoothly with periodicity, half of it surviving when no word is adjacent to itself.
|
| 2960 |
Louder, Longer, Livelier: Acoustic Shortcuts and Underspecified Rationales in Speech LLM Judges
2609.36979
|
cs.SDeess.AS
|
Mingyue Huo, Shivam Mehta, Bhavin Jawade, Yinghong Lan, Haoqi Li |
LLM-as-a-judge is widely used for evaluating text, but extending this paradigm to speech requires models to interpret acoustic as well as linguistic evidence. This introduces a modality-specific risk: a speech judge may treat a perceptually salient cue as evid...LLM-as-a-judge is widely used for evaluating text, but extending this paradigm to speech requires models to interpret acoustic as well as linguistic evidence. This introduces a modality-specific risk: a speech judge may treat a perceptually salient cue as evidence of quality even when that cue is irrelevant to the target criterion or receives more weight than human listeners give it. We call this behavior an acoustic shortcut. To study it, we audit six speech LLM judges using controlled manipulations of intensity, content richness, and emotional delivery. We evaluate both pointwise scoring and pairwise comparison, using human preference calibration to interpret the results. The judges consistently reward louder audio, prefer content-rich speech more strongly than human listeners do, and map emotional delivery into quality preferences. These effects are most visible in pairwise comparison, while pointwise scores often obscure them. More concerningly, the accompanying rationales rarely identify the acoustic cue that changes a judgment and instead repeatedly rely on a limited vocabulary, leaving them acoustically underspecified. Together, these findings show that reliable speech judges must both resist acoustic shortcuts and ground their rationales in the acoustic evidence behind their decisions. To support reproducibility and future audits, we also release SpeechJudgeAudit, the controlled stimuli and evaluation tools used in this study.
|
| 2961 |
Beyond Acoustic Prefixes: Persistent Access to Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition
2603.27205
|
cs.SD
|
Hao Shi, Yuan Gao, Xugang Lu, Tatsuya Kawahara |
Large language models (LLMs) provide strong linguistic priors for serialized output training (SOT), yet LLM-based multi-talker ASR degrades substantially as the number of overlapping talkers increases. Conventional systems expose acoustic evidence primarily th...Large language models (LLMs) provide strong linguistic priors for serialized output training (SOT), yet LLM-based multi-talker ASR degrades substantially as the number of overlapping talkers increases. Conventional systems expose acoustic evidence primarily through an initial projected mixture prefix, requiring the decoder to preserve and recover talker-relevant information indirectly throughout autoregressive generation. We first examine whether this limitation can be resolved by enriching the static prefix using discrete connectionist temporal classification (CTC) tokens, hybrid token--acoustic prompts, and continuous talker-specific representations. The results suggest that acoustic content alone does not fully address the conditioning bottleneck. We therefore extend onset-based serialization from the output target to the acoustic-conditioning pathway and introduce persistent decoder-side access to SOT-aligned serialized acoustic memory. Talker-specific representations are organized in utterance-onset order and retained as external acoustic memory, while the conventional mixture prefix provides a complementary global-conditioning path. LLM layers query this memory throughout generation through gated residual cross-attention. We further introduce a second adaptation stage that jointly applies low-rank updates to the acoustic-retrieval pathway and selected LLM self-attention projections. Experiments on LibriMix show consistent improvements over static-prefix prompting. These results indicate that effective LLM-based multi-talker ASR depends not only on providing richer acoustic representations, but also on maintaining persistent access to acoustic evidence structured according to the serialized output.
|
| 2962 |
AdaptDuplex: from static to adaptive full-duplex spoken dialogue
2609.29217
|
cs.SD
|
Zhiyang Zhou, Yingxin Shang, Zhou Wang, Hongwei Cai, Weixu Wang |
Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mecha...Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which extends Qwen3-Omni with such a mechanism, co-designed across three layers. A compact token-level protocol represents every window as a canonical sequence, trains dual-stream alignment through a bounded text lead over speech, and exposes every behavioral decision as an explicit token for training-free runtime control via logits bias. Adaptive mechanisms dynamically predict among discrete window durations and augment direct response as needed with non-blocking cognitive consolidation and multi-flight external reasoning. A progressive pipeline introduces these behaviors through a three-stage Thinker curriculum, then Talker-only and joint SFT, with GRPO as a preliminary increment. On Full-Duplex-Bench v1 and v1.5, AdaptDuplex outperforms DuplexOmni on 18 of 21 comparable turn-taking, overlap-behavior, and timing metrics and MiniCPM-o 4.5 on 19 of 22, with gains in both interaction decisions and response timing. On human-recorded HumDial-FDBench, it attains the top Final score (69.6) of the compared duplex models.
|
| 2963 |
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
2604.16659
|
cs.SD
|
Jaechul Roh, Virat Shejwalkar, Amir Houmansadr |
Fine-tuning on benign data is known to degrade safety alignment in text and vision LLMs, but whether distinct input properties drive this vulnerability differently remains unclear. Audio introduces a richer problem where benign samples can neighbor harmful con...Fine-tuning on benign data is known to degrade safety alignment in text and vision LLMs, but whether distinct input properties drive this vulnerability differently remains unclear. Audio introduces a richer problem where benign samples can neighbor harmful content through what is said or how it sounds. We present the first systematic study of benign fine-tuning safety in Audio LLMs, evaluating three state-of-the-art models with a proximity-based framework that decomposes embedding-space distance into semantic, acoustic, and mixed axes. We find that the dominant vulnerability axis is architecture-conditioned, determined by how each model's encoder and projector transform audio into the backbone LLM's input space. Across three models, benign fine-tuning elevates Jailbreak Success Rate (JSR) from single digits to as high as 87%, with the most damaging axis shifting from semantic to acoustic proximity depending on encoder design. Mechanistically, fine-tuning selectively suppresses late-layer refusal circuits while frozen encoders preserve upstream representations: the model still detects harmful content but stops refusing, a recognition-refusal dissociation. Two practical defenses, filtering training data to maximize distance from harmful embeddings and a textual system prompt at inference, reduce JSR to near-zero without architectural modification. These findings show that safety evaluation should account for modality and architecture, while highlighting Audio LLMs as a useful testbed for understanding alignment fragility.
|
| 2964 |
HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech
2606.28249
|
cs.SDeess.AS
|
Sihang Nie, Xiaofen Xing, Rui Xing, Haoming Li, Ruitong Xiao |
Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges to statistically averaged prosody, limiting emotional expressiveness. While prefe...Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges to statistically averaged prosody, limiting emotional expressiveness. While preference-driven optimization offers a promising alternative, existing approaches suffer from two structural mismatches: information conflict, where content and emotion in a shared latent space produce conflicting gradients, leading to reward hacking and semantic degradation; and scale gap, where sparse sentence-level rewards struggle to guide dense frame-level generation. To overcome these challenges, we propose HPRO, a hierarchical progressive reward optimization framework. Within HPRO, we introduce the HD-Emo codec as a novel differentiable reward model to mitigate the information conflict. It extracts speech into distinct content and style preference tokens, structurally isolating emotional optimization from semantic content. Building upon this structured preference space, HPRO bridges the scale gap by progressively aligning frame-, word- and sentence-level objectives. Experiments demonstrate that HPRO significantly enhances emotional expressiveness, while effectively preserving linguistic intelligibility. The code and audio samples are publicly available at https://xxh333.github.io/hpro-demo/.
|
| eess.AS 14 papers | ||||
| 2965 |
Model-Guided Design of Low-Context Speech Probes for Cochlear Synaptopathy
2609.36272
|
eess.AS
|
Ahsan J. Cheema, David Meng, Jorge Mejia, Sanna Hou, Sunil Puria |
Cochlear neural degeneration (CND) can impair suprathreshold coding without elevating pure-tone thresholds, complicating its diagnosis when it coexists with hair cell loss. We present a unified comparison of temporal and noise-based probes for CND detection us...Cochlear neural degeneration (CND) can impair suprathreshold coding without elevating pure-tone thresholds, complicating its diagnosis when it coexists with hair cell loss. We present a unified comparison of temporal and noise-based probes for CND detection using low-context vowel-consonant-vowel (VCV) syllables to reduce linguistic and contextual cues. Using a phenomenological auditory nerve model, we simulated responses to 21 VCV tokens under time compression, reverberation, and speech-in-noise conditions across presentation levels and seven CND profiles. We computed mutual information (MI) between inner hair cell potentials and auditory nerve neurograms and quantified information loss relative to a normal-hearing baseline. Time compression and amplitude-modulated (AM) noise produced the largest modeled information losses. We then evaluated these stimuli in a consonant-identification study involving 36 listeners with normal audiograms, 12 of whom reported difficulty understanding speech in noise. Neither 40 percent time compression in quiet nor AM noise alone distinguished listeners with and without these difficulties. However, compressed speech presented in AM noise separated the two groups. This partial agreement between model predictions and behavior supports our MI-based stimulus design framework and motivates further evaluation of combined temporal and noise-based probes for CND detection.
|
| 2966 |
Perception-Inspired Bayesian Causal Fusion for Audiovisual Source Localization
2609.36441
|
eess.AS
|
Kyung Yun Lee, Sungnyun Kim, Sebastian J. Schlecht, Tae-Hyun Oh, Vesa V\"alim\"aki |
Multimodal fusion promises more accurate perception but only when the modalities share a common cause. When they do not, the second modality carries no information about the target, and fusing it can only corrupt the estimate. We cast this whether-to-fuse deci...Multimodal fusion promises more accurate perception but only when the modalities share a common cause. When they do not, the second modality carries no information about the target, and fusing it can only corrupt the estimate. We cast this whether-to-fuse decision as Bayesian causal inference, following the optimal-observer model of human multisensory perception, and implement it as a plug-and-play layer on top of frozen audio and visual models for sound event localization and detection. The model infers a common-cause posterior over visible candidates, then gates precision-weighted fusion accordingly. Fusing unconditionally more than doubles the direction error, whereas the causal gate improves on-screen localization while limiting off-screen degradation, without any joint network retraining.
|
| 2967 |
Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT
2609.36754
|
eess.AS
|
Ki Woong Moon, Daniel Brenner |
Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechani...Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.
|
| 2968 |
WenetSpeech-Min: A Large-Scale Minnan Speech Corpus with Dual Transcriptions for Dialectal Speech Processing
2609.36834
|
eess.AS
|
Haoyu Zhang, Chunjiang He, Hongtao Li, Zeyu Zhu, Qituan Shangguan |
Progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora. For Minnan speech, existing resources remain limited, and few provide paired Minnan and Mandarin transcripts at scale. To address these gaps, we introduce W...Progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora. For Minnan speech, existing resources remain limited, and few provide paired Minnan and Mandarin transcripts at scale. To address these gaps, we introduce WenetSpeech-Min, an open-source corpus comprising around 10,000 hours of Minnan speech collected from diverse online media, with paired Minnan and Mandarin transcripts for every utterance. We further establish an automatic speech recognition (ASR) benchmark covering both Minnan and Mandarin transcripts and a text-to-speech synthesis (TTS) benchmark using Minnan transcripts, with manually verified evaluation sets for both tasks. To assess the effectiveness of the corpus, we train ASR and TTS models on WenetSpeech-Min and compare them with representative systems on the proposed benchmarks. The resulting models outperform the evaluated open-source models on most metrics and achieve competitive performance against commercial systems. We will release the corpus, benchmarks, and models to facilitate reproducible research on Minnan speech technology.
|
| 2969 |
SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding
2609.37601
|
eess.AS
|
Jisoo Park, Seonghak Lee, Hyojin Park, Junseok Kwon |
Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as acoustic reconstruction, optimi...Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as acoustic reconstruction, optimizing waveform fidelity while ignoring whether the generated speech preserves high-level semantic content. In this work, we revisit this formulation and argue that EEG signals carry not only acoustic but also semantic information. We identify two key limitations of prior methods: (1) the neglect of spatial relationships between EEG electrodes, and (2) the failure to exploit the semantic structure of the N400 paradigm, where congruent and incongruent trials reflect distinct semantic processing. We propose SENSE(Semantic-EEG Neural Speech SynthEsis), which combines a graph-based EEG encoder over electrode geometry with EEG Semantic Conditioning (ESC), aligning EEG to a pretrained semantic space using only congruent trials. On the N400 dataset, SENSE consistently outperforms prior methods on both acoustic and semantic metrics, and model-internal channel attribution suggests distributed reliance on auditory, sensorimotor, and centro-parietal regions, consistent with known speech-perception neuroscience. In the unseen-subject setting, SENSE trained on only two subjects already surpasses the strongest baseline trained on all eighteen subjects in word error rate, and matches it on acoustic metrics with as few as eight subjects.
|
| 2970 |
Selective Lookahead for Attention-Based Streaming ASR
2609.37611
|
eess.AS
|
Yichen Jia, Bastiaan Tamm, Hugo Van hamme |
End-to-end attention-based speech recognition is accurate offline but hard to stream: outputs can depend on future audio, and a little future context per layer makes the lookahead grow with the number of layers. We address this with two mechanisms. A bounded-l...End-to-end attention-based speech recognition is accurate offline but hard to stream: outputs can depend on future audio, and a little future context per layer makes the lookahead grow with the number of layers. We address this with two mechanisms. A bounded-lookahead chunk encoder caps every chunk's future receptive field at a constant number of chunks, independent of the number of layers, via one age-selection rule shared by self-attention and the depthwise convolution. On this encoder, dynamic future-chunk decoding lets a per-token trigger commit a token or wait and re-decode it; we propose a learned trigger as the general mechanism, with a simple confidence threshold as an effective fallback. On full LibriSpeech test-clean the dynamic system matches the best static-lookahead accuracy (6.5%) at a median latency of 306 ms versus 860 ms for one-chunk static lookahead, and a wait budget bounds the deferral tail below the static baseline's 90th percentile at 0.1 points more WER.
|
| 2971 |
Signal-Independent and Signal-Dependent Neural Ambisonic Matrix Encoding for Arbitrary Arrays with Variable Microphone Counts
2609.37691
|
eess.AS
|
Shichao Hu, Zhiheng Jin, Chunyang Xu, Mengyao Zhu |
Recent neural Ambisonic encoders accommodate diverse array geometries, yet many existing neural encoders require a fixed microphone count because the number of microphone channels is embedded in the network architecture. This requirement limits deployment acro...Recent neural Ambisonic encoders accommodate diverse array geometries, yet many existing neural encoders require a fixed microphone count because the number of microphone channels is embedded in the network architecture. This requirement limits deployment across devices with different microphone configurations and adaptation to changes in available channels. To address this limitation, we investigate Transformer-based matrix encoding for arbitrary microphone arrays with variable microphone counts. This is achieved through shared microphone-wise processing and masked self-attention that models inter-microphone relationships across variable-size arrays. Within this framework, we consider signal-independent (SI) encoding, which predicts encoding matrices from array transfer functions, and introduce a signal-dependent (SD) extension that additionally incorporates the observed microphone signals. Both models are trained on simulated scenes using LibriSpeech sources and extensively evaluated under changes in source type, unseen microphone counts, and increased source counts beyond those used during training. Both SI and SD outperform conventional least-squares (LS) encoding in aggregate reconstruction performance across the evaluated conditions. SD consistently achieves stronger overall performance than SI. These results demonstrate that the proposed framework enables array-agnostic Ambisonic encoding while retaining generalization across microphone counts and acoustic source conditions.
|
| 2972 |
Acoustic Honeybee Queen-State Detection Under Unseen Conditions
2609.37845
|
eess.AS
|
Mahsa Abdollahi, Nico Coallier, Maxime Fraser Franco, Tiago H. Falk |
Honeybee queen loss is a major threat to colony health, yet queen-status assessment remains largely manual and disruptive. Acoustic monitoring offers a non-invasive alternative by enabling continuous analysis of hive sounds. In this paper, we benchmark convent...Honeybee queen loss is a major threat to colony health, yet queen-status assessment remains largely manual and disruptive. Acoustic monitoring offers a non-invasive alternative by enabling continuous analysis of hive sounds. In this paper, we benchmark conventional and learned acoustic representations for automated detection of queen absence, comparing task-specific convolutional neural networks with pretrained audio transformers. Experiments are performed on 5,129 audio recordings from 3,285 hives across 47 apiaries collected between October 2024 and August 2026, evaluated using hive- and apiary-independent splits. Results show that models relying on modulation spectrograms achieve the best performance, reaching an Area Under the Receiver Operating Characteristic (AUROC) of 0.81 and Area Under the Precision-Recall Curve (AUPRC) of 0.37 on unseen hives; performance decreases under unseen-apiary evaluation. Overall, our results highlight the promise of modulation-based audio representations for non-invasive queen-status monitoring and highlight the challenge in cross-apiary model generalization.
|
| 2973 |
QK-GCC: Learnable Query-Key Spectral Matching for Robust Time Delay Estimation
2609.38000
|
eess.AS
|
Jinkai Zhang, Weiye Chen, Yue Huang, Xiaotong Tu, Xinghao Ding |
Time delay estimation (TDE) is a fundamental component of microphone-array sound source localization. Generalized cross-correlation (GCC) is widely used because it is efficient and interpretable, but its handcrafted spectral matching and predefined frequency w...Time delay estimation (TDE) is a fundamental component of microphone-array sound source localization. Generalized cross-correlation (GCC) is widely used because it is efficient and interpretable, but its handcrafted spectral matching and predefined frequency weighting are vulnerable to noise and reverberation. Existing neural GCC variants mainly improve robustness by enhancing input signals or modeling GCC responses, while the cross-channel spectral matching step itself remains handcrafted. We propose QK-GCC, a learnable GCC-like framework that replaces handcrafted weighted spectral matching in GCC with Query-Key matching between two microphone signals. The two microphone signals are encoded as magnitude-phase frequency tokens and mapped to Query and Key representations, respectively, enabling frequency reliability learning and local spectral evidence aggregation for delay estimation. Experiments in simulated reverberant rooms across diverse SNR and reverberation conditions show that QK-GCC improves TDE accuracy over GCC-PHAT and learning-based GCC variants, while remaining lightweight and generalizing to unseen source types. The code is available at https://github.com/zhangjinkai33-ui/QK-GCC.
|
| 2974 |
FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech
2609.35791
|
eess.AS
|
Puneet Mathur, Dinesh Manocha |
Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based ...Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set $0.853$ (at FP<=0.10) in a zero-shot setting. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.
|
| 2975 |
Zephyr: An Efficient Audio Denoising System Using Spiking Neural Networks Enabled With A Sparsity-Aware Flexible FPGA PE Array
2609.37711
|
eess.AS
|
Cheng-En Chang, Chi-Wei Kao, Chung-Lun Yang, Yan-Lin Jiang, Yi-Chen Huang |
In this work we look to neuromorphic computing to solve the power consumption problem that audio denoising neural networks face on edge devices like smartphones, wireless headphones and hearing aids. Spiking neural networks (SNNs) have the potential to solve t...In this work we look to neuromorphic computing to solve the power consumption problem that audio denoising neural networks face on edge devices like smartphones, wireless headphones and hearing aids. Spiking neural networks (SNNs) have the potential to solve this problem due to their high activation sparsity and low complexity, however many SOTA SNNs require hardware that supports a mixture of operations to be able to fully perform inference. To solve this problem, we convert SOTA audio denoising neural network Spiking-FullSubNet to a hardware friendly version showing that via QAT and activation function simplification we can achieve $\approx28\times$ improvement in power consumption to 52.9nJ per 32ms audio frame when calculated for custom digital hardware in a 45nm process node. We then propose a digital circuit which by means of a sparsity-aware flexible PE array can perform inference of the heterogeneous compute load of Spiking-FullSubNet, and validate this circuit on a PYNQ-Z1 FPGA achieving a real-time factor of 0.727 at 100MHz.
|
| 2976 |
Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios
2603.08249
|
eess.AS
|
Pol Buitrago, Javier Hernando |
Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training. Syntheti...Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training. Synthetic visual data have been shown to be an effective augmentation strategy for addressing AV data scarcity. However, a more challenging scenario arises for languages such as Catalan, where no real audiovisual data are available for training. In this study, we investigate whether AVSR can be bootstrapped in such a zero-AV-resource setting, using synthetic visual data as the sole source of visual supervision. We synthesize over 700 hours of talking-head video and fine-tune a pre-trained AV-HuBERT model. On a manually annotated Catalan benchmark, our model achieves near state-of-the-art (SOTA) performance with much fewer parameters and training data than SOTA ASR systems such as Whisper-large-v3, outperforms an identically trained audio-only baseline, and preserves multimodal advantages under acoustic degradation. Scalable synthetic video thus offers a viable substitute for real recordings in zero-AV-resource AVSR.
|
| 2977 |
Teacher-Free Self-Distilled Consistency Trajectory Learning for Fast Speech Enhancement
2609.10392
|
eess.AS
|
Shuubham Ojha, Carol Espy-Wilson |
Consistency trajectory models offer a route to fast, high- quality speech enhancement, collapsing the many reverse steps of diffusion-based enhancers into a handful. When instantiated on a Schr\"odinger bridge (SB), which pins the generative process to fixed c...Consistency trajectory models offer a route to fast, high- quality speech enhancement, collapsing the many reverse steps of diffusion-based enhancers into a handful. When instantiated on a Schr\"odinger bridge (SB), which pins the generative process to fixed clean and noisy endpoints, exist- ing consistency-trajectory enhancers (SBCTMs) still require a pretrained teacher to supply trajectory supervision, which raises training cost and ties the final quality to that of the teacher. We propose a teacher-free, self-distilled consistency- trajectory framework that removes the external teacher result- ing in a 5X reduction in per epoch training time. Our model is trained with a three-stage curriculum of clean speech pre- diction, a self-distilled shortcut objective, and perceptual fine-tuning with a multi-resolution short-time Fourier trans- form (MR-STFT) loss. Using the same NCSN++ backbone as SBCTM, our model attains a wide-band PESQ of 3.01, ES- TOI 0.87, and SI-SDR 19.07 dB on VoiceBank+DEMAND compared to 3.57, 0.87 and 12.8 dB for the teacher based model. Further, we find that a geometric schedule at low reverse step count maximizes perceptual quality, while a higher-step uniform schedule favors signal fidelity.
|
| 2978 |
Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
2609.33645
|
eess.AS
|
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang |
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LL...Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.
|