← Back to Index
Daily Research Digest

arXiv Papers

2026-07-28
499
Papers
4
Categories
499
Translated
收藏清单 0
机器人学 (Robotics)
48
cs.RO / 1 / 2607.22858

A Replay-Constrained Simulation Framework for Personalization of Powered Knee--Ankle Prosthesis Controllers

一种用于个性化动力膝踝假肢控制器的重放约束仿真框架
Le, Duong, Posh, Ryan, Cheng, Shihao, Ghaffari, Maani, Gregg, Robert D.
Abstract
Personalization of impedance controllers for powered prosthetic legs is critical to accommodating individual gait biomechanics but remains challenging. Existing methods rely on time-intensive human-in-the-loop exploration and/or constrain optimization to low-dimensional, single-joint parameter subspaces. Sim-to-real transfer has enabled high-dimensional locomotion control for legged robots, but in assistive device control the human partner remains un-modelable. We present a replay-constrained simulation framework: a MuJoCo-based simulator reproduces prosthetic knee-ankle dynamics while replaying recorded hip kinematics and feedback-based ground reaction forces from individual walking data, bypassing the need to model complex human neuromuscular control mechanisms. We demonstrate the framework with a deep reinforcement learning policy that personalizes phase-dependent stiffness, damping, and equilibrium angle at both joints simultaneously, maximizing a biomimicry-based reward computed solely from onboard prosthesis measurements. Experiments with three participants with transfemoral amputation during level-ground walking at 0.8~m/s demonstrate strong simulation-to-hardware predictive validity (Pearson $r=0.96$--$0.997$). The best-performing policy on hardware was consistently predicted within the top five simulation policies for all participants. The learned controllers improved overall biomimicry rewards by 42--59\% relative to the unpersonalized baseline. The framework supports scalable high-dimensional personalization of powered prosthetic legs and is amenable to extension to higher-dimensional controller parameterizations such as neural-network controllers.
Chinese Translation
动力假肢腿的阻抗控制器个性化对于适应个体步态生物力学至关重要,但仍然面临挑战。现有方法依赖于耗时的人机交互探索和/或将优化限制在低维度的单关节参数子空间。模拟到现实的转移使得腿部机器人能够进行高维度的运动控制,但在辅助设备控制中,人类合作伙伴仍然无法建模。我们提出了一种重放约束的仿真框架:基于MuJoCo的仿真器再现了假肢膝踝的动力学,同时重放来自个体步态数据的记录的髋关节运动学和基于反馈的地面反作用力,绕过了建模复杂人类神经肌肉控制机制的需求。我们通过一种深度强化学习策略展示了该框架,该策略同时个性化两个关节的相位依赖性刚度、阻尼和静力平衡角,最大化仅基于假肢内部测量计算的仿生奖励。在以0.8 m/s的速度进行平地行走的三名股骨上截肢参与者的实验中,显示出强大的仿真到硬件的预测有效性(Pearson $r=0.96$--$0.997$)。在所有参与者中,硬件上表现最佳的策略在仿真中始终被预测为前五名策略。与未个性化基线相比,学习到的控制器将整体仿生奖励提高了42%至59%。该框架支持动力假肢腿的可扩展高维个性化,并且适合扩展到更高维度的控制器参数化,例如神经网络控制器。
cs.RO / 2 / 2607.22964

Pose-Aware Modeling to Mitigate Pose-Related Artifacts in Tactile Gloves

姿态感知建模以减轻触觉手套中的姿态相关伪影
Yu, Tianhong Catherine, Kou, Ziyi, Huang, Mia, Niehues, Taylor, Luo, Yiyue, Guan, Li, Zhang, Dingtian
Abstract
Tactile gloves digitize contact and force during hand-object interactions, enabling robotics applications in dexterous manipulation, teleoperation, and learning from demonstration. To preserve hand dexterity and capture the nuances of natural interactions, these gloves and the integrated tactile sensors are designed to be soft, flexible, and comfortable. However, such flexible sensors are sensitive not only to contact forces but also unavoidably to hand pose changes, resulting in pose-related artifacts (PRAs). PRAs are especially problematic in the low-force range, resulting in misdetections or late-onset detections of contact, which raises the minimum detectable force (MDF) of the glove. In this work, we characterize the PRAs in relation to pose and force. Building on these insights, we introduce a glove-agnostic algorithmic framework that leverages hand pose information, which is increasingly available, to mitigate PRAs without glove modifications. Our pose-aware force estimation model augments tactile-to-force pipelines with a residual prediction branch that explicitly accounts for pose-induced sensor deformations. We validate our approach across 3 glove designs and 15 users, reducing MDF by 10.4%, 12.2%, and 18.3%, with consistent improvements across all evaluated metrics. This method provides a practical path to improving the usability of tactile gloves in data collection and diverse robotic applications.
Chinese Translation
触觉手套在手与物体交互过程中数字化接触和力,使得在灵巧操作、远程操作和示范学习等机器人应用中得以实现。为了保持手的灵巧性并捕捉自然交互的细微差别,这些手套及其集成的触觉传感器被设计为柔软、灵活和舒适。然而,这种灵活的传感器不仅对接触力敏感,而且不可避免地对手的姿态变化敏感,从而导致姿态相关伪影(PRAs)。在低力范围内,PRAs尤其成问题,导致接触的误检测或延迟检测,从而提高了手套的最小可检测力(MDF)。在本研究中,我们将PRAs与姿态和力的关系进行了表征。在此基础上,我们提出了一种与手套无关的算法框架,利用越来越多可用的手姿态信息来减轻PRAs,而无需对手套进行修改。我们的姿态感知力估计模型通过一个残差预测分支增强了触觉到力的管道,明确考虑了姿态引起的传感器变形。我们在3种手套设计和15名用户中验证了我们的方法,MDF分别降低了10.4%、12.2%和18.3%,在所有评估指标上均表现出一致的改善。这种方法为提高触觉手套在数据收集和多样化机器人应用中的可用性提供了一条实用路径。
cs.RO / 3 / 2607.22997

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

面向视觉-语言-动作操作的Real2Sim2Real:基于AMD ROCm的管道
Yang, Qing, Wang, Xun, Wang, Ziguan, Li, Zhenjiang, Wang, Hongqiang, Weng, Dongdong
Abstract
Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed by industry leaders such as Jensen Huang (``the next big thing is Physical AI, AI with a body,'' GTC Paris, June 2025) and Dr. Lisa Su (`we're entering the world of Physical AI ... this is where AI enters the real world,' CES 2026). This paper presents an end-to-end, fully AMD-accelerated technology stack for embodied manipulation, spanning data-center training silicon, Radeon PRO simulation/rendering GPUs, and Ryzen AI edge compute, unified by the open ROCm software stack. We demonstrate that training and deploying VLA-based manipulation policies does not require a CUDA-locked ecosystem. Four progressive demonstrations are presented: (1) a Sim-to-Real manipulation pipeline trained with SmolVLA and deployed on a physical Franka arm; (2) a semantic, language-grounded object-selection task (`one-of-three'); (3) a Real2Sim synthetic-data generation pipeline that fuses 3D Gaussian Splatting (3DGS) reconstructions of real scenes with the Genesis physics engine; and (4) large-scale reinforcement learning for quadruped and humanoid locomotion benchmarked across multiple hardware platforms. All pipelines run natively on ROCm + PyTorch on RDNA4 (Radeon AI PRO R9700) and RDNA3.5 (Radeon PRO W7900) hardware and are reproducible on the free Radeon Cloud Platform.
Chinese Translation
物理人工智能——将大型视觉-语言-动作(VLA)模型与在现实世界中行动的具身代理相结合——已成为人工智能的下一个主要前沿,这一观点得到了行业领袖的认可,如Jensen Huang(“下一个重大突破是物理人工智能,即有身体的人工智能,”GTC巴黎,2025年6月)和Dr. Lisa Su(“我们正在进入物理人工智能的世界……这就是人工智能进入现实世界的地方,”CES 2026)。本文提出了一种端到端、完全基于AMD加速的技术栈,用于具身操作,涵盖数据中心训练硅片、Radeon PRO模拟/渲染GPU和Ryzen AI边缘计算,统一于开放的ROCm软件栈。我们展示了基于VLA的操作策略的训练和部署并不需要一个被CUDA锁定的生态系统。我们展示了四个渐进的示例:(1)一个使用SmolVLA训练并部署在物理Franka臂上的Sim-to-Real操作管道;(2)一个语义、语言基础的物体选择任务(“三选一”);(3)一个Real2Sim合成数据生成管道,将真实场景的3D高斯点云重建与Genesis物理引擎融合;以及(4)在多个硬件平台上进行的大规模强化学习,用于四足和类人运动的基准测试。所有管道均在RDNA4(Radeon AI PRO R9700)和RDNA3.5(Radeon PRO W7900)硬件上原生运行于ROCm + PyTorch,并可在免费的Radeon Cloud Platform上复现。
cs.RO / 4 / 2607.22999

WCM: World-Cognition Model for Generalizable Human-Robot Interaction

WCM:通用人机交互的世界认知模型
Chen, Yuzhen, Zhou, KC
Abstract
Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks. Current robot-control paradigms, including vision-language-action policies and world-model-based planners, are mainly optimized for instruction execution, leaving users with little visibility into why an action is chosen and few mechanisms to redirect, correct, or teach the robot through interaction. To solve this problem, we present the World-Cognition Model (WCM), a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime. SLAK separates perception, reasoning, control, and memory, while the runtime allows reasoning, dialogue, and execution to proceed concurrently. WCM further introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks. Teaching episodes and autonomous task rollouts are refined into chain-of-thought supervision to continually improve the model. WCM achieves a 73.8% average success rate across nine real-world human-robot interaction tasks, including tasks held out from CoT fine-tuning and a long-horizon task learned through teaching.
Chinese Translation
语言代理现在可以与软件中的用户流畅互动,但机器人在物理任务中仍然难以实现类似的交互。目前的机器人控制范式,包括视觉-语言-动作策略和基于世界模型的规划器,主要优化用于指令执行,导致用户对为何选择某个动作几乎没有可见性,也缺乏通过交互重定向、纠正或教导机器人的机制。为了解决这个问题,我们提出了世界认知模型(World-Cognition Model, WCM),这是一个基于SLAK架构(感知、逻辑、动作和知识)和异步运行时的人本化具身代理。SLAK将感知、推理、控制和记忆分开,而运行时允许推理、对话和执行并行进行。WCM进一步引入了一种人机互动教学模式,使用户能够互动地教导机器人处理困难或长期任务。教学过程和自主任务展开被精炼为思维链监督,以持续改进模型。WCM在九个真实世界的人机交互任务中实现了73.8%的平均成功率,包括从CoT微调中保留的任务和通过教学学习的长期任务。
cs.RO / 5 / 2607.23108

The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation

精度的诅咒:高精度机器人操作的数据缩放法则
Xu, Cuijie, Xu, Yuanfan, Xue, Min, Lin, Jianjie, Wang, Jian, Zhang, Xudong, Wang, Yu, Yu, Jincheng
Abstract
While scaling laws for imitation learning have primarily focused on generalization in open-world settings, the relationship between data and precision in closed-world tasks like robotic assembly remains largely unexplored. This paper systematically investigates this relationship and introduces a novel scaling law. We find that to achieve a fixed success rate, the required number of demonstrations $N$ grows super-exponentially as the target precision $P$ approaches a limit $c$. This relationship is accurately captured by the model $\log(N) \propto 1/(P-c)$. Crucially, we reveal that the limit precision $c$ is not a static physical constant of the task but an emergent property of the entire agent system, including its sensors and expert policy. Through experiments on canonical manipulation tasks, we validate this law and demonstrate that improving system components, such as adding a wrist camera or using a more effective expert, measurably lowers $c$, thus expanding the system's achievable precision. Our work provides a new theoretical framework for precision in robotics and a quantitative metric to evaluate system capabilities. Furthermore, these findings provide a practical methodology for guiding the development and debugging of high-precision manipulation systems.
Chinese Translation
尽管模仿学习的缩放法则主要集中在开放世界环境中的泛化,但在机器人组装等封闭世界任务中,数据与精度之间的关系仍然 largely 未被探索。本文系统地研究了这一关系,并引入了一种新颖的缩放法则。我们发现,为了实现固定的成功率,所需的演示次数 $N$ 随着目标精度 $P$ 接近极限 $c$ 而超指数级增长。这个关系可以通过模型 $ ext{log}(N) ext{propto} rac{1}{(P-c)}$ 准确捕捉。重要的是,我们揭示了极限精度 $c$ 并不是任务的静态物理常数,而是整个代理系统的一个涌现属性,包括其传感器和专家策略。通过对典型操作任务的实验,我们验证了这一法则,并展示了改善系统组件(如添加腕部摄像头或使用更有效的专家)可以显著降低 $c$,从而扩展系统可实现的精度。我们的工作为机器人精度提供了新的理论框架和评估系统能力的定量指标。此外,这些发现为指导高精度操作系统的开发和调试提供了一种实用的方法论。
cs.RO / 6 / 2607.23204

Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot

通过上下文感知的前言生成实现低延迟的轮流对话在现实世界对话机器人中的应用
Okafuji, Yuki, Inoue, Koji, Ohira, Yoshiki
Abstract
Large language model (LLM)-based dialogue systems suffer response delays because generation begins only after final speech recognition. While fixed fillers are a workaround, they become unnatural over time. We propose a two-stage incremental framework that decouples prefatory-response preparation from speech onset. Once user intent becomes predictable, an intent readiness detector triggers LLM-based generation of a short prefatory response. Concurrently, a voice activity projection (VAP) model determines when to deliver it. Through a field experiment with a route-guidance robot in a shopping mall, we evaluated three conditions: no-filler, fixed-filler, and contextual-preface. Both fixed-filler and contextual-preface significantly reduced initial response latency relative to no-filler. Relative to fixed-filler, contextual-preface had significantly longer initial response latency but a significantly shorter initial-to-main gap. Exploratory ratings showed no significant differences. These results indicate a timing trade-off.
Chinese Translation
基于大型语言模型(LLM)的对话系统由于响应延迟而受到影响,因为生成过程仅在最终语音识别完成后才开始。虽然固定填充词是一种解决方法,但随着时间的推移,它们变得不自然。我们提出了一种两阶段增量框架,将前言响应的准备与语音开始解耦。当用户意图变得可预测时,意图准备检测器触发基于LLM的短前言响应生成。同时,语音活动预测(VAP)模型确定何时传递该响应。通过在购物中心与路线引导机器人进行的实地实验,我们评估了三种条件:无填充词、固定填充词和上下文前言。相较于无填充词,固定填充词和上下文前言显著减少了初始响应延迟。相较于固定填充词,上下文前言的初始响应延迟显著更长,但初始到主要响应的间隔显著更短。探索性评分显示没有显著差异。这些结果表明存在时间上的权衡。
cs.RO / 7 / 2607.23268

Sling2Sim2Real: One-Shot Elastic System Identification for Non-Destructive Slingshot Policy Learning

Sling2Sim2Real:用于非破坏性弹弓策略学习的一次性弹性系统识别
Kang, Wonjae, Kim, Geonwoo, Song, Minseok, Park, Daehyung
Abstract
Elastic object manipulation (EOM) involves highdimensional, nonlinear, and elastic deformations. The diverse deformation properties of elastic objects substantially expand the relevant state space, requiring extensive exploration to learn accurate manipulation policies for tasks such as slingshot manipulation. While simulation enables large-scale and safe exploration compared to costly and potentially destructive real-world trials (e.g., repeated projectile launches), accurately calibrating elastic behavior between the real world and simulation remains challenging since elastic properties are largely indistinguishable from visual observations alone. To address these challenges, we propose Sling2Sim2Real, a one-shot Real2Sim2Real framework that identifies elastic parameters from a single non-destructive interaction and enables policy learning in simulation. The framework consists of two stages: 1) a multi-start Real2Sim system identification method that exploits parameter covariance to estimate elastic properties, and 2) simulation-based policy learning followed by zero-shot Sim2Real transfer using the calibrated simulator. We evaluate Sling2Sim2Real on a slingshot manipulation task using a Franka Emika Panda arm and elastic bands with diverse physical properties across varying target distances. Experimental results demonstrate that Sling2Sim2Real achieves accurate policy learning and robust generalization while significantly reducing the amount of required real-world interaction.
Chinese Translation
弹性物体操控(EOM)涉及高维度、非线性和弹性变形。弹性物体的多样化变形特性显著扩展了相关状态空间,要求进行广泛探索以学习准确的操控策略,例如弹弓操控。虽然与昂贵且可能具有破坏性的现实世界试验(例如,重复发射弹丸)相比,模拟能够进行大规模且安全的探索,但准确校准现实世界与模拟之间的弹性行为仍然具有挑战性,因为仅凭视觉观察很难区分弹性特性。为了解决这些挑战,我们提出了Sling2Sim2Real,一个一次性Real2Sim2Real框架,该框架通过单次非破坏性交互识别弹性参数,并在模拟中实现策略学习。该框架由两个阶段组成:1)一种多起始的Real2Sim系统识别方法,利用参数协方差来估计弹性特性;2)基于模拟的策略学习,随后使用校准的模拟器进行零-shot Sim2Real转移。我们在使用Franka Emika Panda臂和具有不同物理特性的弹性带进行的弹弓操控任务上评估了Sling2Sim2Real。实验结果表明,Sling2Sim2Real实现了准确的策略学习和稳健的泛化,同时显著减少了所需的现实世界交互量。
cs.RO / 8 / 2607.23384

Semantic Semi-Incremental Data-Association-Free Object SLAM

语义半增量无数据关联对象SLAM
Zhang, Yihao, Hong, Jungseok, Leonard, John J.
Abstract
Data association between landmark measurements and landmark variables has long been a central challenge in SLAM, as estimation accuracy depends critically on associating measurements with the correct landmark variables. Recent advances in deep learning have created new opportunities for the problem; data association can now leverage not only positional measurements but also semantic information about object landmarks, such as class labels from neural object detectors and feature vectors from visual foundation models. In this paper, we present a generalized data-association-free SLAM framework that jointly estimates data associations, robot poses, landmark positions, and landmark semantics from odometry, and positional and semantic measurements of landmarks. The proposed framework (i) creates a synergy between data association and landmark semantics estimation; (ii) adopts a semi-incremental estimation scheme for improved accuracy and computational efficiency; and (iii) provides a principled justification, guidelines, and heuristics for landmark-number estimation, improving the interpretability and practical usability of the framework. The proposed framework and algorithms are evaluated on synthetic and real-world datasets with two types of semantic information, class labels and real-valued feature vectors, and demonstrate superior performance compared to strong baselines.
Chinese Translation
地标测量与地标变量之间的数据关联长期以来一直是SLAM中的一个核心挑战,因为估计精度在很大程度上依赖于将测量与正确的地标变量关联起来。最近在深度学习方面的进展为这一问题创造了新的机会;数据关联现在不仅可以利用位置测量,还可以利用关于对象地标的语义信息,例如来自神经对象检测器的类别标签和来自视觉基础模型的特征向量。在本文中,我们提出了一种广义的无数据关联SLAM框架,该框架从里程计以及地标的位置信息和语义测量中联合估计数据关联、机器人姿态、地标位置和地标语义。所提出的框架 (i) 在数据关联和地标语义估计之间创造了协同效应;(ii) 采用半增量估计方案以提高准确性和计算效率;(iii) 为地标数量估计提供了原则性依据、指导方针和启发式方法,提高了框架的可解释性和实际可用性。所提出的框架和算法在合成和真实世界数据集上进行了评估,使用了两种类型的语义信息,即类别标签和实值特征向量,并且与强基线相比展示了优越的性能。
cs.RO / 9 / 2607.23473

PRISM: Polynomial Representations for Interaction-Structured Motor Control

PRISM:交互结构化运动控制的多项式表示
Lee, Seung Hyun, Yu, Stella X.
Abstract
Robot policies are typically MLPs mapping observations to actions. Yet robot observations are physical variables, and many action-relevant cues arise not from individual variables but from their interactions; power, inertial effects, contact, slip, and compliance depend on products among observable signals. We introduce PRISM, a policy representation that makes polynomial interactions among observable physical variables explicit, learnable, and compact. Rather than listing all polynomial terms, PRISM uses a factorized polynomial module to expose higher-order interaction features efficiently. In reinforcement learning, it keeps the standard MLP backbone but applies a gradually activated element-wise polynomial function after it. In imitation learning, it replaces linear proprioceptive conditioning in Diffusion Policy with a polynomial layer trained end-to-end. Across humanoid locomotion and contact-rich manipulation, PRISM improves performance over standard MLP policies and larger MLPs with matched capacity, showing that interaction structure cannot be replaced by capacity alone. It also yields sensorless compliant behavior without force, wrench, tactile input, contact labels, or admittance control. These results suggest that polynomial representations should become a standard architectural choice for embodied motor control. The project page is available at https://lsh3163.github.io/prism/
Chinese Translation
机器人策略通常是将观察映射到动作的多层感知器(MLP)。然而,机器人观察的是物理变量,许多与动作相关的线索并非来自单个变量,而是来自它们之间的交互;功率、惯性效应、接触、滑移和顺应性依赖于可观察信号之间的乘积。我们提出了PRISM,这是一种政策表示,使可观察物理变量之间的多项式交互变得明确、可学习且紧凑。PRISM并不是列出所有多项式项,而是使用一个因式分解的多项式模块高效地暴露高阶交互特征。在强化学习中,它保持标准的MLP骨干,但在其后应用逐渐激活的逐元素多项式函数。在模仿学习中,它用一个端到端训练的多项式层替代了扩散策略中的线性本体感知条件。在类人步态和接触丰富的操作中,PRISM的表现优于标准的MLP策略和具有匹配能力的大型MLP,表明交互结构不能仅通过能力来替代。它还在没有力、扭矩、触觉输入、接触标签或导纳控制的情况下实现了无传感器的顺应行为。这些结果表明,多项式表示应成为具身运动控制的标准架构选择。项目页面可访问 https://lsh3163.github.io/prism/
cs.RO / 10 / 2607.23515

LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks

LEACL:用于长时间操作任务的强化学习的LLM增强自动课程学习
Heravi, Faraz, Ouyang, James, Xu, Zifan, Kumar, Arjun, Sung, Yoonchang, Stone, Peter
Abstract
Long-horizon manipulation tasks pose significant challenges for reinforcement learning due to sparse reward signals and long horizons. Automatic curriculum learning (ACL) has been proposed to tackle these challenges by progressively training agents on a sequence of tasks, from easier to more difficult. However, the success of ACL depends heavily on task-dependent specifications-such as well-defined task parameter spaces and difficulty measures-which are often manually crafted and difficult to generalize across diverse tasks. Recent advances in large language models (LLMs) offer a promising alternative by enabling the decomposition of complex tasks into meaningful subtasks using the LLMs' web-scale common-sense knowledge. This decomposition can provide a natural curriculum structure for efficient learning of long-horizon tasks. However, existing LLM-based methods typically rely on hand-designed dense reward functions to learn each subtask, which can introduce bias and still requires significant human supervision. In this work, we propose LLM-enhanced automatic curriculum learning (LEACL), a framework that integrates LLMs and ACL to address these limitations. Specifically, LLMs are used to both decompose tasks into subtasks and to generate task-dependent specifications for each subtask. These specifications are then used by ACL algorithms to guide learning using only sparse reward signals, eliminating the need for dense reward design. We evaluate LEACL on five long-horizon manipulation tasks from the LIBERO benchmark. LEACL achieves better asymptotic performance in terms of the success rates compared to human-designed dense rewards.
Chinese Translation
长时间操作任务由于稀疏的奖励信号和较长的时间跨度,对强化学习提出了重大挑战。自动课程学习(ACL)被提出以应对这些挑战,通过逐步训练代理在一系列任务上,从简单到困难。然而,ACL的成功在很大程度上依赖于任务依赖的规范——例如明确的任务参数空间和难度度量——这些规范通常是手动设计的,且难以在不同任务之间进行推广。最近在大型语言模型(LLMs)方面的进展提供了一种有前景的替代方案,通过利用LLMs的网络规模常识知识,将复杂任务分解为有意义的子任务。这种分解可以为高效学习长时间任务提供自然的课程结构。然而,现有基于LLM的方法通常依赖于手工设计的稠密奖励函数来学习每个子任务,这可能引入偏见,并且仍然需要大量的人类监督。在本研究中,我们提出了LLM增强的自动课程学习(LEACL),这是一个整合LLMs和ACL以解决这些局限性的框架。具体而言,LLMs被用于将任务分解为子任务,并为每个子任务生成任务依赖的规范。这些规范随后被ACL算法用于指导学习,仅使用稀疏奖励信号,从而消除了对稠密奖励设计的需求。我们在LIBERO基准的五个长时间操作任务上评估了LEACL。与人类设计的稠密奖励相比,LEACL在成功率方面实现了更好的渐近性能。
cs.RO / 11 / 2607.23565

Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter

基于预期风险引导的强化学习在动态杂乱环境中的安全飞行
Mei, Yuchao, Zhang, Guohao, Ai, Luxia, Chen, Haopeng, Tao, Wenbing
Abstract
Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided reinforcement learning framework. Leveraging privileged simulator states, we construct a directionally aligned future collision risk map based on the Closest Point of Approach (CPA). Through an asymmetric actor-critic architecture, the network is trained to self-predict this structured risk, which explicitly guides the visual policy during deployment. A lightweight spatio-temporal encoder extracts motion cues directly from onboard depth sequences, bypassing explicit object tracking or optical flow estimation. Extensive simulated and real-world experiments demonstrate that our method effectively improves safety margins and flight efficiency in dense dynamic clutters compared to existing baselines. Furthermore, the learned policy achieves robust zero-shot Sim-to-Real transfer on a physical quadrotor, relying purely on abstracted spatio-temporal depth sequences and its self-predicted risk priors, validating the effectiveness of our approach and its robust generalization from simulation to reality.
Chinese Translation
安全的四旋翼导航在杂乱和动态环境中不仅依赖于瞬时几何感知,更关键的是要预测由相对运动引起的碰撞风险。传统的模块化流程常常受到感知延迟的困扰,而依赖于隐式标量奖励的端到端学习方法在没有物理基础监督的情况下,往往难以提取可靠的时空特征。为了解决这个问题,我们提出了一种基于预期风险引导的强化学习框架。利用特权模拟器状态,我们构建了一个基于最近接触点(Closest Point of Approach, CPA)的方向对齐未来碰撞风险图。通过不对称的演员-评论家架构,网络被训练以自我预测这一结构化风险,从而在部署过程中明确引导视觉策略。一个轻量级的时空编码器直接从机载深度序列中提取运动线索,绕过了显式的物体跟踪或光流估计。大量的模拟和现实世界实验表明,与现有基线相比,我们的方法有效地提高了在密集动态杂乱环境中的安全边际和飞行效率。此外,所学习的策略在物理四旋翼上实现了稳健的零-shot模拟到现实转移,完全依赖于抽象的时空深度序列及其自我预测的风险先验,验证了我们方法的有效性及其从模拟到现实的稳健泛化能力。
cs.RO / 12 / 2607.23602

Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models

物理空间中相邻集合的动作优于世界模型中的最佳预测
Li, Liangyu, Liu, Qingwen, Liu, Mingqing
Abstract
Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action sequence, choose the minimum, execute its first action block, and replan. This rule can fail even when the terminal cost perfectly and accurately reflects the true task objective in the physical world. Residual prediction error can give an infeasible sequence an anomalously low cost, and a larger proposal pool gives such errors more chances to outrank feasible alternatives. We call this conditional failure proposal overgeneration. In Cube candidate execution audits, increasing the total proposal budget from 72 to 288 reduces the feasibility of selection by minimum latent cost from .375 to .062 for position targets and from .344 to .031 for targets defined by position and yaw, although every larger pool contains a feasible sequence. We introduce Adjacent Set Action Reconstruction (ASAR). Among proposals with low cost, ASAR measures density from standardized early action prefixes and reconstructs a full sequence from an adjacent set with a light anchor from the sequence with minimum cost. On a Carry and Release evaluation set of 75 queries, Kernel ASAR improves event completion success over matching selection by 28.0, 24.0, and 18.7 percentage points under latent cost and by 18.7, 20.0, and 17.3 points under a trajectory reachability cost at 72, 144, and 288 proposals. Analysis of finite proposal pools characterizes selection risk from the lower tail, separation by a related radius support statistic, and sequence containment under an explicit local feasibility condition.
Chinese Translation
基于采样和潜在世界模型的控制器为每个候选动作序列分配一个预测的终端成本,选择最低的成本,执行其第一个动作块,然后重新规划。即使终端成本完美且准确地反映了物理世界中的真实任务目标,这一规则也可能失败。残余预测误差可能导致不可行的序列具有异常低的成本,而更大的提案池则使这些错误有更多机会超越可行的替代方案。我们称这种条件性失败为提案过生成。在 Cube 候选执行审核中,将总提案预算从 72 增加到 288,导致通过最低潜在成本选择的可行性从 0.375 降低到 0.062(针对位置目标),以及从 0.344 降低到 0.031(针对由位置和偏航定义的目标),尽管每个更大的池中都包含一个可行序列。我们引入了相邻集合动作重构(Adjacent Set Action Reconstruction,ASAR)。在低成本的提案中,ASAR 从标准化的早期动作前缀中测量密度,并从具有最低成本序列的相邻集合中重构完整序列。在一个包含 75 个查询的 Carry and Release 评估集中,Kernel ASAR 在 72、144 和 288 个提案下,分别提高了事件完成成功率 28.0、24.0 和 18.7 个百分点(在潜在成本下),以及 18.7、20.0 和 17.3 个百分点(在轨迹可达性成本下)。对有限提案池的分析表征了来自下尾的选择风险,通过相关半径支持统计量的分离,以及在明确的局部可行性条件下的序列包含。
cs.RO / 13 / 2607.23684

Towards Ultrafast Depth Sensing Via Active Event-based Stereo Vision

基于主动事件的立体视觉实现超快深度感知的研究
Li, Jianing, Zhang, Yunjian, Han, Haiqian, Huang, Kangyao, Ji, Xiangyang
Abstract
Conventional frame-based imaging for active stereo systems has encountered major challenges in fast-motion scenarios. However, how to design a novel paradigm for ultrafast depth sensing remains an open issue. In this paper, we propose a novel problem setting, namely active event-based stereo vision, which attempts to integrate binocular event cameras and an infrared 2D pattern projector for high-speed dense depth sensing. Technically, we first build a stereo camera prototype system and present a real-world dataset with over 21.5k spatiotemporal synchronized labels at 15 Hz, while also establishing a realistic synthetic dataset with stereo event streams and 23.8k synchronized labels at 20 Hz. Then, we propose ActiveEventNet+, a lightweight yet effective event-based stereo matching neural network that learns to generate high-quality dense disparity maps from stereo event streams with low latency. Our ActiveEventNet+ mainly involves three innovations: incorporating lightweight blocks into event-based stereo matching frameworks, designing a novel cost volume with dynamic interactions between stereo pairs, and presenting an effective temporal consistency architecture to fully use rich temporal cues in event streams. The results show that our ActiveEventNet+ outperforms state-of-the-art methods while significantly reducing computational complexity. Our solution offers superior depth sensing performance compared to conventional frame-based stereo cameras in high-speed scenes. In particular, the lightweight ActiveEventNet enables the prototype system to achieve real-time processing at speeds up to 150 FPS. We believe that this novel active event-based stereo vision paradigm can provide new insights into the design of future high-speed depth sensing camera systems. Our dataset and code can be available at https://github.com/jianing-li/active_event_based_stereo.
Chinese Translation
传统的基于帧的成像在主动立体系统中在快速运动场景下遇到了重大挑战。然而,如何设计一种新颖的超快深度感知范式仍然是一个未解决的问题。本文提出了一种新颖的问题设置,即基于主动事件的立体视觉,旨在将双目事件相机与红外二维模式投影仪结合,以实现高速密集深度感知。从技术上讲,我们首先构建了一个立体相机原型系统,并提供了一个现实世界的数据集,其中包含超过21.5k个时空同步标签,频率为15 Hz,同时还建立了一个具有立体事件流和23.8k个同步标签的现实合成数据集,频率为20 Hz。然后,我们提出了ActiveEventNet+,这是一种轻量级但有效的基于事件的立体匹配神经网络,能够从低延迟的立体事件流中生成高质量的密集视差图。我们的ActiveEventNet+主要涉及三项创新:将轻量级模块纳入基于事件的立体匹配框架,设计具有立体对之间动态交互的新型代价体积,以及提出一种有效的时间一致性架构,以充分利用事件流中的丰富时间线索。结果表明,我们的ActiveEventNet+在显著降低计算复杂度的同时,超越了最先进的方法。与传统的基于帧的立体相机相比,我们的解决方案在高速场景中提供了卓越的深度感知性能。特别是,轻量级的ActiveEventNet使得原型系统能够以高达150 FPS的速度实现实时处理。我们相信,这种新颖的基于主动事件的立体视觉范式可以为未来高速深度感知相机系统的设计提供新的见解。我们的数据集和代码可在 https://github.com/jianing-li/active_event_based_stereo 获取。
cs.RO / 14 / 2607.23702

Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization

尝试一次,然后最优:跨情节探索摊销的去冗余程序记忆
Ge, Haizhou, Ouyang, Haochen, Chen, Zhixing, Jia, Yufei, Li, Yue, Shi, Lu, Han, Lei, Zhou, Guyue, Huang, Ruqi
Abstract
Manipulating objects with hidden internal state, such as a latched microwave, forces a robot to probe before it can act. Yet a robot that has solved an instance once re-runs the same probes whenever it encounters that instance again, because existing cross-episode memories target task success and organize reuse around states, not the object or the cost of re-exploring it. We present Instance-Oriented Memory (IOM), an object-centric framework that amortizes this exploration: from a single encounter that uncovers the hidden state, whether or not it succeeds, IOM records a short procedure for manipulating that instance, keys it on the object's identifiable features, and injects it as a soft bias on a procedure-conditioned policy. A later encounter recognizes the object and recalls its procedure instead of re-exploring. We instantiate this distillation with an off-the-shelf vision-language model (VLM) that parses each encounter into the procedure without task-specific training. Across four articulated-object tasks, two in simulation (microwave, door) and two on a real robot (bottle, cabinet), an oracle procedure memory cuts manipulation operations by 16-30% over re-exploration at non-regressing success, and the VLM instantiation recovers 69-88% of that saving out of the box. Because the procedure is a soft bias on a feedback-driven policy, an incorrect memory is recovered from rather than obeyed: success holds even when a retrieved procedure is wrong, as for $\approx$12% of door instances. Across all tasks the benefit is purely one of efficiency: success never regresses, and on the real robot even improves. Code will be released upon acceptance.
Chinese Translation
操控具有隐藏内部状态的物体,例如带锁的微波炉,迫使机器人在行动之前进行探测。然而,一旦机器人解决了某个实例,它在再次遇到该实例时会重新执行相同的探测,因为现有的跨情节记忆主要关注任务成功,并围绕状态组织重用,而不是物体本身或重新探索的成本。我们提出了实例导向记忆(Instance-Oriented Memory, IOM),这是一个以物体为中心的框架,用于摊销这种探索:从一次揭示隐藏状态的单次遭遇中,无论成功与否,IOM记录一个短程序以操控该实例,将其键入物体的可识别特征,并将其注入作为程序条件策略上的软偏置。后续的遭遇识别该物体并回忆其程序,而不是重新探索。我们用一个现成的视觉-语言模型(Vision-Language Model, VLM)来实现这一提炼,该模型将每次遭遇解析为程序,而无需特定任务的训练。在四个关节物体任务中,两个在仿真环境中(微波炉、门),两个在真实机器人上(瓶子、橱柜),一个理想程序记忆在非退化成功的情况下将操控操作减少了16-30%的重新探索,而VLM实例化在开箱即用的情况下恢复了69-88%的节省。由于该程序是反馈驱动策略上的软偏置,因此错误的记忆是被恢复而非遵循的:即使检索到的程序是错误的,成功依然存在,例如在约12%的门实例中。在所有任务中,收益纯粹是效率上的:成功从未退化,且在真实机器人上甚至有所改善。代码将在接受后发布。
cs.RO / 15 / 2607.23704

LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratories

LabRobFail:化学自驾实验室中机器人故障分析的基准测试
Wang, Haobo, Sun, Baoli, Zou, Anqi, Huang, Dongsheng, Lv, Zelin, Wang, Ning, Li, Rui, Zhou, Dongzhan, Guo, Weiyu, Wang, Zhihui, Ouyang, Wanli
Abstract
The deployment of embodied agents in self-driving laboratories could accelerate scientific discovery, yet their reliability is constrained by the irreversible and safety-critical nature of chemical experiments. Progress is further hindered by scarce failure data and the lack of fine-grained evaluation protocols. To address these challenges, we introduce LabRobFail, a failure-centric framework for learning and evaluating robotic failure analysis in chemical laboratories. LabRobFail-Sim injects controllable failures at the control, physics, and semantic levels, enabling the construction of LabRobFail-Data, which contains over 20,000 trajectories across 70+ task scenarios, five failure categories, and 11 fine-grained failure types. LabRobFail-Bench evaluates six capabilities spanning task understanding, failure detection, temporal localization, severity assessment, failure classification, and actionable correction. We further develop LabRobFail-VLM, a domain-specialized vision-language model that generates structured failure diagnoses and recovery instructions. On seen environments, it achieves 92.58% failure-detection accuracy and 85.58% temporal-localization accuracy, substantially outperforming general-purpose VLMs. When integrated as a real-time supervisor, it improves downstream VLA task success rates by 10-20 percentage points, demonstrating the value of fine-grained failure understanding for closed-loop recovery and reliable laboratory autonomy. Our code and data are available at https://github.com/Su-ISE-2001/SciRobo
Chinese Translation
在自驾实验室中部署具身智能体可以加速科学发现,但其可靠性受到化学实验不可逆和安全关键性质的限制。进展进一步受到稀缺故障数据和缺乏细粒度评估协议的阻碍。为了解决这些挑战,我们提出了LabRobFail,这是一个以故障为中心的框架,用于学习和评估化学实验室中的机器人故障分析。LabRobFail-Sim在控制、物理和语义层面注入可控故障,从而构建了LabRobFail-Data,该数据集包含超过20,000条轨迹,涵盖70多个任务场景、五类故障和11种细粒度故障类型。LabRobFail-Bench评估六种能力,包括任务理解、故障检测、时间定位、严重性评估、故障分类和可操作的纠正。我们进一步开发了LabRobFail-VLM,一个领域专用的视觉-语言模型,能够生成结构化的故障诊断和恢复指令。在已知环境中,它实现了92.58%的故障检测准确率和85.58%的时间定位准确率,显著优于通用视觉-语言模型。当作为实时监督者集成时,它将下游视觉-语言任务的成功率提高了10-20个百分点,证明了细粒度故障理解在闭环恢复和可靠实验室自主性中的价值。我们的代码和数据可在 https://github.com/Su-ISE-2001/SciRobo 获取。
cs.RO / 16 / 2607.23726

Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning

用于稀疏奖励长时间跨度强化学习的层次化软演员-评论家
Elashaal, Zahra Abdalla, Hfaiedh, Afef, Khraief, Nahla, Ellabib, Issmail, Cirrincione, Giansalvo
Abstract
Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To address these challenges, we propose a two-level Hierarchical Reinforcement Learning (HRL) framework. The first level handles high-level strategic planning, while the low-level uses the continuous-control Soft Actor-Critic (SAC) algorithm, and they utilize entropy-regularized policy optimization. The proposed framework was trained and evaluated using the Search-and-Rescue-2 (SAR-2) dataset. HRL-SAC effectively addresses sparse-reward long-horizon search problems characterized by delayed rewards and continuous control, and its outperforming the flat SAC baseline reinforcement learning in terms of success rates, coverage efficiency, and convergence. These findings indicate that hierarchical entropy-regularized policies are a promising solution to tackle long-horizon sparse-reward reinforcement learning tasks.
Chinese Translation
在稀疏奖励的长时间跨度任务中,探索面临着显著的挑战。为了解决这些挑战,我们提出了一种两级层次化强化学习(HRL)框架。第一级处理高层次的战略规划,而低级则使用连续控制的软演员-评论家(Soft Actor-Critic, SAC)算法,并利用熵正则化的策略优化。所提出的框架使用搜索与救援-2(Search-and-Rescue-2, SAR-2)数据集进行了训练和评估。HRL-SAC有效解决了稀疏奖励长时间跨度搜索问题,这些问题的特点是延迟奖励和连续控制,并且在成功率、覆盖效率和收敛性方面优于平面SAC基线强化学习。这些发现表明,层次化的熵正则化策略是解决长时间跨度稀疏奖励强化学习任务的一个有前景的解决方案。
cs.RO / 17 / 2607.23743

Learning Traversability-Aware Global Planners for Long Horizon Off-Road Navigation

学习考虑可通行性的全局规划器以实现长距离越野导航
Viswanath, Kasi, Gregory, Jason M., Kolhe, Shaunak, Saripalli, Srikanth
Abstract
Autonomous navigation across large off-road environments remains a challenging problem. Onboard sensors perceive only the immediate surroundings, yet safe and efficient routes depend on terrain features that extend well beyond the sensor horizon. Geo-spatial data sources such as satellite imagery, aerial LiDAR, and vector maps can close this gap, but learning traversability from them is difficult: dense labels are unavailable at scale, and existing methods rely on short-range sensing. We propose an efficient formulation that learns a continuous traversability map from overhead data, supervised directly by human-driven GPS trajectories and shaped by self-supervised geometric priors from LiDAR. Alongside the model, we release a public dataset of 299 scenes spanning $\sim\!1{,}244\,\mathrm{km}^{2}$ of diverse terrain, paired with $1{,}130\,\mathrm{km}$ of human driving. In field trials on a Clearpath Warthog across seven routes at two sites, our method achieves trajectories within $3.66\%$ of human path length and reduces operator interventions by $\sim\!85\%$ compared to local-planner-only autonomy.
Chinese Translation
在广阔的越野环境中进行自主导航仍然是一个具有挑战性的问题。机载传感器只能感知周围的即时环境,而安全高效的路线依赖于超出传感器视野的地形特征。地理空间数据源,如卫星图像、航空激光雷达(LiDAR)和矢量地图,可以弥补这一差距,但从中学习可通行性是困难的:大规模的密集标签不可用,现有方法依赖于短距离感知。我们提出了一种高效的公式,通过直接由人类驱动的GPS轨迹监督,并通过来自LiDAR的自监督几何先验进行塑造,从上方数据中学习连续的可通行性地图。我们还发布了一个公共数据集,包含299个场景,覆盖约1,244平方公里的多样地形,并配有1,130公里的人类驾驶数据。在对Clearpath Warthog进行的现场试验中,在两个地点的七条路线中,我们的方法实现了与人类路径长度相差3.66%的轨迹,并将操作员干预减少了约85%,相比于仅依赖局部规划器的自主性。
cs.RO / 18 / 2607.23782

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

$N_0$-VTLA:利用潜在触觉标记扩展视觉-触觉-语言-动作模型
NeoteAI Team, Fudan TEAI Team
Abstract
We present $N_0$-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, $N_0$-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, $N_0$-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. $N_0$-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.
Chinese Translation
我们提出了 $N_0$-VTLA,这是一种视觉-触觉-语言-动作(VTLA)基础模型,能够实现(1)具有触觉感知和触觉反馈控制的细粒度接触丰富操控,以及(2)基于存储的部署数据进行离线策略改进。在现有的基于视觉的骨干网络基础上,我们提出了一种触觉集成的训练方案,包括视觉-触觉预训练、分阶段触觉通路集成和优势条件的离线策略改进。在预训练阶段,策略从我们的规模庞大的视觉-触觉机器人数据集NeoData中学习广泛的接触先验;据我们所知,$N_0$-VTLA是第一个在大规模触觉数据上进行预训练的VTLA模型。在后训练阶段,我们通过一个预测触觉通路增强策略,该通路将大规模学习到的接触模式提炼为下游触觉中心操控所需的细微运动调整。对于离线策略改进,我们引入了ALTER,这是一种优势条件的离线强化学习方法,将相对进展和轨迹事件比较转换为用于固定部署语料库的策略训练的二元优势标签,进一步改善了在接触丰富技能(如可变形物体操控)上的任务特定学习。在接触丰富的基准测试中,$N_0$-VTLA以较大优势超越了强基线:在九个真实机器人NeoReal任务中全部获胜,并在二十任务模拟套件中达到63.8%的平均成功率,而最强基线为44.0%。使用ALTER训练的$N_0$-VTLA策略在三个长时间跨度的真实机器人任务中达到75-95%的成功率。这些结果为多功能触觉驱动的操控策略奠定了基础。
cs.RO / 19 / 2607.23783

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

$N_0$-TWAM:用于接触丰富操作的触觉原生世界-动作模型的扩展
NeoteAI Team, Fudan TEAI Team
Abstract
We present $N_0$-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on contact-rich tasks. We pre-train $N_0$-TWAM at large scale with visuo-tactile joint training over tactile-rich demonstrations spanning six embodiments and 450 tasks. We use NeoForce, a unified force-based tactile representation, to form a physically grounded contact signal that conditions action generation. To improve long-horizon and multi-stage manipulation, we introduce tactile contact events for task staging and advance through them during execution. For real-time efficiency, we adopt an asymmetric Mixture-of-Transformers architecture that pairs a full-width expert for video prediction with slim experts for downstream action and tactile prediction. Evaluations on both real and simulated benchmarks justify the capabilities of $N_0$-TWAM across a range of contact-rich tasks, and demonstrate the benefit of data scaling for precise tactile and action prediction. In summary, $N_0$-TWAM endows a world-action model with predictive capabilities to foresee vision, touch and action, building a solid foundation for fine-grained manipulation on open contact-rich tasks. The codebase and model checkpoints will be made publicly available to foster further research and development in tactile-enabled robotic manipulation.
Chinese Translation
我们提出了 $N_0$-TWAM,这是一种用于接触丰富操作的触觉原生世界-动作模型,能够预测未来的视觉和未来的接触。据我们所知,这是第一个在大规模下训练的触觉世界-动作模型,并且在接触丰富的任务上表现出强大的能力。我们通过对跨越六种体现和450个任务的触觉丰富演示进行视觉-触觉联合训练,在大规模上预训练了 $N_0$-TWAM。我们使用 NeoForce,这是一种统一的基于力的触觉表示,形成一个物理基础的接触信号,以指导动作生成。为了改善长时间和多阶段的操作,我们引入了触觉接触事件用于任务分阶段,并在执行过程中逐步推进。为了实现实时效率,我们采用了一种不对称的混合变换器架构,将用于视频预测的全宽专家与用于下游动作和触觉预测的精简专家相结合。在真实和模拟基准上的评估证明了 $N_0$-TWAM 在一系列接触丰富任务中的能力,并展示了数据扩展在精确触觉和动作预测中的优势。总之,$N_0$-TWAM 赋予了世界-动作模型预测视觉、触觉和动作的能力,为开放接触丰富任务上的精细操作奠定了坚实的基础。代码库和模型检查点将公开发布,以促进触觉驱动的机器人操作的进一步研究和开发。
cs.RO / 20 / 2607.23784

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

言语引导的机器人策略合成:小小的言语能产生巨大的影响
Chen, Daphne, Jain, Archit Ritesh, Goossen, Eric, Romig, Emma, Murray, Michael, Walker, Nick, Cakmak, Maya
Abstract
While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail. In this work, we propose ARCHITECT, a framework that treats robot policy acquisition as an interactive program synthesis task. ARCHITECT leverages the reasoning capabilities of LLM coding agents to synthesize modular robot programs that utilize a suite of perception and control tools. Unlike end-to-end models where distribution shift leads to unpredictable, cascading failures, our modular architecture allows users to isolate failures and localize feedback at the level of abstraction required. We introduce an iterative process where a human supervisor provides natural language corrections to steer the policy. These corrections are grounded in the policy code by program execution traces and distilled into a persistent skill library, a form of long-term in-context learning which enables the agent to accumulate a repertoire of reusable, interpretable behaviors. In a benchmark evaluation on a Franka Panda robot, ARCHITECT outperforms state-of-the-art VLA models and program synthesis baselines on complex, long-horizon tasks, including articulated object manipulation and cloth folding. Our results demonstrate that the synthesized skill library enables the system to transfer to novel tasks with decreasing human intervention, providing a steerable and data-efficient alternative to black-box robot learning. Website: https://robo-architect.github.io/
Chinese Translation
尽管视觉-语言-动作模型展示了令人印象深刻的零-shot操作能力,但它们仍然是根本上难以解释、适应或纠正的黑箱策略,在不可避免地失败时尤为如此。在本研究中,我们提出了ARCHITECT,一个将机器人策略获取视为交互式程序合成任务的框架。ARCHITECT利用大规模语言模型(LLM)编码代理的推理能力,合成利用一系列感知和控制工具的模块化机器人程序。与端到端模型不同,后者在分布转移时会导致不可预测的级联失败,我们的模块化架构允许用户在所需的抽象层次上隔离故障并本地化反馈。我们引入了一个迭代过程,其中人类监督者提供自然语言的修正以引导策略。这些修正通过程序执行轨迹与策略代码相结合,并提炼成一个持久的技能库,这是一种长期的上下文学习形式,使代理能够积累可重用、可解释行为的 repertory。在对Franka Panda机器人的基准评估中,ARCHITECT在复杂的长时间任务上超越了最先进的视觉-语言-动作模型和程序合成基线,包括关节物体操作和布料折叠。我们的结果表明,合成的技能库使系统能够在减少人类干预的情况下转移到新任务,提供了一种可引导和数据高效的替代黑箱机器人学习的方法。网站:https://robo-architect.github.io/
cs.RO / 21 / 2607.23797

Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map

注意力记忆:基于视觉-语言-运动图的语言条件再感知
Ghosh, Dibyendu
Abstract
A robot carrying a persistent, behavior-annotated map faces two planning questions, and its memory answers only one well. The \emph{spatial-navigation} question -- how to walk around a room -- we address first and report a negative: building on Vision--Language--Motion Maps (VLMM), a behavior-aware planner cost cuts a planning-time objective by $\sim$35\% over 28 AI2-THOR scenes, but under closed-loop execution the real benefit nearly vanishes ($\sim$4\%) and an on-demand vision--language model (VLM) does as well. The \emph{resource-allocation} question differs: under a limited perception budget, what should the robot re-observe now to keep its map fresh? Framing re-perception as this attention decision, we show a persistent map's memory (change-history, or even just recency of last sighting) yields the best schedule (held-out), matching an oracle, while the memoryless VLM prior is poor. Because the schedule reallocates budget toward what matters, memory's benefit concentrates on the important objects ($\sim$1.6$\times$ the mean), and a downstream fetch task confirms fewer wasted trips; the gain grows with per-instance heterogeneity exactly as a Cauchy--Schwarz bound predicts -- it equals $\mathrm{Var}(\sqrt\lambda)$, the variance of root-volatility. With a real CLIP prior on rendered objects the advantage is $+21$--$26\%$. The map's distinctive value appears when the task is \emph{language-conditioned}: told what to track, VLMM grounds the relevant objects (open-vocabulary) and tracks their change (memory), beating even a strong relevance-weighted recency baseline ($+2.5\%$) -- so its motion channel adds value beyond a last-seen timestamp -- and an on-demand VLM ($+8.9\%$); neither language nor dynamics alone suffices. The map earns its keep not by telling the robot how to walk around a room, but by telling it what to pay attention to.
Chinese Translation
一台携带持久行为注释地图的机器人面临两个规划问题,而其记忆仅能较好地回答其中一个。 extit{空间导航}问题——如何在房间内行走——我们首先进行探讨,并报告一个负面结果:基于视觉-语言-运动图(VLMM)的行为感知规划器在28个AI2-THOR场景中将规划时间目标减少了约35%,但在闭环执行下,实际收益几乎消失(约4%),而按需视觉-语言模型(VLM)也表现相似。 extit{资源分配}问题则有所不同:在有限的感知预算下,机器人现在应该重新观察什么以保持其地图的新鲜度?将再感知框架视为这一注意力决策,我们展示了持久地图的记忆(变化历史,甚至仅是最近观察的时间)产生了最佳调度(保留集),与一个神谕相匹配,而无记忆的VLM先验表现较差。由于调度将预算重新分配到重要对象上,记忆的益处集中在重要对象上(约为均值的1.6倍),而下游取物任务确认了更少的浪费行程;收益随着每个实例的异质性增长,正如柯西-施瓦茨界限所预测的那样——它等于 ext{Var}( ext{sqrt} ext{λ}),即根波动的方差。当使用真实的CLIP先验对渲染对象进行处理时,优势为21%至26%。当任务是 extit{语言条件}时,地图的独特价值显现:在被告知跟踪什么时,VLMM将相关对象(开放词汇)与其变化(记忆)结合,超越了即便是强相关性加权的近期基线(+2.5%)——因此其运动通道的价值超出了最后观察时间戳——以及按需VLM(+8.9%);仅有语言或动态单独使用都不足以满足需求。地图的价值并不在于告诉机器人如何在房间内行走,而在于告诉它应关注什么。
cs.RO / 22 / 2607.23867

BC-NMPC: Battery-Constrained NMPC with Propulsion Prediction and Replanning for High-Speed Flight

BC-NMPC:具有推进预测和重新规划的电池约束非线性模型预测控制用于高速飞行
Gupta, Parakh M., Mihulka, Matej, Novosad, Matej, Penicka, Robert, Saska, Martin
Abstract
Trajectory tracking performance of Uncrewed Aerial Vehicles (UAVs) degrades during high-speed and agile flight due to the depletion of the battery and subsequent loss of maximum available thrust. In applications such as drone racing, the consequent trajectory tracking error leads to a collision with obstacles and a subsequent failure to complete the race. In this paper, we present a novel method for integrating battery and propulsion system models into a Nonlinear Model Predictive Controller (NMPC) framework to enable real-time prediction of the voltage, consumed current, power, and maximum available thrust of the platform. This enables our approach to account for the dynamic variations in the maximum available thrust of the UAV caused by battery discharge, allowing it to plan for the depleting thrust and improve trajectory tracking performance. A trajectory planning algorithm is implemented to replan the trajectory in-flight based on evolving thrust limits. The accuracy of the proposed model is verified in real-world flight experiments, while the effectiveness of the replanning algorithm is evaluated in simulation. Compared to an uncompensated flight, our novel approach demonstrates achieves a collision-free flight to achieve a 6-fold decrease in tracking Root Mean Square Error (RMSE), a 46 % increase in flight distance, and a 100 % increase in flight time in an obstacle-ridden environment.
Chinese Translation
在高速和灵活飞行过程中,无人驾驶飞行器(UAV)的轨迹跟踪性能因电池耗尽而下降,导致最大可用推力的丧失。在无人机竞速等应用中,随之而来的轨迹跟踪误差会导致与障碍物碰撞,进而无法完成比赛。本文提出了一种新颖的方法,将电池和推进系统模型集成到非线性模型预测控制(NMPC)框架中,以实现对平台电压、消耗电流、功率和最大可用推力的实时预测。这使得我们的方法能够考虑由于电池放电引起的UAV最大可用推力的动态变化,从而能够为逐渐减少的推力进行规划,并改善轨迹跟踪性能。我们实现了一种轨迹规划算法,以根据不断变化的推力限制在飞行中重新规划轨迹。通过实际飞行实验验证了所提模型的准确性,同时在仿真中评估了重新规划算法的有效性。与未补偿飞行相比,我们的新方法实现了无碰撞飞行,使轨迹均方根误差(RMSE)降低了6倍,飞行距离增加了46%,飞行时间在障碍物环境中增加了100%。
cs.RO / 23 / 2607.23899

Embodied GPT-5.1: Evidence of a World Model?

具身的 GPT-5.1:世界模型的证据?
Spinelli, Roberto, Martins, Thiago C.
Abstract
This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot despite having no prior embodiment, no training in simulated environments, and no exposure to sensorimotor experience. Using only low-resolution first-person images and a discrete action set, the model was tasked with navigation and object-directed behaviors such as locating and contacting a target toy. Across multiple trials, GPT-5.1 demonstrated emergent capabilities that suggest elements of spatial reasoning and physical understanding. These included maintaining short-term memory of object locations after they left the camera frame, inferring the physical consequences of its own movements, and executing coherent action sequences such as colliding with an object and reversing to visually verify the outcome. At the same time, the model displayed inefficiencies and perceptual limitations, including imprecise alignment strategies and occasional misidentification of distant distractors. Overall, the results indicate that GPT-5.1 exhibits signs of world-model-like behavior in an embodied setting, despite the absence of any embodiment-related training, a finding that challenges long-standing views in cognitive science and robotics which hold that a physical body is a necessary prerequisite for developing such forms of intelligence. The findings motivate deeper investigation into the emergence, limits, and robustness of physical understanding in large language models.
Chinese Translation
本探索性研究考察了大型多模态语言模型 GPT-5.1 是否能够作为物理移动机器人高层控制器,尽管其没有先前的具身经验、未在模拟环境中训练,也未接触传感器运动经验。仅使用低分辨率的第一人称图像和离散的动作集,该模型被赋予导航和目标导向行为的任务,例如定位和接触目标玩具。在多次试验中,GPT-5.1 展现出一些新兴能力,暗示其具备空间推理和物理理解的元素。这些能力包括在物体离开摄像机画面后保持短期记忆、推断自身运动的物理后果,以及执行连贯的动作序列,如与物体碰撞并倒退以视觉验证结果。同时,该模型也表现出一些低效和感知限制,包括不精确的对齐策略和偶尔错误识别远处的干扰物。总体而言,结果表明,尽管缺乏任何与具身相关的训练,GPT-5.1 在具身环境中展现出类似世界模型的行为,这一发现挑战了认知科学和机器人学中长期以来的观点,即物理身体是发展此类智能的必要前提。这些发现激励我们深入研究大型语言模型中物理理解的出现、局限性和稳健性。
cs.RO / 24 / 2607.23969

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

LeapBot-WA:通过预测潜在对齐实现世界锚定动作模型
Liu, Pei, Zheng, Nan, Zhang, Lang, Peng, Daojie, Zhang, Yanan, Kong, Feilong, Feng, Mingyue, Liu, Jiachao, Wang, Yaonong, Chen, Qifeng, Ma, Jun
Abstract
World Action Models (WAMs) have emerged as a powerful paradigm for embodied intelligence, yet the prevailing reliance on pixel-level video generation creates a fundamental bottleneck. Forcing models to reconstruct task-irrelevant visual details dissipates representational capacity and renders policies vulnerable to visual distractors. In this paper, we propose LeapBot-WA, which establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor. Departing from the traditional reliance on visual synthesis, LeapBot-WA shifts the core of world modeling to Predictive Semantic Alignment, extracting abstract physical dynamics directly within a latent foundation space. To bridge the modality gap between non-Gaussian predictive features and diffusion priors, we introduce the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift. Furthermore, we design an Asymmetric Mixture-of-Transformers (MoT) architecture. During training, an Anchor Diffusion Transformer acts as a privileged dynamics expert to guide the Action Diffusion Transformer; at inference, this heavy dynamics branch is pruned, enabling zero-overhead execution. LeapBot-WA achieves state-of-the-art performance among predictive models on LIBERO and matches top-tier generative WAMs on RoboTwin 2.0 without requiring large-scale trajectory pre-training. It further demonstrates superior zero-shot robustness to unseen environments and successful real-world transfer, establishing a highly efficient and robust latent-centric paradigm for scalable robotic control. Code: https://github.com/LeapWM/leapbot-wa.
Chinese Translation
世界动作模型(WAMs)已成为具身智能的强大范式,但对像素级视频生成的依赖造成了根本性瓶颈。强迫模型重建与任务无关的视觉细节消耗了表征能力,并使策略易受视觉干扰。在本文中,我们提出了LeapBot-WA,它通过将联合嵌入预测架构(Joint-Embedding Predictive Architecture, JEPA)作为世界锚,建立了一种新颖的预测潜在范式。与传统的视觉合成依赖不同,LeapBot-WA将世界建模的核心转向预测语义对齐,直接在潜在基础空间中提取抽象的物理动态。为了弥合非高斯预测特征与扩散先验之间的模态差距,我们引入了各向同性语义自编码器(Isotropic Semantic Autoencoder, ISAE),它将锚的潜在空间重塑为适合扩散的流形,以防止偏离流形。此外,我们设计了一种非对称混合变换器(Asymmetric Mixture-of-Transformers, MoT)架构。在训练过程中,锚扩散变换器作为特权动态专家指导动作扩散变换器;在推理时,这一重动态分支被修剪,实现零开销执行。LeapBot-WA在LIBERO数据集上实现了预测模型的最先进性能,并在RoboTwin 2.0上与顶级生成WAMs相匹配,而无需大规模轨迹预训练。它进一步展示了对未见环境的优越零-shot鲁棒性和成功的现实世界迁移,为可扩展机器人控制建立了一种高效且稳健的潜在中心范式。代码:https://github.com/LeapWM/leapbot-wa。
cs.RO / 25 / 2607.23989

Co-planning of Flight Corridors and Communication Infrastructure for Urban Drone Logistics Networks

城市无人机物流网络的飞行走廊与通信基础设施的协同规划
He, Yingjie, Wang, Yikang, Gao, Zhenyu
Abstract
Reliable wireless connectivity is essential for urban air mobility (UAM) networks in dense urban environments. It is therefore imperative to carefully plan the supporting communication infrastructure for UAM flight corridors. Most existing works optimize communication infrastructure and UAV flight paths independently, often leading to unnecessary base station (BS) deployment or excessive flight detours. This paper studies the joint optimization of BS deployment and UAV flight corridors in complex urban environments, aiming to minimize both infrastructure investment and flight distance while satisfying communication quality constraints. We propose CR-CMAB, a channel reciprocity-guided combinatorial multi-armed bandit framework. The framework constructs high-fidelity radio maps using 3D ray tracing, selects BS combinations via coverage-aware CMAB search, and dynamically expands the search space by identifying promising BS locations through channel reciprocity. Experimental results from a detailed case study demonstrate that CR-CMAB outperforms baseline methods with moderate computational time, yielding more strategically positioned BSs and shorter flight corridors. This study offers a practical planning perspective for cost-effective and communication-reliable UAM deployment in future smart cities.
Chinese Translation
在密集城市环境中,可靠的无线连接对于城市空中出行(UAM)网络至关重要。因此,仔细规划UAM飞行走廊所需的通信基础设施是非常必要的。现有的大多数研究独立优化通信基础设施和无人机(UAV)飞行路径,这往往导致不必要的基站(BS)部署或过度的飞行绕行。本文研究了在复杂城市环境中基站部署与无人机飞行走廊的联合优化,旨在在满足通信质量约束的同时,最小化基础设施投资和飞行距离。我们提出了CR-CMAB,这是一种基于信道互易性的组合多臂老虎机框架。该框架利用三维光线追踪构建高保真无线电地图,通过覆盖感知的CMAB搜索选择基站组合,并通过识别有前景的基站位置动态扩展搜索空间。详细案例研究的实验结果表明,CR-CMAB在适度的计算时间内优于基线方法,能够提供更具战略性的基站位置和更短的飞行走廊。本研究为未来智能城市中经济高效且通信可靠的UAM部署提供了实用的规划视角。
cs.RO / 26 / 2607.24008

FutureRTC: Real-Time Robot Execution with Anticipatory-Conditioned Action Chunking

FutureRTC:基于预期条件的实时机器人执行与动作分块
Jiang, Hai, Zou, Yixian, Liang, Binbin, Liu, Boqian, Meng, Fanman, Liu, Shuaicheng
Abstract
Real-time deployment of Vision-Language-Action (VLA) policies necessitates asynchronous execution, wherein subsequent action chunks are computed concurrently with the execution of the current chunk, leading to prediction-execution misalignment and manifesting as inter-chunk discontinuities. Existing methods either superficially smooth chunk boundaries, require costly policy optimization, or exclusively forward-predict proprioceptive states yet neglect critical visual observations. In this paper, we propose \textbf{FutureRTC}, a plug-and-play adaptation framework that predicts execution-time observations and states for asynchronous VLA control without modifying the underlying policy. Specifically, FutureRTC features a state correction module to compensate for the discrepancy between rolled-forward and actual execution-time proprioceptive states and an observation prediction module that forecasts execution-time visual representations by leveraging robot motion as an explicit physical prior through motion-aware feature transport and reconstruction. Furthermore, we introduce a policy consistency loss to align the action chunks generated from predicted contexts with those produced under the expected execution-time inputs of the VLA policy. Extensive experiments across simulated and real-world environments demonstrate that FutureRTC achieves superior robustness to inference delays, resulting in smoother trajectories, faster execution, and consistently higher task success rates.
Chinese Translation
视觉-语言-动作(VLA)策略的实时部署需要异步执行,其中后续动作分块与当前分块的执行并行计算,这导致预测与执行之间的不一致,并表现为分块间的断裂。现有方法要么仅仅表面上平滑分块边界,要么需要昂贵的策略优化,或者仅仅向前预测本体状态而忽视关键的视觉观测。本文提出了 extbf{FutureRTC},一种即插即用的适应框架,能够在不修改基础策略的情况下预测异步VLA控制的执行时观测和状态。具体而言,FutureRTC具有一个状态校正模块,用于补偿滚动前推与实际执行时本体状态之间的差异,以及一个观测预测模块,通过运动感知特征传输和重建,利用机器人运动作为显式的物理先验,预测执行时的视觉表示。此外,我们引入了一种策略一致性损失,以使从预测上下文生成的动作分块与在VLA策略的预期执行时输入下产生的动作分块对齐。在模拟和真实环境中的大量实验表明,FutureRTC在推理延迟方面表现出更强的鲁棒性,从而实现了更平滑的轨迹、更快的执行和始终更高的任务成功率。
cs.RO / 27 / 2607.24029

Moving-Horizon Estimation and Nonlinear Model Predictive Control of Cable-Driven Soft Manipulators

基于移动视野估计和非线性模型预测控制的电缆驱动软操纵器
Xun, Lingxiao, Li, Haihong, Zheng, Gang
Abstract
Precise control of soft manipulators remains challenging due to the difficulty of developing accurate yet computationally tractable models for model-based estimation and control. Reduced Cosserat-rod models provide a physics-based and control-oriented description of soft-robot dynamics, offering an explicit alternative to purely data-driven input-output representations. In this paper, we propose a moving-horizon estimation (MHE) and nonlinear model predictive control (NMPC) framework for cable-driven soft manipulators based on reduced Cosserat dynamics. A smooth cable-length-driven modeling formulation is developed by approximating the complementarity relationship between cable tension and cable slackness, enabling cable-length control without direct tension sensing. Based on this formulation, an MHE method is introduced to estimate the reduced state and reconstruct the manipulator configuration from end-effector pose measurements and cable-length information. An NMPC controller is then formulated to achieve task-space control under cable-length and cable-rate constraints. The proposed framework is validated through numerical simulations and experiments. Simulation results demonstrate the effectiveness of the estimator and controller for pose and strain-related regulation on a multi-cable soft manipulator. Experimental results on a four-cable prototype further show that the proposed MHE-NMPC scheme can be implemented in real time and enables accurate end-effector position tracking through cable-length control.
Chinese Translation
由于开发既准确又计算上可行的模型以进行基于模型的估计和控制的困难,精确控制软操纵器仍然具有挑战性。简化的Cosserat杆模型提供了一种基于物理的、面向控制的软机器人动力学描述,为纯数据驱动的输入输出表示提供了明确的替代方案。本文提出了一种基于简化Cosserat动力学的电缆驱动软操纵器的移动视野估计(MHE)和非线性模型预测控制(NMPC)框架。通过近似电缆张力与电缆松弛之间的互补关系,开发了一种平滑的电缆长度驱动建模形式,从而实现无需直接张力传感的电缆长度控制。基于该建模形式,提出了一种MHE方法,用于估计简化状态并根据末端执行器姿态测量和电缆长度信息重构操纵器配置。然后,制定了一种NMPC控制器,以在电缆长度和电缆速率约束下实现任务空间控制。通过数值仿真和实验验证了所提出的框架。仿真结果表明,该估计器和控制器在多电缆软操纵器的姿态和应变相关调节方面的有效性。对四电缆原型的实验结果进一步表明,所提出的MHE-NMPC方案可以实时实施,并通过电缆长度控制实现准确的末端执行器位置跟踪。
cs.RO / 28 / 2607.24036

WARL: Wrench-Augmented Reinforcement Learning for Task-Agnostic Learning in Legged Robots

WARL:用于腿部机器人任务无关学习的扭矩增强强化学习
Yoneda, Keita, Kawaharazuka, Kento, Okada, Kei
Abstract
While reinforcement learning for legged robots has achieved high motor performance, it has been constrained by the limited exploration capability of actions confined to the joint space. To address this issue, this study proposes a new method, Wrench-Augmented Reinforcement Learning (WARL), which introduces a wrenche (force and torque) into the action space. The proposed method combines wrench-guided exploration with a success rate-based curriculum mechanism to expand exploration capabilities in the early stages of learning, with the ultimate goal of acquiring behaviors based solely on joint control. Experiments using a quadruped robot demonstrated that WARL can learn robustly across diverse terrains and motor tasks without requiring terrain-specific reward adjustments or complex curriculum designs. Furthermore, an ablation study verified the effectiveness of the Switching Curriculum, which gradually eliminates the wrench. On the other hand, we also show that introducing a wrench can encourage behaviors that do not sufficiently exploit the robot's physical embodiment. These findings suggest that while wrench-based exploration enhancement is effective for improving learning efficiency, designing it in a way that is consistent with the robot's physical structure is a critical future challenge.
Chinese Translation
尽管腿部机器人的强化学习已实现高水平的运动性能,但其探索能力受到限制,主要局限于关节空间内的动作。为了解决这一问题,本研究提出了一种新方法——扭矩增强强化学习(Wrench-Augmented Reinforcement Learning,WARL),该方法将扭矩(力和扭矩)引入动作空间。所提方法结合了扭矩引导的探索与基于成功率的课程机制,以扩展学习早期阶段的探索能力,最终目标是基于关节控制获取行为。使用四足机器人进行的实验表明,WARL能够在不同的地形和运动任务中稳健学习,而无需进行特定地形的奖励调整或复杂的课程设计。此外,消融研究验证了逐步消除扭矩的切换课程的有效性。另一方面,我们还表明,引入扭矩可能会促使一些行为未能充分利用机器人的物理表现。这些发现表明,尽管基于扭矩的探索增强在提高学习效率方面是有效的,但以与机器人物理结构一致的方式进行设计是未来的一个关键挑战。
cs.RO / 29 / 2607.24079

Effective Parameters, Real Behavior: Renormalization for Robotics -- From Infinite Electron Mass to Sim-to-Real Gap

有效参数与真实行为:机器人学中的重整化——从无限电子质量到模拟与现实之间的差距
Sun, Youran, Guo, Jiaxuan, Ren, Xingyu, Yi, Chugang, Yang, Haizhao
Abstract
Bridging the sim-to-real gap is a central problem in robotics, and the prevailing approach is to build increasingly accurate simulators. Here, we propose another approach based on renormalization: using effective, resolution-dependent parameters to absorb details omitted by the simulator and reproduce real behavior. These parameters may differ from measured physical values because they compensate for what the simulator leaves out. We demonstrate this mechanism analytically for proportional--derivative (PD) control at finite simulation frequency, where proportional feedback changes the effective derivative gain and derivative feedback changes the effective inertia. We then interpret dynamic rope manipulation and underwater swimming through the same perspective. Finally, we present a practical procedure for choosing observables, identifying omitted physics, and determining effective parameters. Renormalization offers robotics a complementary path across the sim-to-real gap: effective parameters, real behavior.
Chinese Translation
弥合模拟与现实之间的差距是机器人学中的一个核心问题,当前的主流方法是构建越来越精确的模拟器。在这里,我们提出了一种基于重整化的另一种方法:使用有效的、依赖于分辨率的参数来吸收模拟器省略的细节,并重现真实行为。这些参数可能与测量的物理值不同,因为它们补偿了模拟器所遗漏的内容。我们在有限模拟频率下对比例-导数(PD)控制进行了分析,展示了这一机制,其中比例反馈改变了有效导数增益,而导数反馈则改变了有效惯性。接着,我们通过相同的视角解释了动态绳索操控和水下游泳。最后,我们提出了一种实用的程序,用于选择可观测量、识别遗漏的物理现象以及确定有效参数。重整化为机器人学提供了一条跨越模拟与现实之间差距的补充路径:有效参数,真实行为。
cs.RO / 30 / 2607.24113

A Case Study on the Acceptance of a Humanoid Robotic Head Employed in Three Public Spaces

关于在三个公共场所使用类人机器人头部的接受度案例研究
Heisler, Marcel, Randecker, Luca, Becker-Asano, Christian
Abstract
Previous research has shown that a human-like robot's acceptance heavily depends on the setting in which it operates and its ability to perform relevant tasks. This paper, first, reports on how our robot processes natural language to generate a multimodal, verbal response integrating emotional expressions based on an emotion simulation backend. Then, it describes how visitors were invited to speak with our robot in their own language at three different, public locations, where the robot was running continuously for several days. The TAM2 questionnaire results reveal that on average users were motivated to use the robot and found it rather useful and easy to use regardless of the specific location. However, public spaces like the tourist information and the city library seem to be a better fit for our interactive, robotic head than an office environment such as the building authority, where the willingness to interact was lower. Overall, the robot's multi-lingual responses were very much appreciated, but every fifth user found the response time too slow impeding the dialog flow, which remains to be improved in future work.
Chinese Translation
以往研究表明,类人机器人被接受的程度在很大程度上依赖于其操作的环境以及其执行相关任务的能力。本文首先报告了我们的机器人如何处理自然语言,以生成基于情感模拟后端的多模态、口头响应,整合情感表达。接着,描述了如何邀请访客在三个不同的公共场所与我们的机器人用他们自己的语言交流,机器人在这些地方连续运行了数天。TAM2问卷的结果显示,平均而言,用户使用机器人的动机较强,认为其相当有用且易于使用,无论具体位置如何。然而,像旅游信息中心和城市图书馆这样的公共空间似乎更适合我们的互动机器人头部,而建筑管理局等办公室环境的互动意愿较低。总体而言,机器人的多语言响应受到高度赞赏,但每五位用户中就有一位认为响应时间过慢,妨碍了对话的流畅性,这在未来的工作中仍需改进。
cs.RO / 31 / 2607.24159

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

DeVA:具有物理指导的解耦视频-动作模型用于机器人策略学习
Zhang, Mengqi, Khose, Sahil, Kareer, Simar, Song, Yuchen, Jain, Unnat, Hoffman, Judy
Abstract
Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations. Video generative models offer a promising foundation by encoding rich spatiotemporal priors through future predictions. However, existing Video-Action Models either couple video and action prediction in a shared backbone, making policy adaptation harder to optimize, or under-utilize video information when guiding the action branch. In this work, we introduce DeVA, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance. DeVA transfers representations from multiple video layers to the action expert, enabling rich information exchange while making policy learning more tractable. It further supervises intermediate video features and the action stream with physically salient guidance (affordance/depth). Experiments on both simulation benchmarks and real-world deployment demonstrate strong performance with limited data, faster convergence than a unified architecture, and clear performance gains from physical guidance.
Chinese Translation
通用机器人操作需要能够在执行语言指令时预测视觉场景如何演变的策略。尽管最近的视觉-语言-动作模型受益于大规模的预训练,但它们主要静态的预训练目标对物理动态和时间因果关系的监督有限,导致控制相关知识需要从下游机器人演示中学习。视频生成模型通过未来预测提供了一个有前景的基础,能够编码丰富的时空先验。然而,现有的视频-动作模型要么在共享主干中耦合视频和动作预测,使得策略适应更难以优化,要么在指导动作分支时未充分利用视频信息。在本研究中,我们提出了DeVA,一种具有专门视频和动作专家、多个层次特征转移和物理显著指导的解耦视频-动作模型。DeVA将来自多个视频层的表示转移到动作专家,促进了丰富的信息交换,同时使策略学习变得更加可行。它进一步用物理显著指导(可用性/深度)对中间视频特征和动作流进行监督。在模拟基准和真实世界部署的实验中,DeVA展示了在有限数据下的强大性能、比统一架构更快的收敛速度,以及来自物理指导的明显性能提升。
cs.RO / 32 / 2607.24190

Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim

不被遗忘:人形机器人头部Kim的个性化情节记忆的实现与评估
Aschenbrenner, Steve, Heisler, Marcel, Sievers, Thomas, Becker-Asano, Christian
Abstract
Social robots that rely on large language models for conversation are unable to retain information across sessions. This absence of memory violates social expectations, potentially preventing the formation of persistent relationships. This paper presents a lightweight episodic memory module that integrates vector-based semantic retrieval with an LLM-controlled dialog system, deployed on the humanoid robot head Kim. The module employs a hybrid scoring function combining cosine similarity with a memory strength metric to retrieve contextually relevant past interactions and inject them into the generation prompt. The system was evaluated in a within-subjects video-based online study (N = 43) using the Human-Robot Interaction Evaluation Scale (HRIES). Results show that episodic memory significantly increased perceived sociability (d = 0.60, p < .001), with the strongest effects on perceived trustworthiness (d = 0.62) and warmth (d = 0.56). Perceived disturbance remained unchanged (d = 0.00), indicating that the implemented approach to personalized recall did not trigger privacy-related discomfort or uncanny valley effects. These findings suggest that episodic memory serves as a social lubricant in embodied Human-Robot Interaction, enhancing relational quality without eliciting negative affective responses.
Chinese Translation
依赖于大型语言模型进行对话的社交机器人无法在会话之间保留信息。这种缺乏记忆的情况违反了社交期望,可能阻碍持久关系的形成。本文提出了一种轻量级情节记忆模块,该模块将基于向量的语义检索与由LLM(大型语言模型)控制的对话系统相结合,部署在人形机器人头部Kim上。该模块采用混合评分函数,将余弦相似度与记忆强度指标相结合,以检索上下文相关的过去互动并将其注入生成提示中。该系统在一项基于视频的在线研究中进行了评估(N = 43),使用了人机交互评估量表(HRIES)。结果显示,情节记忆显著提高了感知的社交性(d = 0.60,p < .001),对感知的可信度(d = 0.62)和温暖感(d = 0.56)影响最为显著。感知的干扰保持不变(d = 0.00),表明实施的个性化回忆方法并未引发隐私相关的不适或恐怖谷效应。这些发现表明,情节记忆在具身人机交互中充当了社交润滑剂,提升了关系质量而未引发负面情感反应。
cs.RO / 33 / 2607.24206

Surgical Re-enactment for Operating Room Workflow Datasets

手术重演用于手术室工作流程数据集
Friedrich, Jana Nina, Ross, Andrea Karin Maria, Henriques, Angelo, Weisser, Mario Peter Martin, Zhang, Ling, Nasseri, Mohammad Ali
Abstract
The introduction of new technologies, such as surgical robots, is driving the vision of a connected, smart operating room (OR). However, realizing this vision requires a deep understanding of surgical workflows, which relies on realistic datasets capturing the actions of all OR personnel from both full room and surgical field perspectives. Acquiring such data in real ORs is prohibitively challenging due to factors such as ethics committee approvals, limited space for camera installation, and sterility regulations preventing the use of tracking markers. We present a step-by-step methodology for re-enacting complete surgical procedures in a reconstructed OR. This approach enables the creation of repeatable and annotatable workflow datasets for training activity recognition models, generating scene graphs, and formalizing surgical process models. Developed for robot-assisted ophthalmic surgery, our methodology combines expert consultation, structured workflow formalization, OR reconstruction, role-based training, real OR observation, and iterative recording with post-take debriefing. We provide concrete recommendations to allow other research groups to seamlessly adopt this methodology for their own surgical domains.
Chinese Translation
新技术的引入,如手术机器人,推动了连接智能手术室(OR)的愿景。然而,实现这一愿景需要深入理解手术工作流程,这依赖于能够从全室和手术区域视角捕捉所有手术室人员动作的真实数据集。在真实手术室中获取此类数据面临极大的挑战,原因包括伦理委员会的批准、摄像头安装空间有限以及无菌规定限制了追踪标记的使用。我们提出了一种逐步的方法论,用于在重建的手术室中重演完整的手术过程。这种方法使得能够创建可重复和可注释的工作流程数据集,以训练活动识别模型、生成场景图和形式化手术过程模型。我们的研究方法专为机器人辅助眼科手术开发,结合了专家咨询、结构化工作流程形式化、手术室重建、基于角色的培训、真实手术室观察以及迭代录制与后期总结。我们提供具体建议,以便其他研究团队能够无缝地将此方法论应用于他们自己的手术领域。
cs.RO / 34 / 2607.24207

FloAff-Kitchen: Bridging Navigation and Manipulation via Canonical and Progressive Floor Affordance Learning

FloAff-Kitchen:通过典范和渐进的地面可供性学习桥接导航与操控
Zhong, Ping, Teng, Manling, Wu, Tao, Chen, Bolei, Xia, Jiazhi, Wang, Jianxin
Abstract
Mobile manipulation requires robots to identify Floor Affordance (FloAff) that maximizes downstream manipulation success rather than merely ensuring navigation feasibility. FloAff prediction is a target-conditioned local spatial reasoning problem, yet existing methods suffer from representation ambiguity caused by irrelevant spatial context and arbitrary object orientations, while entangling shared and task-specific knowledge across heterogeneous manipulation skills. To address these challenges, we propose a unified framework for FloAff prediction from egocentric multimodal perception, consisting of canonical representation learning and progressive affordance prior learning. Specifically, we introduce a Canonical Floor Affordance Representation (CFAR), which learns canonical interaction geometry by preserving affordance-relevant local structure while eliminating nuisance spatial variations unrelated to robot base placement. We further propose Progressive Floor Affordance Learning (PFAL), which learns transferable FloAff priors from a foundation manipulation task and progressively adapts them to heterogeneous downstream manipulation skills. To facilitate systematic evaluation, we establish the first cross-scene, multi-view FloAff-Kitchen benchmark covering diverse manipulation skills, scene layouts, furniture styles, and viewpoints. Extensive experiments on three benchmark settings demonstrate that our method consistently outperforms strong baselines, while ablation studies validate the contribution of each proposed component. Project page: https://csu-hero-lab.github.io/FloAff-Kitchen_Web/
Chinese Translation
移动操控要求机器人识别最大化下游操控成功的地面可供性(Floor Affordance, FloAff),而不仅仅是确保导航的可行性。FloAff 预测是一个目标条件的局部空间推理问题,但现有方法受到无关空间上下文和任意物体方向引起的表征模糊的影响,同时在异质操控技能中纠缠共享和特定任务的知识。为了解决这些挑战,我们提出了一个统一的 FloAff 预测框架,该框架基于自我中心的多模态感知,包含典范表征学习和渐进的可供性先验学习。具体而言,我们引入了典范地面可供性表征(Canonical Floor Affordance Representation, CFAR),该表征通过保留与可供性相关的局部结构,同时消除与机器人基座位置无关的干扰空间变异,学习典范交互几何。我们进一步提出了渐进地面可供性学习(Progressive Floor Affordance Learning, PFAL),该方法从基础操控任务中学习可转移的 FloAff 先验,并逐步将其适应于异质下游操控技能。为了便于系统评估,我们建立了第一个跨场景、多视角的 FloAff-Kitchen 基准,涵盖多样的操控技能、场景布局、家具风格和视点。在三个基准设置上的大量实验表明,我们的方法始终优于强基线,而消融研究验证了每个提出组件的贡献。项目页面:https://csu-hero-lab.github.io/FloAff-Kitchen_Web/
cs.RO / 35 / 2607.24233

Quality-Adaptive Multi-UAV 3D Reconstruction with Sparse Workload Redistribution

质量自适应的多无人机三维重建与稀疏工作负载重分配
Sportich, Benjamin, Boubakri, Kenza, Simonin, Olivier, Renzaglia, Alessandro
Abstract
3D reconstruction of unknown environments is a key application in robotics but is severely limited by the computational and energy capabilities of current aerial platforms. Deploying multiple UAVs and providing efficient and scalable path planning strategies are common approaches, but effective online coordination among UAVs remains a significant challenge. To address this problem, we propose a quality-adaptive decentralized decision-making strategy to build a 3D map with user-defined degrees of fidelity. The approach integrates a quality-oriented criterion based on TSDF confidence into view generation and information gain estimation to produce viewpoints consistent with the desired fidelity target. Additionally, we employ two levels of coordination: a penalty factor in the viewpoint evaluation to encourage local dispersion among the UAVs and a global imbalance correction mechanism. The latter, based on regularized clustering and optimal task assignment, is only triggered when an unbalanced configuration relative to high-information regions is detected. Simulation results demonstrate that the proposed method improves path efficiency compared to state-of-the-art multi-UAV exploration approaches, while also achieving higher-fidelity reconstructions in terms of coverage and accuracy. We make our code publicly available to the community.
Chinese Translation
未知环境的三维重建是机器人技术中的一个关键应用,但受到当前空中平台计算和能量能力的严重限制。部署多个无人机并提供高效且可扩展的路径规划策略是常见的方法,但无人机之间的有效在线协调仍然是一个重大挑战。为了解决这个问题,我们提出了一种质量自适应的去中心化决策策略,以构建具有用户定义保真度的三维地图。该方法将基于TSDF(Truncated Signed Distance Function)置信度的质量导向标准整合到视图生成和信息增益估计中,以产生与所需保真度目标一致的视点。此外,我们采用了两级协调机制:在视点评估中引入惩罚因子以鼓励无人机之间的局部分散,以及一个全球不平衡校正机制。后者基于正则化聚类和最优任务分配,仅在检测到相对于高信息区域的不平衡配置时触发。仿真结果表明,所提方法在路径效率上优于最先进的多无人机探索方法,同时在覆盖率和准确性方面实现了更高保真的重建。我们将代码公开供社区使用。
cs.RO / 36 / 2607.24267

FeelWorld: Visuo-Tactile World Model for Hierarchical Contact Prediction and Planning

FeelWorld:用于层次接触预测和规划的视觉-触觉世界模型
Ma, Wenxuan, Zhang, Chaofan, Xue, Chao, Cai, Yinghao, Yao, Guocai, Cui, Shaowei, Wang, Shuo
Abstract
Humans plan physical interactions by imagining the possible outcomes of candidate actions. However, existing visual world models primarily capture appearance dynamics while overlooking the tactile states that govern contact-rich interactions, potentially producing imagined futures that appear visually plausible but violate physical dynamics. We introduce FeelWorld, a hierarchical visuo-tactile world model that jointly predicts future visual latents and three tactile states. FeelWorld organizes these states hierarchically as contact state, a 3D tactile latent that encodes force-related information, and slip state. These states are jointly predicted by a shared latent dynamics model with explicit supervision. To prevent irrelevant tactile signals during free-space motion from degrading visual prediction, we introduce a contact-gated asymmetric attention mechanism that maintains a visual-only prediction pathway before contact and enables joint visuo-tactile dynamics prediction during contact. The model is further trained with autoregressive rollouts and context noise injection to improve robustness to compounding errors. The predicted contact and slip states also support contact-aware CEM planning. Experiments on chip grasping, fruit grasping, and USB insertion show that FeelWorld reduces 10-step LPIPS from 0.084 to 0.058 and maintains an LPIPS that is 61% lower than that of the visual baseline after an 80-step autoregressive rollout. FeelWorld also achieves an average zero-shot planning success rate of 81.7%, providing an effective approach for incorporating tactile sensing into world models.
Chinese Translation
人类通过想象候选动作的可能结果来规划物理交互。然而,现有的视觉世界模型主要捕捉外观动态,而忽视了支配接触丰富交互的触觉状态,这可能导致想象的未来在视觉上看似合理但违反物理动态。我们提出了FeelWorld,一种层次化的视觉-触觉世界模型,能够共同预测未来的视觉潜变量和三种触觉状态。FeelWorld将这些状态层次化组织为接触状态、编码与力相关信息的三维触觉潜变量和滑动状态。这些状态由一个共享的潜变量动态模型共同预测,并进行显式监督。为了防止在自由空间运动中无关的触觉信号降低视觉预测的准确性,我们引入了一种接触门控的非对称注意机制,该机制在接触之前维持一个仅视觉的预测路径,并在接触期间启用联合的视觉-触觉动态预测。该模型还通过自回归展开和上下文噪声注入进行进一步训练,以提高对累积误差的鲁棒性。预测的接触和滑动状态也支持接触感知的CEM规划。在芯片抓取、水果抓取和USB插入的实验中,FeelWorld将10步LPIPS从0.084降低到0.058,并在80步自回归展开后保持比视觉基线低61%的LPIPS。FeelWorld还实现了81.7%的平均零-shot规划成功率,为将触觉感知纳入世界模型提供了一种有效的方法。
cs.RO / 37 / 2607.24292

Learning Adaptive Multi-Task Guidance, Navigation, and Control via Hypernetworks

通过超网络学习自适应多任务引导、导航与控制
Castan, Ricard Marsal I, Arora, Aman, Richard, Antoine, Orsula, Andrej, Pradalier, Cédric, Olivares-Méndez, Miguel A.
Abstract
Autonomous free-flying robots in orbital environments require controllers that are both versatile and resource-efficient, yet maintaining a separate, task-specific policy for each mission profile is architecturally brittle and limits operational flexibility as requirements evolve. We introduce HYPER-GNC, a multi-task reinforcement learning framework in which a hypernetwork maps physics-informed task embeddings to the weights of a shared actor-critic policy, enabling a single compact controller to master four distinct GNC tasks: velocity tracking, docking, inspection, and navigation with obstacle avoidance. The continuous embedding space allows the controller to generalize to novel mission configurations at deployment time without any retraining. Extensive experiments demonstrate that HYPER-GNC achieves sample efficiency comparable to single-task specialists while maintaining stability under significant inertial perturbations and external body wrenches. We further validate the framework on a physical satellite emulator, successfully bridging the simulation-to-reality gap across all mission profiles. Code, trained models, and deployment scripts are made publicly available to support reproducibility.
Chinese Translation
在轨道环境中,自主自由飞行机器人需要既灵活又资源高效的控制器,但为每个任务配置维护单独的特定策略在架构上是脆弱的,并限制了随着需求变化而演变的操作灵活性。我们提出了HYPER-GNC,一个多任务强化学习框架,其中超网络将物理信息任务嵌入映射到共享的演员-评论家策略的权重,使得一个紧凑的控制器能够掌握四个不同的引导、导航与控制(GNC)任务:速度跟踪、对接、检查和带障碍物避让的导航。连续的嵌入空间使得控制器能够在部署时对新任务配置进行泛化,而无需任何重新训练。大量实验表明,HYPER-GNC在样本效率上与单任务专家相当,同时在显著的惯性扰动和外部力矩下保持稳定。我们进一步在物理卫星仿真器上验证了该框架,成功地跨越了所有任务配置的仿真与现实之间的差距。代码、训练模型和部署脚本已公开,以支持可重复性。
cs.RO / 38 / 2607.24296

PAC-DP: PAC-Bayesian Diffusion Policy Learning

PAC-DP:PAC-贝叶斯扩散策略学习
Yeganegi, Mohammad Hasan, Yu, Dian, Del Prete, Andrea, Khadiv, Majid, Saveriano, Matteo
Abstract
Diffusion Policies (DPs) are able to perform complex manipulation tasks. However, DPs are typically trained by minimizing a denoising objective, which provides limited control over generalization in the finite-data regimes common in robotics. In this letter, we propose PAC-DP, an approach that increases the performance of DPs in robotic manipulation tasks. By modeling the DP as a Bayesian neural network, and defining a PAC-Bayes generalization bound, we derive a novel training objective that augments the standard denoising loss with a Kullback-Leibler divergence regularizer between the posterior and prior parameter distributions. From the theoretical perspective, our approach provides a principled approach to regularize the training of DPs without significantly increasing the training time. From the practical point of view, experimental results demonstrate improved denoising performance, lower variational negative log-likelihood, and higher success rates across multiple robotic manipulation benchmarks. Crucially, the largest improvements are observed in low-data training regimes and complex tasks, establishing PAC-DP as a theoretically grounded framework for robot policy learning.
Chinese Translation
扩散策略(Diffusion Policies, DPs)能够执行复杂的操作任务。然而,DPs 通常通过最小化去噪目标进行训练,这在机器人领域常见的有限数据环境中对泛化能力的控制有限。在本文中,我们提出了 PAC-DP,这是一种提高 DPs 在机器人操作任务中表现的方法。通过将 DP 建模为贝叶斯神经网络,并定义 PAC-Bayes 泛化界限,我们推导出一种新的训练目标,该目标通过在后验和先验参数分布之间引入 Kullback-Leibler 散度正则化项来增强标准去噪损失。从理论角度来看,我们的方法为在不显著增加训练时间的情况下对 DPs 的训练进行正则化提供了一种有原则的方法。从实践角度来看,实验结果表明,在多个机器人操作基准测试中,去噪性能得到改善,变分负对数似然降低,成功率提高。重要的是,在低数据训练环境和复杂任务中观察到最大的改进,确立了 PAC-DP 作为机器人策略学习的理论基础框架。
cs.RO / 39 / 2607.24317

A note on the motion representation and configuration update in time stepping schemes for the constrained rigid body

关于受约束刚体的时间步进方案中的运动表示和配置更新的说明
Müller, A.
Abstract
The dynamics of a holonomically constrained rigid body can be modeled by Newton-Euler equations subjected to geometric constraints. This is frequently formulated as a differential-algebraic equation (DAE) system of index 1. Inmultibody system (MBS) dynamics it is common (1) to numerically solve this system by means of integration schemes for ordinary differential equations, and (2) to treat the rigid body motion on the direct product Lie group SO (3)R3, although rigid body motions form the semidirect product Lie group SE (3). It is has been observed that the constraint satisfaction depends on which Lie group is used as configuration space (c-space). In this paper the problem is considered from a geometric perspective. It is shown that the constraints are exactly satisfied by a numerical integration scheme if they define a subgroup of the c-space. The subgroups of SE (3) have a significance for modeling mechanical systems, including lower kinematic (Reuleaux) pairs and are implicitly used in MBS modeling. It is concluded that SE (3) is the appropriate cspace for numerical DAE modeling of a constrained rigid body. This result does not immediately apply to MBS, however.
Chinese Translation
一个全约束刚体的动力学可以通过受几何约束的牛顿-欧拉方程进行建模。这通常被表述为一个指数为1的微分代数方程(DAE)系统。在多体系统(MBS)动力学中,通常(1)通过常微分方程的积分方案数值求解该系统,以及(2)将刚体运动视为直接乘积李群 SO(3)R3,尽管刚体运动形成半直接乘积李群 SE(3)。观察到约束的满足程度依赖于所使用的李群作为配置空间(c-space)。本文从几何的角度考虑了这一问题。研究表明,如果约束定义了 c-space 的一个子群,则数值积分方案可以精确满足这些约束。SE(3) 的子群在建模机械系统时具有重要意义,包括较低运动学(Reuleaux)对,并在 MBS 建模中隐含使用。最后得出结论,SE(3) 是受约束刚体的数值 DAE 建模的适当 c-space。然而,这一结果并不立即适用于 MBS。
cs.RO / 40 / 2607.24320

Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform

基于持续强化学习的自主赛车泛化研究:以RoboRacer平台为例
Siegert, Joel, Ghignone, Edoardo, Magno, Michele
Abstract
A key challenge in modern robotics is to adapt to changing environments, a challenge that is exacerbated when simulations cannot encompass every possible real-world configuration, and therefore Reinforcement Learning (RL) in the physical world becomes necessary. Continual Reinforcement Learning provides the tools to address this challenge; however, both the frameworks and the methods remain underexplored. Autonomous Racing and in particular the RoboRacer competition provide a testing ground for such methods, as learning to drive on a new track-floor combination with the least amount of new experience naturally frames a continual learning problem. This work tries to address this gap by proposing a continual RL framework based on Continual Backpropagation that is able, with only real-world data, to train a generalistic policy on a set of tracks and then fine- tune it within 15 minutes to outperform classical controllers. Furthermore, a comparison method based on offline RL is proposed, and a simulation analysis of the plasticity properties of the methods is conducted.
Chinese Translation
现代机器人技术面临的一个关键挑战是适应不断变化的环境,当模拟无法涵盖所有可能的现实世界配置时,这一挑战尤为突出,因此在物理世界中进行强化学习(Reinforcement Learning, RL)变得必要。持续强化学习提供了应对这一挑战的工具;然而,相关框架和方法仍然未得到充分探索。自主赛车,特别是RoboRacer竞赛,为此类方法提供了测试平台,因为在新的赛道-地面组合上以最少的新经验学习驾驶自然构成了一个持续学习问题。本研究试图通过提出一个基于持续反向传播(Continual Backpropagation)的持续强化学习框架来填补这一空白,该框架能够仅利用现实世界数据在一组赛道上训练出通用策略,并在15分钟内进行微调以超越经典控制器。此外,提出了一种基于离线强化学习的比较方法,并对这些方法的可塑性特性进行了模拟分析。
cs.RO / 41 / 2607.24369

Model Predictive Planner for UAV Navigation in Non-Convex Air Corridors

非凸空中走廊中无人机导航的模型预测规划器
Silva Jr., Henrique, Santos, Marcelo A., Raffo, Guilherme V.
Abstract
This work presents a motion planning framework for UAV navigation in non-convex urban air corridors. The planner is based on a mixed-integer tracking model predictive control formulation that enforces corridor feasibility and dynamic consistency within a single optimization problem. To guarantee convergence to the target and mitigate the occurrence of local minima induced by non-convex geometry, a shortest-path-based offset cost with feasibility constraints is embedded directly into the planning problem. Numerical simulations show that the proposed formulation generates dynamically valid trajectories that satisfy the corridor constraints and converge to the target without relying on external global planning stages.
Chinese Translation
本文提出了一种用于无人机在非凸城市空中走廊中导航的运动规划框架。该规划器基于混合整数跟踪模型预测控制(mixed-integer tracking model predictive control)形式,能够在单一优化问题中强制执行走廊的可行性和动态一致性。为了保证收敛到目标并减轻由非凸几何形状引起的局部最小值的出现,最短路径基础的偏移成本与可行性约束直接嵌入到规划问题中。数值仿真表明,所提出的形式生成了动态有效的轨迹,满足走廊约束并收敛到目标,而无需依赖外部全局规划阶段。
cs.RO / 42 / 2607.24481

ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm

ArmnetBench v0.1:低成本机械臂农场上操控策略的并行真实世界评估
Selvaraj, Praveen, Uttini, Lorenzo, Kuosmanen, Ville
Abstract
Real-world evaluation is a bottleneck in developing generalist robot manipulation policies. Each rollout requires physical hardware and an operator to set up, reset, and score it. We introduce ArmnetBench v0.1, a benchmark run on a fleet of low-cost SO-101 cells under light on-site supervision. v0.1 validates this arm farm end to end and compares 7 policies across 12 tasks with both single-arm and bimanual configurations. Each policy is trained or fine-tuned on 50 demonstrations per task; the benchmark contains 2,518 policy rollouts and 600 reference demonstrations. All 3,118 episodes carry a three-way label (successful, suboptimal, or failure). Policy rollouts are human-scored, while demonstrations are successful by construction. Beyond evaluation, its quality-labelled trajectories support downstream learning, from reward and predictive world models to policies trained on mixed-quality data. The leaderboard is an initial comparison under this shared budget. We release the 3,118 core episodes in LeRobot v3.0 and RoboMeter formats.
Chinese Translation
真实世界评估是开发通用机器人操控策略的瓶颈。每次实验都需要物理硬件和操作员进行设置、重置和评分。我们介绍了 ArmnetBench v0.1,这是一个在低成本 SO-101 机械臂群体上运行的基准测试,且在轻度现场监督下进行。v0.1 验证了该机械臂农场的端到端性能,并在 12 个任务中比较了 7 种策略,包括单臂和双臂配置。每种策略在每个任务上经过 50 次演示进行训练或微调;该基准测试包含 2,518 次策略实验和 600 次参考演示。所有 3,118 次实验均带有三种标签(成功、次优或失败)。策略实验由人工评分,而演示则是通过构建保证成功的。除了评估,其质量标记的轨迹还支持下游学习,从奖励和预测世界模型到在混合质量数据上训练的策略。排行榜是这一共享预算下的初步比较。我们以 LeRobot v3.0 和 RoboMeter 格式发布了 3,118 个核心实验。
cs.RO / 43 / 2607.24482

Distributed Coordination for Resilient Multi-UAV Remote Sensing: A Photovoltaic Inspection Case Study

面向韧性的多无人机遥感的分布式协调:光伏检测案例研究
GP-Lenza, Guillermo, Fernandez-Cortizas, Miguel, Molina, Martin, Campoy, Pascual
Abstract
Deploying multiple UAVs for remote sensing enables proportional reductions in mission time, but realizing these benefits requires the fleet to coordinate at runtime: distributing sensing targets, responding to platform failures, and recovering from degraded data quality. In inspection campaigns, where mission value depends on complete coverage and the usability of every capture, a centralized ground-station coordinator is a single point of failure: a lost link or station fault leaves sensing gaps that cannot be filled without operator intervention. We propose the \textbf{SwarmLink}, an inter-agent communication infrastructure that non-invasively extends any existing aerial framework with peer-to-peer coordination capability, without modifying the host system. We apply it to photovoltaic plant inspection as a representative large-scale sensing campaign, extending Aerostack2 with a distributed auction that unifies initial sensing-target allocation, platform-failure recovery, and data-quality-triggered reassignment into a single runtime mechanism. All three disruption scenarios reduce to the same re-auction over remaining targets and active platforms, requiring zero modifications to the Aerostack2 core and no ground-station involvement during the mission.
Chinese Translation
部署多个无人机进行遥感可以有效缩短任务时间,但实现这些收益需要机队在运行时进行协调:分配感知目标、应对平台故障以及恢复降级的数据质量。在检测活动中,任务的价值依赖于全面覆盖和每次捕获的可用性,集中式地面站协调器则成为单点故障:丢失的链接或站点故障会导致无法填补的感知空白,必须依赖操作员的干预。我们提出了 extbf{SwarmLink},一种代理间通信基础设施,它以非侵入的方式扩展任何现有的空中框架,具备点对点协调能力,而无需修改主机系统。我们将其应用于光伏电站检测,作为一个代表性的规模化感知活动,通过分布式拍卖将初始感知目标分配、平台故障恢复和数据质量触发的重新分配统一为一个单一的运行机制。这三种干扰场景都简化为对剩余目标和活跃平台的重新拍卖,要求对Aerostack2核心零修改,并且在任务执行期间无需地面站的参与。
cs.RO / 44 / 2607.24485

{\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision

{\tau}: 从未来视觉监督中学习触觉增强的视觉-语言-动作模型
Cheng, Ning, Xu, Jinan, Li, Wanlin, Chen, Yangzhi, Gao, Jing, Wang, Yiqun, Peng, Kelan, Han, Wenjuan
Abstract
Learning the informative tactile representation while effectively adapting it to pretrained Vision-Language-Action (VLA) models remains challenging at both the data and modeling levels. At the data level, limited task-specific demonstrations constrain representation quality, whereas large-scale pretraining incurs substantial costs. At the modeling level, existing methods either focus on instantaneous contact states or model temporal interaction dynamics using 6D wrench sequences, leaving high-dimensional tactile signals underexplored. To address these challenges, we present {\tau}, a touch-augmented VLA framework that learns an action-conditioned spatiotemporal tactile representation from future visual supervision inspired by the Joint-Embedding Predictive Architecture (JEPA), and fuses it with vision-language features for action generation under limited data. This supervision operates in latent space and is used only during training, adding no deployment overhead. We also introduce TacAura, a dataset of synchronized vision, proprioception, and vision-based tactile signals across four representative contact-rich manipulation tasks. Experiments show that {\tau} outperforms existing models and generalizes to unseen objects and scenes, delivering improved manipulation performance and robustness
Chinese Translation
在数据和建模层面上,有效地适应预训练的视觉-语言-动作(VLA)模型,同时学习信息丰富的触觉表示仍然具有挑战性。在数据层面,有限的任务特定演示限制了表示质量,而大规模预训练则产生了可观的成本。在建模层面,现有方法要么专注于瞬时接触状态,要么使用6D扭矩序列建模时间交互动态,从而使高维触觉信号未得到充分探索。为了解决这些挑战,我们提出了{\tau},一个触觉增强的VLA框架,它从未来视觉监督中学习一个基于动作条件的时空触觉表示,该框架受到联合嵌入预测架构(JEPA)的启发,并将其与视觉-语言特征融合,以在有限数据下生成动作。这种监督在潜在空间中操作,仅在训练期间使用,不增加部署开销。我们还引入了TacAura,一个跨四个代表性接触丰富操控任务的同步视觉、身体感知和基于视觉的触觉信号的数据集。实验表明,{\tau}在性能上优于现有模型,并能推广到未见过的物体和场景,提供了更好的操控性能和鲁棒性。
cs.RO / 45 / 2607.24493

KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation

KAI:一种运动学感知接口用于数据高效的关节物体操控
Li, Yaping, Zhaxizhuoma, Yu, Qiaojun, Zeng, Jia, Lin, Dahua, Pang, Jiangmiao
Abstract
Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects. By embedding interpretable geometric and kinematic priors into policy learning, KAI provides a strong inductive bias aligned with the underlying structure of articulated motion. This design effectively improves sample efficiency, with gains particularly pronounced in low-data regimes: across six simulation tasks, our method achieves an average success rate of 82.9%, matching or surpassing baseline performance while using only half the demonstration data. Our method also exhibits robust generalization to unseen backgrounds and visual distractors, transferring from a single clean training environment to cluttered real-world scenes. KAI's action-agnostic design further enables co-training with human interaction videos to enhance real-world robustness: under diverse visual distractions, our method with video co-training achieves over 70% average success rate.
Chinese Translation
关节物体操控需要理解运动学结构,而仅依靠机器人演示学习这一点既困难又昂贵。我们提出了运动学感知关节接口(KAI),这是一种结构化的中间表示,能够捕捉关节物体的运动学结构。通过将可解释的几何和运动学先验嵌入到策略学习中,KAI 提供了与关节运动的基本结构相一致的强归纳偏置。这一设计有效提高了样本效率,尤其在低数据环境下表现尤为显著:在六个仿真任务中,我们的方法实现了平均成功率82.9%,在仅使用一半演示数据的情况下,与基线性能相匹配或超越。我们的方法还表现出对未见背景和视觉干扰的强健泛化能力,能够从单一干净的训练环境迁移到杂乱的现实场景。KAI 的动作无关设计进一步支持与人类互动视频的共同训练,以增强现实世界的鲁棒性:在多样的视觉干扰下,我们的方法通过视频共同训练实现了超过70%的平均成功率。
cs.RO / 46 / 2607.24538

NEO: NeRF It Once, Edit It Many Times for Continuous Object Manipulation

NEO:一次生成 NeRF,多次编辑以实现连续物体操作
Zieliński, Mikołaj, Hall, David, Belter, Dominik, Moghadam, Peyman
Abstract
In this paper, we present NEO, a unified framework providing language-guided NeRF editing for robotic manipulation. Our paper introduces (i) a language-guided object removal that combines neural field resampling with multiview-consistent progressive inpainting, (ii) a direct NeRF weight editing method utilizing knowledge distillation, composing original and edited NeRFs via a teacher-student model, enabling coherent modeling of future scene states before a robot executes an action, and (iii) the first benchmark (NEO-Dataset) for quantitatively evaluating NeRF scene editing methods suitable for robot manipulation. We show that our approach outperforms state-of-the-art baselines in scene editing tasks, including object removal and pick-and-place robotic experiments, yielding visually coherent and geometrically consistent edits that reduce artifacts commonly introduced by prior methods.
Chinese Translation
在本文中,我们提出了 NEO,一个统一框架,提供基于语言指导的 NeRF 编辑以用于机器人操作。我们的论文介绍了 (i) 一种结合神经场重采样与多视图一致的渐进式修补的语言指导物体移除方法,(ii) 一种利用知识蒸馏的直接 NeRF 权重编辑方法,通过教师-学生模型组合原始和编辑后的 NeRF,使得在机器人执行动作之前能够连贯地建模未来场景状态,以及 (iii) 第一个基准数据集 (NEO-Dataset),用于定量评估适合机器人操作的 NeRF 场景编辑方法。我们展示了我们的方法在场景编辑任务中优于最先进的基线,包括物体移除和抓取-放置的机器人实验,产生视觉上连贯且几何上一致的编辑,减少了先前方法常引入的伪影。
cs.RO / 47 / 2607.24629

Development of a Handheld Actuation Mechanism for a Tendon-driven Robotically Steered Guidewire

一种用于腱驱动机器人引导导丝的手持驱动机制的开发
Chavenet, Saima, Brumfiel, Timothy A., Konda, Revanth, Desai, Jaydev P.
Abstract
An endovascular intervention begins with a skilled clinician manually navigating a long, slender wire, called a guidewire, to the target location within the vasculature. Due to factors, such as vessel tortuosity and lack of steerability at the guidewire tip, manual navigation of a guidewire could be challenging, potentially resulting in vessel damage, perforation, dissection, and occlusion, as well as postsurgical complications, such as thrombosis. This work details the development of a handheld actuation mechanism for tendon-driven robotic guidewires by utilizing the aforementioned spooling mechanism. Incorporating the compact spooling mechanism into a handheld device enables clinicians to leverage both motorized guidewire steering and manual rotation and translation, if desired. The proposed device is designed such that, while holding the device, the clinician can execute feeding, rotation, and bending motions for up to 1.5 m of the robotically steerable guidewire using a joystick and momentary switch on the handle. Furthermore, the proposed device is demonstrated through navigation in an anatomically accurate phantom aorta model.
Chinese Translation
血管内干预始于一名熟练的临床医生手动导航一根称为导丝的细长线材,直至血管内的目标位置。由于血管的扭曲性和导丝尖端缺乏可操控性等因素,手动导航导丝可能会面临挑战,可能导致血管损伤、穿孔、剥离和堵塞,以及术后并发症,如血栓形成。本研究详细介绍了一种用于腱驱动机器人导丝的手持驱动机制的开发,利用上述卷绕机制。将紧凑的卷绕机制整合到手持设备中,使临床医生能够在需要时同时利用电动导丝的操控和手动旋转与移动。所提出的设备设计使得在握持设备的同时,临床医生可以通过手柄上的操纵杆和瞬时开关执行导丝的进给、旋转和弯曲动作,操作长度可达1.5米。此外,所提出的设备通过在解剖准确的幻影主动脉模型中的导航进行了演示。
cs.RO / 48 / 2607.24744

Data Pyramid for Embodied Manipulation

用于具身操作的数据金字塔
Ye, Yifan, Fu, Yankai, Lv, Yaoxu, Hou, Bohan, Cen, Jun, Kong, Lingdong, Zheng, Duo, Chen, Tianxing, Liu, Jiaming, Cao, Ziang, Lou, Yunfan, Chow, Wei, Sun, Xian, Wang, Yingshuo, Ge, Kuangzhi, Chi, Xiaowei, Zhang, Xidong, Pang, Zhibo, Zhong, Yiwu, Han, Sirui, Lu, Zhihe, Yuan, Weihao, Chen, Qifeng, Wang, Michael Yu, Mu, Yao, Liu, Ziwei, Yang, Jianfei, Luo, Ping, Zhang, Shanghang
Abstract
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.
Chinese Translation
多模态基础模型通过消耗整个互联网学习了视觉和语言。具身代理没有这样的捷径,因为它们需要将观察与物理状态和动作结合的数据。这些信号可以由多种数据源提供,程度各异。在本研究中,我们将具身数据生态系统组织为一个涵盖五个互补来源的“金字塔”:真实机器人数据、UMI风格数据、自我中心和外部中心数据、仿真数据以及通用视觉-语言数据。我们围绕可扩展性与机器人对齐之间的张力来组织金字塔,并进一步根据数据质量、多样性、可重用性和物理真实性对每个来源进行特征描述。然后,我们通过数据配方的视角分析近期的具身基础模型,考察在预训练过程中如何选择、对齐和混合不同的数据源。对于具身大脑模型、视觉-语言-动作模型和世界-动作模型,我们将数据组成与感知、推理、规划、动作生成和世界预测的能力相关联。最后,我们讨论了六个开放挑战:构建大规模触觉数据集、收集失败与恢复数据、开发可扩展的数据收集管道、在不同具身之间对齐动作、利用自我中心数据进行灵巧操作,以及为机器人学习设计原则性的数据配方。我们希望这项工作为下一代具身系统的设计奠定基础。
计算机视觉 (Computer Vision)
187
cs.CV / 1 / 2607.22687

DINOv3-MIL: Per-Kidney Multi-Label Tumour and Cyst Detection from Foundation-Model Patch Tokens on KiTS23

DINOv3-MIL:基于基础模型补丁标记的每肾脏多标签肿瘤和囊肿检测(KiTS23)
M, Vishalakshi, Sharma, Sahil, P, Pramod Kumar
Abstract
Foundation vision models trained on natural images transfer to medical tasks without domain pre-training, but volumetric classification requires aggregating tens of thousands of patch tokens per study, and the aggregator constrains how the resulting model can be interpreted. We compare three aggregators on identical frozen DINOv3 ViT-H/16+ features for renal tumour/cyst detection on KiTS23 (966 kidneys; n=97 test): a CLS-token linear probe, gated attention multiple instance learning (MIL) over 55,296 patch tokens, and a prototype head following ProtoViT. Attention MIL achieves the highest AUROC for tumour (0.74, 95% CI 0.64-0.83) and cyst (0.80, 0.70-0.88), with attention enriched 7.5-9.8x over chance within annotated lesions. The prototype head does not transfer to cyst detection (AUROC 0.51), exposing an interpretability-performance trade-off at this token scale.
Chinese Translation
在没有领域预训练的情况下,基于自然图像训练的基础视觉模型可以迁移到医学任务,但体积分类需要对每个研究聚合数万个补丁标记,而聚合器限制了结果模型的可解释性。我们比较了三种聚合器在相同的冻结 DINOv3 ViT-H/16+ 特征上进行肾脏肿瘤/囊肿检测的效果,数据集为 KiTS23(966 个肾脏;n=97 测试):一个 CLS-token 线性探测器,基于 55,296 个补丁标记的门控注意力多实例学习(MIL),以及遵循 ProtoViT 的原型头。注意力 MIL 在肿瘤(AUROC 0.74,95% CI 0.64-0.83)和囊肿(AUROC 0.80,0.70-0.88)检测中取得了最高的 AUROC,在标注病灶内的注意力增强达 7.5-9.8 倍。原型头在囊肿检测中未能转移(AUROC 0.51),暴露了在此标记规模下可解释性与性能之间的权衡。
cs.CV / 2 / 2607.22696

MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion

MegaSlide-DiT:面向高效视频扩散的内存中心适应与可变形局部注意力
Liu, Jiacheng, Liu, Jason
Abstract
High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a single workstation. A 100 billion-plus parameter DiT easily requires over a terabyte of persistent state, while naive spatiotemporal self-attention grows quadratically in sequence length. These two walls -- parameter memory and activation memory -- prevent researchers from adapting massive generative models without large GPU clusters. We revisit this problem from a systems perspective and introduce MegaSlide-DiT, a prototype that demonstrates how a pre-trained 105B DiT can be adapted on a single H200 GPU with 1.5 TB of host RAM. Our key insight is that the GPU need not own the model state: all persistent weights, master weights and optimizer moments remain in host memory, while only transient shards are streamed to the GPU on demand. Simultaneously, we replace quadratic global attention with 3D Deformable Slide Attention (3D-DSA), a motion-adaptive local attention operator that reduces both memory and computational complexity to linear in the sequence length. We report detailed memory accounting, execution traces and evaluation results to substantiate our design. MegaSlide-DiT does not claim to train a 105B model from scratch on a single GPU, nor does it magically solve bandwidth limits; rather, it offers a pragmatic path for full-parameter adaptation of massive video diffusion models on high-end workstations.
Chinese Translation
基于扩散变换器(Diffusion Transformers, DiTs)的高分辨率视频扩散模型在保真度方面表现优异,但很快就会耗尽单个工作站的内存预算。一个超过1000亿参数的DiT模型轻松需要超过一TB的持久状态,而简单的时空自注意力在序列长度上呈平方增长。这两个瓶颈——参数内存和激活内存——阻碍了研究人员在没有大型GPU集群的情况下适应庞大的生成模型。我们从系统的角度重新审视这个问题,并引入了MegaSlide-DiT,一个原型展示了如何在单个H200 GPU上适应一个预训练的105B DiT,配备1.5 TB的主机内存。我们的关键见解是,GPU无需拥有模型状态:所有持久权重、主权重和优化器时刻都保留在主机内存中,而只有瞬态分片在需要时流式传输到GPU。同时,我们用3D可变形滑动注意力(3D Deformable Slide Attention, 3D-DSA)替代了平方复杂度的全局注意力,这是一种运动自适应的局部注意力操作符,将内存和计算复杂度降低到与序列长度线性相关。我们报告了详细的内存核算、执行跟踪和评估结果,以证实我们的设计。MegaSlide-DiT并不声称在单个GPU上从零开始训练一个105B模型,也没有神奇地解决带宽限制;相反,它为在高端工作站上全参数适应庞大的视频扩散模型提供了一条务实的路径。
cs.CV / 3 / 2607.22698

FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog

FogDrive:一种用于分级雾霾下感知的多模态合成驾驶数据集
Panwar, Vansh
Abstract
Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal alignments needed to evaluate robust sensor fusion. Real-world weather datasets suffer from uncontrolled collection and single-level, uncalibrated conditions, while synthetic alternatives either target camera-only restoration or lack the paired clean-and-foggy structure needed to benchmark "defog-then-detect" pipelines. We present FogDrive, a rigorously calibrated, multi-modal autonomous-driving dataset bridging data-centric engineering and robust machine learning. Built with the CARLA simulator, FogDrive contains 660 scenes (~133k fully annotated frames, 50:50 day/night) across four synchronized cameras (RGB, depth, semantic segmentation), a LiDAR and semantic-LiDAR pair, and front radar. Physically consistent fog is modeled independently on camera channels (Koschmieder model) and LiDAR channels (Beer-Lambert law) at three calibrated visibility densities (160m, 100m, 50m). Every scene ships in four matched variants (clean plus three graded fog levels) with cross-calibrated 2D and 3D bounding boxes. A semantic-segmentation-based quality audit over 8k images validates annotations at 95.1% precision and over 99% recall for vehicles within 40m. We establish baseline benchmarks with state-of-the-art architectures (TransFusion, BEVFusion, YOLOv8-m) across two paradigms: 3D multi-modal fusion and 2D image restoration. These yield critical data-centric insights: mixing multi-density fog during training tightens 3D bounding-box geometry without added data-scaling cost, while in 2D pipelines image-quality metrics (PSNR, SSIM) prove poor predictors of downstream detection performance. FogDrive will be fully open-sourced alongside our data-generation framework to accelerate robust, multi-modal research.
Chinese Translation
在恶劣天气条件下的感知仍然是可靠自主驾驶的一个关键瓶颈,然而现有基准缺乏评估稳健传感器融合所需的系统性多模态对齐。现实世界的天气数据集受到不受控的收集和单一层次、未校准条件的影响,而合成替代品要么仅针对相机恢复,要么缺乏用于基准“去雾后检测”的配对清晰与雾霾结构。我们提出了FogDrive,这是一个经过严格校准的多模态自主驾驶数据集,旨在连接数据驱动工程与稳健机器学习。FogDrive基于CARLA模拟器构建,包含660个场景(约13.3万帧完全标注,白天/夜晚各占50%),涵盖四个同步摄像头(RGB、深度、语义分割)、一个激光雷达和一个语义激光雷达对,以及前向雷达。物理一致的雾霾在摄像头通道(Koschmieder模型)和激光雷达通道(Beer-Lambert定律)上独立建模,具有三种校准的能见度密度(160米、100米、50米)。每个场景提供四个匹配变体(清晰加上三个分级雾霾水平),并配有交叉校准的2D和3D边界框。基于语义分割的质量审计覆盖8000张图像,验证了在40米内车辆标注的95.1%精度和超过99%的召回率。我们在两个范式下使用最先进的架构(TransFusion、BEVFusion、YOLOv8-m)建立基准基线:3D多模态融合和2D图像恢复。这些结果提供了关键的数据驱动见解:在训练过程中混合多密度雾霾可以在不增加数据缩放成本的情况下收紧3D边界框几何,而在2D管道中,图像质量指标(PSNR、SSIM)被证明是下游检测性能的差劣预测因子。FogDrive将与我们的数据生成框架一起完全开源,以加速稳健的多模态研究。
cs.CV / 4 / 2607.22702

MIME: Multimodal Interactive Motion Encoder

MIME:多模态互动运动编码器
Zucek, Addison, Gupta, Prerit, Kuatova, Kamila, Bera, Aniket
Abstract
Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the relationships between actors. We introduce the Multimodal Interactive Motion Encoder (MIME), which, to our knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion. MIME captures individual and shared structure using stream based co-attention with explicit interaction features and curriculum based contrastive training. On Inter-X text-motion retrieval, MIME consistently outperforms early and late fusion baselines across gallery sizes, achieving a 12.8% relative improvement in text-to-motion R@1 at a 2,000-sample gallery. We further evaluate MIME as a frozen auxiliary prior within TIMotion and InterMask on the unseen InterHuman dataset. MIME improves semantic alignment metrics while maintaining comparable FID in TIMotion. These results show that interaction aware multimodal encoding improves multi person motion retrieval and transfers across datasets to support downstream motion generation.
Chinese Translation
文本-运动表示学习迅速发展,越来越多的研究关注于动画、增强现实/虚拟现实(AR/VR)和具身人工智能中的多人物交互。这些场景需要将语言与个体演员动态及演员之间关系对齐的表示。我们介绍了多模态互动运动编码器(MIME),据我们所知,这是第一个专门为两人互动运动设计的多模态编码器。MIME通过基于流的共同注意机制捕捉个体和共享结构,并结合显式交互特征和基于课程的对比训练。在Inter-X文本-运动检索任务中,MIME在不同画廊规模下始终优于早期和晚期融合基线,在2,000个样本的画廊中实现了文本到运动R@1的12.8%相对提升。我们进一步在未见的InterHuman数据集上,将MIME作为冻结的辅助先验应用于TIMotion和InterMask。MIME在保持TIMotion中相似FID的同时,提高了语义对齐指标。这些结果表明,关注交互的多模态编码改善了多人物运动检索,并在数据集之间迁移,以支持下游运动生成。
cs.CV / 5 / 2607.22703

Histopathological Spectrum-Guided Prostate Stratification via Segmentation-Assisted Diagnostic Transformer

基于组织病理学谱的前列腺分层通过分割辅助诊断变换器
Li, Leyang, Chen, Lihua, Hu, Huangang, Hao, Tianhang, Cheng, Hao, Zhang, Xin, Sun, Qianru, Lu, Bingxu, Yu, Wenlong, Duan, Feng
Abstract
Prostate cancer diagnosis with multiparametric MRI (mpMRI) is commonly based on PI-RADS assessment or binary classification, which suffer from subjectivity and fail to capture clinically relevant pathological heterogeneity. To address this limitation, we construct a Prostate Cancer Histopathology Spectrum Dataset (PCa-HSD) and formulate a clinically meaningful four-class classification task, addressing the underrepresentation of benign lesions that are easily confounded with prostate cancer in existing datasets. We propose Language-guided Segmentation-assisted Diagnostic Transformer model (LSDT), which leverages zero-shot segmentation to provide anatomical priors and performs effective multi-modal slice fusion for classification. Our proposed method consistently improves accuracy across backbones, achieving the best average accuracy of 0.633 and JointRecall of 0.768 in five-fold cross-validation on a cohort of 344 patients. These results demonstrate that integrating pathology supervision and anatomical priors significantly enhances fine-grained prostate MRI classification and provides a more clinically relevant paradigm for risk stratification. Code will be made publicly available in a future revision.
Chinese Translation
前列腺癌的诊断通常依赖于多参数磁共振成像(mpMRI)的PI-RADS评估或二元分类,这些方法存在主观性,并未能捕捉到临床相关的病理异质性。为了解决这一局限性,我们构建了前列腺癌组织病理学谱数据集(PCa-HSD),并制定了一个具有临床意义的四类分类任务,以解决现有数据集中良性病变的代表性不足,这些良性病变易与前列腺癌混淆。我们提出了一种语言引导的分割辅助诊断变换器模型(LSDT),该模型利用零样本分割提供解剖先验,并进行有效的多模态切片融合以进行分类。我们提出的方法在不同的基础模型上持续提高了准确性,在344名患者的五折交叉验证中,达到了最佳平均准确率0.633和联合召回率0.768。这些结果表明,整合病理监督和解剖先验显著增强了细粒度的前列腺MRI分类,并为风险分层提供了更具临床相关性的范式。代码将在未来的修订中公开。
cs.CV / 6 / 2607.22704

Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model

托卡马克偏滤器中中性粒子发射层析成像的可见光成像诊断:一种高效的基于变压器的替代模型
Wang, Xiao, Si, Hao, Chen, Qiang, Zhang, Yu-Xiang, Zhang, Beihe, Yang, Jianhua, Yang, Qingquan, Sun, Dengdi, Lyu, Wanli, Xu, Guosheng, Tang, Jin
Abstract
Nuclear fusion has made significant progress in recent years and is expected to become one of the most important pathways to addressing global energy challenges. This paper focuses on observing plasma using visible-light cameras, analyzing its spatio-temporal motion cues, and predicting the two-dimensional spatial distribution of light intensity, aiming to provide a foundational basis for future scientific experiments using deep neural networks. Specifically, we propose Delta-InvFormer, a novel backbone network centered on a differential Transformer. The key insight is that by taking consecutive video frames as input, we can better capture the dynamics of the plasma. Moreover, spatial and temporal differential self-attention effectively mitigates interference from noisy signals, ensuring high-quality feature extraction. These features are then fused into a compact and informative representation, which is fed into a decoder network to predict the distribution. Based on real experimental data collected from the Experimental Advanced Superconducting Tokamak (EAST) large-scale scientific facility, our results demonstrate that the proposed model not only significantly accelerates traditional methods for distribution prediction but also achieves competitive reconstruction accuracy. The source code of this paper will be released on https://github.com/Event-AHU/OpenFusion
Chinese Translation
近年来,核聚变取得了显著进展,预计将成为应对全球能源挑战的重要途径之一。本文聚焦于使用可见光相机观察等离子体,分析其时空运动线索,并预测光强的二维空间分布,旨在为未来使用深度神经网络的科学实验提供基础依据。具体而言,我们提出了Delta-InvFormer,这是一种以差分变压器为核心的新型骨干网络。关键的见解在于,通过将连续的视频帧作为输入,我们可以更好地捕捉等离子体的动态。此外,空间和时间的差分自注意力有效减轻了噪声信号的干扰,确保了高质量的特征提取。这些特征随后被融合成紧凑且信息丰富的表示,输入到解码器网络中以预测分布。基于从实验先进超导托卡马克(EAST)大型科学设施收集的真实实验数据,我们的结果表明,所提出的模型不仅显著加速了传统的分布预测方法,还实现了具有竞争力的重建精度。本文的源代码将发布在 https://github.com/Event-AHU/OpenFusion
cs.CV / 7 / 2607.22705

EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations

EditCLEVR:一个用于对象中心表示组合忠实性的配对场景干预基准
Karnam, Anuraag Gadehothur, Sathish, Tarunesh
Abstract
Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations. Existing evaluations usually score segmentation, single-image factor prediction, or downstream accuracy, but these tests do not directly ask whether a per-object representation behaves correctly under a controlled semantic edit. We introduce EditCLEVR, a paired-scene intervention benchmark in which each example contains a before/after pair of CLEVR-style renders with the same object indices and scene layout, and either exactly one known attribute change on one known object or a no-edit re-render for drift measurement. The protocol includes probe-free diagnostics for representation-change localization and stability, together with probe-decoded semantic faithfulness metrics that test whether the predicted scene change matches the intended intervention across in-distribution and compositional out-of-distribution (OOD) suites, allowing code-space movement and decoded object-attribute correctness to be evaluated separately. We introduce the semantic metric Scene-Graph Intervention Accuracy (SGIA), which requires the full after-scene prediction to be correct and the only predicted before-to-after semantic change to be the intended object-factor edit. We also establish Delta-SGIA as a companion diagnostic that checks the single-site change pattern without requiring the full after-scene graph to be correct. Baseline evaluations on ground-truth-mask backbones, learned-slot models, SAM 2 + frozen-ViT models, and one mask-feature hybrid indicate that CoGenT-OOD-core degradation can persist under ground-truth instance masks, that mask source accounts for part but not all of native performance, and that locality or stability alone can overstate semantic faithfulness. Code is available at https://github.com/torux-bughunter/EditCLEVR.
Chinese Translation
对象中心学习旨在将场景表示为可以在新组合中重用其属性的对象。现有的评估通常评分于分割、单图像因子预测或下游准确性,但这些测试并没有直接询问每个对象的表示在受控语义编辑下是否表现正确。我们引入了EditCLEVR,一个配对场景干预基准,其中每个示例包含一对前后CLEVR风格的渲染图,这些渲染具有相同的对象索引和场景布局,并且在一个已知对象上有恰好一个已知属性的变化,或者进行无编辑的重新渲染以测量漂移。该协议包括无探针的表示变化定位和稳定性诊断,以及探针解码的语义忠实性指标,测试预测的场景变化是否与预期的干预相匹配,涵盖了分布内和组合的分布外(OOD)套件,允许代码空间的移动和解码的对象属性正确性被单独评估。我们引入了语义指标场景图干预准确性(SGIA),要求完整的后场景预测必须正确,并且唯一预测的前后语义变化必须是预期的对象因子编辑。我们还建立了Delta-SGIA作为一个伴随诊断,检查单点变化模式而不要求完整的后场景图是正确的。在真实掩膜骨干、学习槽模型、SAM 2 + 冻结的ViT模型和一个掩膜特征混合模型上的基线评估表明,CoGenT-OOD-core降解在真实实例掩膜下仍然存在,掩膜源部分但不是全部解释了原生性能,并且仅依赖局部性或稳定性可能会夸大语义忠实性。代码可在 https://github.com/torux-bughunter/EditCLEVR 获取。
cs.CV / 8 / 2607.22707

LowAux-RDNet: Low-Pass Residual Supervision with Scene-Balanced Real-World Training for Single-Image Reflection Removal

LowAux-RDNet:基于场景平衡的真实世界训练的低通残差监督用于单幅图像反射去除
Li, Jizhong
Abstract
Single-image reflection removal aims to recover a clean transmission layer from one image captured through glass. We study an explicit decomposition pipeline built on RDNet and introduce LowAux, a training-only low-pass reflection auxiliary objective. The original residual target remains the main reflection supervision, while symmetrically filtered prediction and target provide a stable low-frequency constraint. We further incorporate scene-balanced real pairs from RRW to broaden real-scene coverage and improve cross-dataset generalization. To avoid evaluation discrepancies caused by model-specific resizing, padding, output quantization, and metric code, we build a unified public benchmark over CEILNet, Real20, Postcard, Objects, and Wild. Under the same evaluator, the proposed system obtains a five-dataset macro average of 27.546 dB PSNR, 0.9220 SSIM, 0.9751 NCC, and 0.004760 LMSE, achieving the highest macro-average PSNR, SSIM, and NCC and the lowest LMSE among the compared public checkpoints and internal variants. Per-dataset and qualitative analyses show that the main benefit is a more balanced performance across diverse reflection distributions, while clear semantic reflections in Postcard remain challenging.
Chinese Translation
单幅图像反射去除旨在从通过玻璃拍摄的一张图像中恢复干净的透射层。我们研究了一个基于RDNet的显式分解管道,并引入了LowAux,这是一种仅用于训练的低通反射辅助目标。原始的残差目标仍然是主要的反射监督,而对称滤波的预测和目标提供了一个稳定的低频约束。我们进一步结合了来自RRW的场景平衡真实对,以扩大真实场景的覆盖范围并改善跨数据集的泛化能力。为了避免因模型特定的调整大小、填充、输出量化和度量代码导致的评估差异,我们在CEILNet、Real20、Postcard、Objects和Wild上建立了一个统一的公共基准。在相同的评估器下,所提出的系统在五个数据集上获得了27.546 dB的宏平均PSNR、0.9220的SSIM、0.9751的NCC和0.004760的LMSE,达到了比较的公共检查点和内部变体中最高的宏平均PSNR、SSIM和NCC以及最低的LMSE。每个数据集的定量和定性分析表明,主要的好处是对不同反射分布的更平衡性能,而Postcard中的清晰语义反射仍然具有挑战性。
cs.CV / 9 / 2607.22708

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

StepX-Edge:通过架构-训练-部署协同设计实现的设备端用户界面视觉-语言模型
Wang, Yin, Hu, Haotian, Han, Jineng, Qiu, Wentao, Ge, Zhenhua, Tang, Liujian, Wang, Fanyi
Abstract
Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen understanding, visual question answering, and element grounding; on the other is the strict compute, memory, and power budget of mobile chips. Existing work either trades one for the other, or stops at simulation without real-device validation. We present StepX-Edge, a 0.9B-parameter on-device UI vision-language model that resolves this tension through three-layer co-design of architecture, training, and deployment. Architecturally, UI-aware Layered Visual Encoding (ULVE) and a Progressive Dimensionality Projection (PDP) connector target the extreme aspect ratios and fine-grained perception of screens, while standard full attention throughout ensures native compatibility with mainstream mobile NPU operators. For training, the five-stage StepX-Curriculum framework is designed around our observation of mutual-promotion effects among UI subtasks, so that all four capabilities grow synergistically under a tight parameter budget rather than interfering. For deployment, a module-wise differentiated two-stage PTQ-to-QAT quantization scheme keeps the post-quantization accuracy loss within 1%. StepX-Edge achieves the strongest overall UI understanding among <=1B models, surpassing all 2B-2.3B baselines on ScreenQA (88.76 F1) and Chinese OCRBench v2 (57.25), and matching 1.3B-2.3B general VLMs on RefCOCO (92.0%) and OCRBench v1 (831) with far fewer parameters. After W4A16+KV8 quantization, the model runs stably on Snapdragon 8 Gen5 devices with ~0.84 s TTFT, 98 tok/s decode, and 1.4 GB peak memory. We will open-source the training data, the full training recipe, and the quantization deployment pipeline.
Chinese Translation
在终端设备上部署具备全面用户界面理解的视觉-语言模型长期以来面临准确性与效率之间的困境:一方面是光学字符识别(OCR)、屏幕理解、视觉问答和元素定位的准确性要求;另一方面是移动芯片在计算、内存和功耗方面的严格限制。现有工作要么在两者之间进行权衡,要么停留在模拟阶段而未进行真实设备验证。我们提出了StepX-Edge,这是一个拥有9亿参数的设备端用户界面视觉-语言模型,通过架构、训练和部署的三层协同设计解决了这一矛盾。在架构方面,用户界面感知的分层视觉编码(ULVE)和渐进维度投影(PDP)连接器针对极端长宽比和屏幕的细粒度感知,而标准的全注意力机制确保与主流移动NPU运算符的原生兼容性。在训练方面,五阶段的StepX-Curriculum框架围绕我们对用户界面子任务之间相互促进效应的观察进行设计,使得所有四种能力在紧凑的参数预算下协同增长,而不是相互干扰。在部署方面,模块化差异化的两阶段PTQ到QAT量化方案使得量化后的准确性损失保持在1%以内。StepX-Edge在<=10亿参数的模型中实现了最强的整体用户界面理解,超越了所有2亿-2.3亿基线模型,在ScreenQA(88.76 F1)和中文OCRBench v2(57.25)上表现优异,并在RefCOCO(92.0%)和OCRBench v1(831)上与1.3亿-2.3亿的通用视觉-语言模型相匹配,同时参数量远少于后者。在W4A16+KV8量化后,该模型在Snapdragon 8 Gen5设备上稳定运行,TTFT约为0.84秒,解码速度为98个token/s,峰值内存为1.4 GB。我们将开源训练数据、完整的训练方案和量化部署流程。
cs.CV / 10 / 2607.22709

RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus

RMS@CC-MMD 2026:通过几何交互和多视角共识进行多模态厌女检测
Hossain, Md. Ajwad
Abstract
The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Memes often rely on a semantic clash between visual and textual modalities, where hateful intent is implicit and culturally grounded. This paper presents GeoMVC (Geometric Interaction and Multi-View Consensus), developed for the CC-MMD Grand Challenge at ICMI 2026. To address the limitations of static feature concatenation, a Geometric Interaction Layer is proposed that models cross-modal alignment via Hadamard products and cosine similarity between frozen visual and textual embeddings. We further mitigate distribution shifts caused by noisy OCR and code-mixed transliteration through a Multi-View Consensus strategy, aggregating predictions across raw, length-filtered, and English-translated text views. The system achieved Rank 2 in the Malayalam partition (Macro F1: 0.892) and Rank 3 in the Chinese partition (Macro F1: 0.895) on Task A, while securing Rank 5 in the Tamil partition (Macro F1: 0.521). A detailed error analysis on the development partition highlights open challenges in modeling localized transliteration and code-mixed sarcasm across Dravidian and Chinese cultural contexts.
Chinese Translation
互联网迷因的传播为自动内容审核引入了新的复杂性,特别是在检测厌女情绪方面。迷因通常依赖于视觉和文本模态之间的语义冲突,其中仇恨意图是隐含的并且根植于文化背景。本文提出了GeoMVC(几何交互和多视角共识),该模型是为ICMI 2026的CC-MMD大挑战开发的。为了克服静态特征连接的局限性,提出了一种几何交互层,通过Hadamard积和冻结的视觉与文本嵌入之间的余弦相似度来建模跨模态对齐。我们进一步通过多视角共识策略减轻由噪声OCR和混合转写引起的分布偏移,聚合原始、长度过滤和英语翻译文本视图的预测。该系统在马拉雅拉姆语分区中获得了第2名(宏F1: 0.892),在中文分区中获得了第3名(宏F1: 0.895),而在泰米尔语分区中获得了第5名(宏F1: 0.521)。对开发分区的详细错误分析突显了在德拉威文化和中文文化背景下建模本地化转写和混合讽刺的开放挑战。
cs.CV / 11 / 2607.22712

scMIR: a vision-language foundation model for single-cell light microscopy image representation

scMIR:一种用于单细胞光学显微镜图像表示的视觉-语言基础模型
Shang, Yifan, Tan, Jiahui, Zeng, Xiangxiang, Zhou, Renjie
Abstract
Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis. Existing representation learning methods mostly rely on task-oriented modeling, which is limited by specific datasets and predefined tasks, making them difficult to generalize across different cell types and microscopy modalities, and experimental conditions. Although general-purpose methods have improved the generalization ability of image representation in recent years, their limited utilization of experimental background and biological context information still poses challenges in complex phenotypic analysis. Here, we propose scMIR, a vision-language foundation model for single-cell light microscopy image representation. By synergistically combining self-supervised image reconstruction with text-guided cross-modal alignment, scMIR can simultaneously encode morphological and biological semantic information in a unified representation space. scMIR is pre-trained on 207,957 image-text pairs, covering various cell types, microscopy modalities, and perturbation conditions. scMIR outperforms existing general models and task-oriented methods as systematically evaluated on various complex tasks using 16 benchmark datasets, including cell classification, clustering, phenotype inference, and batch effect correction tasks. Furthermore, scMIR shows a strong generalization ability across various tasks without requiring task-specific fine-tuning. With its unique advantages, we envision scMIR may promote the standardization and automation of high-throughput phenotyping workflows through supporting various downstream analysis tasks.
Chinese Translation
单细胞光学显微镜图像已成为表征细胞表型的重要数据来源,但其复杂性和异质性给高通量自动分析带来了挑战。现有的表示学习方法大多依赖于任务导向建模,这受到特定数据集和预定义任务的限制,使其难以在不同细胞类型、显微镜模式和实验条件下进行推广。尽管通用方法近年来提高了图像表示的泛化能力,但它们对实验背景和生物学上下文信息的有限利用仍然在复杂表型分析中带来了挑战。在此,我们提出了scMIR,一种用于单细胞光学显微镜图像表示的视觉-语言基础模型。通过协同结合自监督图像重建与文本引导的跨模态对齐,scMIR能够在统一的表示空间中同时编码形态学和生物学语义信息。scMIR在207,957对图像-文本对上进行了预训练,涵盖了各种细胞类型、显微镜模式和扰动条件。通过在16个基准数据集上对细胞分类、聚类、表型推断和批次效应校正等各种复杂任务进行系统评估,scMIR的表现优于现有的通用模型和任务导向方法。此外,scMIR在各种任务中表现出强大的泛化能力,无需特定任务的微调。凭借其独特优势,我们设想scMIR可以通过支持各种下游分析任务,促进高通量表型工作流程的标准化和自动化。
cs.CV / 12 / 2607.22714

Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems

针对嵌入式汽车系统的优化 RetinaNet 架构的实时语义分割
D, Sai Sidharth
Abstract
Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power. This paper presents an optimized semantic segmentation architecture derived from the RetinaNet detection framework, adapted for dense pixel-wise prediction and tailored for deployment on resource-constrained embedded hardware. The proposed architecture, termed Opt-RetinaSeg, replaces the standard ResNet-50 backbone with a hybrid lightweight feature extractor, restructures the Feature Pyramid Network (FPN) to reduce redundant multi-scale computation, and introduces a compact segmentation head guided by focal-loss-inspired class balancing to address the severe foreground-background imbalance common in road scenes. We further apply a three-stage optimization pipeline consisting of structured channel pruning, post-training INT8 quantization, and knowledge distillation from a high-capacity teacher network. Evaluated on the Cityscapes and BDD100K datasets and deployed on an NVIDIA Jetson Xavier NX and a Qualcomm QCS610 automotive SoC, the proposed model achieves 73.9% mIoU at 70.4 FPS, representing a 7.4x inference speedup and a 4x reduction in model size relative to the ResNet-50 baseline, with less than 3% accuracy degradation. These results indicate that RetinaNet-derived architectures, when systematically optimized, are viable candidates for real-time semantic segmentation in embedded automotive perception pipelines
Chinese Translation
实时感知是高级驾驶辅助系统(ADAS)和自动驾驶车辆的基础要求,但嵌入式汽车平台对计算、内存和功耗施加了严格的限制。本文提出了一种基于 RetinaNet 检测框架的优化语义分割架构,适用于密集的像素级预测,并针对资源受限的嵌入式硬件进行了调整。所提出的架构称为 Opt-RetinaSeg,采用混合轻量级特征提取器替代标准的 ResNet-50 主干,重构特征金字塔网络(FPN)以减少冗余的多尺度计算,并引入一个紧凑的分割头,通过受焦点损失启发的类别平衡来解决道路场景中常见的严重前景-背景不平衡。我们进一步应用了一个三阶段优化流程,包括结构化通道剪枝、后训练 INT8 量化和从高容量教师网络的知识蒸馏。在 Cityscapes 和 BDD100K 数据集上进行评估,并在 NVIDIA Jetson Xavier NX 和 Qualcomm QCS610 汽车 SoC 上部署,所提出的模型在 70.4 FPS 下实现了 73.9% 的 mIoU,相较于 ResNet-50 基线实现了 7.4 倍的推理加速和 4 倍的模型尺寸缩减,且准确率下降不足 3%。这些结果表明,经过系统优化的基于 RetinaNet 的架构是嵌入式汽车感知管道中实时语义分割的可行候选者。
cs.CV / 13 / 2607.22715

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework

面向儿童的AIGC视频风险评估:基准与知识支持的迭代推理框架
Mi, Lewen, Li, Manyi, Sun, Yuling, Zhang, Yufan, Shi, Yuxin, Bian, Yulong, Li, Xiangxian, Liu, Juan
Abstract
The rapid growth of Artificial Intelligence-generated content (AIGC) is reshaping video production and circulation, exposing children to an increasing volume of AIGC videos. Unlike traditionally produced videos, AIGC videos often exhibit greater uncertainty in visual details, narrative coherence, and content expression, which may introduce developmentally inappropriate risks for children. However, existing video safety research is largely designed for general violation detection from an adult perspective and remains insufficient for identifying the fine-grained, implicit, and context-dependent risks that children may encounter when viewing AIGC videos. To address this gap, we study child-oriented AIGC video reviewing, making three contributions. First, we construct CAVSR, a benchmark of 605 real-world videos collected from multiple platforms, and develop a hierarchical risk taxonomy comprising 6 top-level categories and 26 fine-grained labels to support systematic evaluation of children's viewing risks. Second, we propose QVRS-E, a knowledge- and experience-augmented video reviewing framework that combines multi-agent collaboration with expert and experiential knowledge to support targeted evidence acquisition and fact-grounded reviewing decisions. Third, extensive experiments demonstrate that our method significantly enhances the reviewing of child-related risks integrated with vision-language models, and yields more robust review reports.
Chinese Translation
人工智能生成内容(AIGC)的快速增长正在重塑视频制作和传播,使儿童接触到越来越多的AIGC视频。与传统制作的视频不同,AIGC视频通常在视觉细节、叙事连贯性和内容表达上表现出更大的不确定性,这可能给儿童带来发展上不适宜的风险。然而,现有的视频安全研究主要是从成人的角度设计的,针对一般违规行为的检测,尚不足以识别儿童在观看AIGC视频时可能遇到的细粒度、隐性和依赖上下文的风险。为了解决这一问题,我们研究了面向儿童的AIGC视频评估,做出了三项贡献。首先,我们构建了CAVSR,一个由多个平台收集的605个真实视频的基准,并开发了一个包含6个顶级类别和26个细粒度标签的层次风险分类法,以支持对儿童观看风险的系统评估。其次,我们提出了QVRS-E,一个知识和经验增强的视频评估框架,结合了多智能体协作与专家和经验知识,以支持针对性的证据获取和基于事实的评估决策。第三,大量实验表明,我们的方法显著增强了与视觉-语言模型结合的儿童相关风险评估,并产生了更为稳健的评估报告。
cs.CV / 14 / 2607.22716

Visual Token Compression Enhances Robustness of MLLMs

视觉标记压缩增强了多模态大语言模型的鲁棒性
Gu, Shishen, Cui, Jiequan, Hu, Wenbo, Shi, Zenglin, Hu, Zhenzhen, Hong, Richang
Abstract
In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision and language modalities cannot be perfectly aligned, the misaligned visual tokens might act as out-of-distribution (OOD) inputs, leading to unpredictable outputs and introducing potential vulnerabilities. Building on this insight, we aim to enhance model robustness against jailbreaks and hallucinations by reducing OOD visual tokens at robust-pruning layers, while also reducing inference cost as a side benefit. Specifically, we measure the distance between each visual token and the language feature space. Then, visual tokens with large distances are identified as OOD tokens, which can be iteratively pruned. To demonstrate the effectiveness of our method, we evaluate it on seven diverse popular benchmarks. Notably, our method yields an average improvement of 13.29\% in defending jailbreak attacks, consistently achieves competitive performance in mitigating hallucinations, and maintains strong results on general datasets like MME.
Chinese Translation
在本文中,我们首次展示了视觉标记修剪增强了多模态大语言模型(MLLMs)的鲁棒性,减轻了诸如越狱攻击和幻觉等脆弱性。由于视觉和语言模态无法完美对齐,未对齐的视觉标记可能作为分布外(OOD)输入,导致不可预测的输出并引入潜在的脆弱性。基于这一见解,我们旨在通过在鲁棒修剪层减少OOD视觉标记来增强模型对越狱和幻觉的鲁棒性,同时作为附带好处降低推理成本。具体而言,我们测量每个视觉标记与语言特征空间之间的距离。然后,将距离较大的视觉标记识别为OOD标记,这些标记可以被迭代修剪。为了证明我们方法的有效性,我们在七个多样化的流行基准上进行了评估。值得注意的是,我们的方法在防御越狱攻击方面平均提高了13.29\%,在减轻幻觉方面始终保持竞争力,并在MME等通用数据集上保持强劲的结果。
cs.CV / 15 / 2607.22717

TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians

TOM-GS:通过静态3D高斯的时间不透明度调制实现可编辑视频表示
Lisowski, Marek, Smoliński, Łukasz, Howil, Kornel, Biliński, Piotr, Mazur, Marcin, Spurek, Przemysław
Abstract
While Implicit Neural Representations (INRs) and dynamic 3D Gaussian Splatting (3DGS) achieve impressive results in video processing, they often fall short of producing representations that are easily editable. Recent methods address this by introducing complex spatial deformations or folded distributions, which constrain optimization and reduce flexibility for downstream editing. In this paper, we introduce TOM-GS, an editable video representation that forgoes complex deformations in favor of regular 3D Gaussians equipped with a continuous temporal opacity formulation. By assigning a learnable temporal mean and scale to the opacity of each Gaussian, our model enables static 3D spatial components to fade smoothly in and out of the scene. Grounded by robust, off-the-shelf pose estimation, our approach maintains a static spatial geometry that naturally supports a wide range of manual and physics-based edits. TOM-GS outperforms prior editable video representations in visual fidelity, while its reliance on standard 3D Gaussians ensures seamless compatibility with established 3D editing tools.
Chinese Translation
尽管隐式神经表示(Implicit Neural Representations, INRs)和动态3D高斯点云(3D Gaussian Splatting, 3DGS)在视频处理方面取得了令人瞩目的成果,但它们在生成易于编辑的表示方面往往显得不足。近期的方法通过引入复杂的空间变形或折叠分布来解决这一问题,这限制了优化过程并降低了后续编辑的灵活性。在本文中,我们提出了TOM-GS,这是一种可编辑的视频表示,放弃了复杂的变形,采用了配备连续时间不透明度公式的常规3D高斯。通过为每个高斯分配可学习的时间均值和尺度,我们的模型使得静态3D空间组件能够平滑地在场景中淡入淡出。基于强大的现成姿态估计,我们的方法保持了静态空间几何形状,自然支持广泛的手动和基于物理的编辑。TOM-GS在视觉保真度方面超越了之前的可编辑视频表示,同时其对标准3D高斯的依赖确保了与现有3D编辑工具的无缝兼容。
cs.CV / 16 / 2607.22718

DAMamba-UNet3D: A Parameter-Efficient Mamba State Space U-Net with Dynamic Adaptive Scan for 3D Medical Image Segmentation

DAMamba-UNet3D:一种参数高效的Mamba状态空间U-Net,具有动态自适应扫描用于3D医学图像分割
Hussain, Mohammad Arafat, Grant, Ellen, Ou, Yangming
Abstract
We propose parameter-efficient SSM-based U-Net architectures for 3D medical image segmentation. Convolutional U-Nets afford O(n) local mixing per layer but lack explicit global context; transformers provide global reasoning at O(n^2) cost in sequence length $n$. State-space models (SSMs), such as Mamba, offer $O(n)$ global propagation per block. Yet, existing medical SSM segmenters rely on fixed scan patterns and large parameter budgets. Dynamic Adaptive Scan (DAS), which learns data-dependent reordering before selective scan, has not been applied to medical imaging or extended to 3D volumes. We propose DAMamba-UNet3D, a hybrid encoder-decoder that integrates tri-plane 3D-DAS blocks at encoder stages E2-E4 while retaining convolutions elsewhere (~5.3M parameters). On BraTS 2020 five-fold cross-validation, DAMamba-UNet3D achieves mean Dice 0.815+/-0.013 (full-volume per-case evaluation) at ~13x lower parameter cost than SegMamba (0.824+\-0.014, ~70M). At comparable scale, DAMamba-L (~70M), a wide DAS-native variant with encoder-only DAMamba and a convolutional bottleneck, reaches 0.829+\-0.012, surpassing retrained SegMamba by 0.5pt. Component ablations show that encoder-only DAS placement is critical as bottleneck and decoder SSM blocks lower Dice. Together, the results suggest that learned tri-plane DAS in a hybrid U-Net is competitive with, and under our large-scale design may improve upon, SegMamba's fixed Tri-orientated Mamba (ToM) scanning on BraTS 2020. Code: https://github.com/marafathussain/DAMamba-UNet3D.
Chinese Translation
我们提出了一种基于状态空间模型(SSM)的参数高效U-Net架构,用于3D医学图像分割。卷积U-Net在每层提供O(n)的局部混合,但缺乏明确的全局上下文;而变换器在序列长度n下以O(n^2)的成本提供全局推理。状态空间模型(SSMs),如Mamba,每个块提供O(n)的全局传播。然而,现有的医学SSM分割器依赖于固定的扫描模式和较大的参数预算。动态自适应扫描(DAS)在选择性扫描之前学习数据依赖的重排序,但尚未应用于医学成像或扩展到3D体积。我们提出了DAMamba-UNet3D,这是一种混合编码器-解码器,集成了在编码器阶段E2-E4的三平面3D-DAS块,同时在其他地方保留卷积(约5.3M参数)。在BraTS 2020五折交叉验证中,DAMamba-UNet3D在约13倍更低的参数成本下实现了平均Dice 0.815+/-0.013(每例全体积评估),而SegMamba为0.824+/-0.014(约70M)。在可比规模下,DAMamba-L(约70M)是一种宽的DAS原生变体,具有仅编码器的DAMamba和卷积瓶颈,达到了0.829+/-0.012,超越了重新训练的SegMamba 0.5点。组件消融实验表明,仅编码器的DAS放置作为瓶颈至关重要,而解码器SSM块降低了Dice。综合来看,结果表明,在混合U-Net中学习的三平面DAS与SegMamba在BraTS 2020上的固定三向Mamba(ToM)扫描具有竞争力,并且在我们的规模设计下可能会有所改进。代码:https://github.com/marafathussain/DAMamba-UNet3D。
cs.CV / 17 / 2607.22719

$\gamma$-Bridge: A Look-Parametric Diffusion Bridge

$eta$-桥:一种外观参数化的扩散桥
Hu, Xuran, Zhu, Yujie, Wang, Tengxi, Li, Jilong, Zhao, Wufan
Abstract
Multiplicative Gamma noise is a signal-dependent degradation in coherent imaging; synthetic aperture radar (SAR) despeckling is its most prominent real-world instance. Existing diffusion denoisers parameterize their forward process by abstract signal-to-noise schedules rather than by the physical look number $L$, so different deployment scenarios typically require separately trained models, and transfer from synthetic Gamma training to real SAR remains challenging without clean ground truth. We introduce $\gamma$-Bridge, a look-parametric bridge whose schedule $L(t)$ connects the noisy observation at $L_{obs}$ to the clean limit through exact multiplicative Gamma marginals. Its closed-form Gamma--L\'evy reverse posterior admits both stochastic and deterministic processes, while observation conditioning and a two-step consistency loss stabilize multi-step inference in the low-SNR single-look regime. Because bridge time directly represents $L$, one conditioned network can smart-start from any admissible input look and stop at a target look number. These two orthogonal controls enable zero-shot restoration over the full admissible grid after training only at $L_{obs} = 1$ on natural images with synthetic Gamma corruption. Combined with a homogeneous-patch look estimator, $\gamma$-Bridge processes data from six spaceborne and airborne SAR sensors without sensor-specific fine-tuning, achieving leading results on standard synthetic benchmarks while providing physically interpretable input and output controls absent from prior denoisers. Codes are released \href{https://github.com/Teriri1999/GammaBridge}{here}.
Chinese Translation
乘法伽马噪声是在相干成像中依赖于信号的退化现象;合成孔径雷达(SAR)去斑是其最显著的实际实例。现有的扩散去噪器通过抽象的信噪比时间表而非物理外观数 $L$ 来参数化其前向过程,因此不同的部署场景通常需要单独训练的模型,而从合成伽马训练到真实 SAR 的迁移在没有干净的真实值的情况下仍然具有挑战性。我们引入了 $eta$-桥,这是一种外观参数化的桥,其时间表 $L(t)$ 通过精确的乘法伽马边际将噪声观察值 $L_{obs}$ 连接到干净极限。其封闭形式的伽马-莱维逆后验允许随机和确定性过程,同时观察条件和两步一致性损失在低信噪比单视图状态下稳定多步推理。由于桥时间直接表示 $L$,一个条件网络可以从任何可接受的输入外观智能启动,并在目标外观数处停止。这两个正交控制使得在仅在具有合成伽马污染的自然图像上以 $L_{obs} = 1$ 训练后,能够对整个可接受网格进行零-shot 恢复。结合均匀补丁外观估计器,$eta$-桥处理来自六个空间和空中 SAR 传感器的数据,无需特定传感器的微调,在标准合成基准上取得领先结果,同时提供物理可解释的输入和输出控制,而这些在先前的去噪器中是缺失的。代码已在此发布 [here](https://github.com/Teriri1999/GammaBridge)。
cs.CV / 18 / 2607.22721

An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia

用于精神分裂症认知修复的交互式视觉语言平台
Mehdi, Nassira Ait, Temmam, Milissa, Larabi, Slimane
Abstract
Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning. For patients diagnosed with schizophrenia, these tasks are crucial for addressing severe cognitive deficits. However, evaluating the correctness of these physical actions generally relies on manual observation by clinicians, which introduces subjectivity and limits the scalability of therapeutic interventions. In this paper, we propose an automated framework based on Vision-Language Models for action verification in cognitive remediation tasks tailored for schizophrenia rehabilitation. The proposed system relies on a camera-monitored tabletop environment composed of structured miniature scenes including roads, a roundabout, a park, and toy vehicles. Patients receive audio instructions describing goal-oriented spatial actions to perform by manipulating a toy vehicle. These interactive physical activities are specifically designed to stimulate targeted cognitive functions, such as sustained attention, motor coordination, spatial navigation, and cognitive flexibility. To verify the correctness of the performed actions without requiring continuous clinical oversight, the system analyzes the video feed tracking the patient's hand and toy movements. A fine-tuned Vision-Language Model interprets the recorded video sequences and generates semantic descriptions of the observed activities, enabling high-level verification of the executed actions with respect to the initial textual instructions. A dedicated dataset of 4634 tabletop cognitive remediation video scenarios was collected to evaluate the proposed approach. Experimental results demonstrate that our specialized framework effectively bridges low-level physical telemetry with high-level clinical feedback, presenting a scalable and objective solution for advanced cognitive rehabilitation.
Chinese Translation
认知修复任务通常要求患者执行涉及物体操作和顺序推理的结构化动作。对于被诊断为精神分裂症的患者,这些任务对于解决严重的认知缺陷至关重要。然而,评估这些物理动作的正确性通常依赖于临床医生的手动观察,这引入了主观性并限制了治疗干预的可扩展性。本文提出了一种基于视觉-语言模型的自动化框架,用于精神分裂症康复中认知修复任务的动作验证。该系统依赖于一个由相机监控的桌面环境,包含结构化的微型场景,包括道路、环形交叉口、公园和玩具车辆。患者接收音频指令,描述通过操作玩具车辆执行的目标导向空间动作。这些交互式物理活动专门设计用于刺激特定的认知功能,如持续注意力、运动协调、空间导航和认知灵活性。为了在不需要持续临床监督的情况下验证执行动作的正确性,该系统分析跟踪患者手部和玩具运动的视频流。经过微调的视觉-语言模型解释记录的视频序列,并生成观察到的活动的语义描述,从而实现对执行动作与初始文本指令的高层次验证。我们收集了一个包含4634个桌面认知修复视频场景的专用数据集,以评估所提出的方法。实验结果表明,我们的专门框架有效地将低层次的物理遥测与高层次的临床反馈连接起来,提供了一种可扩展且客观的高级认知康复解决方案。
cs.CV / 19 / 2607.22722

A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection

一种新型对抗样本:测量人类与模型之间的差距及其与OOD检测的关系
Borji, Ali
Abstract
Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer recognize the image. Prior work showed such examples can be generated at scale but left three questions untested: whether humans really perform worse than the model, whether standard out-of-distribution (OOD) detection and calibration tools catch it, and whether existing defenses mitigate it. We answer all three on MNIST, CIFAR-10, and ImageNet. (i) An independent recognizer proxy drops to ~49% on CIFAR-10 while the model stays at 100% -- a gap a small human pilot (N=5) corroborates directly and that is not explained by signal loss (a matched-magnitude Gaussian control degrades recognizability faster); a CLIP zero-shot proxy confirms the gap at ImageNet scale too. (ii) Confidence- and energy-based OOD detectors and calibration are structurally blind (0% detection, ECE ~= 0), while a feature-space Mahalanobis detector flags 100% -- but is evaded by an adaptive attacker at no cost to success. (iii) No classical defense, including adversarial training (45% robust accuracy), reduces attack success (correlation with large-epsilon_l resistance r ~= 0). A mechanistic analysis further shows the attack destroys low-level texture far faster than edge/shape structure.
Chinese Translation
几乎所有的对抗攻击都通过添加不可察觉的扰动来欺骗模型。相反,我们研究的是相反的情况:一种大且明显可见的扰动,导致模型保持其原始的正确预测,即使人类将不再识别该图像。先前的研究表明,这种样本可以大规模生成,但留下了三个未测试的问题:人类是否真的表现得比模型差,标准的分布外(OOD)检测和校准工具是否能够捕捉到它,以及现有的防御措施是否能够缓解它。我们在MNIST、CIFAR-10和ImageNet上回答了这三个问题。(i) 一个独立的识别代理在CIFAR-10上的准确率降至约49%,而模型保持在100%——这一差距得到了一个小型人类实验(N=5)的直接证实,并且无法通过信号损失来解释(匹配幅度的高斯控制使得可识别性下降得更快);一个CLIP零-shot代理在ImageNet规模上也确认了这一差距。(ii) 基于置信度和能量的OOD检测器和校准在结构上是盲目的(0%检测率,ECE约为0),而一个特征空间的Mahalanobis检测器则标记了100%——但被一个自适应攻击者以零成本规避。(iii) 没有任何经典防御措施,包括对抗训练(稳健准确率为45%),能够降低攻击成功率(与大ε_l抗性之间的相关性r约为0)。机制分析进一步表明,该攻击比边缘/形状结构更快地破坏低级纹理。
cs.CV / 20 / 2607.22723

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models

通过分类引导的大型视觉-语言模型从文档中提取视觉信息
Li, Huafu, Chen, Guo, Xia, Jia, Wang, Lei, Du, Wei, Yao, Yun, Peng, Weijun, Li, Liming
Abstract
Visual information extraction (VIE) from visually rich documents remains challenging due to high layout variability and real-world impairments. Existing methods typically rely on sequential OCR pipelines or end-to-end models requiring extensive labeled data and layout-specific training, limiting their scalability.We propose a classification-guided large vision-language model (LVLM) framework for multi-type VIE that achieves high accuracy with minimal supervision. The approach decouples document-type classification from content extraction and employs in-context learning (ICL)-based dynamic prompt engineering to inject task-specific knowledge, enabling robust zero-shot inference across diverse layouts. From a theoretical perspective, the proposed method can be viewed as a form of conditional computation that reduces task uncertainty and improves information efficiency during prompt-based inference. Evaluated on a real-world bidding dataset with 16 certificate types, our zero-shot method (based on Qwen2.5-VL-7B) outperforms a strong supervised baseline by 18.35 percentage points in F1-score (86.43\% vs. 68.08\%) and 0.23 in normalized edit distance (0.90 vs. 0.67). Optional domain-specific fine-tuning further improves performance to 93.65\% F1 and 0.93 NED, demonstrating superior robustness against seals, watermarks, and low contrast. The framework offers an efficient, scalable solution for complex document understanding in office automation. Code is available at https://github.com/FairmeHIT/Multi-VIE, and fine-tuned models at https://huggingface.co/fairme/Qwen2.5-VL-7B-SFT.
Chinese Translation
从视觉丰富的文档中提取视觉信息(VIE)仍然面临挑战,主要由于布局的高度变异性和现实世界中的干扰。现有方法通常依赖于顺序的光学字符识别(OCR)管道或需要大量标注数据和特定布局训练的端到端模型,这限制了它们的可扩展性。我们提出了一种分类引导的大型视觉-语言模型(LVLM)框架,用于多类型的VIE,该框架在最小监督下实现了高准确性。该方法将文档类型分类与内容提取解耦,并采用基于上下文学习(ICL)的动态提示工程来注入任务特定知识,从而在不同布局中实现稳健的零-shot推理。从理论角度来看,所提出的方法可以视为一种条件计算形式,减少了任务的不确定性,并在基于提示的推理过程中提高了信息效率。在一个包含16种证书类型的真实世界招标数据集上进行评估,我们的零-shot方法(基于Qwen2.5-VL-7B)在F1-score上比强监督基线高出18.35个百分点(86.43 ext{%}对68.08 ext{%}),在归一化编辑距离上高出0.23(0.90对0.67)。可选的特定领域微调进一步将性能提升至93.65 ext{%} F1和0.93 NED,展示了对印章、水印和低对比度的优越鲁棒性。该框架为办公自动化中的复杂文档理解提供了一种高效、可扩展的解决方案。代码可在https://github.com/FairmeHIT/Multi-VIE获取,微调模型可在https://huggingface.co/fairme/Qwen2.5-VL-7B-SFT获取。
cs.CV / 21 / 2607.22725

Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification

结构保护主导基于深度学习的激光散斑材料分类中的数据增强
Salem, Mohamed Abdallah, Diab, Nourhan Zein
Abstract
Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are poorly matched to coherent imaging. Laser speckle patterns are not generic textures; they arise from coherent interference, and their discriminative content is carried by structured stochastic spatial and frequency statistics. This study examines how controlled augmentation perturbations influence speckle-based material classification on the SensiCut dataset. We train ResNet18 and EfficientNet-B0 under a parametric augmentation framework comprising rotation, Gaussian blur, independent Gaussian noise, spatially correlated speckle-aware noise, intensity jitter, and spatial masking, and evaluate test performance using macro F1-score averaged over three random seeds. Separate ordinary least squares models link augmentation parameters to performance for each architecture. Across both models, Gaussian blur exerts a strong negative effect (p < 0.001), indicating that low-pass filtering suppresses high-frequency structure that is informative for material discrimination. Independent pixel-wise noise is likewise harmful (p = 0.003 for EfficientNet-B0 and p = 0.001 for ResNet18), consistent with disruption of local spatial coherence. In contrast, spatially correlated perturbations yield significant positive coefficients (p = 0.004 for EfficientNet-B0 and p = 0.001 for ResNet18), showing that variability can improve robustness when it preserves speckle organization. The fitted models explain a substantial fraction of performance variation (R2 = 0.796 for EfficientNet-B0 and R2 = 0.879 for ResNet18). These results show that, in laser speckle imaging, augmentation effectiveness is determined primarily by structural preservation rather than perturbation magnitude. The findings motivate physics-aware augmentation design for coherent optical sensing.
Chinese Translation
数据增强通常用于提高图像分类的泛化能力,但标准策略的基本假设与相干成像不匹配。激光散斑图案并不是通用纹理;它们源于相干干涉,其判别内容由结构化的随机空间和频率统计特征所承载。本研究考察了受控增强扰动如何影响在 SensiCut 数据集上的基于散斑的材料分类。我们在一个参数化增强框架下训练了 ResNet18 和 EfficientNet-B0,该框架包括旋转、高斯模糊、独立高斯噪声、空间相关的散斑感知噪声、强度抖动和空间遮蔽,并通过对三个随机种子进行平均的宏 F1 分数评估测试性能。单独的普通最小二乘模型将增强参数与每个架构的性能联系起来。在两个模型中,高斯模糊产生了显著的负面影响(p < 0.001),表明低通滤波抑制了对材料区分有信息量的高频结构。独立的逐像素噪声同样有害(EfficientNet-B0 的 p = 0.003 和 ResNet18 的 p = 0.001),与局部空间相干性的破坏一致。相反,空间相关的扰动产生了显著的正系数(EfficientNet-B0 的 p = 0.004 和 ResNet18 的 p = 0.001),表明当变异性保持散斑组织时,可以提高鲁棒性。拟合模型解释了性能变化的相当一部分(EfficientNet-B0 的 R2 = 0.796 和 ResNet18 的 R2 = 0.879)。这些结果表明,在激光散斑成像中,增强的有效性主要由结构保护决定,而不是扰动幅度。研究结果激励了针对相干光学传感的物理感知增强设计。
cs.CV / 22 / 2607.22726

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

PCA:面向持久性的压缩与聚合用于快速视频大型语言模型
Song, Zihan, Ye, Shuo, Zhao, Bo, Zhang, Ruixin, Zhang, Jiayu, Ding, Shouhong, Yu, Zitong
Abstract
Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindrance to efficient reasoning. This paper introduces a training-free $\mathbf{P}$ersistence-Aware $\mathbf{C}$ompression and $\mathbf{A}$ggregation (PCA) method designed to preserve high-fidelity raw visual information before the encoding stage. PCA can be built on arbitrary VLLMs and consists of two modules: 1) A Dynamic Downsampling (DD) module that adaptively removes redundant frames by analyzing frame-wise similarity. 2) A Persistence-Aware Motion Enhancement (PAME) module that enriches each selected keyframe by aggregating the temporal context of its neighbors, ensuring that essential information is preserved even after aggressive frame reduction. Our approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model. Extensive experiments demonstrate that PCA consistently outperforms existing state-of-the-art approaches in both efficiency and accuracy, achieving a speedup of 1.8$\times$ to 2.5$\times$ compared to the baseline VLLM. The code is open-sourced at https://github.com/Heisenberg10110/PCA.
Chinese Translation
尽管视频大型语言模型(VLLMs)在视频理解方面取得了令人鼓舞的进展,但长时间帧中的冗余仍然阻碍了高效推理。本文提出了一种无训练的$ extbf{P}$ersistence-Aware $ extbf{C}$ompression and $ extbf{A}$ggregation(PCA)方法,旨在在编码阶段之前保留高保真度的原始视觉信息。PCA可以构建在任意VLLMs之上,包含两个模块:1)动态下采样(Dynamic Downsampling, DD)模块,通过分析帧间相似性自适应地去除冗余帧;2)面向持久性的运动增强(Persistence-Aware Motion Enhancement, PAME)模块,通过聚合邻近帧的时间上下文来丰富每个选定的关键帧,确保在激进的帧减少后仍能保留重要信息。我们的方法显著减少了长上下文建模的计算,同时提升了基线模型的性能。大量实验表明,PCA在效率和准确性上始终优于现有的最先进方法,与基线VLLM相比实现了1.8$ imes$到2.5$ imes$的加速。代码已开源,地址为https://github.com/Heisenberg10110/PCA。
cs.CV / 23 / 2607.22727

Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

可信赖的医学分割:临床图像退化下的不确定性感知 U-Net 评估
Kaliaperumal, Pranav, Kaliaperumal, Manisha
Abstract
Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning. We present a reproducible framework for evaluating uncertainty-aware segmentation under con- trolled clinical degradation. Our experiments use a synthetic multimodal brain tumor MRI cohort generated with a biophysical phantom simulator that follows the BraTS protocol. We train U-Net and Attention U-Net baselines for multi-class tumor sub-region segmentation and augment both models with Monte Carlo dropout to estimate per-voxel uncertainty. Across eight clinically motivated corruption types at five severity levels, we measure segmentation accuracy, calibration, failure detection, and selective prediction coverage. On clean data, Attention U-Net achieves a whole-tumor Dice of 0.990; under severe Gaussian noise, its performance falls to 0.089. Predictive uncertainty rises with degradation and tracks segmentation error (Pearson r = 0.53 under severity-3 Gaussian noise), allowing us to flag failures with an AUROC of 0.843. These results argue for uncertainty-aware inference as a practical safety layer in physician-in-the-loop radiology workflows. We release the code, trained models, and evaluation protocol to support direct reproduction.
Chinese Translation
医学图像分割模型在理想成像条件下通常报告高基准准确率,但在临床退化下的失败可能并不明显:传感器噪声、患者运动、低分辨率采集和对比度变化都可能在不发出明显警告的情况下改变模型行为。我们提出了一个可重复的框架,用于在受控的临床退化下评估不确定性感知的分割。我们的实验使用了一个合成的多模态脑肿瘤 MRI 队列,该队列是通过遵循 BraTS 协议的生物物理幻影模拟器生成的。我们训练了 U-Net 和 Attention U-Net 基线模型,用于多类肿瘤子区域分割,并通过蒙特卡洛 dropout 增强这两种模型,以估计每个体素的不确定性。在五个严重程度水平下的八种临床相关退化类型中,我们测量了分割准确性、校准、失败检测和选择性预测覆盖率。在干净数据上,Attention U-Net 实现了整体肿瘤 Dice 为 0.990;在严重高斯噪声下,其性能降至 0.089。预测不确定性随着退化而上升,并与分割错误相关(在严重程度为 3 的高斯噪声下,Pearson r = 0.53),使我们能够以 0.843 的 AUROC 标记失败。这些结果支持在医生参与的放射学工作流程中,将不确定性感知推理作为一种实用的安全层。我们发布了代码、训练模型和评估协议,以支持直接重现。
cs.CV / 24 / 2607.22728

CrossSpine: Multi-scale Cross-sequence Attention with Anatomical Priors for Automated Pfirrmann Grading

CrossSpine:结合解剖先验的多尺度跨序列注意力用于自动化Pfirrmann分级
Nguyen, Hai Son, Vu, Duong Ngoc, Nguyen, Trong-Nghia, Van, Bien Tran, Pham, Van-Dem, Xuan, Trang Mai, Vu, Huan, Van Luong, Thien
Abstract
Automated grading of Lumbar Disc Degeneration is essential for the objective quantification of structural changes associated with low back pain. Observing that baseline models underperformed on our data, we propose a framework designed to overcome these limitations. First, we present the Cross-sequence Attention Spine (CrossSpine) framework, a novel architecture that employs a cross-sequence attention mechanism to adaptively fuse features from different MRI sequences at multiple spa- tial scales. Second, we contribute a meticulously curated dataset aimed at automated Pfirrmann grading. Finally, we introduce an IVD-aware classification technique that integrates anatomical disc-level information, enabling the model to learn level-specific degeneration priors. Our experi- ments demonstrate the superiority of this approach: CrossSpine achieved a relative improvement exceeding 125% in the Macro F1 score, while boosting the Mean AUPRC by 99% and the Mean AUROC by 36% com- pared to the baseline.
Chinese Translation
自动化腰椎间盘退变分级对于客观量化与下背痛相关的结构变化至关重要。我们观察到基线模型在我们的数据上表现不佳,因此提出了一个旨在克服这些局限性的框架。首先,我们提出了跨序列注意力脊柱(CrossSpine)框架,这是一种新颖的架构,采用跨序列注意力机制,在多个空间尺度上自适应融合来自不同MRI序列的特征。其次,我们贡献了一个精心策划的数据集,旨在实现自动化Pfirrmann分级。最后,我们引入了一种考虑椎间盘(IVD)信息的分类技术,整合了解剖学层面信息,使模型能够学习特定层级的退变先验。我们的实验表明该方法的优越性:与基线相比,CrossSpine在宏观F1分数上实现了超过125%的相对提升,同时将平均AUPRC提升了99%,平均AUROC提升了36%。
cs.CV / 25 / 2607.22729

Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction

打开模型的视野:视频和上下文感知的多模态后续反馈预测
Kim, Min-Jae, Moon, Jun-Yeong, Sung, Mujeen, Park, Gyeong-Moon
Abstract
Backchannels, which signal listener states like empathy and understanding, are fundamental to natural human interaction. However, current approaches rely solely on audio and text. This omits crucial visual cues, such as facial expressions and gestures, as well as broader conversational contexts, which are necessary for accurate prediction. In this paper, we introduce Context-Aware Multimodal Alignment for Backchannel Prediction (CAMA-BC), a novel framework that leverages visual information through Multi-Layer Multimodal Alignment (MMA). Our alignment process comprises two stages. First, Context Alignment (MMA-CA) utilizes unlabeled dialogues with videos to capture conversational contexts. Next, Backchannel Alignment (MMA-BA) fine-tunes the representations specifically for backchannel prediction. Experimental results show that CAMA-BC significantly outperforms both existing methods and simple multimodal baselines, with particular effectiveness in recognizing complex backchannels such as empathy.
Chinese Translation
后续反馈是信号传达听众状态(如同理心和理解)的重要组成部分,对于自然的人际互动至关重要。然而,当前的方法仅依赖于音频和文本,这忽略了面部表情、手势等关键的视觉线索,以及进行准确预测所需的更广泛的对话上下文。在本文中,我们提出了一种上下文感知的多模态对齐框架用于后续反馈预测(Context-Aware Multimodal Alignment for Backchannel Prediction, CAMA-BC),该框架通过多层多模态对齐(Multi-Layer Multimodal Alignment, MMA)利用视觉信息。我们的对齐过程分为两个阶段。首先,上下文对齐(MMA-CA)利用带视频的未标记对话来捕捉对话上下文。接下来,后续反馈对齐(MMA-BA)专门针对后续反馈预测微调表示。实验结果表明,CAMA-BC在识别复杂的后续反馈(如同理心)方面显著优于现有方法和简单的多模态基线。
cs.CV / 26 / 2607.22731

Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment

无标定的室内环境3D多摄像头人员追踪
Veng, Ponleur, Vaufreydaz, Dominique, Kong, Phutphalla
Abstract
Multi-Camera People Tracking (MCPT) traditionally relies on precise intrinsic and extrinsic camera calibration to project 2D detections into a unified 3D world coordinate system.However, manual calibration constitutes a major bottleneck in large-scale dataset generation from unconstrained video archives. This work proposes a unified calibration-free 3D MCPT framework that infers geometric structure directly from visual data using deep foundation models. The system integrates anchor-free detection (YOLOX), robust tracking (BoT-SORT), omni-scale appearance embedding (OsNet), pose estimation (HRNet via MMPose), and transformer-based geometric reconstruction using the Visual Geometry Grounded Transformer (VGGT). A pose-guided 3D lifting strategy projects head keypoints onto a reconstructed manifold, eliminating dependence on ground-plane homography. Global identity association is formulated as hierarchical agglomerative clustering under a joint appearance-geometry cost with strict velocity gating. Evaluation on the AI City Challenge 2024 demonstrates a HOTA score of 53.13% without access to ground-truth calibration matrices, establishing a strong baseline for purely vision-based 3D tracking.
Chinese Translation
多摄像头人员追踪(MCPT)传统上依赖于精确的内外部相机标定,将2D检测投影到统一的3D世界坐标系中。然而,手动标定在从不受限的视频档案生成大规模数据集时构成了一个主要瓶颈。本研究提出了一种统一的无标定3D MCPT框架,该框架利用深度基础模型直接从视觉数据中推断几何结构。该系统集成了无锚点检测(YOLOX)、鲁棒追踪(BoT-SORT)、全尺度外观嵌入(OsNet)、姿态估计(HRNet通过MMPose)以及基于变换器的几何重建,使用视觉几何基础变换器(VGGT)。一种姿态引导的3D提升策略将头部关键点投影到重建的流形上,消除了对地面平面单应性的依赖。全球身份关联被构建为在严格速度门控下的联合外观-几何成本下的层次聚合聚类。在2024年AI城市挑战赛上的评估显示,在没有访问真实标定矩阵的情况下,HOTA得分为53.13%,为纯视觉基础的3D追踪建立了一个强有力的基准。
cs.CV / 27 / 2607.22733

Generative Augmentation for EEG Motor Imagery Classification: A Class-Conditional VAE with Cycle-Consistent Decoder Refinement

用于脑电图运动想象分类的生成增强:一种具有循环一致解码器精炼的类别条件变分自编码器
Moldoveanu, Matei, Sirois, Alain, Ali, Claire Ben, Lotte, Fabien, Yger, Florian
Abstract
We investigate whether a generative model can supply useful synthetic motor-imagery (MI) electroencephalography (EEG) trials that improve the accuracy of independent downstream classifiers. We train a class-conditional variational autoencoder (CVAE) with an integrated latent classifier on the Zhou motor-imagery dataset, using the learned per-class prior as a generator: sampling the prior for a given label and decoding it into a synthetic, label-consistent signal. A constraint on the covariance matrix of the generated data encourages preservation of covariance structure, and the model is trained with a schedule that alternates ordinary VAE training with a decoder-focused phase that sharpens the generative pathway used for augmentation. We measure the effect of adding synthetic trials to the training set under two evaluation protocols -- within-user (pooled 60/20/20 split across subjects) and cross-user (leave-one-subject-out, LOSO) -- across four representative EEG classification pipelines: Common Spatial Patterns with Linear Discriminant Analysis (CSP+LDA), tangent-space features with a Support Vector Machine (TGSP+SVM), Minimum Distance to Riemannian Mean (MDM), and a neural network based on EEGNetv4 (henceforth EEGNet). Results are aggregated across independent augmentation draws, random seeds (within-user), or leave-one-subject-out folds (cross-user), with uncertainty reported as 95\% confidence intervals (Student's $t$-distribution) computed over per-seed/per-fold averages. We find that synthetic EEG from the CVAE is most credible as a source of class-structured, covariance-like data rather than as a substitute for real raw EEG: it can raise the point estimate for MDM, but the broader augmentation claim remains conservative -- observed gains are small and classifier-dependent.
Chinese Translation
我们研究了生成模型是否能够提供有用的合成运动想象(MI)脑电图(EEG)试验,从而提高独立下游分类器的准确性。我们在Zhou运动想象数据集上训练了一个集成潜在分类器的类别条件变分自编码器(CVAE),使用学习到的每类先验作为生成器:为给定标签采样先验并将其解码为合成的、标签一致的信号。对生成数据的协方差矩阵施加约束,以鼓励保持协方差结构,并且模型的训练采用交替的调度,结合普通VAE训练与聚焦解码器的阶段,以锐化用于增强的生成路径。我们在两个评估协议下测量将合成试验添加到训练集的效果——用户内(在受试者之间进行60/20/20的汇总分割)和用户间(留一受试者法,LOSO)——在四个代表性的EEG分类管道中进行评估:与线性判别分析的公共空间模式(CSP+LDA)、与支持向量机的切空间特征(TGSP+SVM)、最小距离到黎曼均值(MDM),以及基于EEGNetv4的神经网络(以下简称EEGNet)。结果在独立增强抽样、随机种子(用户内)或留一受试者法折叠(用户间)中进行汇总,报告的不确定性为95\%置信区间(学生$t$分布),计算基于每个种子/每个折叠的平均值。我们发现,来自CVAE的合成EEG作为类结构、类似协方差数据的来源最为可信,而不是作为真实原始EEG的替代品:它可以提高MDM的点估计,但更广泛的增强主张仍然是保守的——观察到的增益较小且依赖于分类器。
cs.CV / 28 / 2607.22734

Fast Fourier Convolutional GAN for 30 m Clear-Sky Land Surface Temperature Gap-Free Reconstruction

快速傅里叶卷积生成对抗网络用于30米清晰天空地表温度无缝重建
Alfouly, Marwa, Halilovic, Smajil, Bochow, Nils, Hamacher, Thomas, Boers, Niklas, Schindler, Konrad
Abstract
Satellite-derived Land Surface Temperature (LST) provides spatially comprehensive data that ground stations cannot match. However, its utility is frequently limited by severe data gaps due to the presence of clouds. As LST is essential for understanding land-atmosphere interactions, numerous methods have been proposed to address this challenge. Yet, the development of a scalable and adaptable pipeline for generating gap-free LST datasets and reconstructing cloud-contaminated pixels remains challenging. Moreover, the reconstruction of extensive missing regions in fine-spatial-resolution observations is particularly difficult. To address this challenge, we propose a Multimodal Fast Fourier Convolutional GAN for reconstructing cloud-contaminated pixels in fine-resolution (30 m) Landsat imagery to generate gap-free clear-sky LST products. The method leverages Fast Fourier Convolution to enable a global receptive field across the image, and is guided by a stack of data consisting of satellite observations and Synthetic Aperture Radar (SAR) data. Across all LST quantiles, the interquartile range of scene-averaged RMSE (computed over reconstructed pixels) is consistently between 0.8 K and 1.8 K. The proposed approach enables the recovery of extensive missing regions, including scenes with more than 70% cloud-induced gaps, while relying on auxiliary data that are readily available at a near-global scale.
Chinese Translation
卫星获取的地表温度(LST)提供了地面站无法匹配的空间全面数据。然而,由于云层的存在,其效用常常受到严重数据缺口的限制。由于LST对于理解陆地-大气相互作用至关重要,已经提出了许多方法来应对这一挑战。然而,开发一个可扩展且适应性强的流程,以生成无缝LST数据集并重建受云污染的像素仍然具有挑战性。此外,在高空间分辨率观测中重建大范围缺失区域尤其困难。为了解决这一挑战,我们提出了一种多模态快速傅里叶卷积生成对抗网络(Multimodal Fast Fourier Convolutional GAN),用于重建高分辨率(30米)Landsat影像中的云污染像素,以生成无缝的清晰天空LST产品。该方法利用快速傅里叶卷积实现图像的全局感受野,并以卫星观测数据和合成孔径雷达(SAR)数据的堆叠数据为指导。在所有LST分位数中,重建像素的场景平均均方根误差(RMSE)的四分位距始终在0.8 K到1.8 K之间。所提出的方法能够恢复大范围缺失区域,包括云层引起的缺口超过70%的场景,同时依赖于在近全球范围内易于获得的辅助数据。
cs.CV / 29 / 2607.22736

Benchmarking the Domain Gap: Model Selection Instability Under Domain Shift in Video Capsule Endoscopy

域间差距的基准测试:视频胶囊内窥镜中域转移下的模型选择不稳定性
Hanson, Dan, Jha, Debesh
Abstract
Video capsule endoscopy (VCE) classification is typically evaluated within a single dataset, yet clinical deployment demands robustness across acquisition sources, labeling policies, and patient populations. We examine this gap using Kvasir-Capsule, Capsule Vision 2024 (CV2024), and a shared-label subset of Galar. We fine-tune a suite of general-domain pretrained backbones on the official Kvasir-Capsule folds under a standardized protocol and evaluate the same checkpoints on two non-source targets within a documented shared-label decision space. We find that the predictive value of in-domain ranking is target-dependent: Kvasir-Capsule ranking aligns more closely with Galar than with CV2024, while the two non-source targets agree only weakly. Consequently, the strongest in-domain backbone leads on one target yet falls to mid-pack on the other, and no single evaluation target reliably predicts the others. A second CV2024-trained configuration set reproduces this target-dependent instability. We conclude that capsule endoscopy model selection should report cross-target ranking stability rather than peak single-dataset performance.
Chinese Translation
视频胶囊内窥镜(VCE)分类通常在单一数据集内进行评估,然而临床应用要求在不同的采集来源、标注政策和患者群体中具备鲁棒性。我们使用 Kvasir-Capsule、Capsule Vision 2024 (CV2024) 和一个共享标签的 Galar 子集来研究这一差距。我们在官方 Kvasir-Capsule 折叠上,按照标准化协议对一系列通用领域的预训练主干网络进行微调,并在两个非源目标上评估相同的检查点,这两个目标位于一个已记录的共享标签决策空间内。我们发现,域内排名的预测价值依赖于目标:Kvasir-Capsule 的排名与 Galar 的一致性高于与 CV2024 的一致性,而这两个非源目标之间的共识则较弱。因此,最强的域内主干网络在一个目标上表现优异,但在另一个目标上却处于中等水平,并且没有单一的评估目标能够可靠地预测其他目标。第二个经过 CV2024 训练的配置集重现了这种依赖目标的不稳定性。我们得出结论,胶囊内窥镜模型选择应报告跨目标排名的稳定性,而非单一数据集的峰值性能。
cs.CV / 30 / 2607.22739

Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features

Cortex:针对《Quake》的紧凑行为克隆与冻结视觉特征
Malyshau, Dzmitry
Abstract
We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learning or explicit memory. Cortex is a compact Quake policy with 10.98 million trainable parameters in a six-layer transformer over a frozen DINOv3 encoder. It is trained on the Quake subset of the public Pixels2Play corpus: 6,849 recordings (about 474.7 hours), represented as 17.09 million cached decision frames with keyboard and mouse actions. One sampled training epoch uses 517,048 four-frame windows and takes 3.3 minutes of policy-head optimization on one RTX 5080, excluding one-time feature extraction. We evaluate two independent batches of 20 stochastic, 120-second episodes on Quake E1M1. Cortex does not complete the level, but every episode reaches the opening door, button room, and gate descent; 19 of 20 episodes in each batch record at least one kill. Under the same time-controlled harness, released P2P-150M and NitroGen checkpoints remain shallower in five matched-duration episodes each. These comparisons are limited by small reference samples and different native interfaces. Ablations show that denser visual tokens improve combat and survival, while longer optimization and naive action history improve offline metrics without consistently improving play. The remaining failures are consistent with covariate shift and motivate targeted corrective data. We release the policy implementation, checkpoint, and a representative rollout.
Chinese Translation
我们研究了一种故意简单的行为克隆策略在视觉丰富的第一人称游戏中能够取得多大进展,直到添加强化学习或显式记忆。Cortex 是一个紧凑的《Quake》策略,具有 1098 万可训练参数,基于一个六层变换器和冻结的 DINOv3 编码器。它在公共 Pixels2Play 数据集的《Quake》子集上进行训练:6849 个录音(约 474.7 小时),表示为 1709 万个缓存决策帧,包含键盘和鼠标操作。一个采样的训练周期使用 517048 个四帧窗口,并在一台 RTX 5080 上进行 3.3 分钟的策略头优化,不包括一次性特征提取。我们在《Quake》E1M1 上评估了两批独立的 20 个随机、120 秒的剧集。Cortex 没有完成关卡,但每个剧集都到达了开门、按钮房间和下降门;每批中的 20 个剧集有 19 个记录至少一次击杀。在相同时间控制的环境下,发布的 P2P-150M 和 NitroGen 检查点在每个匹配时长的剧集中仍然较浅。这些比较受到小参考样本和不同原生接口的限制。消融实验表明,更密集的视觉标记改善了战斗和生存,而更长的优化和天真的动作历史在离线指标上有所改善,但并未始终提高游戏表现。剩余的失败与协变量转移一致,并激励针对性纠正数据。我们发布了策略实现、检查点和一个代表性的回放。
cs.CV / 31 / 2607.22740

A Diagnostic Gap Framework for Evaluating Reconstruction Fidelity in Weakly Supervised Mammography

用于评估弱监督乳腺摄影重建保真度的诊断差距框架
Bertrand, Vinceline, Cardei, Ionut
Abstract
Weakly supervised pipelines for medical imaging have become increasingly popular over the years. These systems often include multiple stages and components, such as reconstruction, generation, and localization, yet standard evaluation metrics provide limited insight into whether clinically relevant information is preserved across each stage. We present the diagnostic gap framework, a practical evaluation tool that measures decision preservation and explanation preservation as a function of measured reconstruction fidelity. To isolate the effect of reconstruction from localization, we evaluate on curated lesion ROI crops using a fidelity ladder of three class-conditional reconstructors---VQ-VAE-GAN, VAE-GAN, and diffusion (SDEdit)---spanning a twenty-fold range in perceptual distance (LPIPS 0.029--0.584). At autoencoder fidelity, both decision and explanation are preserved: AUC changes remain within $\pm$0.005 and attribution similarity (HiResCAM, Grad-CAM++) stays high. At diffusion fidelity, both collapse: pooled AUC drops by 0.253 and mass-pathology AUC falls below chance. The diagnostic gap is thus a measurable function of reconstruction fidelity rather than an intrinsic cost of reconstruction, and the framework provides an architecture-agnostic instrument for identifying when and where multi-stage pipelines lose diagnostic signal.
Chinese Translation
近年来,弱监督医学影像处理管道变得越来越受欢迎。这些系统通常包括多个阶段和组件,如重建、生成和定位,但标准评估指标对每个阶段是否保留临床相关信息提供的洞察有限。我们提出了诊断差距框架,这是一种实用的评估工具,测量决策保留和解释保留作为重建保真度的函数。为了隔离重建与定位的影响,我们在策划的病灶 ROI 裁剪上进行评估,使用三种类别条件重建器的保真度梯度——VQ-VAE-GAN、VAE-GAN 和扩散(SDEdit),其感知距离跨越二十倍范围(LPIPS 0.029--0.584)。在自编码器保真度下,决策和解释均得到保留:AUC 变化保持在 $m{ ext{±0.005}}$ 之内,归因相似性(HiResCAM, Grad-CAM++)保持较高。在扩散保真度下,两者均崩溃:汇总 AUC 降低 0.253,病理 AUC 降至偶然水平以下。因此,诊断差距是重建保真度的可测量函数,而不是重建的内在成本,该框架提供了一种与架构无关的工具,用于识别多阶段管道何时以及何处失去诊断信号。
cs.CV / 32 / 2607.22745

AI-generated Images Challenge Visual Trust in High-risk Scenarios

人工智能生成图像在高风险场景中挑战视觉信任
Wang, Yi-Zhi, Xiao, Yichen, Yue, Linan, Gao, Weibo, Du, Yichao, Fang, Pengfei, Di, Shimin, Zhang, Min-Ling
Abstract
Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation. Yet existing detection benchmarks rarely examine synthetic images in public- and individual-safety contexts, where misleading visual content may carry substantial risks. Here we introduce SafeIMG, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2. Unlike benchmarks centred on generic imagery and image-level labels, SafeIMG evaluates not only whether detectors recognise synthetic images, but also whether their decisions reflect human-identified anomalies. To this end, SafeIMG provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies. We evaluate specialized synthetic-image detectors and vision-language models (VLMs), and find that neither provides reliable detection. The strongest VLM identifies only 49.5% of generated images, whereas the best specialised detector identifies 33.1%, compared with 81.7% accuracy for human evaluators. Model explanations cover only 29.8\% of human-annotated anomalies and predominantly capture local defects in text, faces and hands. Their coverage falls to 15.0% for commonsense conflicts and 12.0% for physical inconsistencies, while detection performance deteriorates further after dissemination-induced image degradation. These findings show that current detectors lack the accuracy, explanatory alignment and robustness needed to evaluate AI-generated images reliably across public- and individual-safety settings.
Chinese Translation
图像生成的快速进展正在侵蚀视觉内容在真实性可能影响公共安全和个人声誉的环境中的证据价值。然而,现有的检测基准很少在公共和个人安全的背景下考察合成图像,而这些误导性的视觉内容可能带来重大风险。在此,我们介绍了 SafeIMG,这是一个以安全为导向的基准,涵盖了 12 种使用 GPT Image 2 生成的公共和个人安全场景。与以通用图像和图像级标签为中心的基准不同,SafeIMG 不仅评估检测器是否识别合成图像,还评估其决策是否反映人类识别的异常。为此,SafeIMG 提供了人类注释,定位可疑区域并解释局部伪影以及更高层次的常识或物理不一致性。我们评估了专门的合成图像检测器和视觉-语言模型(VLMs),发现两者均未提供可靠的检测。最强的 VLM 仅识别出 49.5% 的生成图像,而最佳的专门检测器识别出 33.1%,相比之下,人类评估者的准确率为 81.7%。模型解释仅覆盖了 29.8% 的人类注释异常,并主要捕捉文本、面部和手部的局部缺陷。对于常识冲突,其覆盖率降至 15.0%,对于物理不一致性降至 12.0%,而在传播引起的图像退化后,检测性能进一步恶化。这些发现表明,当前的检测器缺乏在公共和个人安全环境中可靠评估 AI 生成图像所需的准确性、解释一致性和鲁棒性。
cs.CV / 33 / 2607.22746

Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge

将全天气建筑损伤映射推进至实例级别:2026年Bright挑战的结果与启示
Chen, Hongruixuan, Huang, He, Wang, Haifeng, Song, Jian, Wang, Junjue, Xuan, Weihao, Mitchell, Hamish, Li, Jiepan, He, Wei, Zhang, Liangpei, Wang, Zijie, Zhong, Chen, Zhao, Jiazhen, Hu, Lei, Hu, Ting, Zhang, Hongyan, Angelides, Gregory, Cha, Miriam, Broni-Bediako, Clifford, Xia, Junshi, Perron, Taylor, Yokoya, Naoto
Abstract
Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed. Post-event optical imagery, however, may be unavailable because of cloud, smoke, or darkness. The Bright Challenge evaluated all-weather building damage mapping from a submeter-resolution pre-event optical image and a post-event SAR image. Participants were required to detect and delineate each building and assign exactly one of three mutually exclusive damage labels. The challenge extended the globally distributed \textsc{Bright} dataset with instance-level annotations for about 291,000 buildings across 16 disaster events spanning seven disaster types. The final phase was evaluated exclusively on two 2025 events absent from training: a wildfire event in California and a hurricane in Jamaica. A total of 157 participants made 1,289 submissions, and 46 teams entered the final phase. The two winning solutions achieved test mAPs of 0.182 and 0.181, approximately 8.7 times the public baseline of 0.021, but remained far below the best in-domain holdout score of 0.513. Across teams ranked in both phases, performance declined sharply and the rank order changed substantially. The two leading solutions independently favored modality-specific encoding, staged or late optical--SAR fusion, and an optical-dominant separation of building localization from damage recognition. The winning method additionally used scene-aware threshold adjustment and pseudo-label adaptation. These results identify cross-event generalization and stable severity discrimination as the principal remaining challenges. All data, annotations, baseline code, and winning solutions are publicly available at https://github.com/ChenHongruixuan/BRIGHT.
Chinese Translation
快速的灾后响应需要及时获得建筑级别的信息,以判断结构是否完好、受损或被毁。然而,灾后光学影像可能因云层、烟雾或黑暗而无法获取。Bright挑战评估了基于亚米分辨率的灾前光学影像和灾后合成孔径雷达(SAR)影像的全天气建筑损伤映射。参与者需检测并勾画出每栋建筑,并准确分配三种互斥损伤标签中的一种。该挑战扩展了全球分布的 extsc{Bright}数据集,为约291,000栋建筑提供了实例级别的注释,涵盖了16个灾难事件和七种灾难类型。最终阶段的评估仅基于两个2025年的事件,这两个事件未在训练中出现:加利福尼亚的野火事件和牙买加的飓风事件。共有157名参与者提交了1,289份方案,46个团队进入了最终阶段。两个获胜方案的测试平均精度(mAP)分别为0.182和0.181,约为公共基线0.021的8.7倍,但仍远低于最佳领域保留分数0.513。在两个阶段中排名的团队表现急剧下降,排名顺序发生了显著变化。两个领先方案独立地倾向于特定模态编码、分阶段或后期光学-合成孔径雷达融合,以及将建筑定位与损伤识别分开的光学主导方法。获胜的方法还使用了场景感知的阈值调整和伪标签适应。这些结果表明跨事件泛化和稳定的严重性区分是主要的剩余挑战。所有数据、注释、基线代码和获胜方案均可在 https://github.com/ChenHongruixuan/BRIGHT 上公开获取。
cs.CV / 34 / 2607.22749

Post-Operative Glioma Segmentation via Loss Stabilization, Normalization and Subspace Attention

通过损失稳定化、归一化和子空间注意力进行术后胶质瘤分割
Crişan, Alexandru, Borza, Diana
Abstract
Tracking residual tumor after surgery is essential for catching recurrence early, but automating post-operative glioma segmentation remains a difficult task. Although transformer-based architectures, such as SwinUNETR, achieved impressive results, few studies test how well they generalize across clinical protocols. In this paper, we conduct an ablation study on the MU-GLIOMA-POST and UCSF-ALPTDG datasets and show that the standard Generalized Dice Loss (GDL) is unstable under domain shift: the Whole Lesion (WL) Dice drops from 0.88 on the internal validation set to 0.73 on the external UCSF test set. To address this, we pair brain-masked percentile normalization with voxel-level contrastive learning. We also propose a Subspace-Aware Class Attention (SACA) module that re-calibrates the bottleneck features and raises Enhancing Tumor (ET) sensitivity by 8% (9.1% relative improvement) on internal validation. Ensembling these refinements with nnU-Net brings every stable configuration to a WL Dice of 0.94, and the SACA variant ensemble achieves the best boundary error (HD95) of 2.92 mm on MU-GLIOMA-POST.
Chinese Translation
术后追踪残余肿瘤对于早期发现复发至关重要,但自动化术后胶质瘤分割仍然是一项困难的任务。尽管基于变换器的架构,如SwinUNETR,取得了令人印象深刻的结果,但很少有研究测试它们在不同临床协议中的泛化能力。在本文中,我们对MU-GLIOMA-POST和UCSF-ALPTDG数据集进行了消融研究,并显示标准的广义骰子损失(Generalized Dice Loss, GDL)在领域转移下不稳定:整体病灶(Whole Lesion, WL)骰子从内部验证集的0.88下降到外部UCSF测试集的0.73。为了解决这个问题,我们将脑部掩膜百分位归一化与体素级对比学习相结合。我们还提出了一个子空间感知类注意力(Subspace-Aware Class Attention, SACA)模块,该模块重新校准瓶颈特征,并在内部验证中将增强肿瘤(Enhancing Tumor, ET)的敏感性提高了8%(相对改善9.1%)。将这些改进与nnU-Net集成,使每个稳定配置的WL骰子达到0.94,而SACA变体集成在MU-GLIOMA-POST上实现了最佳边界误差(HD95)为2.92毫米。
cs.CV / 35 / 2607.22752

Beyond Error-vs-Discard Characteristic: Toward Stable and Reliable Evaluation for Face Image Quality Assessment

超越错误与丢弃特征:面部图像质量评估的稳定可靠评估方法
Wani, Bhavesh, Babnik, Žiga, Štruc, Vitomir, Terhörst, Philipp
Abstract
Face Image Quality Assessment (FIQA) aims to estimate the utility of facial images for reliable recognition. The evaluation of FIQA methods is predominantly based on the Error-versus-Discard Characteristic (EDC), which evaluates performance by progressively discarding low-quality samples and measuring recognition error on the retained subset. In this work, we demonstrate that the widely used EDC protocol has fundamental limitations: Test-Set Divergence and Threshold Drift, which together limit the reliability and comparability of FIQA methods. To address this, we propose discard-based EDC variants and a rank-based Rank Consistency Evaluation (RCE) metric that operates on the entire test set without discarding samples, using a fixed decision threshold. Extensive experiments on five datasets, four face recognition models, and 15 state-of-the-art FIQA methods demonstrate both the limitations of EDC and the effectiveness of the proposed approaches in enabling a more reliable and comparable evaluation. Despite evaluated on face images only, the limitations arise from the EDC protocol rather than the biometric modality, suggesting a broader applicability to biometric quality assessment in general.
Chinese Translation
面部图像质量评估(FIQA)旨在估计面部图像在可靠识别中的实用性。FIQA方法的评估主要基于错误与丢弃特征(EDC),该方法通过逐步丢弃低质量样本并测量保留子集上的识别错误来评估性能。在本研究中,我们展示了广泛使用的EDC协议存在根本性局限性:测试集分歧和阈值漂移,这两者共同限制了FIQA方法的可靠性和可比性。为了解决这个问题,我们提出了基于丢弃的EDC变体和一种基于排名的一致性评估(RCE)指标,该指标在整个测试集上操作而不丢弃样本,并使用固定的决策阈值。在五个数据集、四个面部识别模型和15种最先进的FIQA方法上进行的广泛实验展示了EDC的局限性以及所提方法在实现更可靠和可比评估方面的有效性。尽管仅在面部图像上进行评估,但这些局限性源于EDC协议而非生物特征模态,这表明其在生物特征质量评估中的更广泛适用性。
cs.CV / 36 / 2607.22753

Real-time Reconstruction of Human Visual Perception from fMRI

基于功能性磁共振成像的人类视觉感知实时重建
Iyer, Rishab S., Tu, Jiaxin Cindy, Villanueva, Cesar Kadir Torrico, Mahishi, Anish, Kempner, Ross P., Prince, Jacob S., Lo, Ernest W., Bhowmick, Akash, Arasu, Hritik, Chughtai, Amaar, McDevitt, Elizabeth A., Scotti, Paul S., Norman, Kenneth A.
Abstract
Real-time closed-loop neurofeedback based on functional magnetic resonance imaging (fMRI) has led to important scientific and clinical advances. However, the sophistication of the analysis methods used in real-time fMRI lags behind the state-of-the-art in fMRI decoding, largely due to computational factors: Most advanced decoding pipelines do not fit within the envelope of real-time processing, where the analysis needs to be conducted in a matter of seconds and without leveraging data acquired later in the session. Here, we present a real-time compatible adaptation of a computationally intensive state-of-the-art pipeline for reconstructing perceived natural images (MindEye2), and we demonstrate that reliable fine-grained decoding is still achievable in this setting. Using RT-Cloud, an open-source, scalable cloud-based platform, we performed a real-time scan where we decoded single-trial visual perception within seconds after an image was shown to the participant. Finally, we use simulated analyses to document the factors driving changes in performance from offline to real-time analysis. This work serves as a proof-of-concept that it is feasible to deploy these powerful fMRI decoding pipelines in real-time analysis, paving the way for their use in brain-computer interfaces for scientific discovery and clinical treatment.
Chinese Translation
基于功能性磁共振成像(fMRI)的实时闭环神经反馈已带来了重要的科学和临床进展。然而,实时fMRI中使用的分析方法的复杂性落后于fMRI解码的最新技术,这在很大程度上是由于计算因素:大多数先进的解码流程无法适应实时处理的要求,因为分析需要在几秒钟内完成,并且不能利用会话中后期获得的数据。在此,我们展示了一种与实时兼容的、计算密集型的最新技术流程的改编,用于重建感知的自然图像(MindEye2),并证明在这种设置下仍然可以实现可靠的细粒度解码。我们使用RT-Cloud,一个开源、可扩展的云平台,进行了实时扫描,在图像展示给参与者后几秒内解码单次视觉感知。最后,我们使用模拟分析记录了从离线分析到实时分析性能变化的驱动因素。这项工作作为一个概念验证,证明了在实时分析中部署这些强大的fMRI解码流程是可行的,为其在科学发现和临床治疗中的脑-机接口应用铺平了道路。
cs.CV / 37 / 2607.22771

Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

廉价探针预测3D-CT视觉-语言模型的昂贵训练
Liang, Renjie
Abstract
Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their tokens, and several token budgets, and the combinations grow fast. Comparing them the usual way means fine-tuning a large language model (LLM) on each combination, and running the whole sweep this way needs far more compute than most groups can spend. We ask whether a cheap probe on the encoder's cached embeddings can stand in for that comparison. We build an image-grounded probing benchmark over (encoder $\times$ compression) cells, with clinical attribute families and two validation gates, scale-sanity and probe-separability, that keep each attribute well-scaled and decodable. These gates are the main methodological contribution. On this benchmark we compare a range of read-out heads, and in a preliminary study we pair each probe with its matched downstream task. The early signal is encouraging: the cheap probe orders the candidates in close agreement with expensive fine-tuning, at about $r\approx0.95$ on the cells measured so far. We read this as an ordinal claim, a ranking predictor rather than an exact estimate, and we are explicit about where it stays preliminary. If it holds up, encoder and compression choices can be screened in minutes with frozen-token probes, with full training spent only on the finalists.
Chinese Translation
选择用于3D~CT视觉-语言模型(VLM)的冻结图像编码器,以及其上的令牌压缩方案,是在众多候选者中进行的搜索。存在多种编码器、多种压缩令牌的方法和多种令牌预算,这些组合迅速增加。以通常的方式比较它们意味着需要对每种组合进行大型语言模型(LLM)的微调,而以这种方式进行全面的比较需要的计算资源远超大多数团队的承受能力。我们探讨是否可以用编码器缓存嵌入上的廉价探针来替代这种比较。我们构建了一个基于图像的探针基准,覆盖(编码器 × 压缩)单元,包含临床属性家族和两个验证门:规模合理性和探针可分离性,以保持每个属性的良好规模和可解码性。这些验证门是我们主要的方法论贡献。在这个基准上,我们比较了一系列读取头,并在初步研究中将每个探针与其匹配的下游任务配对。早期信号令人鼓舞:廉价探针对候选者的排序与昂贵的微调结果高度一致,在目前测量的单元中约为 $r ext{≈}0.95$。我们将其视为一种序数声明,即排名预测器而非精确估计,并明确指出其仍处于初步阶段。如果这一结果成立,编码器和压缩选择可以在几分钟内通过冻结令牌探针进行筛选,而完整的训练仅花费在最终候选者上。
cs.CV / 38 / 2607.22808

Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution

混合语义与光谱集成方法用于稳健的合成图像源归属
Hossain, Md. Ajwad
Abstract
The rapid advancement of text-to-image (T2I) models has necessitated robust Synthetic Image Source Attribution (SIA) methodologies. A critical challenge in SIA is the distribution shift between pristine training images and real-world deployed images, which undergo unknown post-processing operations such as JPEG compression and blurring. In this work, proposed for the DLMMDD Challenge at ICANN 2026, we introduce a dual-branch ensemble framework fusing Semantic Deep Learning with Mathematical Forensic Feature Extraction. The semantic branch employs EfficientNet-B0 regularized with Exponential Moving Averaging (EMA) and Label Smoothing. The forensic branch extracts 126 mathematical features -- including SVD spectral profiles and Local Binary Patterns -- from high-pass noise residuals, compressed via Truncated SVD and classified with XGBoost. Evaluated on a dataset of 10 generators where 55% of the test set is degraded, our approach achieves a private leaderboard accuracy of 95.60%. Furthermore, the entire pipeline is highly computationally efficient, requiring no GPU acceleration and executing end-to-end on a standard CPU in under 6.5 hours, highlighting the practicality and scalability of mathematical forensics for real-world deployment.
Chinese Translation
文本到图像(T2I)模型的快速发展使得稳健的合成图像源归属(SIA)方法变得必要。SIA中的一个关键挑战是原始训练图像与实际部署图像之间的分布转变,后者经历了未知的后处理操作,如JPEG压缩和模糊。在本研究中,我们为2026年ICANN的DLMMDD挑战提出了一种双分支集成框架,将语义深度学习与数学取证特征提取相结合。语义分支采用经过指数移动平均(EMA)和标签平滑正则化的EfficientNet-B0。取证分支从高通噪声残差中提取126个数学特征,包括SVD光谱特征和局部二值模式,经过截断SVD压缩并使用XGBoost进行分类。在一个包含10个生成器的数据集上评估,其中55%的测试集被降级,我们的方法在私有排行榜上达到了95.60%的准确率。此外,整个流程具有高度的计算效率,无需GPU加速,并且在标准CPU上端到端执行时间少于6.5小时,突显了数学取证在实际部署中的实用性和可扩展性。
cs.CV / 39 / 2607.22830

ID-V2V: Identity-Preserving Video Restylization

ID-V2V:身份保留的视频风格化
Xu, Yuancheng, He, Mingming, Salamanca, Pablo, Ma, Li, Kant, Yash, Steven, Emmett, Debevec, Paul, Yu, Ning
Abstract
In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and lip synchronization. A key obstacle is the absence of paired training data, as identity-preserving restylized video pairs are rare in real-world settings. To address this, we propose a decoupling of source-grounded identity preservation and edit-driven video synthesis. Our key insight is that facial appearance and expression should remain invariant, with illumination being the primary permissible variation. We therefore cast identity preservation as a video relighting problem, while modeling visual edit propagation as controlled video synthesis guided by the edited keyframe. Building on this formulation, we introduce ID-V2V, a video-to-video generative framework integrating complementary control signals: relit facial regions and facial normal maps tightly constrain facial likeness and performance, while edited keyframes and depth sequences enable flexible and temporally coherent generation. This design enables constructing training pairs from a single video, eliminating the need for scarce paired data. Extensive experiments demonstrate that ID-V2V significantly outperforms existing methods in preserving facial likeness and fine-grained facial performance, supports both single- and multi-subject scenarios, and delivers high visual quality, highlighting its potential as a human-centric tool for real-world content production. The code is available at: https://github.com/Eyeline-Labs/ID-V2V.
Chinese Translation
在视觉叙事中,人类表演是创意意图和叙事意义的核心。然而,在实现灵活的视觉编辑的同时,保留人类身份和表演仍然是生成视频模型面临的挑战。我们将这一挑战形式化为身份保留的视频风格化,该方法在保留面部相似性和表演(包括表情、眼神和嘴唇同步)的同时,传播由编辑关键帧指定的场景、光照和风格变化。一个主要障碍是缺乏配对训练数据,因为在现实世界中,身份保留的风格化视频对非常稀缺。为了解决这个问题,我们提出将源基础的身份保留与编辑驱动的视频合成解耦。我们的关键见解是,面部外观和表情应保持不变,而光照是主要的允许变化。因此,我们将身份保留视为一个视频重光照问题,同时将视觉编辑传播建模为由编辑关键帧引导的受控视频合成。在此基础上,我们引入了ID-V2V,一个视频到视频的生成框架,集成了互补的控制信号:重光照的面部区域和面部法线图紧密约束面部相似性和表演,而编辑关键帧和深度序列则支持灵活且时间一致的生成。这一设计使得从单个视频构建训练对成为可能,消除了对稀缺配对数据的需求。大量实验表明,ID-V2V在保留面部相似性和细致的面部表演方面显著优于现有方法,支持单一和多主体场景,并提供高视觉质量,突显其作为人本工具在现实内容生产中的潜力。代码可在以下链接获取:https://github.com/Eyeline-Labs/ID-V2V。
cs.CV / 40 / 2607.22847

Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling

注视锚定的社交网络:通过联合建模解码隐含关系
Hou, Yuqi, Chen, Zhuo, Hu, Han, Kim, Je Woo, Jiao, Jianbo, Chang, Hyung Jin
Abstract
Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals independently, treating gaze as an i.i.d. quantity or predicting social semantics in isolation. Recent multi-person methods attempt to address this but often treat social relations as rigid, post-hoc classifications decoupled from the gaze estimation process. This oversimplification fails to capture the nuanced nature of social intent, which acts as an underlying driver of gaze behavior rather than a secondary categorical output. We address these limitations by proposing ANCHOR, a target-centric paradigm designed to decode gaze-anchored social intent by modeling the joint distribution of visual attention and latent implicit relations. Our approach surfaces these dependencies as the latent structural scaffolding of gaze behavior. The architecture utilizes a relational attention mechanism to capture fine-grained interpersonal links, leveraging feature-wise modulation for efficient multi-person parsing from a single vision backbone. To stabilize the training of this coupled formulation, we implement an optimization synergy to resolve the inherent conflicts between spatial gaze accuracy and latent social reasoning. This approach ensures robust generalization by seeking stable, flat minima while simultaneously harmonizing competing task gradients. We validate our framework on an extended benchmark featuring dense multi-person annotations and novel social influence rankings. Our results demonstrate state-of-the-art performance and provide the first quantitative evidence that implicit social hierarchies can be robustly disentangled and learned directly from static gaze patterns.
Chinese Translation
人类的注视不仅指向视觉目标;它在静态图像中还作为社交意图的微妙指示。然而,标准模型通常独立处理个体,将注视视为独立同分布(i.i.d.)的量,或孤立地预测社交语义。最近的多人物方法试图解决这一问题,但往往将社交关系视为与注视估计过程脱钩的刚性后验分类。这种过于简化的处理未能捕捉社交意图的细微特征,而社交意图实际上是注视行为的潜在驱动因素,而非次要的类别输出。我们通过提出ANCHOR来解决这些局限性,ANCHOR是一种以目标为中心的范式,旨在通过建模视觉注意力和潜在隐含关系的联合分布来解码注视锚定的社交意图。我们的方法将这些依赖关系呈现为注视行为的潜在结构支架。该架构利用关系注意机制捕捉细粒度的人际联系,利用特征级调制实现从单一视觉骨干网络的高效多人物解析。为了稳定这一耦合模型的训练,我们实施了一种优化协同机制,以解决空间注视精度与潜在社交推理之间的固有冲突。这种方法通过寻求稳定的平坦最小值,同时协调竞争任务的梯度,确保了稳健的泛化能力。我们在一个扩展基准上验证了我们的框架,该基准具有密集的多人物注释和新颖的社交影响排名。我们的结果展示了最先进的性能,并提供了第一个定量证据,证明隐含社交层级可以从静态注视模式中稳健地解开和直接学习。
cs.CV / 41 / 2607.22861

Robustifying pathology foundation models via fine-tuning

通过微调增强病理基础模型的鲁棒性
Filiot, Alexandre, Thaeter, Oskar, Schmauch, Benoit, Guillou, Lionel
Abstract
Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, undermining deployment across laboratories. We develop a novel fine-tuning recipe that improves the robustness of pathology FMs to acquisition factors. Applied to ten different FMs, our fine-tuning strategy consistently improves robustness for every model as well as downstream performance, with no observed trade-off. On average, it raises the PathoROB robustness index by 23% (from 0.72 to 0.87) and increases the overall cross-benchmark performance by 43% on Patho-Bench, HEST and THUNDER combined, with individual gains reaching up to 72% in robustness (Phikon-v2) and 76% in performance (Midnight-12k). We publicly release the fine-tuned versions of Phikon-v2 (Phaet) and Midnight-12k (Mascaret) at https://huggingface.co/wearewaiv/models.
Chinese Translation
病理基础模型(FMs)生成强大的切片级表示,但对扫描仪和染色变异性敏感,这削弱了其在不同实验室的应用。我们开发了一种新颖的微调方案,旨在提高病理 FMs 对采集因素的鲁棒性。该方案应用于十种不同的 FMs,结果显示我们的微调策略在每个模型上均能持续提高鲁棒性和下游性能,且没有观察到任何权衡。平均而言,PathoROB 鲁棒性指数提高了 23%(从 0.72 提升至 0.87),在 Patho-Bench、HEST 和 THUNDER 的综合评估中,整体跨基准性能提高了 43%,个别模型的鲁棒性提升最高可达 72%(Phikon-v2)和性能提升最高可达 76%(Midnight-12k)。我们已在 https://huggingface.co/wearewaiv/models 上公开发布了微调后的 Phikon-v2(Phaet)和 Midnight-12k(Mascaret)版本。
cs.CV / 42 / 2607.22864

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

Spatial-IQ:通过层次能力测试解构空间智能
Rim, Patrick, Long, Tom, Prashnani, Ekta, Rosenholtz, Ruth, Boudaoud, Ben, Xenopoulos, Peter, Wong, Alex, Kim, Joohwan, Jung, Jae-Hyun
Abstract
Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as black boxes, limiting their ability to identify the underlying causes of lower performance: when a model fails a spatial reasoning task, it remains difficult to ascertain whether the hurdle is perceptual, such as recognizing object boundaries, or cognitive, such as reasoning about occlusion to infer hidden geometry. We introduce Spatial-IQ, a hierarchical diagnostic framework that decomposes object counting in stacked 3D structures into 9 perceptual and cognitive sub-tasks organized by the developmental stages of human spatial cognition, with mental rotation as an additional target probe. Using NVIDIA Isaac Sim, we procedurally generated a diverse dataset of roughly 80,000 stacked 3D structures with per-task ground truth. We evaluate models across three output formats (free-response text, multiple-choice images, and image editing) alongside a human baseline. The Spatial-IQ framework shows that top-performing models often succeed at the target task (object counting) without succeeding on the lower-level sub-tasks intended to support it, and that models differ in how much of these hierarchical chains they preserve, often revealing shortcut behavior that raw target-task accuracy alone would obscure. Finally, we demonstrate that training models with chain-of-thought (CoT) supervision over our hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition as both a diagnostic tool and a training signal.
Chinese Translation
多模态大型语言模型(MLLMs)在视觉解读方面表现出色,但在空间推理任务上却未能达到人类的可靠水平。现有基准测试将这些模型视为黑箱,限制了它们识别性能较低的根本原因的能力:当模型在空间推理任务中失败时,难以判断障碍是感知性的(例如,识别物体边界)还是认知性的(例如,推理遮挡以推断隐藏几何形状)。我们提出了Spatial-IQ,一个层次诊断框架,将堆叠的3D结构中的物体计数分解为9个感知和认知子任务,这些子任务按照人类空间认知的发展阶段组织,并将心理旋转作为额外的目标探测。使用NVIDIA Isaac Sim,我们程序生成了大约80,000个堆叠3D结构的多样化数据集,并为每个任务提供了真实标签。我们在三种输出格式(自由响应文本、多项选择图像和图像编辑)上评估模型,并与人类基线进行比较。Spatial-IQ框架显示,表现最佳的模型往往在目标任务(物体计数)上成功,但在支持该任务的低层次子任务上却未能成功,并且模型在保留这些层次链的程度上存在差异,通常揭示了仅通过目标任务准确性无法察觉的捷径行为。最后,我们证明,通过对我们的层次子任务进行思维链(CoT)监督训练模型,并结合可验证奖励的强化学习,显著提高了子任务之间的空间一致性和目标任务的准确性,支持了所提出的分解作为诊断工具和训练信号的价值。
cs.CV / 43 / 2607.22891

AdaKAN: A dual-branch adaptive Kolmogorov-Arnold network for medical image segmentation

AdaKAN:一种用于医学图像分割的双分支自适应Kolmogorov-Arnold网络
Alzu'bi, Dalia, Bhattacharyya, Deep, Ayub, Ali, Hamza, A. Ben
Abstract
Medical image segmentation is a fundamental task in computer-aided diagnosis, yet it remains challenging due to the complexity of anatomical structures and the variability across imaging modalities. In this paper, we propose AdaKAN, an Adaptive Kolmogorov-Arnold Network (KAN) that synergistically integrates convolutional operations with a novel efficient KAN (EffiKAN) block, comprised of an efficient attention mechanism and an adaptive KAN (AdaptKAN) module. This module features a dual-branch design: one branch employs a KAN layer with Bernstein polynomial activations for globally smooth and stable function approximation, while the other branch performs channel-wise refinement through projection operations and adaptive scaling. AdaKAN adopts a U-shaped architecture that effectively captures both long-range dependencies and fine-grained local features, overcoming the limitations of conventional convolutional and Transformer-based segmentation models. Skip connections are employed to preserve spatial details during encoding and facilitate accurate reconstruction during decoding. Extensive experiments conducted on diverse medical imaging datasets demonstrate that AdaKAN achieves state-of-the-art performance in segmentation accuracy.
Chinese Translation
医学图像分割是计算机辅助诊断中的一项基础任务,但由于解剖结构的复杂性和成像模态的变异性,这一任务仍然具有挑战性。本文提出了AdaKAN,一种自适应Kolmogorov-Arnold网络(KAN),它将卷积操作与一种新颖的高效KAN(EffiKAN)模块协同集成,该模块由高效注意力机制和自适应KAN(AdaptKAN)模块组成。该模块具有双分支设计:一条分支采用具有Bernstein多项式激活的KAN层,以实现全局平滑和稳定的函数逼近,而另一条分支通过投影操作和自适应缩放进行通道级的细化。AdaKAN采用U形架构,有效捕捉长距离依赖关系和细粒度局部特征,克服了传统卷积和基于Transformer的分割模型的局限性。在编码过程中采用跳跃连接以保持空间细节,并在解码过程中促进准确重建。在多样的医学成像数据集上进行的广泛实验表明,AdaKAN在分割精度方面达到了最先进的性能。
cs.CV / 44 / 2607.22913

Small-Pollinator Detection in Cluttered Field Video

杂乱场景视频中的小型授粉者检测
Onal, Onur, Chen, Chen
Abstract
Detecting pollinators in field video is challenging: targets are small, visually similar, and observed against cluttered vegetation under blur and occlusion. We present a systematic empirical study of small-pollinator detection under a practical single-GPU compute budget. Using the BuzzSpot challenge dataset, we compare YOLO and RF-DETR models across input resolutions and evaluate sliced inference, class-gated fusion, size-routed ensembling, and post-hoc temporal processing. RF-DETR Large at 1344-pixel resolution achieved our best hidden-test result, reaching 0.405 mAP50:95 and outperforming the 1120-pixel model (0.379) and the best single-model YOLO26m baseline (0.366). The strongest gains came from adopting RF-DETR and increasing its input resolution, indicating that detector choice and input resolution were more effective levers than added inference-time complexity; the resolution gain was strongest for small objects and the rarer bumblebee and moth classes. Sliced-inference fusion, size-routed ensembling, and warm-started 1536-pixel continuation did not surpass this result, while post-hoc temporal processing did not improve the leaked diagnostic evaluation. Error analysis identified bee-hoverfly discrimination as the clearest remaining bottleneck: neighboring frames rarely supplied correctly classified hoverfly evidence for post-hoc correction. These findings motivate learned feature-level temporal aggregation before the final classification decision.
Chinese Translation
在场地视频中检测授粉者具有挑战性:目标小,视觉相似,并且在模糊和遮挡的杂乱植被中观察到。我们在实际的单GPU计算预算下,进行了系统的实证研究,探讨小型授粉者的检测。使用BuzzSpot挑战数据集,我们比较了YOLO和RF-DETR模型在不同输入分辨率下的表现,并评估了切片推理、类别门控融合、大小路由集成和后处理时间处理。RF-DETR Large在1344像素分辨率下取得了我们最佳的隐藏测试结果,达到0.405 mAP50:95,超越了1120像素模型(0.379)和最佳单模型YOLO26m基线(0.366)。最大的提升来自于采用RF-DETR并提高其输入分辨率,表明检测器选择和输入分辨率比增加推理时间复杂性更为有效;对于小物体以及较为稀有的大黄蜂和蛾类,分辨率提升最为显著。切片推理融合、大小路由集成和热启动的1536像素继续并未超过这一结果,而后处理时间处理未能改善泄漏的诊断评估。错误分析表明蜜蜂与蜻蜓的区分是最明显的瓶颈:相邻帧很少提供正确分类的蜻蜓证据以进行后处理修正。这些发现促使我们在最终分类决策之前进行学习特征级的时间聚合。
cs.CV / 45 / 2607.22919

Controlling Embedding Spaces with Text-Conditioned Transformations

通过文本条件变换控制嵌入空间
Fioresi, Joseph, Heilbron, Fabian Caba, Nathani, Pankaj, Shah, Mubarak, Kafle, Kushal
Abstract
Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compress high-level semantics into a single vector, which comes at the cost of primarily expressing a dominant semantics like main object while suppressing other important attributes such as camera angle or color tone. We propose a text-conditioned transformation of visual embeddings that makes such attributes explicitly accessible. Given a natural language description of an attribute category (e.g., "color" or "art style"), a network generates an affine transformation that emphasizes the specified attribute. Conditioning on text enables it to learn many attributes simultaneously, accessing them at inference time through an intuitive interface. The network is trained to align transformed embeddings with the frozen latent space, enabling retrieval using existing large-scale embeddings without any re-encoding. When applied to a full set, the same mechanism transforms the latent space for attribute disentanglement tasks such as multi-clustering. By operating directly in latent space, our method provides a unified and efficient framework for controlling embedding spaces, demonstrating state-of-the-art performance across both attribute-based retrieval and multi-attribute organization tasks with near-zero inference cost. Project page: https://joefioresi718.github.io/ControlEmbed_webpage/
Chinese Translation
像 CLIP 这样的多模态嵌入空间使得语义相似性检索和跨模态零样本分类等强大功能成为可能。这些嵌入将高级语义压缩为单一向量,但这也导致主要表达主导语义(如主要物体),而抑制其他重要属性(如摄像机角度或色调)。我们提出了一种视觉嵌入的文本条件变换,使这些属性能够被明确访问。给定一个属性类别的自然语言描述(例如,“颜色”或“艺术风格”),网络生成一个仿射变换,强调指定的属性。基于文本的条件使其能够同时学习多个属性,并通过直观接口在推理时访问它们。该网络经过训练,使变换后的嵌入与冻结的潜在空间对齐,从而能够利用现有的大规模嵌入进行检索,而无需重新编码。当应用于完整集合时,同一机制可用于属性解耦任务,如多聚类。通过直接在潜在空间中操作,我们的方法提供了一个统一且高效的框架来控制嵌入空间,在基于属性的检索和多属性组织任务中展示了最先进的性能,几乎没有推理成本。项目页面:https://joefioresi718.github.io/ControlEmbed_webpage/
cs.CV / 46 / 2607.22924

Layering Virtual Try-On

分层虚拟试穿
Feng, Chun, Chen, Bowei, Shan, Mengyi, Kemelmacher-Shlizerman, Ira
Abstract
In the real world, fashion is about layering: adding a jacket over a shirt, or a sequence of adding and removing layers, rather than just a single-layer swap. This fundamental real-world task remains a challenge in existing Virtual Try-On (VTON) methods, which excel at single-layer replacement but are not designed to layer or de-layer an existing outfit. This paper proposes Layering Virtual Try-On (LVTON), a layering benchmark and method that preserves an existing outfit while enabling sequential layering. We find that current VTON paradigms are fundamentally ill-equipped for LVTON, as their reliance on cloth-agnostic representations and single-item datasets discards essential layering context. Our key insight is that the LVTON challenge must be disentangled into two distinct competencies: (1) General VTON Priors (e.g., deformation, identity preservation) and (2) Specific Layering Knowledge (e.g., layering order and occlusion reasoning). First, our model obtains general VTON priors by being trained on data produced by an automatic data generation pipeline that synthesizes samples from fashion videos via segmentation and inpainting. Second, the model is fine-tuned on a small, dedicated LVTON dataset to learn the layering logic. Our method achieves state-of-the-art results on our LVTON benchmark and demonstrates superior generalizability on traditional VTON benchmarks, setting new state-of-the-art results when fine-tuned and exhibiting zero-shot capabilities.
Chinese Translation
在现实世界中,时尚是关于分层的:在衬衫上加一件夹克,或是通过添加和去除层次的顺序,而不仅仅是单层的替换。这一基本的现实任务在现有的虚拟试穿(Virtual Try-On, VTON)方法中仍然是一个挑战,这些方法擅长于单层替换,但并未设计用于对现有服装进行分层或去层。本文提出了分层虚拟试穿(Layering Virtual Try-On, LVTON),这是一个分层基准和方法,旨在保留现有服装的同时实现顺序分层。我们发现当前的VTON范式在LVTON方面根本不具备能力,因为它们依赖于与布料无关的表示和单一物品数据集,从而忽略了重要的分层上下文。我们的关键见解是,LVTON挑战必须拆解为两个不同的能力:(1)一般VTON先验(例如,变形、身份保持)和(2)特定的分层知识(例如,分层顺序和遮挡推理)。首先,我们的模型通过在一个自动数据生成管道上进行训练来获得一般VTON先验,该管道通过分割和修复从时尚视频中合成样本。其次,该模型在一个小型专用的LVTON数据集上进行微调,以学习分层逻辑。我们的方法在我们的LVTON基准上实现了最先进的结果,并在传统VTON基准上展示了更优的泛化能力,在微调时设定了新的最先进结果,并展现了零样本能力。
cs.CV / 47 / 2607.22959

HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale

HALLELUAI:一种具备幻觉感知的人工智能系统,用于大规模超真实图像到视频生成
Sakpal, Aniket, Jiang, Yang, Davoudi, Rouzbeh, Hassantabar, Shayan, Najmabadi, Mani
Abstract
AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quality control remains a major constraint to scaling production. We present HALLELUAI, an end-to-end system that moderates and regenerates image-to-video outputs to meet expert-level creative standards and deliver ultra-realistic videos with consistent end-user quality of experience (QoE) at scale. The system integrates a video moderation module that evaluates frame-level aesthetics, temporal motion fidelity, and fine-grained hallucination risks relative to the source image, with an agentic regeneration module that iteratively fixes failures through prompt refinement, controlled camera adjustments, targeted model or image switching, and structured retry strategies. The moderation logic is aligned with domain-specific creative guidelines and produces granular, machine-actionable feedback that directly drives regeneration. In human-in-the-loop evaluations with creative experts, HALLELUAI shows strong alignment and reliably outputs ultra-realistic, production-grade videos suitable for product and marketing placements at scale. This framework advances trustworthy AI generated video content by enforcing visual realism, brand safety, and strict input-image fidelity while enabling image-to-video generation at scale.
Chinese Translation
人工智能生成的视频在营销、产品叙事和创意工作流程中越来越多地被使用,但自动化的高精度质量控制仍然是扩大生产的主要限制。我们提出了HALLELUAI,这是一种端到端系统,能够调节和再生图像到视频的输出,以满足专家级的创意标准,并在大规模上提供具有一致终端用户体验(QoE)的超真实视频。该系统集成了一个视频调节模块,该模块评估帧级美学、时间运动保真度以及相对于源图像的细粒度幻觉风险,并配备一个代理再生模块,通过提示优化、受控相机调整、针对性的模型或图像切换以及结构化重试策略,迭代修复失败。调节逻辑与特定领域的创意指南相一致,并生成细致的、可机器操作的反馈,直接驱动再生。在与创意专家的人工参与评估中,HALLELUAI显示出强大的对齐性,并可靠地输出适合大规模产品和营销投放的超真实、生产级视频。该框架通过强制视觉真实感、品牌安全性和严格的输入图像保真度,推动了可信赖的人工智能生成视频内容,同时实现了大规模的图像到视频生成。
cs.CV / 48 / 2607.22973

mmSimPrior: Learning Simulation Priors for Data-Efficient Real-World Generalizable Radar-Based Human Motion Reconstruction

mmSimPrior:用于数据高效的现实世界可泛化雷达基础人类运动重建的仿真先验学习
Guo, Cheng, Cao, Qiming, Xu, Shengkai, Xie, Haoyu, Su, Kaixiang, Wang, Pu, Xue, Hongfei
Abstract
Millimeter-wave (mmWave) radar offers privacy-preserving and lighting-robust sensing for human motion reconstruction, but learning models that generalize across real deployments require diverse paired radar-motion data that are costly to collect. Simulation provides scalable supervision, yet models trained on clean synthetic signals transfer poorly because of multipath, clutter, response statistics, and resolution degradation. We present mmSimPrior, a simulation-pretrained framework that factorizes transferable knowledge into signal, motion, and radar-to-motion mapping priors. A multi-modal signal encoder is pretrained with a physics-informed domain-randomization curriculum that emulates propagation- and acquisition-level variations, while a joint-temporal tokenizer learns a discrete prior over plausible human motion. A shared mapping prior supports classification over a learned motion codebook for constrained zero-shot reconstruction and continuous regression for flexible limited-data adaptation. We further construct a 4.2M-frame, 31K-sequence dataset suite and introduce a No-Overlap Setting that excludes repeated complete subject-environment-location-motion configurations across adaptation and test. Experiments on mmSimPrior-Real and RT-Pose demonstrate consistent gains: with only 24 paired real sequences, mmSimPrior-Reg reduces MPJPE by 24.7% to 39.0% over the strongest baseline across the three environments, while mmSimPrior-Cls reduces zero-shot MPJPE by 8.5% without finetuning.
Chinese Translation
毫米波(mmWave)雷达为人类运动重建提供了隐私保护和抗光照干扰的感知能力,但在真实部署中,学习能够泛化的模型需要多样化的配对雷达-运动数据,而这些数据的收集成本高昂。仿真提供了可扩展的监督,但在干净的合成信号上训练的模型由于多径、杂波、响应统计和分辨率降级而转移效果不佳。我们提出了mmSimPrior,一个仿真预训练框架,将可转移知识分解为信号、运动和雷达-运动映射先验。多模态信号编码器通过物理知识驱动的领域随机化课程进行预训练,该课程模拟传播和采集级别的变化,同时联合时间标记器学习可行人类运动的离散先验。共享映射先验支持在学习的运动代码本上进行分类,以实现受限的零样本重建和灵活的有限数据适应的连续回归。我们进一步构建了一个包含420万帧、31000个序列的数据集,并引入了一个无重叠设置,排除了在适应和测试中重复的完整主体-环境-位置-运动配置。对mmSimPrior-Real和RT-Pose的实验表明了一致的提升:仅使用24个配对的真实序列,mmSimPrior-Reg在三个环境中将MPJPE降低了24.7%至39.0%,而mmSimPrior-Cls在不进行微调的情况下将零样本MPJPE降低了8.5%。
cs.CV / 49 / 2607.22994

Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning

打破合成-真实领域捷径以实现无训练生成重放的类增量学习
Zhang, Tao, Fan, Qixuan, Liang, Yiyuan, Wang, Yanjie, Yan, Song, Tian, Tian, Zhou, Jiahuan, Yan, Luxin, Zhong, Sheng, Zou, Xu
Abstract
Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concerns regarding privacy and storage. Thus, generative replay has emerged as a viable alternative, synthesizing old data using frozen pretrained text-to-image (T2I) models without any extra training. However, we observe that directly mixing synthetic old-class data with real new-class data during incremental training leads to significant performance degradation. This issue stems from a "domain shortcut", where models rely on domain-discriminative features instead of semantic class cues. To address this, we propose DREAM ($\underline{\mathbf{D}}$omain-$\underline{\mathbf{R}}$egularized $\underline{\mathbf{E}}$xemplar-free $\underline{\mathbf{A}}$lignment $\underline{\mathbf{M}}$odel), which uses a training-free generator to synthesize old-class data and eliminates domain shortcut via subspace rectification and orthogonal projection, while reinforcing semantic alignment through real-anchored prototype regularization. Extensive experiments on 4 datasets demonstrate that DREAM outperforms existing exemplar-free CIL methods and achieves state-of-the-art performance. Our source code is available at https://github.com/Light-ZhangTao/DREAM.
Chinese Translation
类增量学习(CIL)要求模型在不断获取新知识的同时避免灾难性遗忘。虽然示例重放有效,但它引发了关于隐私和存储的担忧。因此,生成重放作为一种可行的替代方案应运而生,利用冻结的预训练文本到图像(T2I)模型合成旧数据,而无需额外训练。然而,我们观察到在增量训练过程中,直接将合成的旧类数据与真实的新类数据混合会导致显著的性能下降。这个问题源于“领域捷径”,模型依赖于领域判别特征而不是语义类线索。为了解决这个问题,我们提出了DREAM($ extbf{D}$omain-$ extbf{R}$egularized $ extbf{E}$xemplar-free $ extbf{A}$lignment $ extbf{M}$odel),该模型使用无训练生成器合成旧类数据,并通过子空间校正和正交投影消除领域捷径,同时通过真实锚定原型正则化增强语义对齐。在4个数据集上的大量实验表明,DREAM的表现优于现有的无示例CIL方法,并达到了最先进的性能。我们的源代码可在https://github.com/Light-ZhangTao/DREAM获取。
cs.CV / 50 / 2607.23023

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

OmniMate:用于互动虚拟形象的开放式实时音视频生成
Song, Quanyue, He, Yishan, Ding, Yanbo, He, Zhixiang, Li, Yongxiang, Jiang, Caigui, Guo, Zhi Zhi
Abstract
Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for interactive avatar systems. However, extending these models to real-time interactive streaming remains challenging, as the generation horizon is unknown in advance and cross-modal identity consistency gradually degrades during long-term generation. To address these challenges, we propose OmniMate, a unified framework for open-ended real-time interactive audio-visual avatar generation. OmniMate jointly synthesizes visual content, speech, and audio effects in real time, enabling natural and immersive multi-turn interactions. To achieve adaptive response progression, we introduce a Generation Progress Controller (GPC) that explicitly models the generation progress of each streaming chunk, allowing the model to complete responses according to the desired progress and achieve seamless transitions between execution and listening states. To preserve long-term cross-modal identity consistency, we propose a Multi-Reference Conditioning Module (MRCM), which leverages multiple reference images and a reference speech segment to provide persistent visual and speaker identity cues throughout long-duration streaming interactions. Extensive experiments on an interaction-oriented adaptation of VerseBench demonstrate that OmniMate achieves high-quality, low-latency streaming generation while maintaining strong long-term audio-visual consistency. The results further show that OmniMate supports realistic, coherent, and responsive interactive avatar experiences over extended multi-turn conversations.
Chinese Translation
最近基于扩散的生成模型的进展使得实时音频驱动的虚拟形象生成和统一的音视频合成成为可能,为互动虚拟形象系统提供了有希望的基础。然而,将这些模型扩展到实时互动流媒体仍然面临挑战,因为生成的时间范围无法提前确定,并且在长期生成过程中跨模态身份一致性逐渐降低。为了解决这些挑战,我们提出了OmniMate,一个用于开放式实时互动音视频虚拟形象生成的统一框架。OmniMate实时联合合成视觉内容、语音和音效,实现自然且沉浸的多轮互动。为了实现自适应的响应进展,我们引入了生成进展控制器(Generation Progress Controller, GPC),该控制器明确建模每个流媒体块的生成进展,使模型能够根据所需进展完成响应,并在执行和监听状态之间实现无缝过渡。为了保持长期的跨模态身份一致性,我们提出了多参考条件模块(Multi-Reference Conditioning Module, MRCM),该模块利用多个参考图像和一个参考语音片段,在长期流媒体互动中提供持续的视觉和说话者身份线索。在针对VerseBench的互动导向适应的广泛实验中,结果表明OmniMate实现了高质量、低延迟的流媒体生成,同时保持了强大的长期音视频一致性。结果进一步表明,OmniMate支持在扩展的多轮对话中实现真实、一致和响应迅速的互动虚拟形象体验。
cs.CV / 51 / 2607.23024

When Less Is More: A Controlled Benchmark of Lightweight CNNs for Satellite Land-Cover Segmentation on DeepGlobe

少即是多:在 DeepGlobe 上对轻量级卷积神经网络进行的卫星土地覆盖分割的受控基准测试
Rehman, Atiq Ur, Donovan, Joseph Michael
Abstract
High-resolution satellite imagery is the backbone of good land-cover classification, and without that, environmental monitoring, urban planning, and sustainable resource management all fall short. Deep learning architectures perform well in semantic segmentation, but the efficiency-accuracy trade-off across classical convolutional encoders is not well quantified under controlled, reproducible conditions. This study compares five architectures VGG16, MobileNetV2, InceptionV3, AlexNet, and CNN on the DeepGlobe Land Cover Classification dataset using three progressively optimized iterations to isolate regularisation, transfer learning, and architectural depth. To ensure performance differentials reflect architectural properties, all experiments used identical preprocessing, hyperparameter, and training protocols without data augmentation or class-imbalance correction. At 24.98 MB, MobileNetV2_v1 had the highest overall accuracy (0.7906) and mean Intersection over Union (0.4625), outperforming deeper alternatives like InceptionV3_v2 (125.17 MB, accuracy 0.7610) and VGG16_v2 (71.13 MB, accuracy 0.7653). Class-wise analysis showed strength in urban, agricultural, and water categories, but rangeland-barren confusion showed that architectural optimization alone cannot optimize spectrally similar minority classes. Strong spatial generalization and crisp boundary delineation were confirmed on held-out test imagery, validating operational applicability. These results show that lightweight, transfer-learned models can match or outperform deeper models in resource-constrained remote-sensing environments, enabling scalable land-cover mapping.
Chinese Translation
高分辨率卫星影像是良好土地覆盖分类的基础,缺乏这一点,环境监测、城市规划和可持续资源管理都将受到影响。深度学习架构在语义分割中表现良好,但在受控、可重复的条件下,经典卷积编码器的效率与准确性之间的权衡尚未得到充分量化。本研究比较了五种架构:VGG16、MobileNetV2、InceptionV3、AlexNet 和 CNN,使用 DeepGlobe 土地覆盖分类数据集,通过三个逐步优化的迭代来隔离正则化、迁移学习和架构深度。为了确保性能差异反映架构特性,所有实验均使用相同的预处理、超参数和训练协议,且未进行数据增强或类别不平衡校正。在 24.98 MB 的体积下,MobileNetV2_v1 具有最高的整体准确率(0.7906)和平均交并比(0.4625),优于更深的替代方案,如 InceptionV3_v2(125.17 MB,准确率 0.7610)和 VGG16_v2(71.13 MB,准确率 0.7653)。按类别分析显示,在城市、农业和水体类别中表现出色,但在草原-荒地的混淆中表明,仅靠架构优化无法优化光谱相似的少数类。在保留的测试影像上确认了强大的空间泛化能力和清晰的边界划分,验证了其操作适用性。这些结果表明,轻量级的迁移学习模型可以在资源受限的遥感环境中与更深的模型相匹配或超越,从而实现可扩展的土地覆盖制图。
cs.CV / 52 / 2607.23046

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs

高分辨率多模态大语言模型中高效视觉标记剪枝的结构冗余建模
Song, Jouwon, Kim, Woohyeong, Kong, Kyeongbo
Abstract
Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that introduces severe latency bottlenecks. While token pruning mitigates this issue, state-of-the-art subset-optimization methods typically rely on iterative subset construction to jointly capture visual diversity and instruction relevance. As visual token counts scale, this sequential dependency introduces significant selection overhead, severely limiting the translation of theoretical FLOPs reductions into actual wall-clock speedups. To address this limitation, we propose Single-Forward Pruner (SFPruner), a structural reformulation of visual token pruning that embeds redundancy control directly into the scoring space, bypassing the need for iterative combinatorial optimization. Our non-iterative framework achieves redundancy-aware importance selection in a single forward pass through two complementary mechanisms. First, to attenuate redundancy at the covariance level, we introduce a semantics-guided ridge leverage scheme. By integrating instruction relevance and visual saliency, this mechanism suppresses dominant covariance directions and mitigates representation bias. Second, ranking-based directional masking resolves residual overlap through asymmetric similarity competition, where higher-scoring tokens explicitly suppress redundant lower-scoring alternatives via parallel tensor operations. Extensive evaluations demonstrate that our approach maintains stable selection costs, reducing the token selection process by up to 110 ms, from 112.4 ms to just 2.5 ms at 512 tokens in Qwen2.5-VL. This structural efficiency successfully translates theoretical token reductions into tangible inference speedups while preserving highly competitive performance against state-of-the-art techniques under aggressive compression.
Chinese Translation
近期的高分辨率多模态大语言模型(MLLMs)在每个输入中生成数千个视觉标记,导致视觉标记的爆炸性增长,从而引入严重的延迟瓶颈。虽然标记剪枝可以缓解这一问题,但最先进的子集优化方法通常依赖于迭代子集构建,以共同捕捉视觉多样性和指令相关性。随着视觉标记数量的增加,这种顺序依赖性引入了显著的选择开销,严重限制了理论FLOPs减少转化为实际时钟加速的能力。为了解决这一限制,我们提出了单次前向剪枝器(Single-Forward Pruner,SFPruner),这是一种视觉标记剪枝的结构重构,直接将冗余控制嵌入评分空间,避免了迭代组合优化的需要。我们的非迭代框架通过两种互补机制在一次前向传递中实现了考虑冗余的关键性选择。首先,为了在协方差层面减轻冗余,我们引入了一种语义引导的岭杠杆方案。通过整合指令相关性和视觉显著性,该机制抑制了主导的协方差方向,减轻了表示偏差。其次,基于排名的方向性掩蔽通过不对称相似性竞争解决了残余重叠,其中得分较高的标记通过并行张量操作显式抑制冗余的低得分替代品。广泛的评估表明,我们的方法保持了稳定的选择成本,将标记选择过程的时间从112.4毫秒减少到仅2.5毫秒,减少幅度高达110毫秒(在512个标记的Qwen2.5-VL中)。这种结构效率成功地将理论标记减少转化为切实的推理加速,同时在激进压缩下保持与最先进技术的高度竞争性能。
cs.CV / 53 / 2607.23052

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models

相似性并非逻辑:双编码器视觉-语言模型的分解推理
Alshehri, Sultan, Yang, Zhantao, Zhang, Han, Savvides, Marios
Abstract
Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person" retrieve images containing both, even when concept detection is reliable. We trace this to an interface-level Bag-of-Concepts effect, where similarity scores approximate mean pooling of concept evidence regardless of operators. Although operator-dependent signals exist in text embeddings, they are too weak or misaligned to affect rankings. Fine-tuning does not reliably resolve this failure because the dominant bottleneck is how similarity aggregates evidence rather than what encoders represent. We propose factored inference, which separates evidence extraction from constraint execution, and introduce LCSE (Logic-Constrained Score Editing), a training-free method that executes constraints externally using concept scores from frozen encoders. We also introduce FACTOR-Bench, where LCSE achieves 85.5% accuracy versus 73.2% for the best fine-tuned baseline, 90.7% when applied to SigLIP 2, and improves NegBench COCO MCQ accuracy from 27.2% to 65.2% while preserving retrieval performance.
Chinese Translation
双编码器视觉-语言模型(VLMs)暴露出一种相似性接口,使得零样本检索成为可能,但未能满足组合约束:像“伞和没有人”的查询会检索到同时包含这两者的图像,即使概念检测是可靠的。我们将此归因于接口层面的概念包效应(Bag-of-Concepts effect),在该效应下,相似性得分近似于概念证据的均值池化,而不考虑操作符。尽管文本嵌入中存在依赖于操作符的信号,但它们过于微弱或不对齐,无法影响排名。微调并不能可靠地解决这一问题,因为主要瓶颈在于相似性如何聚合证据,而非编码器所表示的内容。我们提出了分解推理(factored inference),将证据提取与约束执行分开,并引入了逻辑约束得分编辑(Logic-Constrained Score Editing, LCSE),这是一种无训练的方法,利用来自冻结编码器的概念得分在外部执行约束。我们还介绍了FACTOR-Bench,在该基准上,LCSE的准确率达到85.5%,而最佳微调基线为73.2%,在应用于SigLIP 2时达到90.7%,并将NegBench COCO MCQ的准确率从27.2%提高到65.2%,同时保持检索性能。
cs.CV / 54 / 2607.23070

DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding

DishSeg24k:一个用于食品分割的大规模基准测试,采用随机专家解码
Wang, Yilin, Shi, Haochen, Chen, Guanyu, Min, Weiqing, Zheng, Jinkai, Yan, Chenggang, Jiang, Shuqiang
Abstract
Food segmentation is essential for applications such as intelligent catering, dietary assessment, and recommendation. However, existing benchmarks fail to capture the complexity of real-world dining scenes. The challenges of dense inter-dish overlap, fine-grained class similarity, and extreme long-tail class distributions exceed the fidelity of current datasets. To fill this gap, we introduce \textbf{DishSeg24k}, a large-scale dish-level segmentation benchmark with 24,096 images, 112,281 instances, and 278 fine-grained categories in real-world dining environments. Based on DishSeg24k, we further propose \textbf{Food Expert-Adaptive Segmentation Transformers (FEAST)} to address these challenges. FEAST models query-based decoding as a Markov Decision Process (MDP), where each decoder layer update is treated as a sequential decision step that explores uncertainty along dish boundaries. We further redesign the decoder with a reinforcement learning (RL)-guided Mixture-of-Experts (MoE) module, in which a dual-critic decoupled optimization scheme separates task-oriented query refinement from structure-aware expert routing. This design promotes expert specialization and prevents expert collapse under long-tail category distributions. Finally, extensive experiments on DishSeg24k demonstrate the state-of-the-art performance of FEAST, which outperforms previous methods by {+3.21\%} mIoU, {+3.68\%} mDice, and {+4.00\%} mAcc, respectively. We further validate the effectiveness of FEAST on FoodSeg103. The dataset and code will be publicly released.
Chinese Translation
食品分割对于智能餐饮、饮食评估和推荐等应用至关重要。然而,现有的基准测试未能捕捉现实就餐场景的复杂性。菜肴之间的密集重叠、细粒度类别相似性以及极端长尾类别分布的挑战超出了当前数据集的保真度。为填补这一空白,我们引入了 extbf{DishSeg24k},这是一个大规模的菜肴级分割基准,包含24,096张图像、112,281个实例和278个细粒度类别,涵盖现实就餐环境。基于DishSeg24k,我们进一步提出了 extbf{食品专家自适应分割变换器(FEAST)} 来应对这些挑战。FEAST模型将基于查询的解码视为马尔可夫决策过程(MDP),其中每个解码器层的更新被视为探索菜肴边界不确定性的顺序决策步骤。我们进一步重新设计了解码器,采用强化学习(RL)指导的专家混合(MoE)模块,其中双重评论员解耦优化方案将任务导向的查询细化与结构感知的专家路由分开。这一设计促进了专家的专业化,并防止了在长尾类别分布下的专家崩溃。最后,在DishSeg24k上的大量实验表明,FEAST的性能达到了最先进水平,分别比之前的方法提高了{+3.21\%} mIoU、{+3.68\\%} mDice和{+4.00\\%} mAcc。我们进一步在FoodSeg103上验证了FEAST的有效性。数据集和代码将公开发布。
cs.CV / 55 / 2607.23078

Inverse Bayesian Inference for Extracting Lesion Dynamics from Longitudinal Spectral CT

用于从纵向光谱CT提取病变动态的逆贝叶斯推断
Förner, Lukas, Wördehoff, Melina, Steffens, Julian, Schmutz, Maximilian, Claus, Rainer, Decker, Josua, Kröncke, Thomas, Tehlan, Kartikay, Wendler, Thomas
Abstract
Longitudinal medical imaging captures temporal evolution of lesions, yet extracting the underlying dynamical parameters governing this evolution remains challenging. We propose an inverse Bayesian framework for inferring lesion dynamics from longitudinal spectral CT. We decompose spectral feature ($x$) evolution into three components: \begin{equation*} \frac{dx_i}{dt} = A_i x_i + B \cdot n + C \cdot \Delta x_{\text{sat}} \end{equation*} where $A_i$ captures intrinsic dynamics (lesion-autonomous evolution), $B$ captures local environment tumour burden (organ tumour burden through satellite count coupling), and $C$ captures environment/satellite state change (i.e., whether surrounding lesions move similarly or not). We demonstrate the framework on photon-counting NSCLC CT data from metastases, recovering distinct dynamical regimes: lung lesions exhibit significant satellite count coupling ($B=-0.34$, $p<0.05$) suggesting competitive dynamics, while liver lesions show synergistic satellite behaviour coupling ($C\approx+1.0$, $p<0.05$). Synthetic validation confirms parameter recovery, and cross-coupling analysis validates that our method detects non-zero coupling when present. This work establishes inverse dynamical inference as a principled methodology for extracting interpretable parameters from longitudinal imaging, moving beyond static feature extraction toward mechanistic characterisation of lesion behaviour. The code and data are available at: https://github.com/lukasf98/inverse-bayesian-inference
Chinese Translation
纵向医学成像捕捉病变的时间演变,但提取支配这种演变的潜在动态参数仍然具有挑战性。我们提出了一种逆贝叶斯框架,用于从纵向光谱CT推断病变动态。我们将光谱特征($x$)的演变分解为三个组成部分:egin{equation*} rac{dx_i}{dt} = A_i x_i + B ullet n + C ullet riangle x_{ ext{sat}} ext{,} egin{equation*} 其中$A_i$捕捉内在动态(病变自主演变),$B$捕捉局部环境肿瘤负担(通过卫星计数耦合的器官肿瘤负担),而$C$捕捉环境/卫星状态变化(即,周围病变是否以相似方式移动)。我们在转移性非小细胞肺癌(NSCLC)CT数据上演示了该框架,恢复了不同的动态状态:肺部病变表现出显著的卫星计数耦合($B=-0.34$,$p<0.05$),这表明竞争性动态,而肝脏病变则显示出协同卫星行为耦合($C ext{近似}=+1.0$,$p<0.05$)。合成验证确认了参数恢复,而交叉耦合分析验证了我们的方法在存在时能够检测到非零耦合。这项工作确立了逆动态推断作为一种原则性的方法论,用于从纵向成像中提取可解释的参数,超越静态特征提取,朝着病变行为的机制特征化迈进。代码和数据可在以下网址获取:https://github.com/lukasf98/inverse-bayesian-inference
cs.CV / 56 / 2607.23096

SHReg: Strictly Rotation-Equivariant Point Cloud Registration via Spherical Harmonics

SHReg:通过球谐函数实现严格旋转等变的点云配准
Wang, Chongjian, Gao, Junjie
Abstract
Point cloud registration critically depends on local features that are both distinctive and robust to arbitrary 3D rotations. Existing learning-based methods typically approximate rotation invariance via fragile local reference frames or extensive data augmentation, providing only empirical invariance and often degrading under unseen rotational transformations. In this paper, we propose SHReg, a strictly rotation-equivariant point cloud registration framework grounded in the representation theory of $SO(3)$. By representing local geometric features as irreducible representations of $SO(3)$, SHReg guarantees exact equivariance under arbitrary rotations without relying on local reference frames. Built upon a spherical-harmonics-based equivariant backbone, SHReg jointly learns rotation-invariant descriptors for robust correspondence matching and rotation-equivariant features that preserve fine-grained orientation information. The preserved equivariant structure enables each correspondence to directly hypothesize a rigid transformation, reducing reliance on large-scale hypothesis sampling in conventional RANSAC-based pipelines and leading to improved robustness under challenging rotational variations. Extensive experiments on 3DMatch, 3DLoMatch, and KITTI demonstrate that SHReg consistently outperforms state-of-the-art methods in registration accuracy, particularly under large rotational perturbations.
Chinese Translation
点云配准在很大程度上依赖于既具有辨识度又对任意三维旋转具有鲁棒性的局部特征。现有的基于学习的方法通常通过脆弱的局部参考框架或广泛的数据增强来近似旋转不变性,仅提供经验性的等变性,并且在未见过的旋转变换下往往会退化。本文提出了SHReg,一种基于$SO(3)$表示理论的严格旋转等变点云配准框架。通过将局部几何特征表示为$SO(3)$的不可约表示,SHReg保证在任意旋转下的精确等变性,而无需依赖局部参考框架。SHReg基于球谐函数的等变主干网络共同学习鲁棒的旋转不变描述符以实现精确的对应匹配,以及保留细粒度方向信息的旋转等变特征。保留的等变结构使得每个对应关系能够直接假设刚性变换,减少了对传统RANSAC管道中大规模假设采样的依赖,从而在具有挑战性的旋转变化下提高了鲁棒性。在3DMatch、3DLoMatch和KITTI上的大量实验表明,SHReg在配准精度上始终优于最先进的方法,特别是在大幅旋转扰动下。
cs.CV / 57 / 2607.23132

DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video

DispatchRAG:基于交通事故视频的现实协议进行紧急调度决策
Adhipradhana, Muhammad Sulthan, Javanmardi, Ehsan, Bao, Naren, Tsukada, Manabu
Abstract
Assessing the severity of a traffic accident scenario is important to decide which emergency service to dispatch. Missing an ambulance dispatch on a pedestrian accident is a fatal issue that can lead to death. Recently, Vision-Language Models (VLMs) have been a promising tool for accident reasoning, yet many VLMs are not grounded in real-life accident response protocols, making them not usable in accident severity assessment off-the-shelf. We introduced DispatchRAG, an accident assessor and dispatcher framework grounded in real-life Japanese traffic-accident response protocols, designed to enhance VLMs to generate an appropriate emergency response during an emergency scenario. Utilizing a RAG-based retrieval mechanism to retrieve the most relevant accident protocol and an LLM-powered reasoner to suggest the most proper response. To support evaluation, we introduce Accident Dispatch Dataset, a comprehensive dataset of accident assessment and emergency response according to Japanese accident response protocols adapted from the MM-AU dataset. We validate our framework on the Accident Dispatch Dataset, showing strong performance across various accident scenarios compared to the baseline VLM, pointing toward integration in autonomous vehicles that can automatically report both their own and nearby accidents.
Chinese Translation
评估交通事故场景的严重性对于决定调度哪种紧急服务至关重要。在行人事故中错过救护车的调度是一个致命问题,可能导致死亡。近年来,视觉-语言模型(Vision-Language Models, VLMs)成为事故推理的有前景工具,但许多VLM并未基于现实生活中的事故响应协议,导致它们无法直接用于事故严重性评估。我们提出了DispatchRAG,一个基于现实日本交通事故响应协议的事故评估和调度框架,旨在增强VLM在紧急场景中生成适当紧急响应的能力。该框架利用基于RAG的检索机制来获取最相关的事故协议,并使用大型语言模型(LLM)驱动的推理器来建议最合适的响应。为了支持评估,我们引入了事故调度数据集(Accident Dispatch Dataset),这是一个根据日本事故响应协议改编的全面事故评估和紧急响应数据集,源自MM-AU数据集。我们在事故调度数据集上验证了我们的框架,显示出在各种事故场景中相较于基线VLM的强大表现,指向在自动驾驶车辆中集成该框架的潜力,以便自动报告自身及附近的事故。
cs.CV / 58 / 2607.23181

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

面向视觉-语言导航的双脑最小充分表征
Wu, Yihao, Xu, Chenyi, Yan, Liqi, Cai, Chenhuan, Min, Geyong, Lin, Bin, Guan, Fangli, Zhang, Jianhui, Li, Pan
Abstract
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening generalization and increasing computational burden. We propose BrainNav, a navigation framework grounded in the Principle of Minimal Sufficiency. BrainNav consists of three components: a Logical Anchor Model that implements instruction-aware selective perception to suppress environmental noise, a Minimalist Constraint Alignment module that serves as a compact cross-modal bottleneck, efficiently synchronizing discrete linguistic intent with continuous latent dynamics while filtering out redundant information, and a Compression World Model that predicts action-conditioned states within a condensed, low-rank latent space. These modules align semantic intent with spatial perception, enhancing the agent's robustness and efficiency in complex tasks. Experiments show that BrainNav improves over prior SOTA by 2.0 % / 1.0 in SR/SPL on R2R-CE val-unseen and 0.94 % / 0.78 on RxR-CE val-unseen. These results indicate that minimally sufficient world representations provide an effective foundation for robust VLN.
Chinese Translation
在连续环境中的视觉与语言导航(VLN-CE)要求代理在自我中心的观察中将语言与环境结合,并在未见场景中进行规划。尽管最近的多模态大型模型和基于世界模型的方法改善了导航,但它们往往保留了过多与任务无关的细节,削弱了泛化能力并增加了计算负担。我们提出了BrainNav,一个基于最小充分性原则的导航框架。BrainNav由三个组件组成:一个逻辑锚模型(Logical Anchor Model),它实现了基于指令的选择性感知,以抑制环境噪声;一个极简约约束对齐模块(Minimalist Constraint Alignment),作为一个紧凑的跨模态瓶颈,能够高效地将离散的语言意图与连续的潜在动态同步,同时过滤冗余信息;以及一个压缩世界模型(Compression World Model),它在一个浓缩的低秩潜在空间中预测基于动作的状态。这些模块将语义意图与空间感知对齐,增强了代理在复杂任务中的鲁棒性和效率。实验表明,BrainNav在R2R-CE val-unseen上相较于之前的最先进技术(SOTA)提高了2.0% / 1.0的成功率(SR)/成功路径长度(SPL),在RxR-CE val-unseen上提高了0.94% / 0.78。这些结果表明,最小充分的世界表征为鲁棒的视觉-语言导航提供了有效的基础。
cs.CV / 59 / 2607.23189

Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design

Fashion-3DLR:一种基于成对时尚元素的可控3D服装生成方法,用于智能设计
Yang, Shenghao, Zhang, Hongtao, Yi, Yuhan, Tang, Zhihao, Cui, Zihao, Wen, Lian, Yan, Han, Gao, Yuan, Zhao, Mingbo
Abstract
AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry. However, 3D garment generation remains in its nascent stage, where in the realm of fashion, the semantic information of diverse design elements exhibits intricate coupling relationships in 3D representations, posing substantial challenges for generating diverse 3D garments. In this work, to handle the above problem, We introduce Fashion-3DLR, a novel 3D garment generation framework that utilizes diverse design elements to create high-quality, versatile 3D garment assets. Specifically, to bridge the semantic gaps between different fashion elements, we propose a Garment Feature Fusion Diffusion Transformer (GFF-DiT) module to integrate 2D fashion design elements, e.g., sketch and texture, into latent space. Within the latent space, we then employ a rectified flow transformer to generate geometry latents, which can be decoded into various 3D garment representations, including 3D Gaussians and meshes. Furthermore, we integrate Fashion-3DLR into downstream tasks, achieving the 3D Gaussian Splatting (3DGS)-driven cloth physical simulation and mesh-based virtual try-on. Experimental results indicate that Fashion-3DLR surpass the previous state-of-the-art methods, which verify that the proposed work can generate well-structured, non-watertight garments capable of physical simulation and virtual try-on, underscoring its potential as a versatile 3D garment design tool.
Chinese Translation
人工智能生成内容(AIGC)取得了显著进展,2D生成模型已成为数字时尚行业的现成工具。然而,3D服装生成仍处于初级阶段,在时尚领域,各种设计元素的语义信息在3D表现中展现出复杂的耦合关系,这给生成多样化的3D服装带来了重大挑战。在本研究中,为了解决上述问题,我们提出了Fashion-3DLR,这是一种新颖的3D服装生成框架,利用多样的设计元素创建高质量、通用的3D服装资产。具体而言,为了弥合不同时尚元素之间的语义差距,我们提出了一种服装特征融合扩散变换器(Garment Feature Fusion Diffusion Transformer,GFF-DiT)模块,将2D时尚设计元素(如草图和纹理)整合到潜在空间中。在潜在空间中,我们采用了一个校正流变换器生成几何潜变量,这些潜变量可以解码为各种3D服装表现,包括3D高斯分布和网格。此外,我们将Fashion-3DLR集成到下游任务中,实现了基于3D高斯点云(3D Gaussian Splatting,3DGS)的布料物理仿真和基于网格的虚拟试穿。实验结果表明,Fashion-3DLR超越了之前的最先进方法,验证了所提出的工作能够生成结构良好、非密闭的服装,能够进行物理仿真和虚拟试穿,突显了其作为多功能3D服装设计工具的潜力。
cs.CV / 60 / 2607.23193

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

OmniScope:面向全模态大语言模型的模态解耦令牌压缩
Su, Jinsen, Luo, Yongdong, Ma, Yuexiao, Hu, Yibo, Jin, Meiguang, Zheng, Xiaowu
Abstract
Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different moments. This cross-modal salience mismatch makes unidirectional guidance prone to discarding answer-critical cues under aggressive compression. We propose OmniScope, a training-free token compression framework that uses the query as a shared semantic anchor while estimating relevance separately for audio and video. OmniScope allocates modality-specific token budgets, prunes visual tokens with an anchor-delta strategy that preserves both global context and temporal changes, and merges audio tokens within each second to reduce redundancy while maintaining temporal continuity. Across four audio-video benchmarks and two Qwen2.5-Omni model scales, OmniScope achieves the best average accuracy across all compression settings. At 25% overall token retention, it delivers up to 3.53x prefill speedup and more than 15% GPU memory reduction, with only a 0.35-point drop in average accuracy. These results suggest a simple design principle for OmniLLM inference: share the query across modalities, but not the salience estimates. The code is available at https://github.com/MAC-AutoML/OmniScope.
Chinese Translation
现有的全模态大语言模型的令牌压缩方法通常依赖于一种模态来决定在另一种模态中保留什么。我们展示了这一假设往往会失效:对于相同的查询,音频和视频的相关性往往在不同的时刻达到峰值。这种跨模态显著性不匹配使得单向指导在激进压缩下容易丢弃关键答案线索。我们提出了OmniScope,一种无训练的令牌压缩框架,它使用查询作为共享语义锚点,同时分别估计音频和视频的相关性。OmniScope分配特定于模态的令牌预算,采用锚点-增量策略修剪视觉令牌,以保留全局上下文和时间变化,并在每秒内合并音频令牌以减少冗余,同时保持时间连续性。在四个音频-视频基准测试和两个Qwen2.5-Omni模型规模上,OmniScope在所有压缩设置中实现了最佳的平均准确率。在25%的整体令牌保留率下,它实现了高达3.53倍的预填充加速和超过15%的GPU内存减少,平均准确率仅下降0.35点。这些结果表明了一个简单的OmniLLM推理设计原则:在模态间共享查询,但不共享显著性估计。代码可在https://github.com/MAC-AutoML/OmniScope获取。
cs.CV / 61 / 2607.23194

Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix

超长场景文本识别:双轴诊断与无训练修复
Raisi, Zobeir, Zelek, John
Abstract
Scene Text Recognition (STR) models are trained almost exclusively on word crops of at most 25 characters, yet real deployments (signage, product labels, dense captions) require reading much longer text. This paper diagnoses that failure and then closes it. The diagnosis separates out-of-length failure into two simultaneously extrapolating axes (the encoder's width axis and the decoder's time axis) and shows that encoder width, not decoder length, is the dominant failure mode. Representation-side fixes bring only partial relief: training-free rotary rescalings recover at most 2-4 points of character error rate (CER), and a weighted fine-tuning recipe recovers 6-8 points while improving standard-benchmark accuracy, yet word accuracy on the Long Text Benchmark (LTB) stays near zero, because the residual gap lies in the decoding mechanism rather than the representation. We then close that gap at inference time, on an unmodified word-level checkpoint: the long image is sliced into overlapping crops at the model's training width, each decoded independently and in-distribution, and the reads stitched by geometry-anchored edit-distance alignment. This procedure reaches 42.79-43.05% bucket-average word accuracy on LTB across two base checkpoints, matching the published state of the art (41.57%) and beating it by 11-12 points on the hardest bucket, at wall-clock parity with plain decoding; applied unchanged to the public PARSeq checkpoint it reaches 47.11%. Once chunking is applied fine-tuning no longer helps: the decoding-side fix alone matches purpose-built architectures. We release the diagnosis harness and implementation.
Chinese Translation
场景文本识别(STR)模型几乎完全在最多25个字符的单词裁剪上进行训练,而实际应用(如标识、产品标签、密集字幕)则需要读取更长的文本。本文对这一失败进行了诊断并提出了解决方案。诊断将超长失败分解为两个同时外推的轴(编码器的宽度轴和解码器的时间轴),并表明编码器宽度而非解码器长度是主要的失败模式。表示侧的修复仅能部分缓解:无训练的旋转重缩放最多恢复2-4个字符错误率(CER)点,而加权微调方案恢复6-8个点,同时提高标准基准的准确性,但在长文本基准(LTB)上的单词准确率仍接近零,因为剩余的差距在于解码机制而非表示。我们随后在推理时填补了这一差距,使用未修改的单词级检查点:将长图像切割为与模型训练宽度重叠的裁剪,每个裁剪独立解码并在分布内进行,读取通过几何锚定的编辑距离对齐进行拼接。该过程在两个基础检查点上达到了42.79-43.05%的LTB桶平均单词准确率,匹配已发布的最新成果(41.57%),并在最难的桶上超越了11-12个点,且与普通解码的时间开销相当;在公共PARSeq检查点上应用无变化时,准确率达到了47.11%。一旦应用了分块,微调便不再有帮助:仅解码侧的修复便可与专门构建的架构相匹配。我们发布了诊断工具和实现代码。
cs.CV / 62 / 2607.23209

Counterfactual Motion Reliability Learning for Robust UAV Tracking

反事实运动可靠性学习用于鲁棒性无人机跟踪
Chen, Yuehai
Abstract
Infrared unmanned aerial vehicle (UAV) tracking is challenging because the target is often small, low-contrast, and easily confused with thermal distractors or cluttered backgrounds. Recent Transformer-based trackers have achieved promising performance by learning strong appearance representations, but their responses can still be dominated by background structures when the target appearance is weak or ambiguous. A natural solution is to introduce temporal motion cues. However, in infrared UAV tracking, motion cues are not always reliable: camera jitter, dynamic backgrounds, sensor noise, and target disappearance may produce temporal variations that are stronger than the true target motion. Therefore, the key challenge is not simply how to use motion, but how to distinguish target-consistent motion from background-induced pseudo motion. To this end, we propose CMRTrack, a counterfactual motion reliability learning framework for robust infrared UAV tracking. CMRTrack first extracts temporal evidence from adjacent search regions using a lightweight motion evidence encoder. During training, a counterfactual target-erased history branch is introduced to construct hard motion references, encouraging the motion encoder to learn reliable target-consistent motion rather than arbitrary temporal changes. The learned motion evidence is then incorporated into a one-stream tracking framework through motion-guided token modulation and reliability-aware score fusion, enabling adaptive feature enhancement and response refinement. Extensive experiments on Anti-UAV410 demonstrate that CMRTrack consistently outperforms representative state-of-the-art trackers and significantly improves the OSTrack baseline, with ablation studies and qualitative analysis verifying the effectiveness of the proposed counterfactual motion reliability learning.
Chinese Translation
红外无人机(UAV)跟踪面临挑战,因为目标通常较小、对比度低,并且容易与热干扰物或杂乱背景混淆。最近基于Transformer的跟踪器通过学习强大的外观表示取得了良好的性能,但当目标外观较弱或模糊时,它们的响应仍可能受到背景结构的主导。一个自然的解决方案是引入时间运动线索。然而,在红外无人机跟踪中,运动线索并不总是可靠:相机抖动、动态背景、传感器噪声和目标消失可能产生的时间变化可能强于真实目标运动。因此,关键挑战不仅在于如何使用运动,而在于如何区分目标一致的运动与背景诱导的伪运动。为此,我们提出了CMRTrack,一种用于鲁棒性红外无人机跟踪的反事实运动可靠性学习框架。CMRTrack首先使用轻量级运动证据编码器从相邻搜索区域提取时间证据。在训练过程中,引入了一个反事实目标擦除历史分支,以构建困难的运动参考,鼓励运动编码器学习可靠的目标一致运动,而不是任意的时间变化。然后,通过运动引导的令牌调制和可靠性感知的得分融合,将学习到的运动证据纳入单流跟踪框架,从而实现自适应特征增强和响应优化。在Anti-UAV410上的大量实验表明,CMRTrack始终优于具有代表性的最先进跟踪器,并显著改善了OSTrack基线,消融研究和定性分析验证了所提出的反事实运动可靠性学习的有效性。
cs.CV / 63 / 2607.23224

BoneAgeTW2: Automated Skeletal Maturation Assessment via the Tanner-Whitehouse 2 Method, Deep Learning, and Clinical Report Generation with Distribution Curves

BoneAgeTW2:通过Tanner-Whitehouse 2方法、深度学习和分布曲线生成临床报告的自动化骨骼成熟评估
Pinto, Juan Manuel Castillo
Abstract
We present BoneAgeTW2, the first fully open-source system to automate the complete Tanner-Whitehouse 2 (TW2) clinical protocol for skeletal maturity assessment end-to-end. The system employs YOLOv8 for precise detection and localization of the 20 TW2 hand bones from radiographic images, and an EfficientNet-B3 backbone with 20 independent classification heads to assign maturation stages (A-I) to each bone simultaneously. From these predictions, the system automatically generates clinical PDF reports including interactive Gaussian distribution curves for all 20 bones, enabling direct comparison with population norms. The model is trained on the public RSNA Pediatric Bone Age Challenge dataset (12,611 hand radiographs) using a pseudo-labeling strategy to derive per-bone stage labels from global bone age annotations. The full codebase is publicly available at https://github.com/jmmana/BoneAgeTW2.
Chinese Translation
我们提出了BoneAgeTW2,这是第一个完全开源的系统,能够端到端自动化完成Tanner-Whitehouse 2 (TW2) 骨骼成熟评估的临床协议。该系统采用YOLOv8精确检测和定位来自放射影像的20块TW2手骨,并使用EfficientNet-B3作为主干网络,配备20个独立的分类头,能够同时为每块骨骼分配成熟阶段(A-I)。基于这些预测,该系统自动生成临床PDF报告,包括所有20块骨骼的交互式高斯分布曲线,便于与人群标准进行直接比较。该模型在公共的RSNA儿童骨龄挑战数据集(12,611张手部放射影像)上进行训练,采用伪标签策略从全局骨龄注释中推导每块骨骼的阶段标签。完整的代码库可在https://github.com/jmmana/BoneAgeTW2获取。
cs.CV / 64 / 2607.23235

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

基于重建的超越参考标题的标题评估框架
Tang, Zhijiang, Qi, Jiaxin, Tang, Kaihua, Zheng, Yuhua, Huang, Jianqiang
Abstract
Image captioning is a primary task in vision--language research, yet assessing how faithfully a caption preserves image semantics without relying on reference captions remains unsettled. Prevailing evaluations rely on human-annotated references, whose content reflects annotator intent and captioning proficiency. In this paper, we study a reconstruction-based principle for caption evaluation: a caption is as good as its capacity to enable reconstruction of the original image. However, because captioning inherently compresses visual information, it is impossible to recover all details, and pixel-wise comparison between reconstructed and source images is neither feasible nor meaningful. Through our in-depth analysis of the nature of captions, whose fundamental purpose is to transmit the semantic content of an image, we propose a revised principle: a caption is as good as its capacity to enable a reconstruction that is semantically equivalent to the original. To assess semantic equivalence, we test whether the reconstruction matches the original image across a suite of downstream vision--language tasks, yielding a reference-free, task-conditioned caption score. We characterize component-dependent limitations and introduce the lower-cost Captioning Turing Test Dataset (CTTD) surrogate.
Chinese Translation
图像标题生成是视觉-语言研究中的一项主要任务,但如何评估标题在多大程度上忠实地保留图像语义而不依赖于参考标题仍未得到解决。现有评估依赖于人工标注的参考,其内容反映了标注者的意图和标题生成的熟练程度。本文研究了一种基于重建的标题评估原则:一个标题的好坏取决于其是否能够有效地重建原始图像。然而,由于标题生成本质上压缩了视觉信息,因此不可能恢复所有细节,重建图像与源图像之间的逐像素比较既不可行也没有意义。通过对标题本质的深入分析,我们提出了一种修订后的原则:一个标题的好坏取决于其是否能够实现与原始图像语义等价的重建。为了评估语义等价性,我们测试重建是否在一系列下游视觉-语言任务中与原始图像匹配,从而产生一个无参考、任务条件的标题评分。我们描述了组件依赖的局限性,并引入了低成本的标题图灵测试数据集(CTTD)替代品。
cs.CV / 65 / 2607.23238

SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models

SARATR-X-v2:面向尺度的SAR基础模型结构预训练
Li, Weijie, Song, Yafei, Liu, Yongxiang, Peng, Bowen, Zhou, Jie, Xia, Jingyuan, Yang, Wei, Liu, Tianpeng, Liu, Zhen, Liu, Li
Abstract
Masked image modeling has become a dominant paradigm for SAR pre-training, yet the design of the reconstruction target remains fundamentally unsettled. This article argues that a SAR pre-training target should satisfy two conditions to produce transferable representations: (i) physics-grounded stability, i.e., approximate invariance of the target operator to multiplicative speckle inherent in coherent imaging; and (ii) semantic scale compatibility, i.e., coverage of the heterogeneous spatial scales that downstream tasks demand. These two conditions are individually achievable but jointly difficult: physics-grounded stability favors fixed operators, while semantic scale compatibility favors data-driven composition. To this end, SARATR-X-v2 reconciles both within a single design. The target is constructed through fixed structural extractors spanning six receptive fields, from blind-spot local aggregation to directional log-ratio region contrast, and fused via learnable weights into one unified supervision signal for masked reconstruction. On twelve SAR benchmarks across classification, detection, and segmentation, SARATR-X-v2 achieves state-of-the-art transfer performance. Under synthetic speckle variation, the proposed target reduces perturbation drift in the learned representation by nearly two orders of magnitude relative to pixel-space supervision. Taken together, these results establish physics-grounded stability and semantic scale compatibility as a principled framework for pre-training target design under coherent imaging, and suggest that effective SAR pre-training is not about reconstructing more signal, but about reconstructing the right structural target.
Chinese Translation
掩蔽图像建模已成为SAR预训练的主导范式,但重建目标的设计仍然存在根本性的不确定性。本文认为,SAR预训练目标应满足两个条件,以产生可迁移的表示:(i)基于物理的稳定性,即目标算子对相干成像中固有的乘法斑点的近似不变性;(ii)语义尺度兼容性,即覆盖下游任务所需的异构空间尺度。这两个条件可以单独实现,但共同实现则较为困难:基于物理的稳定性倾向于固定算子,而语义尺度兼容性则倾向于数据驱动的组合。为此,SARATR-X-v2在单一设计中调和了这两者。该目标通过固定的结构提取器构建,涵盖六个感受野,从盲点局部聚合到方向性对数比区域对比,并通过可学习的权重融合为一个统一的监督信号,用于掩蔽重建。在分类、检测和分割的十二个SAR基准测试中,SARATR-X-v2实现了最先进的迁移性能。在合成斑点变化下,所提出的目标将学习表示中的扰动漂移减少了近两个数量级,相较于像素空间监督。综合来看,这些结果确立了基于物理的稳定性和语义尺度兼容性作为相干成像下预训练目标设计的原则框架,并表明有效的SAR预训练并非是重建更多信号,而是重建正确的结构目标。
cs.CV / 66 / 2607.23265

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

WaveZip:基于小波驱动的时空解耦用于视频令牌压缩
Zeng, Yuhui, Chen, Wang, Huang, Jinfa, Xie, Tianyu, Luo, Yongdong, Ji, Jiayi, Zheng, Xiawu, Luo, jiebo
Abstract
Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. While recent efficient methods attempt to compress tokens via hard pruning or uniform merging, they operate strictly in the spatial feature domain, where robust structural context and discriminative semantic details are inherently entangled. In this work, we propose WaveZip, a joint signal-frequency-domain framework for efficient video inference. Driven by the insight that temporal redundancy resides in low-pass approximation scales while spatial saliency strongly correlates with high-frequency components, WaveZip leverages Discrete Wavelet Transforms (DWT) to disentangle these signals. Temporally, it employs 1D DWT to analyze query-frame relevance, and the resulting high-frequency coefficients are further gated by inter-frame differences, with both signals jointly driving the dynamic allocation of a precise frame-level token budget. Spatially, a 2D DWT decomposes features into low-frequency approximations and high-frequency detail components, where the high-frequency coefficients are modulated within query-salient regions to regulate spatial reconstruction. Importantly, WaveZip requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency. Extensive experiments on long video understanding benchmarks demonstrate that WaveZip retains 99.6% of the full performance under an extreme 10x compression ratio, consistently outperforming state-of-the-art methods.
Chinese Translation
现有的大型视觉语言模型(LVLMs)在长视频理解方面面临挑战,主要是由于视觉令牌的二次计算成本。尽管最近一些高效方法试图通过硬剪枝或均匀合并来压缩令牌,但它们严格在空间特征域内操作,而在该域中,稳健的结构上下文和区分性的语义细节本质上是交织在一起的。在本研究中,我们提出了WaveZip,这是一种用于高效视频推理的联合信号频域框架。WaveZip的核心思想是,时间冗余存在于低通近似尺度中,而空间显著性与高频成分强相关。WaveZip利用离散小波变换(DWT)来解耦这些信号。在时间上,它采用一维DWT来分析查询帧的相关性,得到的高频系数进一步通过帧间差异进行门控,这两种信号共同驱动精确的帧级令牌预算的动态分配。在空间上,二维DWT将特征分解为低频近似和高频细节成分,其中高频系数在查询显著区域内进行调制,以调节空间重建。重要的是,WaveZip不需要特定任务的训练,并且可以无缝集成到现成的LVLM中,以提高推理效率。在长视频理解基准上的大量实验表明,WaveZip在极端的10倍压缩比下保留了99.6%的完整性能,始终优于最先进的方法。
cs.CV / 67 / 2607.23271

What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

CLIP所知但无法表达的内容:从冻结的中间特征中恢复否定信息
Lu, Chen-Yi, Chen, Yueh-Shao, Chaterji, Somali
Abstract
Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to negation. We attribute this failure to a phenomenon we call Representational Collapse: by tracking compositional divergence and visual alignment across the CLIP text encoder, we show that middle layers build compositional syntax, but the final layers collapse this structure as visual alignment rises, producing a syntax-blind final representation. To recover the lost negation signal without altering pretrained weights, we propose PeakPatch, a lightweight post-hoc correction system that intercepts the encoder at its compositional peak. An Embedding Correction Network (ECN) uses cross-attention to extract a negation-specific signal from the peak layer, anchored to a stable baseline, and predicts a deviation vector that re-injects the lost syntax into the final-layer embedding space. A complementary Score Correction Network (SCN) predicts bounded scalar score offsets for discriminative tasks. Both modules are trained jointly end-to-end while all CLIP parameters remain frozen, adding only 5.2M parameters (3.5% of the backbone) and preserving the standard cosine similarity interface. On NegBench, PeakPatch achieves 74.3% on COCO MCQ (+35.1 over CLIP, +17.8 over the best encoder fine-tuning method) and 65.5% on VOC MCQ, while outperforming all fine-tuning baselines on fully out-of-distribution negation retrieval despite training only 3.5% of the parameters. The corrected embeddings also transfer to text-to-image generation (+18.4 negation score) and generalize across ViT-B/32, ViT-L/14, and SigLIP backbones. Project URL: https://stevencylu.github.io/PeakPatch/.
Chinese Translation
对比视觉-语言模型如CLIP将语义相反的短语(例如,“一只狗”与“不是一只狗”)映射到几乎相同的嵌入中,从而使其对否定不敏感。我们将这一失败归因于我们称之为表示崩溃(Representational Collapse)的现象:通过跟踪CLIP文本编码器中的组合发散和视觉对齐,我们表明中间层构建了组合语法,但最终层在视觉对齐增强时崩溃了这一结构,产生了对语法失去敏感性的最终表示。为了在不改变预训练权重的情况下恢复丢失的否定信号,我们提出了PeakPatch,这是一种轻量级的后期修正系统,能够在编码器的组合峰值处拦截信号。嵌入修正网络(Embedding Correction Network, ECN)利用交叉注意力从峰层提取特定于否定的信号,并以稳定基线为锚点,预测一个偏差向量,将丢失的语法重新注入到最终层的嵌入空间。一个补充的评分修正网络(Score Correction Network, SCN)为判别任务预测有界的标量评分偏移。两个模块在所有CLIP参数保持冻结的情况下进行联合端到端训练,仅增加5.2M参数(占主干的3.5%),并保持标准的余弦相似性接口。在NegBench上,PeakPatch在COCO MCQ上达到了74.3%(比CLIP提高了35.1,比最佳编码器微调方法提高了17.8),在VOC MCQ上达到了65.5%,同时在完全超出分布的否定检索中超越了所有微调基线,尽管只训练了3.5%的参数。修正后的嵌入在文本到图像生成中也表现良好(否定评分提高了18.4),并在ViT-B/32、ViT-L/14和SigLIP主干上具有良好的泛化能力。项目网址:https://stevencylu.github.io/PeakPatch/
cs.CV / 68 / 2607.23335

The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

门始终关闭:向冻结的视觉-语言模型注入辅助信号
Farazi, Moshiur, Ramasinghe, Sameera, Ciftler, Bekir Sait, Turza, Mahbub Ahmed, Rahman, Shafin
Abstract
Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that the optimiser almost always decides on zero: across five injection designs, every gated pathway becomes behaviourally closed, with accuracy invariant to ablating the pathway at inference even when the gate parameter would nominally pass 30-45% of the signal. We attribute this suppression phenomenon to two regimes, a dead-gradient regime formalised through the caption-invariance of image-derived signals, and a negative-utility regime in which the auxiliary signal actively hurts the loss. Rather than fight suppression, we exploit it: we regularise LoRA fine-tuning with geometric auxiliary losses from hyperbolic visual relational graphs (IoA-driven entailment cones and angular repulsion on the Lorentz manifold), coupled only through the forward pass at training time and dropped at inference. Disaggregating GQA by question type exposes a clean dissociation. Three configurations without geometric losses at inference lose 2.85-3.39pp on relational questions while gaining ~1pp on attribute questions; a fourth that trains with the losses but infers through a soft prompt loses 5.14pp on rel for only +0.23pp on attr, so training-time regularisation alone does not protect relational accuracy without a geometric inference pathway. Configurations that keep the geometric pathway at inference preserve vanilla-level relational accuracy and match the attribute gain. Out of distribution on VSR, the RMS-prefix recipe preserves the spatial signal; stripping the geometric losses (G2) collapses VSR by 4.6pp, isolating them as the OOD source. A secondary result: embedding-norm alignment is necessary for generation-safe prefix injection, and learnable gates should be replaced with fixed, non-optional injection at matched scales.
Chinese Translation
视觉-语言模型(VLMs)中的辅助信号通道通常配备可学习的门,以便优化器可以决定接受多少信号。我们发现优化器几乎总是选择零:在五种注入设计中,每个带门通道的行为都变得封闭,准确率在推理时对去除该通道不变,即使门参数名义上会传递30-45%的信号。我们将这种抑制现象归因于两个机制,一是通过图像衍生信号的标题不变性形式化的死梯度机制,二是辅助信号在损失中积极造成伤害的负效用机制。我们并不试图对抗抑制,而是利用它:我们通过来自双曲视觉关系图的几何辅助损失(基于信息的蕴含锥和洛伦兹流形上的角度排斥)来正则化LoRA微调,这些损失仅在训练时通过前向传播连接,并在推理时被丢弃。通过问题类型对GQA进行解构,揭示了一个清晰的解离。在推理时,没有几何损失的三种配置在关系问题上损失了2.85-3.39个百分点,而在属性问题上获得了约1个百分点;第四种配置在训练时使用损失但通过软提示进行推理,在关系问题上损失了5.14个百分点,仅在属性问题上获得了+0.23个百分点,因此仅靠训练时的正则化无法在没有几何推理通道的情况下保护关系准确性。保持几何通道在推理时的配置保持了原始水平的关系准确性,并匹配了属性增益。在VSR的分布外,RMS-prefix方法保持了空间信号;去除几何损失(G2)使VSR下降了4.6个百分点,将其孤立为OOD来源。一个次要结果是:嵌入范数对齐对于生成安全的前缀注入是必要的,且可学习的门应被固定的、不可选的匹配尺度注入所替代。
cs.CV / 69 / 2607.23368

Explaining BiomedCLIP with Weighted Banzhaf Interactions Supported by Tree-Gram Parsing

用加权Banzhaf交互和树语法解析解释BiomedCLIP
Rymarski, Jakub, Rempała, Adam, Sobieski, Bartłomiej, Biecek, Przemysław
Abstract
Vision-Language Models (VLMs) are demonstrating significant capabilities in medical tasks like radiology analysis, yet providing faithful and interpretable explanations remains a key consideration for their responsible deployment in clinical settings. However, existing explanation methods, such as the widely used FIxLIP framework, often struggle with the fine-grained nature of modern tokenizers. The tokenization problem fragments clinical concepts---splitting terms like "saddle embolus" into scattered, meaningless subwords---which leads to noisy, semantically incoherent cross-modal attributions. Such fragmentation also results in a combinatorial explosion of interaction possibilities, obscuring the model's true reasoning. To address this, we introduce ParseFIxLIP, an extension that incorporates the Tree-Gram Parsing into the Banzhaf interaction game used by FIxLIP. This semantically informed strategy utilizes dependency parsing trees to define explanation players by grouping related text tokens into semantically coherent units. Our smart_depth grouping strategy, merging tokens according to spaCy token dependency tree, successfully mitigates concept fragmentation, yielding substantially more interpretable cross-modal interactions by unifying complex medical concepts. Quantitatively, while baselines struggled with the high dimensionality of long captions, our parsing approach maintained statistical robustness and semantic parsimony. Qualitative analysis on BiomedCLIP, validated on medical imagery (ROCOv2) and general examples, confirms that the approach accurately captures the synergistic influence of grouped words on model predictions. In conclusion, our work offers intuitive and clinically relevant insights into VLM decision-making, fulfilling the critical need for coherent explanations in the medical domain.
Chinese Translation
视觉-语言模型(VLMs)在放射学分析等医疗任务中展现出显著的能力,但提供可信且可解释的解释仍然是其在临床环境中负责任部署的关键考虑。然而,现有的解释方法,如广泛使用的FIxLIP框架,往往难以应对现代分词器的细粒度特性。分词问题将临床概念碎片化——将“鞍状栓子”等术语拆分为零散且无意义的子词——这导致了嘈杂且语义不连贯的跨模态归因。这种碎片化还导致交互可能性的组合爆炸,掩盖了模型的真实推理。为了解决这一问题,我们引入了ParseFIxLIP,这是一个扩展,结合了树语法解析到FIxLIP使用的Banzhaf交互游戏中。这种语义信息驱动的策略利用依赖解析树通过将相关文本标记分组为语义连贯的单元来定义解释参与者。我们的smart_depth分组策略根据spaCy的标记依赖树合并标记,成功减轻了概念碎片化,统一复杂的医疗概念,从而产生更具可解释性的跨模态交互。在定量方面,尽管基线方法在长标题的高维性上遇到困难,但我们的解析方法保持了统计稳健性和语义简约性。对BiomedCLIP的定性分析,在医疗影像(ROCOv2)和一般示例上验证,确认该方法准确捕捉了分组词对模型预测的协同影响。总之,我们的工作为VLM决策提供了直观且与临床相关的见解,满足了医疗领域对连贯解释的迫切需求。
cs.CV / 70 / 2607.23373

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

UltraViT:针对大型视觉语言模型的延迟优化设备端视觉编码器
Metaxas, Ioannis Maniadis, Bulat, Adrian, Baldrati, Alberto, Zaganidis, Anestis, Ouali, Yassine, Kim, Hyeonuk, Tzimiropoulos, Georgios
Abstract
Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or smaller language models, the vision encoder is largely overlooked, typically deployed as a monolithic, computationally heavy feature extractor. Moreover, there is no previous effort that designs a vision encoder for LVLMs directly optimized for on-device latency. In this paper, we present UltraViT, a vision encoder for LVLMs, explicitly designed and optimized for on-device performance. Specifically, by taking into account real on-device latencies, we systematically design a pyramidal architecture that strategically integrates and adapts heterogeneous spatial mixers at the macro-block level. Furthermore, to pre-train UltraViT, we propose a novel two-stage generative pre-training strategy: cultivating rich spatial features via dense distillation, followed by direct generative supervision from a capacity-mixed frozen LLM. Compared to standard contrastive and SSL, we show that our pre-training is much more effective for achieving high-level semantic grounding for UltraViT needed for the subsequent generative multimodal alignment of LVLM training. Extensive experiments demonstrate that our on-device latency-informed design combined with our tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.
Chinese Translation
大型视觉语言模型(LVLMs)由于其庞大的计算需求,仍然面临瓶颈,限制了它们在资源受限的边缘设备上的部署。尽管对LVLMs的压缩工作主要集中在视觉令牌的减少或更小的语言模型上,但视觉编码器往往被忽视,通常作为一个单一的、计算量大的特征提取器进行部署。此外,之前没有任何工作直接为LVLMs设计针对设备端延迟优化的视觉编码器。在本文中,我们提出了UltraViT,一种专为LVLMs设计并优化的视觉编码器,明确考虑了设备端性能。具体而言,通过考虑实际的设备端延迟,我们系统地设计了一种金字塔架构,在宏块级别上战略性地集成和适应异构空间混合器。此外,为了预训练UltraViT,我们提出了一种新颖的两阶段生成预训练策略:通过密集蒸馏培养丰富的空间特征,随后通过容量混合的冻结语言模型(LLM)进行直接生成监督。与标准对比学习和自监督学习(SSL)相比,我们展示了我们的预训练在实现UltraViT所需的高层次语义基础方面更为有效,这对于后续的LVLM训练中的生成多模态对齐至关重要。大量实验表明,我们基于设备端延迟的信息设计结合我们量身定制的训练策略,为高效的LVLM编码建立了新的最先进水平,在设备端以近1.7倍的速度显著超越现有的以编码器为中心的基线。
cs.CV / 71 / 2607.23445

Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models

Omni-Prune:面向查询的统一令牌剪枝以提高全模态大型语言模型的效率
Zhong, Yiming, Nie, Chang, Shan, Caifeng
Abstract
Omnimodal large language models (OmniLLMs) are rapidly extending multimodal reasoning to cover synchronized audio and video. However, the resulting audio-video token sequences are long, leading to high prefill latency and GPU memory usage at inference time. Existing token pruning methods, designed mainly for vision-only inputs, miss both the cross-modal links between audio and video and the user query that decides which content matters. To bridge this gap, we present Omni-Prune, a training-free, query-aware audio-visual token pruning framework that jointly removes redundancy from both modalities while keeping task-relevant cross-modal evidence. Specifically, Omni-Prune first splits the token sequence into adaptive time windows placed at audio saliency peaks, then scores audio and video tokens on a single scale that combines encoder attention with text-query relevance, and pairs related audio-video tokens so that they are kept together. Within each window, a final K-medoids step then selects a few representative tokens, adding diverse cues that score-based selection alone would miss. Extensive experiments demonstrate that Omni-Prune outperforms established baseline methods, delivering up to 3.25x prefill speedup and 1.3x memory reduction while retaining over 99% of full-model performance.
Chinese Translation
全模态大型语言模型(OmniLLMs)正在迅速扩展多模态推理,以覆盖同步音频和视频。然而,生成的音频-视频令牌序列较长,导致推理时的预填充延迟和GPU内存使用量高。现有的令牌剪枝方法主要针对仅视觉输入,未能考虑音频和视频之间的跨模态联系以及决定内容重要性的用户查询。为了解决这一问题,我们提出了Omni-Prune,这是一种无训练、面向查询的音频-视觉令牌剪枝框架,能够同时去除两个模态中的冗余,同时保留与任务相关的跨模态证据。具体而言,Omni-Prune首先将令牌序列分割成放置在音频显著性峰值处的自适应时间窗口,然后在一个结合编码器注意力和文本查询相关性的单一尺度上对音频和视频令牌进行评分,并将相关的音频-视频令牌配对,以便它们一起保留。在每个窗口内,最终的K-中位数步骤选择少量代表性令牌,增加基于评分的选择所遗漏的多样化线索。大量实验表明,Omni-Prune的性能优于既有基线方法,实现了高达3.25倍的预填充加速和1.3倍的内存减少,同时保留了超过99%的全模型性能。
cs.CV / 72 / 2607.23451

Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance

基于Prompt-S6和语义感知知识引导的多模态物体重识别
Zhou, Weixiang, Zuo, Jiabei, Wang, Yuhao, Wang, Cong, Lu, Huchuan, Su, Zhixun
Abstract
Multi-modal object Re-Identification (ReID) aims to retrieve specific objects by integrating complementary information from multiple modalities. However, existing multi-modal ReID methods do not effectively address background interference suppression or achieve tri-modal alignment, instead focusing on pairwise feature fusion. Moreover, many current aggregation approaches suffer from high computational complexity. To address these limitations, we propose PRISM, a novel multi-modal ReID framework built upon Prompt-S6 (PS6) and semantic-aware knowledge guidance. PS6 maintains the linear complexity and strong sequence modeling capability of Mamba while enabling efficient cross-modal interaction. Leveraging these advantages, we design two key components: Semantic-Driven Token Pruning (SDTP) and Progressive Fusion Network (PFN). Parsing semantic priors from the segmentation foundation models, the SDTP then leverages these priors and applies dynamic token pruning to suppress background noise and refine feature representations. The PFN progressively aggregates multi-modal features to achieve tri-modal alignment and fully exploit modality complementarity. With the proposed modules, PRISM generates more robust multi-modal representations under complex scenarios. Extensive experiments on four multi-modal object ReID benchmarks demonstrate the effectiveness and efficiency of our approach. The source code is available at https://github.com/zw-absin/PRISM.
Chinese Translation
多模态物体重识别(ReID)旨在通过整合来自多个模态的互补信息来检索特定物体。然而,现有的多模态ReID方法未能有效解决背景干扰抑制或实现三模态对齐,而是专注于成对特征融合。此外,许多当前的聚合方法面临着高计算复杂度的问题。为了解决这些局限性,我们提出了PRISM,一个基于Prompt-S6(PS6)和语义感知知识引导的新型多模态ReID框架。PS6保持了Mamba的线性复杂度和强大的序列建模能力,同时实现了高效的跨模态交互。利用这些优势,我们设计了两个关键组件:语义驱动的令牌修剪(SDTP)和渐进融合网络(PFN)。SDTP从分割基础模型中解析语义先验,然后利用这些先验并应用动态令牌修剪来抑制背景噪声并优化特征表示。PFN逐步聚合多模态特征,以实现三模态对齐并充分利用模态互补性。通过所提出的模块,PRISM在复杂场景下生成更强健的多模态表示。在四个多模态物体ReID基准上的大量实验表明了我们方法的有效性和效率。源代码可在 https://github.com/zw-absin/PRISM 获取。
cs.CV / 73 / 2607.23468

Robust 6-DoF Object Pose Tracking with Built-In Recovery under Occlusions and Rapid Object Motions

具有内置恢复机制的鲁棒6自由度物体姿态跟踪:应对遮挡和快速物体运动
Opra, Balázs, Ghafari, Léo, Stewart, Thomas, Stachniss, Cyrill
Abstract
Real-time 6-DoF object pose tracking is essential for many robotics applications, and several approaches exist. Yet even today's approaches remain unreliable under temporary full occlusions and rapid object motions. Once tracking is lost, most methods struggle to detect the failure and recover automatically, often requiring manual re-initialization. In this paper, we address the problem of robust model-based 6-DoF tracking of unseen objects from RGB-D data, especially in scenarios with occlusion and fast motion. We propose a novel method that combines efficient learning-based keypoint matching with optimization-based alignment and introduces a novel failure detection and recovery module. Our system monitors pose reliability, detects tracking divergence or occlusions, and performs a global re-detection and pose estimation step that robustly verifies recovery candidates before resuming tracking. Our evaluation on standard tracking benchmarks and on a new dataset of occluded and fast-moving scenes shows that our method matches state-of-the-art accuracy on easy tracking sequences, maintains high tracking speed at 57.6 frames per second, and provides the most robust tracking performance under challenging conditions. Thus, we believe that our approach is a relevant step forward in robust 6-DoF object tracking from RGB-D data.
Chinese Translation
实时6自由度物体姿态跟踪对于许多机器人应用至关重要,现有多种方法可供选择。然而,即使是今天的技术在面对暂时的完全遮挡和快速物体运动时仍然不够可靠。一旦跟踪丢失,大多数方法难以检测失败并自动恢复,通常需要手动重新初始化。本文针对从RGB-D数据中对未见物体进行鲁棒的基于模型的6自由度跟踪的问题,特别是在遮挡和快速运动的场景中。我们提出了一种新方法,将高效的基于学习的关键点匹配与基于优化的对齐相结合,并引入了一种新颖的失败检测与恢复模块。我们的系统监控姿态的可靠性,检测跟踪偏差或遮挡,并执行全局重新检测和姿态估计步骤,稳健地验证恢复候选者,然后再继续跟踪。我们在标准跟踪基准和一个新的遮挡与快速移动场景数据集上的评估表明,我们的方法在简单跟踪序列中达到了最先进的准确性,保持了每秒57.6帧的高跟踪速度,并在具有挑战性的条件下提供了最鲁棒的跟踪性能。因此,我们相信我们的方法是从RGB-D数据中进行鲁棒6自由度物体跟踪的重要进展。
cs.CV / 74 / 2607.23472

VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

VIPER:上下文视觉物理推理用于物理上合理的视频生成
Chen, Tianxiao, Chen, Hanmo, Chen, Huajin, Li, Bo, Ye, Qi, Jiang, Peng-Tao
Abstract
Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficult with standard text and image conditions. The core challenge is a conditioning bottleneck: material response, contact interaction, deformation, and motion trajectory are continuous and relational physical cues that are hard to specify exhaustively in language but can be demonstrated naturally by video. We propose VIPER, a Visual In-Context Physics Reasoning framework for reference-guided image-to-video generation. Given a target image, a brief target prompt, and a reference video, VIPER treats the reference as a dense visual demonstration of the desired physical process rather than an appearance template. It uses a Multimodal Large Language Model (MLLM) to extract reference-derived physical cues and guide a pretrained image-to-video generator through a hierarchical training strategy, enabling physical behavior transfer while preserving the visual prior of the base generator. To support this setting, we construct VIPER-19K, a curated dataset with material, trajectory, and physical-impact annotations, together with filtered reference-target pairs. Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality. Qualitative results further demonstrate that VIPER can transfer reference-derived physical behavior to new target scenes without requiring carefully engineered prompts.
Chinese Translation
现代视频生成模型能够合成视觉上引人注目且时间上连贯的片段,但在标准文本和图像条件下控制其物理行为仍然困难。核心挑战在于条件瓶颈:材料响应、接触交互、变形和运动轨迹是连续的、关系性的物理线索,难以用语言详尽地指定,但可以通过视频自然展示。我们提出了VIPER,一个用于参考引导的图像到视频生成的上下文视觉物理推理框架。给定一个目标图像、一个简短的目标提示和一个参考视频,VIPER将参考视为所需物理过程的密集视觉演示,而不是外观模板。它使用多模态大型语言模型(Multimodal Large Language Model, MLLM)提取参考派生的物理线索,并通过分层训练策略引导预训练的图像到视频生成器,从而实现物理行为转移,同时保留基础生成器的视觉先验。为了支持这一设置,我们构建了VIPER-19K,一个经过精心策划的数据集,包含材料、轨迹和物理影响的注释,以及过滤后的参考-目标对。在未见验证集上的实验表明,VIPER在参考视频的物理相似性和人类偏好上优于代表性的视频生成和视频作为提示的基线,同时保持竞争力的整体视频质量。定性结果进一步表明,VIPER能够将参考派生的物理行为转移到新的目标场景,而无需精心设计的提示。
cs.CV / 75 / 2607.23491

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

PlanCraft:为建筑师启发的渐进式三维住宅场景生成进行草图绘制、细化和家具布置
Zeng, Pengyu, Dai, Yuqin, Yin, Jun, Han, Ziyang, Hei, Ng Cheuk, Zhong, Jing, Shi, Chaoyang, Jin, ZhanXiang, Jiang, Maowei, Han, Yuxing, Lu, Shuai
Abstract
Two structural insights have been overlooked in automated residential floor plan generation. First, design is inherently progressive. Architects begin with rough strokes and refine them over time, whereas existing methods typically require their conditioning representation to be fully specified before generation, a fundamental mismatch with how design actually works. Second, the 2D floor plan is not an optional intermediate but an irreplaceable spatial contract. Once room boundaries, doors, and windows are fixed, furnishing reduces from open-ended spatial reasoning to bounded constraint satisfaction. Bypassing this contract, as existing 3D systems do by delegating layout to language models, yields overlapping rooms and implausible proportions; directly calling general-purpose language models likewise produces geometrically invalid layouts. Guided by these insights, we present PlanCraft. SketchPlan supplies the missing training signal by replaying the architect's drawing process on 80K real floor plans, producing partial sketches at every completeness level. PlanCraft-Diff progressively sharpens an incomplete sketch into a geometrically precise, vectorizable floor plan through a coarse-to-fine strategy. With the spatial contract established, PlanCraft-Agent then furnishes the scene within well-defined room boundaries. Experiments show that PlanCraft achieves a 61.1\% lower FID than the best existing 2D method and surpasses existing 3D systems by 15 points in expert-rated spatial rationality, with a sketch at only 25\% completion already outperforming all fully specified baselines.
Chinese Translation
在自动化住宅平面图生成中,有两个结构性见解被忽视。首先,设计本质上是渐进的。建筑师从粗略的草图开始,并随着时间的推移进行细化,而现有方法通常要求其条件表示在生成之前完全指定,这与设计的实际工作方式存在根本不匹配。其次,二维平面图不是可选的中间步骤,而是不可替代的空间契约。一旦房间边界、门和窗户被固定,家具布置就从开放式空间推理转变为有限约束满足。现有的三维系统通过将布局委托给语言模型来绕过这一契约,导致重叠的房间和不合理的比例;直接调用通用语言模型同样会产生几何上无效的布局。在这些见解的指导下,我们提出了PlanCraft。SketchPlan通过在80K真实平面图上重放建筑师的绘图过程,提供了缺失的训练信号,在每个完整性水平上生成部分草图。PlanCraft-Diff通过粗到细的策略,逐步将不完整的草图锐化为几何精确、可向量化的平面图。在建立空间契约后,PlanCraft-Agent在明确的房间边界内进行场景布置。实验表明,PlanCraft的FID比现有最佳二维方法低61.1%,在专家评估的空间合理性上超越现有三维系统15分,且仅在25%的完成度下的草图已经超越了所有完全指定的基线。
cs.CV / 76 / 2607.23492

To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion

抹去,还是不抹去:具有保留意识的自适应排名子空间扩展的稳健训练无关概念抹除
Saha, Shaswati, Anguluri, Rajasekhar, Gaur, Manas
Abstract
Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while preserving model utility on benign concepts. Current CETs face a trade-off between erasure robustness and utility: stronger edits erase the target more reliably but degrade utility on non-target concepts, and vice versa. This stems from how existing methods define what to erase and what to preserve. Many CETs rely on static concept banks specified manually, generated by LLMs, or selected by CLIP image-text similarity. Such banks do not model how prompts steer the model during denoising, leaving it vulnerable to triggers that reintroduce the target while suppressing nearby benign concepts. We present Preservation-aware Adaptive Ranked Subspace Expansion (PARSE), a training-free framework for robust concept erasure in latent diffusion models. Given a target, PARSE queries the diffusion model with classifier-free guidance to dynamically discover target-inducing erase concepts and nearby retain concepts in the model vocabulary. It then edits the cross-attention value space with a preservation-aware projection that removes target directions while leaving retain directions intact. For triggers beyond this vocabulary-indexed space, PARSE iteratively searches for re-emergence triggers by textual inversion and adaptively expands the erased subspace only when a new trigger direction does not conflict with retain semantics. We also introduce the Balanced Erasure Utility Score (BEUS), which combines robustness (ASR under multiple attacks) and utility preservation (FID) via bounded monotone transforms and harmonic mean aggregation. Experiments on NSFW, artistic style, and object erasure, with a large-scale robustness-utility analysis over many CET baselines, show that PARSE erases multiple concepts robustly without sacrificing post-edit utility.
Chinese Translation
概念抹除技术(CETs)对文本到图像的扩散模型进行编辑,以抹去不希望出现的目标,例如不适宜内容或受版权保护的风格,同时保留模型在良性概念上的效用。目前的CET面临抹除稳健性与效用之间的权衡:更强的编辑可以更可靠地抹去目标,但会降低非目标概念的效用,反之亦然。这源于现有方法如何定义要抹去的内容和要保留的内容。许多CET依赖于手动指定的静态概念库、由大型语言模型(LLMs)生成或通过CLIP图像-文本相似性选择的概念库。这些库未能建模提示在去噪过程中如何引导模型,使其容易受到重新引入目标的触发器的影响,同时抑制附近的良性概念。我们提出了具有保留意识的自适应排名子空间扩展(PARSE),这是一个针对潜在扩散模型的稳健概念抹除的无训练框架。给定一个目标,PARSE通过无分类器引导查询扩散模型,动态发现诱发目标的抹除概念和模型词汇中附近的保留概念。然后,它通过一种保留意识的投影编辑交叉注意力值空间,去除目标方向,同时保持保留方向不变。对于超出该词汇索引空间的触发器,PARSE通过文本反演迭代搜索重新出现的触发器,并仅在新的触发器方向与保留语义不冲突时自适应扩展抹除子空间。我们还引入了平衡抹除效用评分(BEUS),它通过有界单调变换和调和平均聚合结合了稳健性(在多次攻击下的ASR)和效用保留(FID)。在不适宜内容、艺术风格和物体抹除的实验中,以及对多个CET基线的大规模稳健性-效用分析,结果表明PARSE能够稳健地抹去多个概念,而不牺牲后编辑效用。
cs.CV / 77 / 2607.23493

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

基于标记-区域引导的跨注意力融合用于多模态情感解读
Farazi, Musa Tur, Reza, Nufayer Jahan
Abstract
Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computationally challenging due to the intricate interplay between visual cues and embedded, often stylized, text, particularly in low-resource languages like Bengali Language. This paper addresses the detection of political intent in Bengali memes by introducing Multimodal Cross-Attention Fusion framework. We first leverage a Vision-Language Model to extract high-fidelity OCR text from noisy meme images. Subsequently, we encode visual and textual features and synthesize them through a cross-modal multi-head attention mechanism that aligns semantic tokens with visual regions. We also investigate the integration of a domain-specific political lexicon as a knowledge prior. Experimental evaluation on the PoliMemeDecode1 dataset shows that our attention-based fusion significantly outperforms unimodal baselines and standard concatenation methods, achieving a state-of-the-art Macro-F1 of approximately 0.94. Interpretability analyzes further confirm that the model effectively learns to ground textual semantics in visual evidence.
Chinese Translation
社交网络上多模态内容的自动分析已成为理解公众情感和信息传播的重要任务。然而,由于视觉线索与嵌入的、通常是风格化的文本之间复杂的相互作用,特别是在像孟加拉语这样的低资源语言中,分类互联网迷因仍然具有计算挑战性。本文通过引入多模态跨注意力融合框架,解决了孟加拉迷因中政治意图的检测问题。我们首先利用视觉-语言模型从嘈杂的迷因图像中提取高保真度的OCR文本。随后,我们对视觉和文本特征进行编码,并通过跨模态多头注意力机制将语义标记与视觉区域对齐进行合成。我们还探讨了将特定领域的政治词汇作为知识先验的整合。在PoliMemeDecode1数据集上的实验评估表明,我们基于注意力的融合显著优于单模态基线和标准拼接方法,达到了约0.94的最先进的宏观F1值。可解释性分析进一步确认模型有效地学习了如何将文本语义与视觉证据相结合。
cs.CV / 78 / 2607.23504

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation

MemVLN:用于视觉与语言导航的情节记忆和程序记忆
Liu, Yuqi, Qian, Shengju, Qu, Tianyuan, Lin, Mingxian, Wang, Zixuan, Wang, Xin, Yu, Bei, Jia, Jiaya
Abstract
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to maintain long-horizon visual history for trajectory consistency while executing actions with low latency. Existing video-based VLN approaches typically struggle to satisfy both demands simultaneously. To address these challenges, we propose MemVLN, a novel VLN framework that achieves state-of-the-art performance with real-time inference efficiency (14 FPS). MemVLN utilizes a visual encoder to process continuous observations and a Large Language Model (LLM) to interpret instructions and generate actions. Central to our approach is an Episodic Memory management that applies pyramidal resolutions. This mechanism concentrates computation on immediate percepts while retaining compressed long-term history. Complementing to this design, we introduce Procedural Memory for fast action with a compact vocabulary of atomic mid-level actions to bypass auto-regressive decoding latency. Experiments on VLN-CE show that MemVLN-4B surpasses the baseline Qwen3-VL-4B architecture by 5.8\% SR in R2R and 9.7\% SR in RxR, while achieving a 7$\times$ speedup in inference latency.
Chinese Translation
在连续环境中的视觉与语言导航(VLN-CE)要求代理在执行低延迟动作的同时,保持长时间的视觉历史以确保轨迹一致性。现有的视频基础VLN方法通常难以同时满足这两项要求。为了解决这些挑战,我们提出了MemVLN,这是一种新颖的VLN框架,能够以实时推理效率(14 FPS)实现最先进的性能。MemVLN利用视觉编码器处理连续观察,并使用大型语言模型(LLM)来解释指令和生成动作。我们方法的核心是应用金字塔分辨率的情节记忆管理。该机制将计算集中在即时感知上,同时保留压缩的长期历史。为了补充这一设计,我们引入了程序记忆,通过紧凑的原子中级动作词汇实现快速动作,以绕过自回归解码延迟。在VLN-CE上的实验表明,MemVLN-4B在R2R中超越基线Qwen3-VL-4B架构5.8\%的成功率(SR),在RxR中超越9.7\\%的成功率,同时实现推理延迟的7倍加速。
cs.CV / 79 / 2607.23511

MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

MOJITO:统一端到端自主驾驶的模态联合学习
Cheng, Zhijing, Zhang, Xuancheng, Di, Donglin, Fan, Lei, Ma, Baorui, Li, Hao, Yang, Xun
Abstract
End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories conditioned on this context. We argue that this one-way perception-to-planning interface forces sensor inputs into a compact representation, losing the fine-grained details critical for planning. Moreover, by constraining the planner to this compressed context, it is difficult to leverage the rich representations offered by modern vision foundation models. To address these issues, we propose MOJITO, a unified sensor-to-action framework for end-to-end autonomous driving built on modal joint learning. MOJITO removes the cascaded interface and instead performs block-wise Modal Joint Attention that simultaneously updates action, image, and LiDAR features, allowing the planner to directly access multi-modal features during action generation. MOJITO achieves 88.9 PDMS on the NAVSIM v1 dataset and 88.4 EPDMS on the more challenging NAVSIM v2 dataset, setting a new state-of-the-art. Extensive experiments further demonstrate strong scalability, instruction following, and diverse trajectory generation. Code and models are available at https://github.com/mumucc01/MOJITO.
Chinese Translation
端到端自主驾驶系统通常遵循级联的两阶段流程,其中感知阶段将多模态传感器输入压缩为紧凑的上下文,而下游规划器则基于该上下文预测轨迹。我们认为,这种单向的感知到规划接口迫使传感器输入转化为紧凑的表示,导致规划所需的细粒度细节丢失。此外,限制规划器使用这种压缩上下文,使得难以利用现代视觉基础模型提供的丰富表示。为了解决这些问题,我们提出了MOJITO,一种基于模态联合学习的端到端自主驾驶统一传感器到动作框架。MOJITO去除了级联接口,而是执行块级模态联合注意力,同时更新动作、图像和激光雷达特征,使规划器在动作生成过程中能够直接访问多模态特征。MOJITO在NAVSIM v1数据集上达到了88.9的PDMS,在更具挑战性的NAVSIM v2数据集上达到了88.4的EPDMS,创造了新的最先进水平。大量实验进一步证明了其强大的可扩展性、指令跟随能力和多样化轨迹生成。代码和模型可在https://github.com/mumucc01/MOJITO获取。
cs.CV / 80 / 2607.23517

Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction

实时人本世界建模用于上半身人机交互
Ji, Chaonan, Qi, Jinwei, Zhang, Peng, Zhang, Bang
Abstract
We present a real-time human-centric world model for upper-body interactive generation, aiming to synthesize coherent local world dynamics centered on a person, where coordinated body, hand, and facial motions evolve jointly with controllable human-object discrete interaction. To this end, we adopt a continuous-discrete joint control scheme with two complementary components: a continuous human state and a discrete interaction state. For continuous human-state control, we introduce a unified implicit representation based on multi-scale motion encoding, in which motion latents from the upper body, hands, and face are fused into a shared latent space. This multi-scale design improves expressiveness across different spatial scales, captures fine-grained human dynamics more effectively, and enables direct control without explicit retargeting. For discrete object interaction-state control, we represent object contact using a small set of language-encoded discrete interaction states, where text serves as an explicit interaction-state command, such as \emph{no contact} or \emph{grasp}, rather than an open-ended generation prompt, and we further construct a dedicated rendering pipeline for human-object interaction data to supervise such discrete interaction states. By combining continuous implicit human-state control with discrete interaction-state control, our model enables precise modeling of how a person moves and interacts with the local environment, including controllable changes to nearby scene states. Finally, we distill the model for efficient streaming real-time inference, achieving 25 FPS on two H100 GPUs. Experiments demonstrate improved fine-grained motion fidelity, more realistic hand-object coordination, and effective real-time interaction, establishing a practical step beyond motion reproduction toward real-time human-centric world modeling.
Chinese Translation
我们提出了一种实时人本世界模型,用于上半身的交互生成,旨在合成以人为中心的连贯局部世界动态,其中协调的身体、手和面部动作与可控的人机离散交互共同演变。为此,我们采用了一种连续-离散联合控制方案,包含两个互补组件:连续的人体状态和离散的交互状态。对于连续的人体状态控制,我们引入了一种基于多尺度运动编码的统一隐式表示,其中来自上半身、手和面部的运动潜变量融合到一个共享的潜在空间中。这种多尺度设计提高了不同空间尺度上的表现力,更有效地捕捉细粒度的人体动态,并实现了无需显式重定向的直接控制。对于离散物体交互状态控制,我们使用一小组语言编码的离散交互状态来表示物体接触,其中文本作为显式的交互状态命令,例如“无接触”或“抓取”,而不是开放式生成提示,并进一步构建了一个专门的渲染管道,以监督这种离散交互状态。通过将连续隐式人体状态控制与离散交互状态控制相结合,我们的模型能够精确建模一个人如何移动并与局部环境互动,包括对附近场景状态的可控变化。最后,我们对模型进行了蒸馏,以实现高效的流媒体实时推断,在两台 H100 GPU 上达到了 25 FPS。实验表明,细粒度运动保真度得到了改善,手-物体协调更加真实,实时交互效果显著,标志着在运动再现向实时人本世界建模的实际进展。
cs.CV / 81 / 2607.23522

ATCNet-CIAM for Multi-Session Motor Imagery EEG Signal Classification

用于多会话运动想象脑电信号分类的ATCNet-CIAM
Hai, Le Huu Son, Hai, Nguyen Chi, Vu, Truong Viet, Nguyen, Nguyen Phuc, Anh, Nguyen Thai, Tu, Ngo Hoang
Abstract
Motor imagery (MI)-based electroencephalography is widely used in non-invasive brain--computer interfaces (BCIs), but robust decoding remains challenging due to inter-subject variability and cross-session non-stationarity. This work proposes ATCNet-CIAM, an enhanced attention temporal convolutional network that integrates a lightweight channel-integrated attention module (CIAM) into the ATCNet framework to improve channel-spatial feature representation for MI decoding. The proposed model is evaluated on BCI Competition IV-2a, BCI Competition IV-2b, and the multi-day WBCIC-MI dataset under standard, within-session, and cross-session protocols. Experimental results show that ATCNet-CIAM achieves 86.32% accuracy on BCI IV-2a and 87.96% on BCI IV-2b under the standard protocol, while reaching 89.46% and 83.64% in the within-session WBCIC-MI on 2C and 3C, respectively. The proposed framework consistently improves classification stability and robustness under session-varying conditions, and ablation study confirms the complementary contribution of the proposed architectural components.
Chinese Translation
基于运动想象(MI)的脑电图(EEG)在非侵入式脑-计算机接口(BCI)中被广泛应用,但由于个体间的差异性和跨会话的非平稳性,稳健解码仍然具有挑战性。本研究提出了ATCNet-CIAM,一种增强的注意力时间卷积网络,该网络将轻量级通道集成注意力模块(CIAM)集成到ATCNet框架中,以改善运动想象解码的通道-空间特征表示。所提出的模型在BCI竞赛IV-2a、BCI竞赛IV-2b以及多日WBCIC-MI数据集上进行了评估,采用标准、会话内和会话间协议。实验结果表明,ATCNet-CIAM在BCI IV-2a上达到了86.32%的准确率,在BCI IV-2b上达到了87.96%的准确率,而在会话内的WBCIC-MI中,2C和3C分别达到了89.46%和83.64%。所提出的框架在会话变化条件下始终提高了分类的稳定性和鲁棒性,消融研究确认了所提出的架构组件的互补贡献。
cs.CV / 82 / 2607.23530

Geometry Meets Semantics: Fractional Gradient Stabilization for Semantic-Driven Bounding Box Optimization in Visual Detection Tasks

几何与语义的结合:用于视觉检测任务中语义驱动的边界框优化的分数梯度稳定化
Ming, Qi, Yang, Haitian, Zhao, Xudong, Zhao, Mingjing, Wang, Liuqian, Liu, Nanqing
Abstract
Bounding boxes are fundamental for object localization in visual detection tasks. Among them, oriented bounding boxes are widely used in visual detection tasks, which provide a more precise directional representation. Generally, IoU-based losses are widely adopted to optimize box regression. However, we observed that IoU-driven box optimization suffers from two key issues: (1) it relies solely on geometric properties while ignoring semantic cues; (2) orientation optimization suffers from unstable gradients, causing oscillations in orientation convergence. In this paper, we propose a Fractional Semantic IoU loss to achieve unified semantic-geometric learning with gradient stabilization. First, we design a semantic similarity metric to guide IoU optimization, building a Semantic IoU loss (SIoU loss) with an adaptive gradient gating mechanism. Then, we revisit the gradient instability issue in oriented box optimization and extend the SIoU loss to a fractional-order formulation to build the \textbf{Fr}actional \textbf{S}emantic \textbf{IoU} \textbf{loss} (FrSIoU loss). The FrSIoU loss accumulates historical IoU states to regularize abnormal gradients during bounding box optimization process. Extensive experiments demonstrate that our approach achieves stable performance gains across different bounding box formulations and diverse visual detection tasks. The code will be available on GitHub.
Chinese Translation
边界框是视觉检测任务中物体定位的基础。其中,定向边界框在视觉检测任务中被广泛使用,因为它提供了更精确的方向表示。通常,基于IoU(Intersection over Union)的损失被广泛采用来优化框回归。然而,我们观察到,基于IoU的框优化存在两个关键问题:(1)它仅依赖几何属性而忽略语义线索;(2)定向优化面临不稳定的梯度,导致方向收敛时的振荡。在本文中,我们提出了一种分数语义IoU损失,以实现统一的语义-几何学习和梯度稳定化。首先,我们设计了一种语义相似度度量来指导IoU优化,构建了具有自适应梯度门控机制的语义IoU损失(SIoU损失)。然后,我们重新审视了定向框优化中的梯度不稳定性问题,并将SIoU损失扩展为分数阶形式,以构建分数语义IoU损失(FrSIoU损失)。FrSIoU损失在边界框优化过程中累积历史IoU状态,以规范异常梯度。大量实验表明,我们的方法在不同的边界框形式和多样的视觉检测任务中实现了稳定的性能提升。代码将会在GitHub上发布。
cs.CV / 83 / 2607.23542

GaitFace: A Multimodal Dataset for Long-Range Person Identification

GaitFace:用于远程人员识别的多模态数据集
Komaty, Alain, Luevano, Luis S., Vidit, Vidit, George, Anjith, Amine, Zeina Al, Marcel, Sébastien
Abstract
Efficient border control is becoming a significant global challenge, mainly due to severe congestion and extended passenger waiting times. To mitigate these bottlenecks and facilitate passenger flow, biometric technologies are increasingly deployed to streamline identity verification and enhance crossing efficiency. Technical limitations frequently impede biometric identification, particularly in long-range surveillance, where systems must deal with adverse atmospheric conditions and degraded image quality. While high-quality frameworks like BRIAR exist, they are frequently restricted to specific government agencies. This paper introduces GaitFace, a new public dataset that contains face and gait data captured at long distances. To ensure that the research reflects authentic border scenarios, we use Pre-Enrollment data, where a traveler registers via a mobile device, and "In-the-Wild" captures, which records individuals at a distance across multiple viewing angles and different cameras. Benchmarking SOTA face and gait models reveals that current architectures fail under low-resolution and elevated viewpoints despite success with optical assistance. GaitFace exposes these critical vulnerabilities, providing a rigorous public benchmark to drive more robust, unconstrained biometric research.
Chinese Translation
高效的边境控制正成为一个重大的全球挑战,主要由于严重的拥堵和乘客等待时间的延长。为了缓解这些瓶颈并促进乘客流动,生物识别技术正日益被应用于简化身份验证并提高通行效率。然而,技术限制常常妨碍生物识别识别,特别是在远程监控中,系统必须应对不利的气象条件和降级的图像质量。尽管存在像BRIAR这样的高质量框架,但它们通常仅限于特定的政府机构。本文介绍了GaitFace,一个新的公共数据集,包含在远距离捕获的面部和步态数据。为了确保研究反映真实的边境场景,我们使用了预登记数据,旅行者通过移动设备进行注册,以及“野外”捕获,记录在多个视角和不同摄像头下的个体。对当前最先进的面部和步态模型进行基准测试显示,尽管在光学辅助下取得成功,但现有架构在低分辨率和高视角下表现不佳。GaitFace揭示了这些关键脆弱性,提供了一个严格的公共基准,以推动更强大、无约束的生物识别研究。
cs.CV / 84 / 2607.23575

D3O: Dynamic Distribution Distillation for Ordinal Regression

D3O:用于序数回归的动态分布蒸馏
Dong, Chunlai, Hu, Yaojun, Xu, Yuyang, Ying, Haochao, Wu, Jian
Abstract
Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered. In practice, however, ordinal labels are often obtained by discretizing underlying continuous semantics through subjective human judgment, resulting in ambiguous boundaries and annotation noise. Such uncertainty challenges existing methods that rely on fixed supervision targets, which may reinforce biased ordering under subjective annotations. To address this limitation, we propose D3O, a dynamic distribution distillation framework that replaces static supervision with training-driven evolution of ordinal label distributions via self-distillation. Specifically, we introduce a contrastive ordinal-aware label enhancement module that leverages vision-language alignment to recover refined label distributions capturing both inter-class ambiguity and instance-level uncertainty. Furthermore, we design a CDF-based cross-layer interaction distillation mechanism to propagate cumulative ordinal structure across network hierarchy, ensuring consistent ordinal geometry in intermediate representations. Extensive experiments on four general ordinal regression tasks demonstrate that our proposed D3O consistently outperforms existing approaches, particularly under severe class imbalance and noisy supervision. These results highlight the effectiveness of dynamic supervision in learning robust ordinal representations beyond fixed targets. The code will be publicly available.
Chinese Translation
序数回归广泛应用于标签离散但本质上有序的场景。然而,在实际应用中,序数标签通常是通过主观的人类判断将潜在的连续语义离散化而获得的,这导致了模糊的边界和标注噪声。这种不确定性对依赖固定监督目标的现有方法提出了挑战,这可能在主观标注下强化偏见排序。为了解决这一局限性,我们提出了D3O,一个动态分布蒸馏框架,它通过自蒸馏将静态监督替换为基于训练驱动的序数标签分布演变。具体而言,我们引入了一种对比序数感知标签增强模块,该模块利用视觉-语言对齐来恢复精细的标签分布,以捕捉类间模糊性和实例级不确定性。此外,我们设计了一种基于CDF的跨层交互蒸馏机制,以在网络层次结构中传播累积的序数结构,确保中间表示中的一致序数几何。对四个通用序数回归任务的广泛实验表明,我们提出的D3O在现有方法中始终表现优异,特别是在严重类别不平衡和噪声监督的情况下。这些结果突显了动态监督在学习超越固定目标的稳健序数表示中的有效性。代码将公开发布。
cs.CV / 85 / 2607.23576

Neuromorphic Object Detection: An In-Depth Study and Future Directions

神经形态物体检测:深入研究与未来方向
Li, Jianing, Li, Dianze, Glover, Arren, Fan, Xiaopeng, Li, Guoqi, Bartolozzi, Chiara, Benosman, Ryad B., Tian, Yonghong
Abstract
Conventional frame-based cameras face significant challenges in detecting objects under high-speed motion blur or in low-light environments. Neuromorphic cameras provide asynchronous visual streams with high temporal resolution and a wide dynamic range, offering a promising solution for object detection under challenging conditions. Despite the development of numerous models and the emergence of various applications in neuromorphic object detection, there is still a lack of deep understanding and standardized benchmarks to assess progress and address key challenges. In this paper, we provide a comprehensive survey and benchmark of existing neuromorphic object detection algorithms. Specifically, we first present a problem description, review the available datasets, and revisit the evaluation metrics. We then explore existing neuromorphic object detection approaches from various perspectives, including event representation, temporal modeling, multimodal fusion, asynchronous processing, low-latency processing, and energy-efficient computing. Furthermore, we evaluate a wide range of representative neuromorphic object detection models and offer detailed analyses of the comparative results. Finally, we discuss unresolved issues in neuromorphic object detection and propose potential future research directions. We hope this survey and benchmark will be a valuable resource for researchers and provide guidance for future advancements in neuromorphic object detection.
Chinese Translation
传统的基于帧的摄像头在高速运动模糊或低光环境下检测物体面临重大挑战。神经形态摄像头提供高时间分辨率和宽动态范围的异步视觉流,为在困难条件下的物体检测提供了有希望的解决方案。尽管在神经形态物体检测领域已经开发了许多模型并出现了各种应用,但仍然缺乏对进展的深入理解和标准化基准来解决关键挑战。在本文中,我们提供了现有神经形态物体检测算法的全面调查和基准评估。具体而言,我们首先呈现问题描述,回顾可用的数据集,并重新审视评估指标。然后,我们从多个角度探讨现有的神经形态物体检测方法,包括事件表示、时间建模、多模态融合、异步处理、低延迟处理和节能计算。此外,我们评估了一系列具有代表性的神经形态物体检测模型,并对比较结果进行了详细分析。最后,我们讨论了神经形态物体检测中尚未解决的问题,并提出潜在的未来研究方向。我们希望这项调查和基准能够成为研究人员的宝贵资源,并为神经形态物体检测的未来进展提供指导。
cs.CV / 86 / 2607.23580

SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion

SketchMamba:一种轻量级状态空间模型用于联合渐进式草图分类和笔画自动补全
Jhaveri, Kavish, Shah, Arya
Abstract
Existing vector-sketch models treat recognition and generation as separate tasks, leaving a gap for streaming interfaces that must understand a drawing as it is being made. We present SketchMamba, a single causal sequence model that continuously classifies a sketch from any partial prefix while simultaneously generating its continuation. We achieve this by applying a dense per-step classification loss to a selective state-space backbone. Evaluated on a 58-class subset of the Quick, Draw! dataset, SketchMamba yields 94.93% final-step accuracy and a progressive-accuracy Area Under the Curve (AUC) of 0.706, crossing 90% of its final accuracy by the time 70% of the strokes are drawn. In a matched-budget comparison, the 1.55 million-parameter backbone ties a causal Transformer while outperforming recurrent and convolutional baselines. Ablations confirm that the dense supervision regime, rather than the architecture alone, drives the early-prediction capability. The results demonstrate that a single causal hidden state can unify progressive recognition and autoregressive generation without auxiliary encoders or task-specific branching.
Chinese Translation
现有的矢量草图模型将识别和生成视为两个独立的任务,这给必须理解正在绘制的图形的流媒体接口留下了空白。我们提出了SketchMamba,这是一种单一因果序列模型,可以从任何部分前缀中持续分类草图,同时生成其延续。我们通过对选择性状态空间骨干网络应用密集的逐步分类损失来实现这一目标。在Quick, Draw!数据集的58类子集上进行评估,SketchMamba实现了94.93%的最终步骤准确率和0.706的渐进准确率曲线下面积(AUC),在70%的笔画绘制完成时就达到了最终准确率的90%。在匹配预算的比较中,具有155万参数的骨干网络与因果Transformer持平,同时超越了递归和卷积基线。消融实验确认,密集监督机制而非单纯的架构驱动了早期预测能力。结果表明,单一因果隐状态可以统一渐进识别和自回归生成,而无需辅助编码器或特定任务的分支。
cs.CV / 87 / 2607.23588

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

JarvisHub:一个开放的画布原生多模态创意代理平台
Lin, Yunlong, Lin, Zixu, Xing, Zhaohu, Li, Biqiang, Li, Chenxin, Wang, Haonan, Wu, Haitao, Liu, Hengyu, Chen, Jianghai, Feng, Kaituo, Li, Kaixin, Chen, Shawn, Huang, Shijue, Chen, Sixiang, Ho, Tsung-Yi, Huang, Wenxuan, Liu, Xiangyan, Hu, Xiaomeng, He, Xuanhua, Sun, Yan, Zhao, Yunqing, Yang, Zhiqin, Wang, Zehan, Tang, Zhengyang, Pang, Tianyu, Yue, Xiangyu
Abstract
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.
Chinese Translation
创意人工智能正从单步资产生成转向长时间跨度的多模态生产。尽管近期的生成模型能够合成高质量的图像、视频、音频片段、用户界面元素、故事板、幻灯片及其他创意资产,但现实世界的创意工作需要的不仅仅是孤立的提示-输出交互。它涉及参考资料、草稿、替代方案、编辑、失败尝试、版本关系、工具操作、评估信号和人类反馈,这些共同构成了一个不断演变的项目状态。现有的基于提示、基于对话和基于节点的生成系统仅部分支持这一状态,因为它们通常会丢弃中间上下文,依赖线性对话,或需要手动指定工作流程。近期的商业系统表明,创意生产正在向代理辅助的方向转变,但它们封闭的架构使得研究代理如何表示上下文、选择工具、修订工件、从失败中恢复以及保持一致性变得困难。为了解决这一问题,我们引入了JarvisHub,一个用于长时间跨度多模态创作的画布原生创意代理平台。JarvisHub将可编辑的画布视为用户工作空间、代理的外部记忆、行动空间和共享项目状态,将多模态工件、依赖关系、版本和反馈表示为类型化的画布节点和链接。通过画布状态、协议桥接和代理运行时的三层架构,JarvisHub使代理能够在可检查和可编辑的创意状态中行动。这一设计使创意代理超越孤立工具使用,朝着持续的、可由人类引导的创意自动化发展,代理能够逐步规划、生成、修订和组织多模态项目,同时用户能够在整个过程中进行检查、指导和干预。
cs.CV / 88 / 2607.23594

Weakly Supervised Instance-Level Gleason Pattern Estimation Using Primary and Secondary Labels

基于主标签和副标签的弱监督实例级格里森模式估计
Sugeta, Nao, Shiku, Kaito, Matsuo, Shinnosuke, Bise, Ryoma
Abstract
In prostate cancer histopathology, the Gleason Score is determined by the most frequent (Primary) and second most frequent (Secondary) Gleason patterns within a whole-slide image. Although these slide-level labels are routinely available in clinical practice, instance-level Gleason annotations are rarely provided, making patch-level learning challenging. We propose a Multiple Instance Learning (MIL) framework that estimates instance-level Gleason patterns from slide-level Primary and Secondary labels. The proposed method formulates instance-level learning according to the clinical definition of the Gleason Score by aggregating instance predictions into class counts and explicitly modeling the Primary pattern, Secondary pattern, and their dominance. Experimental results demonstrate that the proposed formulation enables effective instance-level learning and outperforms existing MIL approaches on the SICAP-MIL dataset.
Chinese Translation
在前列腺癌组织病理学中,格里森评分是通过全切片图像中最频繁(主)和第二频繁(副)格里森模式来确定的。尽管这些切片级标签在临床实践中通常是可用的,但实例级格里森注释却很少提供,这使得基于图块的学习变得具有挑战性。我们提出了一种多实例学习(Multiple Instance Learning, MIL)框架,该框架从切片级主标签和副标签中估计实例级格里森模式。所提方法根据格里森评分的临床定义,通过将实例预测聚合为类别计数,并明确建模主模式、副模式及其主导关系,来构建实例级学习。实验结果表明,所提的公式化方法能够有效地进行实例级学习,并在SICAP-MIL数据集上优于现有的MIL方法。
cs.CV / 89 / 2607.23600

ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion

ConFusion:用于细粒度可控红外与可见光图像融合的连续融合空间学习
Yurong, Guo, Yufei, He, Yonghao, Li, Dongliang, Chang, Ke, Zhang, Zhanyu, Ma
Abstract
Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, producing fused images that adapt to diverse user requirements and downstream tasks. However, existing methods typically rely on predefined discrete control conditions, leading to a sparse space that fails to support fine-grained modulation demands. To address this, we propose ConFusion, a novel framework that learns the continuous fusion space via Gaussian-conditioned spatial-aware modulation, enabling instance-level fine-grained controllable infrared and visible image fusion. ConFusion employs a dual-branch architecture to disentangle modality-invariant and modality-specific representations under joint reconstruction and text-guided semantic alignment. Gaussian-conditioned instance modulation variables coupled with Grounded SAM-based instance masks guide instance-level fine-grained modulation through the Mask-Guided Specific Feature Modulator, while the Text-Driven Invariant Feature Enhancer improves semantic consistency and enhances fusion. During inference, the multimodal large language model parses user intents into instance-level modulation variables to guide image fusion. Extensive experiments show that ConFusion achieves state-of-the-art performance across multiple metrics in both fusion quality and downstream tasks, while supporting fine-grained controllable image fusion. Our code is available at https://github.com/HeyufeiAnto/Confusion
Chinese Translation
可控的红外-可见光图像融合旨在整合互补的热信息和结构信息,并通过灵活的区域感知调制,生成适应多样用户需求和下游任务的融合图像。然而,现有方法通常依赖于预定义的离散控制条件,导致稀疏空间无法满足细粒度调制需求。为此,我们提出了ConFusion,一个通过高斯条件空间感知调制学习连续融合空间的新框架,使得实例级细粒度可控红外与可见光图像融合成为可能。ConFusion采用双分支架构,在联合重建和文本引导的语义对齐下,解耦模态不变和模态特定的表示。高斯条件实例调制变量结合基于Grounded SAM的实例掩膜,通过掩膜引导的特定特征调制器引导实例级细粒度调制,而文本驱动的不变特征增强器则提高语义一致性并增强融合效果。在推理过程中,多模态大型语言模型将用户意图解析为实例级调制变量,以指导图像融合。大量实验表明,ConFusion在融合质量和下游任务的多个指标上均实现了最先进的性能,同时支持细粒度可控图像融合。我们的代码可在 https://github.com/HeyufeiAnto/Confusion 获取。
cs.CV / 90 / 2607.23608

Markerless Motion Capture in Routine Clinical Upper Limb Assessments: Validity and Insights Beyond Ordinal Scoring

无标记运动捕捉在常规临床上肢评估中的应用:有效性及超越序数评分的见解
Unger, Tim, Lambercy, Olivier, Gassert, Roger, Luft, Andreas R., Cotton, R. James, Awai, Chris Easthope
Abstract
The Action Research Arm Test (ARAT) is a widely-used upper limb outcome measure in neurorehabilitation, but its ordinal scoring is subjective and suffers from limited sensitivity and specificity. We evaluated whether artificial-intelligence (AI)-based markerless motion capture (MMC), embedded into ARAT assessments during clinical routine, accurately reconstructs upper limb movement and yields valid, objective kinematic metrics carrying clinically meaningful information beyond the ordinal score. Across 47 sessions from 20 mixed-neurological patients (1,174 ARAT tasks), biomechanical reconstruction was accurate and robust across impairment levels, and kinematic metrics showed the discrimination pattern expected of a construct-valid measure. In longitudinal case studies, the metrics added the specificity and sensitivity the ordinal score lacks: a domain decomposition exposed patient-specific recovery profiles underlying equal ARAT gains (specificity), and kinematic improvement continued to be detected after the ARAT had saturated (sensitivity). MMC in clinical routine can thus provide valid, objective, sensitive, and specific kinematic measurement complementing ordinal scoring.
Chinese Translation
行动研究手臂测试(Action Research Arm Test, ARAT)是神经康复中广泛使用的上肢结果测量工具,但其序数评分具有主观性,并且在敏感性和特异性方面存在局限性。我们评估了基于人工智能(Artificial Intelligence, AI)的无标记运动捕捉(Markerless Motion Capture, MMC)在临床常规的ARAT评估中是否能够准确重建上肢运动,并提供超越序数评分的有效、客观的运动学指标,这些指标具有临床意义的信息。在来自20名混合神经病患者的47个会话(1,174个ARAT任务)中,生物力学重建在各类损伤水平上均表现出准确性和稳健性,运动学指标显示出预期的构念效度测量的区分模式。在纵向案例研究中,这些指标增加了序数评分所缺乏的特异性和敏感性:领域分解揭示了在相同ARAT增益下患者特定的恢复模式(特异性),而运动学改善在ARAT达到饱和后仍然被检测到(敏感性)。因此,MMC在临床常规中可以提供有效、客观、敏感和特异的运动学测量,补充序数评分。
cs.CV / 91 / 2607.23631

PathSelect: Sequential Token Selection for Whole Slide Pathology

PathSelect:全幻灯片病理的顺序标记选择
Chen, Jingzhi, He, Landi, Chen, Zehong, Wu, Peihang, Xu, Lijian
Abstract
Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominantly rely on spatial sampling or training-free pruning, which risk diluting weak but informative signals, leading to the loss of critical diagnostic evidence due to the spatially diffuse nature of pathological cues. We reformulate WSI token pruning as a sequential selection process, enabling the model to autonomously learn an optimal routing strategy rather than relying on static heuristics. We herein propose a decoupled routing framework integrated as an active plugin into the fully pre-trained SlideChat base model, leaving both the slide encoder and large language model frozen. To provide continuous gradients for the non-differentiable pruning operation during training, we introduce PathSelect. PathSelect employs a variance-preserving noise gate to modulate each patch's information flow via a differentiable Soft Top-K operator, paired with a diagonal-attention Denoiser that recovers the perturbed representations without semantic leakage. At inference, the PathSelect module is entirely detached. Relying solely on the trained Scorer, a deterministic Hard Top-K operator executes adaptive, data-dependent trajectory termination, significantly accelerating downstream generative processing with exceptionally low sequential token selection latency. Driven by an empirical average of only 44.86 tokens under a maximum constraint of K = 128, our framework achieves 74.00% overall accuracy on SlideBench (TCGA), representing an approximate 36.6x spatial token reduction relative to the uncompressed baseline average while consistently outperforming sampling-based counterparts.
Chinese Translation
千兆像素全幻灯片图像(WSIs)由于极长的序列长度,给视觉语言模型(VLMs)带来了根本性的计算瓶颈。现有的方法主要依赖于空间采样或无训练修剪,这可能会稀释微弱但信息丰富的信号,导致由于病理线索的空间分散特性而丧失关键的诊断证据。我们将WSI标记修剪重新表述为一个顺序选择过程,使模型能够自主学习最佳路由策略,而不是依赖静态启发式方法。我们在此提出一个解耦的路由框架,作为一个主动插件集成到完全预训练的SlideChat基础模型中,同时保持幻灯片编码器和大型语言模型不变。为了在训练过程中为不可微分的修剪操作提供连续的梯度,我们引入了PathSelect。PathSelect采用保持方差的噪声门,通过可微分的Soft Top-K操作调节每个补丁的信息流,并配备一个对角注意力去噪器,恢复扰动的表示而不泄漏语义。在推理时,PathSelect模块完全脱离。仅依赖训练好的评分器,确定性的Hard Top-K操作执行自适应的数据依赖轨迹终止,显著加速下游生成处理,同时具有极低的顺序标记选择延迟。在最大约束K = 128下,基于经验的平均仅为44.86个标记,我们的框架在SlideBench(TCGA)上实现了74.00%的整体准确率,相较于未压缩基线平均值,空间标记减少约36.6倍,同时始终优于基于采样的对手。
cs.CV / 92 / 2607.23638

WGDnet: Wishart-guided Geometric-aware Deep Network for PolSAR Image Classification

WGDnet:基于Wishart指导的几何感知深度网络用于极化合成孔径雷达图像分类
Shi, Junfei, Zhang, Haojia, Cheng, Yu, Li, Yuke
Abstract
Polarimetric Synthetic Aperture Radar (PolSAR) classification underpins all-weather Earth observation. Conventional Wishart methods depend on rigid handcrafted operators with limited adaptability, while mainstream deep networks ignore PolSAR native Wishart scattering statistics. Additionally, fixed convolution windows fail to capture multi-scale, multi-directional terrain patterns, harming boundary detection and small-object characterization. To mitigate these drawbacks, we propose WGDNet, a Wishart-guided geometric-aware deep network. It integrates three core designs: (1) learnable Wishart convolutions with directional kernels for multi-scale statistical edge feature extraction; (2) an orientation-prior aggregation module that estimates dominant local directions and confidences to refine directional Wishart outputs adaptively; (3) GAnet, a scale-direction adaptive geometric-aware convolution that dynamically reshapes sampling grids to model anisotropic terrain and retain fine details. Our contributions lie in learnable Wishart statistical modeling, orientation-prior feature aggregation, and geometry-adaptive convolution. Evaluations across four real PolSAR datasets verify WGDNet surpasses existing state-of-the-art approaches in classification accuracy and boundary fidelity.
Chinese Translation
极化合成孔径雷达(PolSAR)分类是全天候地球观测的基础。传统的Wishart方法依赖于刚性手工设计的算子,适应性有限,而主流深度网络忽视了PolSAR固有的Wishart散射统计特性。此外,固定的卷积窗口无法捕捉多尺度、多方向的地形模式,影响边界检测和小物体特征提取。为了解决这些缺陷,我们提出了WGDNet,一种基于Wishart指导的几何感知深度网络。它整合了三个核心设计:(1)具有方向性核的可学习Wishart卷积,用于多尺度统计边缘特征提取;(2)一个方向优先聚合模块,估计主导局部方向和置信度,以自适应地细化方向性Wishart输出;(3)GAnet,一种尺度-方向自适应的几何感知卷积,动态重塑采样网格以建模各向异性地形并保留细节。我们的贡献在于可学习的Wishart统计建模、方向优先特征聚合和几何自适应卷积。对四个真实PolSAR数据集的评估验证了WGDNet在分类精度和边界保真度上超越了现有的最先进方法。
cs.CV / 93 / 2607.23657

GRAPE: Graduated Routing for Articulated Portrait mesh Estimation

GRAPE:用于关节肖像网格估计的渐进路由
Liu, Yunfei, Lin, Lijian, Zhu, Ye, Li, Yu
Abstract
Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rely on 3D Morphable Models (3DMMs). However, face-centric models suffer from the "floating head" assumption, conflating head pose with global rotation due to the lack of neck kinematics. Conversely, body-centric models lack high-fidelity facial expression capabilities. Furthermore, current methods struggle to disentangle jaw articulation from expression blendshapes, often over-relying on expressions for mouth opening. These limitations make monocular portrait recovery difficult across representation, supervision, and anatomical parameter estimation. To address these limitations, we introduce GRAPE(Graduated Routing for Articulated Portrait mesh Estimation). We build a Portrait Parametric Model (PPM) with an explicit torso-to-head kinematic chain and a canonical injection step to merge FLAME and the SMPL-X torso. We propose a Progressive Anatomical Alignment (PAA) network, which is composed of a pretrained portrait encoder, a Graduated-Mask Router, and coarse-to-fine experts that follow the portrait anatomical prior. We then train this network with multi-source supervision that combines sparse anatomical keypoints, feature distillation, foreground mask constraints, and relative geometry constraints. Experiments show that GRAPE improves portrait mesh recovery quality, pose alignment, and jaw--expression disentanglement over prior methods. We also demonstrate that our method can benefit the downstream tasks of audio-driven talking-head generation and 3D portrait generation.
Chinese Translation
关节肖像网格估计是3D理解、虚拟形象生成和沉浸式交互的基础。现有方法主要依赖于3D可变形模型(3DMMs)。然而,以面部为中心的模型受到“漂浮头部”假设的影响,由于缺乏颈部运动学,将头部姿态与全局旋转混淆。相反,以身体为中心的模型缺乏高保真的面部表情能力。此外,当前方法在将下颌运动与表情混合形状区分开时面临困难,往往过度依赖表情来表示嘴部张开。这些局限性使得单目肖像恢复在表示、监督和解剖参数估计方面变得困难。为了解决这些问题,我们提出了GRAPE(用于关节肖像网格估计的渐进路由)。我们构建了一个具有明确躯干到头部运动链的肖像参数模型(PPM),并通过规范注入步骤将FLAME与SMPL-X躯干合并。我们提出了一种渐进解剖对齐(PAA)网络,该网络由一个预训练的肖像编码器、一个渐进掩码路由器和遵循肖像解剖先验的粗到细专家组成。然后,我们使用结合稀疏解剖关键点、特征蒸馏、前景掩码约束和相对几何约束的多源监督来训练该网络。实验表明,GRAPE在肖像网格恢复质量、姿态对齐和下颌-表情解耦方面优于先前的方法。我们还展示了我们的方法可以为音频驱动的说话头生成和3D肖像生成等下游任务带来益处。
cs.CV / 94 / 2607.23658

XMatchAD: A Cross-Modal Matching Perspective on Reconstruction-based Anomaly Detection

XMatchAD:基于重建的异常检测的跨模态匹配视角
Cai, Mingxiu, Zhang, Zhe, Wu, Gaochang, Chai, Tianyou
Abstract
The remarkable success of reconstruction-based methods in Unsupervised Anomaly Detection (UAD) lies in their ability to identify and localize anomalies by modeling discrepancies between input images and their reconstructed counterparts. However, these approaches often struggle to capture subtle anomalies and tend to produce blurred anomaly boundaries, which significantly limits their effectiveness, particularly in complex multi-class scenarios. To address these issues, we present XMatchAD, a novel UAD framework that reinterprets the task from a pseudo cross-modal matching perspective. Specifically, the input and reconstructed images are treated as two complementary modalities and their matching relationships are precisely exploited for anomaly detection. First, a pre-trained feature extractor is employed to encode discriminative representations. Second, an attention-guided cross-modal matching mechanism is introduced to match local inter-modal anomaly-related patterns while mutually refining the features. This enhances the sensitivity to anomalies with diverse shapes and subtle deviations and significantly improves the precision of anomaly detection and localization. Third, we design an adaptive frequency-aware fusion module that further delineates sharp anomaly boundaries through the coupling of high-frequency components from cross-modal multi-scale representations. Comprehensive evaluations on MVTec-AD, VisA, and MPDD benchmarks demonstrate that our method consistently achieves superior performance, outperforming state-of-the-art methods in multi-class anomaly detection and localization. The code will be released at https://github.com/Mingxiu-Cai/XMatchAD.
Chinese Translation
基于重建的方法在无监督异常检测(UAD)中的显著成功在于其通过建模输入图像与重建图像之间的差异来识别和定位异常。然而,这些方法往往难以捕捉细微的异常,并且倾向于产生模糊的异常边界,这显著限制了它们的有效性,尤其是在复杂的多类场景中。为了解决这些问题,我们提出了XMatchAD,这是一种新颖的UAD框架,从伪跨模态匹配的角度重新解释了这一任务。具体而言,输入图像和重建图像被视为两种互补的模态,并精确利用它们的匹配关系进行异常检测。首先,采用预训练的特征提取器来编码具有区分性的表示。其次,引入了一种基于注意力的跨模态匹配机制,以匹配局部的跨模态异常相关模式,同时相互优化特征。这增强了对具有多样形状和细微偏差的异常的敏感性,并显著提高了异常检测和定位的精度。第三,我们设计了一个自适应频率感知融合模块,通过结合来自跨模态多尺度表示的高频成分,进一步勾勒出清晰的异常边界。在MVTec-AD、VisA和MPDD基准上的全面评估表明,我们的方法始终实现了卓越的性能,在多类异常检测和定位中超越了最先进的方法。代码将发布在https://github.com/Mingxiu-Cai/XMatchAD。
cs.CV / 95 / 2607.23669

RRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes

RRTrack:动态场景下稳健且可恢复的物体六维姿态跟踪
Li, Junyue, Zheng, Ye, Chen, Yifan, Sun, Zhe, Li, Xuelong
Abstract
Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computationally expensive, while current trackers struggle with fast motion and complete occlusion due to their reliance on continuous visibility. To address these challenges, we present RRTrack, an efficient, recoverable object 6D pose tracker that enables robust tracking through fast motion and target disappearance--reappearance. RRTrack introduces a 2D--6D closed-loop tracking strategy that integrates memory-based video object segmentation (VOS) with 6D pose refinement. The 2D branch maintains target localization, and the 6D branch verifies geometric consistency before memory updates. In addition, a DINOv2-based dual-bank template matching module is developed to recover lost targets by jointly exploiting offline synthetic templates and online observation anchors while maintaining real-time efficiency. We also introduce a synthetic RGB-D benchmark comprising three robotic scenarios with fast motion and full occlusion. Experimental results on the synthetic benchmark demonstrate that RRTrack improves equal-subset mean ADD-S AR by 66.3\% and ADD-S AUC by 65.7\% over FoundationPose while achieving 55.2 FPS. Real-world experiments further validate the robustness of RRTrack under noisy sensing conditions. Project page: https://github.com/7kevin24/RRTrack
Chinese Translation
稳健的物体六维姿态跟踪对于在动态和遮挡场景中操作的机器人系统至关重要。每帧估计器虽然准确,但计算开销较大,而当前的跟踪器由于依赖于持续可见性,在快速运动和完全遮挡的情况下表现不佳。为了解决这些挑战,我们提出了RRTrack,一种高效且可恢复的物体六维姿态跟踪器,能够在快速运动和目标消失-重现的情况下实现稳健跟踪。RRTrack引入了一种二维-六维闭环跟踪策略,将基于记忆的视频物体分割(VOS)与六维姿态细化相结合。二维分支负责维持目标定位,而六维分支在更新记忆之前验证几何一致性。此外,我们开发了一种基于DINOv2的双库模板匹配模块,通过联合利用离线合成模板和在线观察锚点来恢复丢失的目标,同时保持实时效率。我们还引入了一个合成RGB-D基准,包含三个具有快速运动和完全遮挡的机器人场景。在合成基准上的实验结果表明,RRTrack在均等子集平均ADD-S AR上比FoundationPose提高了66.3%,在ADD-S AUC上提高了65.7%,同时实现了55.2 FPS的帧率。现实世界的实验进一步验证了RRTrack在噪声感知条件下的稳健性。项目页面:https://github.com/7kevin24/RRTrack
cs.CV / 96 / 2607.23673

Contrastive Parameter Disentanglement for Multi-modal Remote Sensing Image Generation

多模态遥感图像生成的对比参数解耦
Zhang, Yu, Zhao, Wenda, Tang, Haojun, Wang, Haipeng
Abstract
Existing remote sensing image generation methods are largely confined to single-modality synthesis and therefore fail to exploit the complementary information inherent in multimodal imagery. To address this limitation, we propose a contrastive parameter disentanglement framework for multimodal remote sensing image generation, which generates semantically consistent and structurally aligned images across multiple modalities, including optical, infrared, and synthetic aperture radar (SAR), from a single text prompt. Specifically, we introduce a contrastive parameter disentanglement module that disentangles shared semantics from modality-specific attributes at the parameter level within an orthogonal core subspace. Based on this module, we develop a disentangled optimization strategy that first constrains the parameter matrix A of the LoRA adapter to capture modality-invariant semantics through a multimodal contrastive objective and then guides multiple parameter matrices B to learn modality-specific attributes under text conditioning. This strategy enables the simultaneous generation of multimodal images with consistent semantic content and distinct modality characteristics. Furthermore, to ensure structural alignment across the generated images, we devise a query-key structure transfer mechanism that jointly models multimodal sampling trajectories during inference by transferring structural correlation priors from an anchor modality to the remaining modalities. Extensive experiments demonstrate that our method outperforms state-of-the-art remote sensing image generation approaches in terms of generation quality, semantic consistency, and structural alignment, while also achieving superior performance in the downstream object classification task.
Chinese Translation
现有的遥感图像生成方法主要局限于单一模态合成,因此未能充分利用多模态图像中固有的互补信息。为了解决这一限制,我们提出了一种用于多模态遥感图像生成的对比参数解耦框架,该框架能够从单一文本提示生成语义一致且结构对齐的多模态图像,包括光学、红外和合成孔径雷达(SAR)图像。具体而言,我们引入了一个对比参数解耦模块,该模块在正交核心子空间内从参数层面解耦共享语义与模态特定属性。基于该模块,我们开发了一种解耦优化策略,该策略首先通过多模态对比目标约束LoRA适配器的参数矩阵A,以捕捉模态不变的语义,然后引导多个参数矩阵B在文本条件下学习模态特定属性。这一策略使得能够同时生成具有一致语义内容和独特模态特征的多模态图像。此外,为了确保生成图像之间的结构对齐,我们设计了一种查询-键结构转移机制,该机制在推理过程中通过将结构相关先验从锚模态转移到其余模态,共同建模多模态采样轨迹。大量实验表明,我们的方法在生成质量、语义一致性和结构对齐方面优于最先进的遥感图像生成方法,同时在下游目标分类任务中也表现出更优的性能。
cs.CV / 97 / 2607.23680

Perturbation-Aware Diffusion-Guided Hybrid Segmentation for Robust and Annotation-Efficient Plant Stress Phenotyping

考虑扰动的扩散引导混合分割用于稳健且高效标注的植物应激表型分析
Chaurakoti, Gurbhit, Kar, Soumyashree
Abstract
Semantic segmentation in agricultural imagery is often evaluated under in-domain protocols, yet practical deployment requires robustness to appearance perturbations, limited annotations, and cross domain shift. This paper presents a diffusion-guided hybrid segmentation framework in which U-Net, DeepLabV3+, and SegFormer backbones generate coarse masks that are refined by Denoising Diffusion Probabilistic Models (DDPM), latent diffusion, or semantic-guided diffusion. The framework is evaluated through a 3x3 architectural screening study on PlantSegV3, followed by boundary-constrained optimization, perturbation-guided retraining, low-data evaluation, constrained hyperparameter screening, and controlled cross-domain adaptation. On PlantSegV3, the best selected hybrid model achieves 71.83% refined mean Intersection-over-Union (mIoU) and 26.10% refined Boundary-F1, and the selected models remain stable under substantially reduced supervision, demonstrating strong annotation efficiency. Perturbation analysis identifies grayscale conversion, fog, coarse dropout, and shadow as the most disruptive appearance shifts, and the resulting augmentation policy substantially improves robustness during retraining. The adapted models further show effective transfer to external agricultural datasets under limited target supervision, indicating that diffusion refinement and boundary-aware optimization provide transferable structural priors. Overall, the results show that carefully matched backbone-refiner pairings, combined with perturbation-aware retraining, can improve structural delineation and robustness under realistic resource and distribution constraints.
Chinese Translation
农业图像中的语义分割通常在领域内协议下进行评估,但实际应用需要对外观扰动、有限标注和跨领域转移具有稳健性。本文提出了一种扩散引导的混合分割框架,其中 U-Net、DeepLabV3+ 和 SegFormer 主干生成粗略掩膜,这些掩膜通过去噪扩散概率模型(Denoising Diffusion Probabilistic Models, DDPM)、潜在扩散或语义引导扩散进行精细化。该框架通过在 PlantSegV3 上进行 3x3 架构筛选研究进行评估,随后进行边界约束优化、扰动引导再训练、低数据评估、约束超参数筛选和受控跨领域适应。在 PlantSegV3 上,最佳选择的混合模型实现了 71.83% 的精细化平均交并比(mean Intersection-over-Union, mIoU)和 26.10% 的精细化边界 F1 值,所选模型在显著减少监督的情况下仍保持稳定,显示出强大的标注效率。扰动分析确定了灰度转换、雾、粗略丢失和阴影为最具干扰性的外观变化,所得到的增强策略在再训练过程中显著提高了稳健性。适应后的模型在有限目标监督下进一步有效转移到外部农业数据集,表明扩散精细化和边界感知优化提供了可转移的结构先验。总体而言,结果表明,精心匹配的主干-精细化配对结合扰动感知再训练可以在现实资源和分布约束下改善结构描绘和稳健性。
cs.CV / 98 / 2607.23687

GNM Head: A Generative aNthropometric Model of the human head

GNM 头部:一种人类头部的生成性人类测量模型
Ploumpis, Stylianos, Bednarik, Jan, Zoss, Gaspard, Guseinov, Ruslan, Prasso, Luca, Chandran, Prashanth, Boyne, Oliver, Choutas, Vasileios, Bolkart, Timo, Wang, Daoye, Chai, Menglei, Qiu, Di, Winberg, Sebastian, Rainer, Gilles, Bridgeman, Lewis, Vicini, Delio, Riviere, Jérémy, Boetzel, Yannick, Koumis, Alexander, Busch, Jay, Herrera, Cynthia, Still, Jacob, Ysebert, Scott, Lincoln, Peter, Escolano, Sergio Orts, Rhemann, Christoph, Wood, Erroll, Beeler, Thabo, Zafeiriou, Stefanos
Abstract
Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight spatial control of generated imagery. However, existing publicly available models are typically limited in anatomical scope, modeling only outer geometry while ignoring intra-oral and ocular structures, and frequently suffer from reduced geometric quality stemming from low-fidelity input datasets. In this report we introduce a new parametric model dubbed Generative aNthropometric Model (GNM), named as a homophone of the human genome. GNM encompasses the head, face, neck, eyeballs, teeth, and tongue, and it is built on an extensive database of high-resolution 3D scans combined with high-quality anatomy specific artist-made samples. This report details the data provenance, the model architecture including the specialized sub-models for the ocular and intra-oral structures, and shows its SotA performance on fitting target 3D face scans. To foster community innovation, the complete GNM framework is made publicly available.
Chinese Translation
人类头部的参数模型是计算机视觉和图形学中传统上用于动画、渲染和重建的重要工具。近年来,它们在生成性大型视觉模型中作为关键的条件信号,允许对生成图像进行精确的空间控制。然而,现有的公开可用模型通常在解剖范围上受到限制,仅建模外部几何形状,而忽略了口腔和眼部结构,并且常常由于低保真度输入数据集而导致几何质量下降。在本报告中,我们介绍了一种新的参数模型,称为生成性人类测量模型(Generative aNthropometric Model,GNM),其名称与人类基因组同音。GNM 包括头部、面部、颈部、眼球、牙齿和舌头,并基于一个广泛的高分辨率 3D 扫描数据库,结合高质量的特定解剖结构艺术家制作的样本。本文详细介绍了数据来源、模型架构,包括眼部和口腔内结构的专门子模型,并展示了其在拟合目标 3D 面部扫描方面的最先进表现。为了促进社区创新,完整的 GNM 框架已公开提供。
cs.CV / 99 / 2607.23694

Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation

SAM3的参数高效适应用于基于提示的外科概念分割
Liu, Changjing, Huang, Yiming, Cui, Beilei, Shao, Liangjing, Bai, Long, Li, Yanheng, Che, Haoxuan, Ren, Hongliang
Abstract
Efficient surgical segmentation empowers clinical diagnosis, intraoperative monitoring, and downstream robotic pipelines for reconstruction and simulation. Although prompt-driven foundation models like Segment Anything Model 3 (SAM3) achieve strong segmentation performance on natural images, surgical data exhibits domain gaps against its pre-training data, resulting in degraded segmentation accuracy. Furthermore, existing medical SAM methods require full-parameter fine-tuning, incurring heavy computational consumption and low efficiency. To address these limitations, this work proposes a parameter-efficient Low-Rank Adaptation (LoRA) adaptation of SAM3 for surgical concept segmentation. We inject low-rank adapters into the prompt encoder, detector and tracker while fully freezing the vision backbone, which only optimizes 0.98% of the total model parameters and supports training on a single consumer GPU. Comprehensive experiments demonstrate that our method consistently outperforms zero-shot SAM3 and other mainstream baselines, and the generated segmentation results can be directly deployed to support downstream robotic surgical scene reconstruction and physical simulation pipelines.
Chinese Translation
高效的外科分割能够增强临床诊断、术中监测以及后续的机器人重建和仿真流程。尽管基于提示的基础模型如Segment Anything Model 3 (SAM3)在自然图像上实现了强大的分割性能,但外科数据与其预训练数据之间存在领域差距,导致分割准确性下降。此外,现有的医学SAM方法需要对所有参数进行微调,造成了巨大的计算消耗和低效率。为了解决这些局限性,本研究提出了一种针对外科概念分割的SAM3参数高效低秩适应(Low-Rank Adaptation, LoRA)。我们在提示编码器、检测器和跟踪器中注入低秩适配器,同时完全冻结视觉主干,仅优化0.98%的总模型参数,并支持在单个消费级GPU上进行训练。全面的实验表明,我们的方法在零样本SAM3和其他主流基线中始终表现优越,生成的分割结果可以直接部署以支持后续的机器人外科场景重建和物理仿真流程。
cs.CV / 100 / 2607.23755

DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation

DAP-Pose:深度时间对齐与物理感知跨模态传感器融合用于鲁棒姿态估计
Lin, Jianhan, Qin, Yuchu, Yuan, Jiateng, Zhang, Wenbo, Gao, Shuai
Abstract
Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.
Chinese Translation
在复杂环境中,使用多模态传感器进行鲁棒且准确的姿态估计对于自主车辆和移动机器人系统至关重要。本文提出了DAP-Pose,一个统一的端到端模型,用于鲁棒的多模态姿态估计。DAP-Pose引入了一个双层跨模态融合(Bi-level Cross-modal Fusion, BCF)模块,该模块从视觉、惯性和全球导航卫星系统(GNSS)测量中捕获互补的语义和几何运动线索。为了处理时间偏移,我们设计了一个深度时间对齐(Deep Temporal Alignment, DTA)模块,该模块在潜在空间中显式对齐异步流,从而实现一致的运动建模,而无需严格的硬件同步。此外,我们通过流形几何和GNSS引导的绝对度量尺度引入了物理感知约束,强制运动一致性并减轻漂移。在公共KITTI基准数据集上进行了实验,以评估DAP-Pose相对于现有方法的性能。DAP-Pose达到了最先进的性能,平均平移误差($t_{rel}$)最低为1.31%,旋转误差($r_{rel}$)为0.46$^{ ext{°}}$。此外,它能够准确估计姿态,并在严重人为注入的时间错位下保持鲁棒性能。
cs.CV / 101 / 2607.23758

RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction

RoadVGGT:道路结构感知前馈道路表面重建
Jiao, Han, Liu, Chen, Sun, Jiakai, Zhang, Zhanjie, Yang, Mengyuan, Li, Yimeng, Zhou, Mofan, Zhan, Kun, Zhao, Lei
Abstract
Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-specialized optimization methods can produce high-quality road representations, but they typically require per-scene training and scene-dependent coverage design around the driving trajectory, limiting scalable reconstruction over newly collected roads. To address these limitations, we introduce RoadVGGT, a road-structure-aware feed-forward framework that reconstructs compact Gaussian road surfaces without test-time per-scene optimization. RoadVGGT uses a geometric foundation model to exploit multi-view images together with provided pose and depth observations, and predicts dense pixel-aligned Gaussian attributes through a learned Gaussian head. To make these dense predictions usable for large road surfaces, we align them into a consistent metric world coordinate system and fuse redundant Gaussians on the road-aligned XY plane through confidence-weighted grid fusion. Category-aware grouping and road--sidewalk junction protection further control fusion around vulnerable road structures. The resulting representation supports RGB and semantic bird's-eye-view maps, elevation estimation, and novel view synthesis. RoadVGGT eliminates the need for per-scene optimization in prior methods, reconstructs complete road surfaces with a compact Gaussian representation, and improves image quality, semantic mapping, and elevation accuracy. Extensive experiments demonstrate the potential of geometric foundation models for scalable feed-forward road surface reconstruction.
Chinese Translation
大规模道路表面重建支持高清地图、自动驾驶感知、标注和仿真。现有的道路专用优化方法能够生成高质量的道路表示,但通常需要针对每个场景进行训练,并围绕驾驶轨迹设计场景依赖的覆盖,这限制了对新收集道路的可扩展重建。为了解决这些限制,我们提出了RoadVGGT,一种道路结构感知的前馈框架,能够在不进行测试时每个场景优化的情况下重建紧凑的高斯道路表面。RoadVGGT利用几何基础模型,结合多视角图像以及提供的姿态和深度观测,预测通过学习的高斯头生成的密集像素对齐高斯属性。为了使这些密集预测可用于大规模道路表面,我们将其对齐到一致的度量世界坐标系,并通过基于置信度加权的网格融合在道路对齐的XY平面上融合冗余高斯。类别感知分组和道路-人行道交界保护进一步控制了脆弱道路结构周围的融合。最终生成的表示支持RGB和语义鸟瞰图、海拔估计以及新视图合成。RoadVGGT消除了先前方法中每个场景优化的需求,以紧凑的高斯表示重建完整的道路表面,并提高了图像质量、语义映射和海拔精度。大量实验表明,几何基础模型在可扩展前馈道路表面重建中的潜力。
cs.CV / 102 / 2607.23794

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

PathScale-R1:病理图像分析中的跨尺度推理
Phan, Chi, Zhang, Tianyi, Wu, Yufeng, Xue, Qiaochu, Zhang, Jiajie, Cai, Linghan, Liu, Zeyu, Wang, Sudong, Jin, Yueming, Hu, Dan
Abstract
Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification. However, existing pathology benchmarks and vision-language models (VLMs) are still largely developed under single-scale settings, limiting their ability to learn clinically meaningful multi-magnification reasoning. Moreover, naively constructed visual question answering (VQA) tasks may be susceptible to text-only or superficial visual shortcuts, leading to unreliable assessments of visual understanding. To address these limitations, we introduce a benchmark and training framework for shortcut-resistant cross-scale pathology reasoning. We design an Adversarial Text-only Screening strategy for semantic reasoning questions and a Structure-controlled Distractor Sampling strategy for visual grounding questions, encouraging models to rely on cross-scale visual evidence. Based on this pipeline, we construct PathScale-VQA, a high-quality cross-scale pathology VQA benchmark with 10,373 multiple-choice questions grounded in 1,368 diagnostic paths across multiple magnification levels. Building on the semantic reasoning set, PathScale-R1 is optimized through Difficulty-driven Reasoning Distillation supervised fine-tuning followed by reinforcement learning with a Scale-aware Reasoning Structure reward, which encourages the use of evidence across magnifications. Extensive experiments demonstrate state-of-the-art performance of PathScale-R1 on cross-scale reasoning tasks and effective transfer to conventional single-scale pathology VQA. Our code is available at https://github.com/iMVR-PL/PathScale-R1.
Chinese Translation
病理诊断本质上是多尺度的,需要将低倍放大下的全局组织结构与高倍放大下的细胞形态进行整合。然而,现有的病理基准和视觉语言模型(VLMs)仍然主要是在单尺度设置下开发的,这限制了它们学习临床上有意义的多倍放大推理的能力。此外,简单构建的视觉问答(VQA)任务可能容易受到仅依赖文本或表面视觉捷径的影响,从而导致对视觉理解的不可靠评估。为了解决这些局限性,我们引入了一个抗捷径的跨尺度病理推理基准和训练框架。我们设计了一种针对语义推理问题的对抗性仅文本筛选策略和一种针对视觉定位问题的结构控制干扰样本采样策略,鼓励模型依赖于跨尺度的视觉证据。在此基础上,我们构建了PathScale-VQA,这是一个高质量的跨尺度病理VQA基准,包含10,373个多项选择问题,基于1,368条跨多个放大级别的诊断路径。基于语义推理集,PathScale-R1通过基于难度的推理蒸馏监督微调进行优化,随后进行带有尺度感知推理结构奖励的强化学习,鼓励在不同放大倍数间使用证据。大量实验表明,PathScale-R1在跨尺度推理任务上表现出最先进的性能,并有效转移到传统的单尺度病理VQA。我们的代码可在 https://github.com/iMVR-PL/PathScale-R1 获取。
cs.CV / 103 / 2607.23803

Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms

超越外观:一种多线索框架及其在移动空地平台上行人关联与跟踪的大规模基准
Wu, Ruiqi, Jiao, Bingliang, Han, Ruize, Yu, Hangzheng, Jiang, Xunkai, Wang, Shining, Hu, Yuanqi, Wang, Wenxuan, Wang, Peng
Abstract
Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory reconstruction in multi-platform cooperative perception. Unlike conventional multiple object tracking, MvMoAT faces frequent viewpoint shifts that distort appearance and undermine cross-view association and temporal tracking. We propose FUSION, a viewpoint-robust Feature Unification framework for multi-view aSsociation and IdentificatiON. Its Multi-cue Adaptive Combination (MAC) module adaptively integrates viewpoint-invariant cues with appearance features to improve cross-view association, while Online Multi-view Feature Synchronization (OMFS) aggregates pedestrian features across historical and cross-view frames for temporally consistent tracking. We also introduce RealMvMoAT, a large-scale benchmark featuring substantial inter- and intra-camera viewpoint variation. It contains 504.9K frames from 7 cameras (5 UAV and 2 ground views) across 10 scenes, with over 7.3M identity-labeled bounding boxes. All cameras exhibit random and substantial motion. To the best of our knowledge, RealMvMoAT is the largest MvMoAT dataset to date. Its scale, viewpoint diversity, complex platform motion, and realistic trajectories provide a comprehensive resource for future research. Experiments on RealMvMoAT and six public benchmarks show that FUSION achieves state-of-the-art performance.
Chinese Translation
多视角多目标关联与跟踪(MvMoAT)在多个摄像头视角之间关联对象并随时间跟踪它们,支持多平台协作感知中的身份持久性和法医轨迹重建。与传统的多目标跟踪不同,MvMoAT面临频繁的视角变化,这会扭曲外观并削弱跨视角关联和时间跟踪。我们提出了FUSION,一种视角鲁棒的特征统一框架,用于多视角关联和识别。其多线索自适应组合(MAC)模块自适应地将视角不变线索与外观特征整合,以改善跨视角关联,而在线多视角特征同步(OMFS)则在历史和跨视角帧中聚合行人特征,以实现时间一致的跟踪。我们还引入了RealMvMoAT,这是一个大规模基准,具有显著的摄像头间和摄像头内视角变化。它包含来自7个摄像头(5个无人机视角和2个地面视角)的504.9K帧,涵盖10个场景,超过7.3M个带身份标签的边界框。所有摄像头均表现出随机且显著的运动。据我们所知,RealMvMoAT是迄今为止最大的MvMoAT数据集。其规模、视角多样性、复杂的平台运动和真实的轨迹为未来的研究提供了全面的资源。在RealMvMoAT和六个公共基准上的实验表明,FUSION达到了最先进的性能。
cs.CV / 104 / 2607.23835

Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations

一致的证据,稳健的识别:几何变换下的忠实归因正则化
Jiao, Xianghao, Chen, Ruoyu, Wang, Wei, Hu, Jiazi, Liang, Jiawei, Sun, Shangquan, Liu, Shiming, Zhang, Qunli, Cao, Xiaochun
Abstract
Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains underexplored. Attribution inconsistency under label-preserving geometric transformations may indicate transformation-sensitive evidence reliance, motivating attribution regularization. However, such supervision is valid only when attribution faithfully reflects the evidence driving predictions. Existing self-supervised methods typically align gradient-based maps such as Grad-CAM, whose limited faithfulness means that attribution consistency need not imply consistency of the underlying decision process, leaving transformation robustness unresolved. We propose an annotation-free attribution regularization framework based on submodular search over image regions. By measuring how candidate subsets affect model outputs, the search extracts compact, class-discriminative evidence as search-derived supervision. We further introduce a submodular ranking loss with path-consistency and termination-alignment terms that respectively align spatially corresponding candidate rankings along paired search trajectories and encourage the transformed trajectory to satisfy the stopping criterion at the target terminal step. The loss provides a differentiable surrogate for regularizing both final attributions and the otherwise discrete evidence-selection process. Experiments on ImageNet-100 show that our method substantially improves attribution stability, Insertion, and Deletion on ViT-B/16 with only a 0.28-point accuracy drop, with similar gains on ViT-L/16. On ImageNet-1K, it improves transformed-input accuracy on ResNet-50 and ConvNeXt-B while limiting the clean-accuracy drop to 0.30 points, demonstrating more consistent evidence reliance with minimal performance loss. Code will be released soon.
Chinese Translation
归因方法广泛用于表征模型预测背后的证据,但其改善模型行为的潜力仍未得到充分探索。在保持标签的几何变换下,归因不一致性可能表明对变换敏感的证据依赖,从而激发了归因正则化的需求。然而,这种监督只有在归因忠实地反映推动预测的证据时才是有效的。现有的自监督方法通常对基于梯度的图(如 Grad-CAM)进行对齐,但其有限的忠实性意味着归因一致性不一定意味着潜在决策过程的一致性,从而使变换鲁棒性问题未得到解决。我们提出了一种基于图像区域的子模搜索的无注释归因正则化框架。通过测量候选子集对模型输出的影响,该搜索提取紧凑的、类别区分的证据作为搜索导出的监督。我们进一步引入了一种具有路径一致性和终止对齐项的子模排序损失,分别对齐配对搜索轨迹上的空间对应候选排序,并鼓励变换后的轨迹在目标终止步骤满足停止标准。该损失为正则化最终归因和原本离散的证据选择过程提供了可微的替代方案。在 ImageNet-100 上的实验表明,我们的方法显著提高了 ViT-B/16 的归因稳定性、插入和删除,仅以 0.28 分的准确率下降为代价,在 ViT-L/16 上也取得了类似的提升。在 ImageNet-1K 上,它提高了 ResNet-50 和 ConvNeXt-B 的变换输入准确率,同时将干净准确率的下降限制在 0.30 分,展示了在最小性能损失下更一致的证据依赖。代码将很快发布。
cs.CV / 105 / 2607.23840

STEER: Steerable Dyadic Head Avatars

STEER:可操控的双人头部虚拟形象
Teotia, Kartik, Rhodin, Helge, Kim, Hyeongwoo, Habermann, Marc, Theobalt, Christian
Abstract
Facial movement and expression are central to face-to-face communication, conveying turn-taking, attention, agreement, and engagement alongside speech. While speech-driven facial animation has made strong progress in lip synchronization and audio-conditioned motion generation, most methods treat conversational behavior as an emergent byproduct of audio, or expose only coarse sequence-level affect control. As a result, key non-verbal channels such as gaze contact and aversion, rhythmic head motion, and emotion remain difficult to explicitly control. We present STEER, a controllable 3D dyadic motion prior for reactive conversational head avatars. STEER factorizes conversational behavior into explicit controls for gaze, head rhythm, and emotion, allowing users to steer how an avatar listens, reacts, and engages with a conversation partner. Since temporally aligned annotations for these behaviors are not available in public dyadic corpora, we introduce a tracking and annotation pipeline that recovers behavioral pseudo-labels from in-the-wild dyadic video. A causal flow-matching transformer then learns partner-aware target motion conditioned on audio, partner motion, emotion and the proposed behavioral controls. We further embed STEER in a photorealistic avatar pipeline by extending a Universal Gaussian Head-Avatar Prior with a learned mapping from tracked parametric motion into its avatar-driving space. This enables controllable animation of high-fidelity Gaussian head avatars without re-training the underlying avatar model. STEER outperforms recent dyadic motion baselines on motion quality, dynamics, and diversity, remains competitive on partner coupling, and enables gaze, head-rhythm, and emotion edits together with an interactive live deployment. We make our code and dataset annotations available at our webpage.
Chinese Translation
面部运动和表情在面对面交流中至关重要,传达轮流发言、注意力、同意和参与感等信息,同时伴随语言表达。尽管基于语音的面部动画在唇部同步和音频条件下的运动生成方面取得了显著进展,但大多数方法将对话行为视为音频的自发副产品,或者仅暴露出粗略的序列级情感控制。因此,诸如目光接触和回避、节奏性头部运动以及情感等关键非语言通道仍然难以明确控制。我们提出了STEER,一种用于反应性对话头部虚拟形象的可控3D双人运动先验。STEER将对话行为分解为目光、头部节奏和情感的显式控制,使用户能够引导虚拟形象如何倾听、反应和与对话伙伴互动。由于公共双人语料库中缺乏这些行为的时间对齐标注,我们引入了一种跟踪和标注流程,从自然场景中的双人视频中恢复行为伪标签。然后,一个因果流匹配变换器学习基于音频、伙伴运动、情感和所提出的行为控制的伙伴感知目标运动。我们进一步将STEER嵌入到一个照片级真实感的虚拟形象流程中,通过扩展一个通用高斯头部虚拟形象先验,学习从跟踪的参数化运动到其虚拟形象驱动空间的映射。这使得在不重新训练基础虚拟形象模型的情况下,实现高保真高斯头部虚拟形象的可控动画。STEER在运动质量、动态性和多样性方面超越了近期的双人运动基准,在伙伴耦合方面保持竞争力,并实现了目光、头部节奏和情感的编辑,同时支持交互式实时部署。我们将在网页上提供我们的代码和数据集标注。
cs.CV / 106 / 2607.23844

OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion Models

OmniCache:用于扩散模型的多维层次特征缓存
He, Zhaoyuan, Muaz, Muhammad, Qiu, Lili
Abstract
High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative quality but remain expensive at inference time because they repeatedly evaluate attention-heavy denoisers over many sampling steps. We address this inefficiency by exploiting redundancy in intermediate diffusion features rather than changing model weights or retraining. We identify four complementary redundancy sources in image and video generation: intra-frame, inter-frame, motion, and denoising-step redundancy. Based on this analysis, we propose OmniCache, a unified hierarchical caching framework that performs multidimensional feature reuse through Token Cache, Frame Cache, Block Cache, and Layered Cache. Unlike token-merging baselines that average matched features, OmniCache uses similarity matching to select cacheable features, skips redundant computation, and restores positionally consistent cached activations, preserving feature order and spatial-temporal structure. The resulting framework reuses spatial features in temporal layers and temporal features in spatial layers, while Layered Cache captures cross-step redundancy at the model-layer level. Across SD3, SVD-XT, and Latte, OmniCache reduces inference latency by up to 35%, 25%, and 28%, respectively, while maintaining visual fidelity and motion coherence in a training-free setting.
Chinese Translation
高分辨率图像和视频扩散模型,包括SD3、FLUX以及最近的视频扩散变换器,显著提高了生成质量,但在推理时仍然昂贵,因为它们在多个采样步骤中反复评估计算量大的去噪器。我们通过利用中间扩散特征中的冗余性来解决这一低效问题,而不是改变模型权重或重新训练。我们在图像和视频生成中识别出四种互补的冗余来源:帧内冗余、帧间冗余、运动冗余和去噪步骤冗余。基于这一分析,我们提出了OmniCache,一个统一的层次缓存框架,通过Token Cache、Frame Cache、Block Cache和Layered Cache实现多维特征重用。与平均匹配特征的token-merging基线不同,OmniCache使用相似性匹配选择可缓存特征,跳过冗余计算,并恢复位置一致的缓存激活,保持特征顺序和时空结构。最终框架在时间层中重用空间特征,在空间层中重用时间特征,而Layered Cache则在模型层级捕获跨步骤冗余。在SD3、SVD-XT和Latte中,OmniCache分别将推理延迟减少了多达35%、25%和28%,同时在无训练的情况下保持视觉保真度和运动一致性。
cs.CV / 107 / 2607.23861

Head Avatars with Dynamic Explicit Hair

具有动态显式头发的头部虚拟形象
Sklyarova, Vanessa, Chen, Haonan, Kabadayi, Berna, Kirschstein, Tobias, Fan, Zicong, Wang, Xi, Pons-Moll, Gerard, Nießner, Matthias, Pollefeys, Marc, Black, Michael J., Thies, Justus
Abstract
We present DynHair, a novel method for tracking and modeling dynamic hair for human head avatars. From video input, we reconstruct a dynamic head avatar with an explicit strand-based hair representation using structured 3D Gaussian Splatting. In contrast to the face region of human head avatars, which can be modeled with 3D Gaussians that are attached or generated with respect to some expressive 3D head model, hair is particularly challenging as it exhibits dynamic motion effects. Therefore, we present a novel method that models the dynamic deformations of the hair strands using a temporal network that is conditioned on angular velocity and acceleration of the head, as well as relative gravity. Specifically, an LSTM encodes the motion history and modulates per-point strand features via FiLM conditioning which further used by MLP to produce physically plausible displacements to canonical hairstyle. We jointly optimize this motion and appearance representation of the hair, with a 3DGS-based representation of the face-region, via differentiable Gaussian splatting with photometric, geometric, and physics-based supervision. As a result of our method, we retrieve hair tracking of the training video data and an animatable head avatar with controllable hair dynamics. In our experiments, we demonstrate state-of-the-art performance in terms of hair dynamics, temporal consistency, and generalization across subjects.
Chinese Translation
我们提出了一种名为DynHair的新方法,用于跟踪和建模人类头部虚拟形象的动态头发。通过视频输入,我们利用结构化的三维高斯点云重建了一个具有显式基于发丝的头发表示的动态头部虚拟形象。与人类头部虚拟形象的面部区域可以通过附加或生成于某些表现性三维头部模型的三维高斯进行建模不同,头发的建模尤其具有挑战性,因为它表现出动态运动效果。因此,我们提出了一种新方法,通过一个时间网络来建模头发发丝的动态变形,该网络以头部的角速度和加速度以及相对重力为条件。具体而言,长短期记忆网络(LSTM)编码运动历史,并通过FiLM条件调节每个点的发丝特征,这些特征随后由多层感知器(MLP)用于产生物理上合理的位移到标准发型。我们通过可微分的高斯点云与光度、几何和基于物理的监督,联合优化头发的运动和外观表示,以及面部区域的三维高斯基表示。通过我们的方法,我们从训练视频数据中恢复了头发跟踪,并获得了一个可动画的头部虚拟形象,具有可控的头发动态。在我们的实验中,我们展示了在头发动态、时间一致性和跨对象的泛化方面的最先进性能。
cs.CV / 108 / 2607.23883

Long-Tailed Medical Image Classification

长尾医学图像分类
Ren, Nathanael, Arya, Saagar
Abstract
In this paper, we examine the difficulties of using standard techniques for medical image classification due to long-tailed distributions (wherein rarer conditions have very few samples) resulting in bias towards diagnosing common diseases and away from rarer diseases. We then discuss and implement deep learning models with techniques such as augmentation to minimize error, especially from rarer diseases. We evaluate various different models with AP, F1 score, AUROC, and loss (all on the validation set). We conclude with the promising results from our best model, and potential applications in the healthcare space.
Chinese Translation
在本文中,我们探讨了由于长尾分布(即稀有病症样本极少)导致的医学图像分类中使用标准技术的困难,这种情况使得对常见疾病的诊断偏向,而对稀有疾病的诊断则相对不足。随后,我们讨论并实施了深度学习模型,并采用数据增强等技术来最小化错误,特别是来自稀有疾病的错误。我们使用平均精度(AP)、F1分数、受试者工作特征曲线下面积(AUROC)和损失(均在验证集上)评估了多种不同的模型。最后,我们总结了最佳模型的良好结果及其在医疗领域的潜在应用。
cs.CV / 109 / 2607.23908

Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets

基于嵌入的异常检测用于清理全球作物类型参考数据集
Shah, Syed Roshaan Ali, Van Tricht, Kristof, Butsko, Christina, Degerickx, Jeroen, Szantoi, Zoltan
Abstract
High quality reference data remain a critical bottleneck for crop-type mapping at any spatial and temporal scale. Operational systems such as WorldCereal aggregate labels from heterogeneous sources such as parcel registers, national databases, field surveys, and map-derived products, each with their own biases, coverage gaps and unknown label noise. Simple global rules are inadequate, since crop phenology and observation conditions vary strongly across regions and seasons. In this study, we focus on a single, operationally relevant question: whether embeddings produced through geospatial foundation models are a viable basis for cleaning the reference data. We propose a practical, locality-aware, embedding-based anomaly (EBA) detection framework that operates on the embeddings of a pretrained Earth-observation encoder. We score each labelled sample against other samples of the same crop in the same area using a pretrained embedding, flag the ones that stand out, and test whether removing or down-weighting them before training yields a better model. We establish that the flagged points are genuinely mislabelled or misplaced in two independent ways: against synthetic ground truth, the detector concentrates injected label errors 2.5-5x above chance in its flagged set (detection AUROC up to 0.84); and on real data, a model-independent test shows that removing or confidence-weighting the flagged held-out points raises measured accuracy in trained models, for both crop type and land cover. Acting on the flags then improves the WorldCereal crop-type model across five macro-regions, evaluated on a fixed held-out split under three views. We find conservative cleaning helps while over-cleaning hurts. The EBA detector approach is designed to be reproducible and extensible, and can serve as a template for cleaning large, noisy Earth observation reference datasets beyond crop mapping.
Chinese Translation
高质量的参考数据仍然是任何空间和时间尺度上作物类型制图的关键瓶颈。诸如WorldCereal等操作系统从异构来源汇总标签,这些来源包括地块登记册、国家数据库、田野调查和地图衍生产品,每个来源都有其自身的偏差、覆盖缺口和未知标签噪声。简单的全球规则不足以应对,因为作物的物候和观测条件在不同地区和季节之间差异显著。在本研究中,我们专注于一个单一的、与操作相关的问题:通过地理空间基础模型生成的嵌入是否可以作为清理参考数据的可行基础。我们提出了一种实用的、具有地方意识的基于嵌入的异常(EBA)检测框架,该框架在预训练的地球观测编码器的嵌入上运行。我们使用预训练的嵌入对同一地区的同一作物的每个标记样本进行评分,标记出那些突出的样本,并测试在训练之前去除或降低它们的权重是否能产生更好的模型。我们通过两种独立的方法确认被标记的点确实是错误标记或错误放置的:在合成真实值的对比中,检测器在其标记集中的注入标签错误浓度是随机情况的2.5-5倍(检测AUROC高达0.84);在真实数据上,模型无关的测试显示,去除或信心加权被标记的保留点提高了训练模型的测量准确性,无论是作物类型还是土地覆盖。对标记进行处理后,WorldCereal作物类型模型在五个宏观区域的表现得到了改善,这些区域在三个视角下进行了固定的保留分割评估。我们发现,保守的清理有帮助,而过度清理则有害。EBA检测器方法旨在可重复和可扩展,并可以作为清理大规模、嘈杂的地球观测参考数据集的模板,超越作物制图的范围。
cs.CV / 110 / 2607.23910

SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception

SimBEV2X:用于多任务车对一切协同感知的大规模数据集和数据生成工具
Mehr, Goodarz, Gohari, Sepideh, Abbas, Montasir, Eskandarian, Azim
Abstract
Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range. However, the development of robust V2X algorithms, particularly those relying on unified spatial representations like bird's-eye view (BEV) representation, is hampered by the lack of large-scale, multi-modal, multi-task datasets. Moreover, collecting and annotating a large set of synchronized, real-world multi-agent data is prohibitively expensive. This has resulted in a landscape where existing V2X datasets are notably limited in both size and scope. To overcome this, we introduce SimBEV2X, an advanced synthetic data generation tool built on the CARLA simulator. SimBEV2X automatically creates randomized driving scenarios to collect multi-modal sensor data alongside various types of ground truth including 3D bounding boxes with unique track IDs, HD map information, BEV segmentation maps, and semantic occupancy voxel grids from both vehicles and RSUs. We also present the SimBEV2X dataset, the largest V2X perception dataset to date. The dataset comprises 258 scenes, each involving up to 8 connected vehicles and up to 4 RSUs across a variety of road networks. The SimBEV2X dataset is an order of magnitude larger than existing V2X datasets and contains 102,200 frames, 588,520 lidar point clouds, more than 3 million images, over 27 million bounding boxes, and a comprehensive set of other annotations. Finally, we establish a strong baseline on the SimBEV2X dataset using CoopDet3D and propose CoBEVFusion, a novel architecture that combines CoopDet3D with fused axial attention (FAX) for context-aware multi-agent feature aggregation, resulting in superior performance. SimBEV2X, the SimBEV2X dataset, and CoBEVFusion are available at https://simbev2x.org and https://github.com/GoodarzMehr/SimBEV2X.
Chinese Translation
通过车对一切(V2X)通信实现的协同感知可以克服单个自主车辆固有的物理限制,例如遮挡和传感器范围有限。然而,开发强健的V2X算法,特别是那些依赖于统一空间表示(如鸟瞰视图(BEV)表示)的算法,受到缺乏大规模、多模态、多任务数据集的制约。此外,收集和标注大量同步的真实世界多智能体数据的成本极高。这导致现有的V2X数据集在规模和范围上都显得相当有限。为了解决这个问题,我们推出了SimBEV2X,这是一个基于CARLA模拟器的先进合成数据生成工具。SimBEV2X自动创建随机驾驶场景,以收集多模态传感器数据以及各种类型的真实数据,包括带有唯一轨迹ID的3D边界框、高精度地图信息、BEV分割图和来自车辆及路边单元(RSUs)的语义占用体素网格。我们还提出了SimBEV2X数据集,这是迄今为止最大的V2X感知数据集。该数据集包含258个场景,每个场景涉及多达8辆连接车辆和多达4个RSUs,涵盖多种道路网络。SimBEV2X数据集的规模比现有的V2X数据集大一个数量级,包含102,200帧、588,520个激光雷达点云、超过300万张图像、超过2700万个边界框以及一整套其他注释。最后,我们在SimBEV2X数据集上建立了一个强基线,使用了CoopDet3D,并提出了CoBEVFusion,这是一种将CoopDet3D与融合轴向注意力(FAX)结合的新架构,用于上下文感知的多智能体特征聚合,从而实现了卓越的性能。SimBEV2X、SimBEV2X数据集和CoBEVFusion可在https://simbev2x.org和https://github.com/GoodarzMehr/SimBEV2X获取。
cs.CV / 111 / 2607.23917

Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention

注视到文本生成:超越人类注意力的类别解码
Mondal, Sounak, Samaras, Dimitris, Zelinsky, Gregory, Hoai, Minh
Abstract
We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as a discriminative task over predefined categories, we formulate it as a generative learning problem: training a model to produce free-form descriptions that capture the rich nuances and open-ended nature of human intentions beyond fixed labels. To this end, we introduce Gazette, the first gaze-to-text decoding framework. Based on multimodal large language models (MLLMs), Gazette learns to decode gaze scanpaths into natural language for goals that may extend beyond categorical labels and require articulation in natural language. To help Gazette filter out individual differences in gaze behavior and learn the goal-specific spatiotemporal dynamics crucial for generating accurate natural language goal descriptions, we propose a novel strategy that leverages the encyclopedic knowledge and reasoning abilities of a large language model to synthesize natural language explanations of goal-directed attentional behavior called think-aloud transcripts. Instruction tuning on these synthetic narratives allows Gazette to achieve state-of-the-art performance in gaze decoding across multiple tasks, demonstrating its generalizability and versatility, thereby enabling gaze to serve as a powerful, non-intrusive cue for inferring human goals and intentions in diverse scenarios.
Chinese Translation
我们提出了一种新颖的学习问题:将注视解码为自然语言描述的人类目标,涵盖多种视觉任务。与之前的研究将注视解码视为在预定义类别上的判别任务不同,我们将其表述为生成学习问题:训练一个模型生成自由形式的描述,以捕捉人类意图的丰富细微差别和开放性特征,而不仅限于固定标签。为此,我们推出了Gazette,这是第一个注视到文本的解码框架。Gazette基于多模态大语言模型(MLLMs),学习将注视扫描路径解码为自然语言,以描述可能超出类别标签并需要用自然语言表达的目标。为了帮助Gazette过滤个体在注视行为上的差异,并学习对生成准确的自然语言目标描述至关重要的目标特定时空动态,我们提出了一种新策略,利用大语言模型的百科知识和推理能力,合成目标导向注意行为的自然语言解释,称为“思考大声转录”(think-aloud transcripts)。对这些合成叙述进行指令调优使得Gazette在多个任务中实现了最先进的注视解码性能,展示了其可推广性和多样性,从而使注视成为推断人类目标和意图的强大且非侵入性的线索,适用于多种场景。
cs.CV / 112 / 2607.23920

What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape

我能编辑什么?开放式策略发现与情感可编辑性景观
Li, Qing, Dong, Zeyu, Cui, Yin, Yan, Chuan, Peng, Xiaojiang
Abstract
Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a particular image can afford for a target emotion. Existing affective image manipulation methods, including recent agentic variants, largely operate within bounded strategy spaces based on predefined factor taxonomies, knowledge libraries, or conventional editing templates, and therefore often miss image-specific, context-grounded strategies. We introduce EmoScope, a multi-agent framework that reframes the task from "how should I edit?" to "what can I edit?" EmoScope first discovers an image-specific editable space through emotion-conditioned affordance reasoning, then uses a semantic hierarchy of anchors, variables, and context to balance content consistency and emotional expressiveness before executing and verifying the edit. Because its plans are expressed as image-specific affordances rather than retrieved templates, EmoScope also exposes the editing strategy as an interactive surface for user refinement at the plan level. In a large-scale human evaluation covering all eight Mikels emotion categories, with 4,693 valid responses across 1,824 pairwise questions, participants preferred EmoScope over two competitive baselines by 88.1% on average. Attribution analysis further shows that EmoScope selects target-emotion-adaptive strategies rather than applying a uniform template. The same affordance-level plan also supports lightweight user refinement in an interactive pilot. Finally, we show that classifier-based metrics exhibit emotion-conditional blind spots toward non-stereotypical, context-grounded edits, and present a relative content-emotion preference-affinity landscape showing that EmoScope's advantage varies systematically across image-emotion combinations.
Chinese Translation
情感图像编辑不仅仅是应用情感滤镜或修改预定义的视觉因素:有效的编辑必须识别特定图像在目标情感下的可编辑性。现有的情感图像处理方法,包括最近的代理变体,主要在基于预定义因素分类法、知识库或传统编辑模板的有限策略空间内运作,因此往往忽视了图像特定的、基于上下文的策略。我们提出了EmoScope,一个多代理框架,将任务从“我应该如何编辑?”重新定义为“我可以编辑什么?”EmoScope首先通过情感条件的可供性推理发现图像特定的可编辑空间,然后利用锚点、变量和上下文的语义层次结构,在执行和验证编辑之前平衡内容一致性和情感表现力。由于其计划以图像特定的可供性而非检索模板的形式表达,EmoScope还将编辑策略呈现为用户在计划层面进行细化的交互界面。在覆盖所有八个Mikels情感类别的大规模人类评估中,参与者在1,824对问题中提供了4,693个有效响应,平均偏好EmoScope超过两个竞争基线的比例为88.1%。归因分析进一步表明,EmoScope选择适应目标情感的策略,而不是应用统一的模板。相同的可供性层级计划还支持在交互式试点中进行轻量级用户细化。最后,我们展示了基于分类器的指标在非刻板、基于上下文的编辑方面存在情感条件盲点,并呈现了一个相对内容-情感偏好亲和力的景观,显示EmoScope的优势在不同的图像-情感组合中系统性变化。
cs.CV / 113 / 2607.23921

DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering

DDVT:用于视觉问答中答案定位的动态双层视觉变换器融合网络
Zhang, Yue, Li, Xiangyu, Fan, Wanshu, Yang, Xin, Zhou, Dongsheng
Abstract
Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propose a question-guided dynamic regional-level module (QGDR) that combines complementary image context through ROI Align and text content, enabling precise localization of text-related visual content. Moreover, we present a cross-modal multi-scale aggregation module (CMA) that enhances feature fusion between pixel-level and region-level features, facilitating the effective localization of visual content associated with grounded answers. Furthermore, we fuse the located visual content with text features to locate the region and provide answers to questions posed about the image. Experimental results demonstrate that our DDVT outperforms state-of-the-art methods on several widely-used benchmarks.
Chinese Translation
视觉问答中的答案定位旨在从给定的自然语言问题中定位与图像视觉内容相关的区域,这一研究因其实际应用而备受关注。本文介绍了一种用于视觉问答中答案定位的动态双层视觉变换器融合网络(DDVT)。具体而言,我们提出了一种基于问题引导的动态区域级模块(QGDR),该模块通过ROI Align结合互补的图像上下文和文本内容,从而实现文本相关视觉内容的精确定位。此外,我们还提出了一种跨模态多尺度聚合模块(CMA),该模块增强了像素级和区域级特征之间的特征融合,促进了与定位答案相关的视觉内容的有效定位。此外,我们将定位的视觉内容与文本特征融合,以定位区域并提供关于图像的问题的答案。实验结果表明,我们的DDVT在多个广泛使用的基准测试中优于最先进的方法。
cs.CV / 114 / 2607.23924

DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection

DuoAD:利用 [CLS] 双重特性实现无训练的少样本异常检测
Tang, Jyun-Ze, Huang, Po-Han, Chang, Ming-Ching, Hsu, Chih-Fan, Li, Jeng-Lin
Abstract
Vision foundation models have enabled strong training-free anomaly detection (AD). However, most existing approaches rely primarily on independent local patch features, leaving the global contextual information encoded by Vision Transformers (ViTs) underexploited. In this work, we identify the dual characteristics of the ViT [CLS] token: its embedding provides anomaly-invariant global semantic representation, while its attention maps implicitly highlight spatially abnormal regions. Building on this observation, we propose a fully automated AD framework leveraging global context to remove manual tunings. Our framework introduces (1) an automatic augmentation selection strategy driven by [CLS]-level semantic consistency, and (2) an attention-guided feature reweighting mechanism that dynamically adjusts patch contributions according to [CLS] attention saliency. By integrating these components over multi-level features, our method achieves stable anomaly scoring and precise localization without training or parameter tuning. Under the one-shot setting, it achieves Image-AUC scores of 97.7%, 93.2%, and 84.5% on MVTec-AD, VisA, and Real-IAD. Using a single fixed configuration across categories, backbones, and datasets, the method establishes a new state-of-the-art for plug-and-play, training-free anomaly detection while maintaining strong robustness and practical scalability.
Chinese Translation
视觉基础模型使得强大的无训练异常检测(AD)成为可能。然而,大多数现有方法主要依赖于独立的局部补丁特征,导致视觉变换器(ViTs)所编码的全局上下文信息未得到充分利用。在本研究中,我们识别了ViT [CLS] 令牌的双重特性:其嵌入提供了异常不变的全局语义表示,而其注意力图则隐含地突出显示了空间异常区域。基于这一观察,我们提出了一种完全自动化的AD框架,利用全局上下文来消除手动调节。我们的框架引入了(1)一种由[CLS]级别语义一致性驱动的自动增强选择策略,以及(2)一种注意力引导的特征重加权机制,动态调整补丁的贡献,依据[CLS]注意力显著性。通过在多层特征上整合这些组件,我们的方法在无训练或参数调节的情况下,实现了稳定的异常评分和精确的定位。在一次性设置下,它在MVTec-AD、VisA和Real-IAD上分别达到了97.7%、93.2%和84.5%的Image-AUC分数。该方法在类别、骨干网络和数据集之间使用单一固定配置,建立了即插即用、无训练异常检测的新最先进水平,同时保持了强大的鲁棒性和实用的可扩展性。
cs.CV / 115 / 2607.23937

Manifold-Constrained Noise Optimization for Diverse Diffusion Sampling

多样性扩散采样的流形约束噪声优化
Shi, Qitan, Jin, Cheng, Liu, Ziyuan, Gu, Yuantao
Abstract
Few-step distilled diffusion models generate high-quality images quickly, but often lose per-prompt diversity, producing near-identical samples across random seeds. Optimizing the initial noise at inference time offers an appealing way to recover this diversity, yet existing methods directly update the initial noise in an unconstrained Euclidean space, ignoring both the geometry of the Gaussian prior and the model's sensitivity to noise frequencies. They therefore introduce auxiliary quality-control objectives to maintain generation fidelity, adding compute and weighting hyperparameters while still requiring conservative updates to prevent degradation. In this work, we propose MoNO, a training-free method that performs Manifold-constrained Noise Optimization on a low-dimensional, quality-stabilizing noise manifold. MoNO sequentially optimizes each new initial noise so that its predicted visual feature complements previous generations, while Riemannian updates on an affine low-frequency sphere preserve prior likelihood and fix unstable high-frequency components by construction. This enables large geodesic steps, removes the need for auxiliary quality-control objectives, and converges in far fewer iterations than prior noise-optimization methods. Experiments with multiple distilled text-to-image diffusion models show that MoNO consistently improves per-prompt diversity while maintaining image quality.
Chinese Translation
少步提炼的扩散模型能够快速生成高质量图像,但往往会失去每个提示的多样性,在随机种子下产生几乎相同的样本。在推理时优化初始噪声提供了一种恢复这种多样性的诱人方式,但现有方法直接在无约束的欧几里得空间中更新初始噪声,忽视了高斯先验的几何特性和模型对噪声频率的敏感性。因此,它们引入了辅助质量控制目标以维持生成的保真度,增加了计算和权重超参数,同时仍需要保守的更新以防止退化。在本研究中,我们提出了 MoNO,一种无训练的方法,在低维、质量稳定的噪声流形上执行流形约束噪声优化。MoNO 依次优化每个新的初始噪声,使其预测的视觉特征与先前生成的内容互补,同时在仿射低频球面上的黎曼更新通过构造保持先验似然并修复不稳定的高频成分。这使得大地质步长成为可能,消除了对辅助质量控制目标的需求,并且在迭代次数上远少于先前的噪声优化方法。对多个提炼的文本到图像扩散模型的实验表明,MoNO 在保持图像质量的同时,持续改善每个提示的多样性。
cs.CV / 116 / 2607.23951

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

TimePLE:重新思考视频时间定位的时间表示
Zeng, Yuhui, Mao, Xinyu, Liu, Xiaokun, Tao, Xin, Huang, Jinfa, Ji, Jiayi, Zheng, Xiawu
Abstract
Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this interval indirectly through two endpoint outputs, represented either as discrete timestamp tokens or continuous boundary coordinates. These formulations differ in how endpoints are encoded, but not in what is predicted: the event interval remains a derived object, while interval validity, duration, and interval-level similarity are handled only implicitly. We propose TimePLE, which reformulates VTG from endpoint prediction to interval-native grounding by predicting a single joint distribution over valid temporal intervals. TimePLE maps each interval to a point in a canonical position-duration square, where every support point corresponds to a valid span and neighboring points represent geometrically similar intervals. Given a video and query, the VLM generates a single latent <|TIMESPAN|> token whose hidden state is decoded into a joint interval distribution, refined through duration-aware coordinate correction, and converted into continuous boundaries. The same interval representation is used to encode input temporal anchors, aligning video-side temporal evidence with output-side span prediction. To reliably align the latent span representation with complete event intervals, we curate 90K-scale grounded samples and human-verify 3K-scale benchmark annotations. Experiments across four VTG benchmarks show that TimePLE consistently outperforms endpoint prediction baselines, achieving an average mIoU of 58.9, with clear gains on short-duration and medium-duration events.
Chinese Translation
视频时间定位(VTG)旨在根据自然语言查询定位连续的视频时间区间。然而,当前基于视觉语言模型(VLM)的方法通常通过两个端点输出间接产生该区间,这些端点要么表示为离散的时间戳标记,要么表示为连续的边界坐标。这些表述在端点编码方式上有所不同,但在预测内容上并无区别:事件区间仍然是一个派生对象,而区间的有效性、持续时间和区间级相似性仅被隐式处理。我们提出了TimePLE,它将VTG从端点预测重新构建为区间本地化,通过预测一个有效时间区间的单一联合分布。TimePLE将每个区间映射到一个规范的位置-持续时间方形中,其中每个支持点对应一个有效的跨度,邻近点则表示几何上相似的区间。给定视频和查询,VLM生成一个单一的潜在<|TIMESPAN|>标记,其隐藏状态被解码为联合区间分布,通过考虑持续时间的坐标校正进行优化,并转换为连续边界。相同的区间表示用于编码输入时间锚点,将视频侧的时间证据与输出侧的跨度预测对齐。为了可靠地将潜在跨度表示与完整的事件区间对齐,我们策划了90K规模的基础样本,并对3K规模的基准注释进行了人工验证。在四个VTG基准测试中的实验表明,TimePLE始终优于端点预测基线,平均mIoU达到58.9,并在短时和中等时事件上取得明显提升。
cs.CV / 117 / 2607.23958

RODR: Riemannian Orthogonally Decoupled Regularization for Disentangled Manifold Representation

RODR:用于解耦流形表示的黎曼正交解耦正则化
Zhu, Jiayu, Zhao, Wenlai
Abstract
Point cloud denoising is essentially a geometric recovery task that aims to reconstruct the intrinsic structure of a smooth 2D Riemannian manifold embedded in R^3 from noisy, discrete ambient-space samples. Despite the remarkable progress of modern manifold-aware encoders and generative transport models in geometric representation learning, a fundamental objective-geometry mismatch remains underexplored. Theoretically, we identified that this mismatched coupling leads to geometric gradient interference, where conflicting optimization objectives result in structural degradation and point clustering. We introduce Riemannian Orthogonally Decoupled Regularization (RODR) to reformulate the optimization trajectory by disentangling the normal (fitting) and tangential (distribution) components. Guided by a vector-attention and entropy-aware adaptive strategy, RODR effectively preserves high-fidelity geometric details while maintaining sampling uniformity. Experiments demonstrate that RODR reaches performance comparable to state-of-the-art baselines and suggests improved distribution regularity and reduced local aggregation effectively. Our work establishes a generic and interpretable framework for disentangled geometric optimization in point cloud processing.
Chinese Translation
点云去噪本质上是一项几何恢复任务,旨在从噪声干扰的离散环境空间样本中重建嵌入在 R^3 中的光滑二维黎曼流形的内在结构。尽管现代流形感知编码器和生成传输模型在几何表示学习方面取得了显著进展,但一个基本的目标-几何不匹配问题仍然未得到充分探索。从理论上讲,我们发现这种不匹配的耦合导致了几何梯度干扰,其中相互冲突的优化目标导致结构退化和点聚集。我们引入了黎曼正交解耦正则化(RODR),通过解耦法线(拟合)和切向(分布)分量来重新制定优化轨迹。在向量注意力和熵感知自适应策略的指导下,RODR 有效地保留了高保真几何细节,同时保持了采样均匀性。实验表明,RODR 的性能可与最先进的基线相媲美,并有效地改善了分布的规则性和减少了局部聚集。我们的工作为点云处理中的解耦几何优化建立了一个通用且可解释的框架。
cs.CV / 118 / 2607.23962

Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation

基于视觉-语言模型的全球导航卫星系统欺骗检测在自主车辆导航中的应用研究
Aldeen, Mohammed, Irfan, Muhammad Sami, Dasgupta, Sagar, Cheng, Long, Rahman, Mizanur, Chowdhury, Mashrur
Abstract
Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attacks that can covertly redirect vehicles or induce unsafe maneuvers. In this paper, we develop the first Vision-Language Model (VLM)-based framework for GNSS spoofing detection for autonomous vehicles by fusing front-camera visual data with in-vehicle sensor readings (e.g., speed, acceleration, yaw rate) against GNSS-derived maneuvers. Our approach introduces a three-stage fine-tuning process that first grounds visual cues, and then calibrates sensor data within a shared semantic space to detect discrepancies between predicted and GNSS-derived maneuvers across three attack scenarios. We also generated an independent real-world dataset by driving an instrumented vehicle on public roads in Tuscaloosa, Alabama, equipped with time-synchronized GNSS, IMU, and camera logs to validate cross-regional generalization of our fine-tuned model on unseen data from training data. On this dataset, we then generated intelligent spoofing attacks, including trajectory mirroring with road-network snapping for wrong-turn attacks, position freezing for overshoot scenarios, and drift generation for stop attacks. On this validation dataset, the zero-shot VLMs baseline F1-score ranges from 23% to 32%, whereas our fine-tuned model achieves an F1-score ranging from 94% to 95%. Results show that our VLM-based approach correctly classified every wrong-turn and stop attacks, and attains 88%-93% accuracy for overshoot attacks. Furthermore, we introduce an adaptive inference policy that reduces VLM invocations to 14% (~86% computational reduction) and yields 65ms-73ms per 4s window. These results point to a practical, on-road layer of defense that complements signal-level integrity checks with the use of VLMs.
Chinese Translation
自主车辆(AVs)依赖全球导航卫星系统(GNSS)进行定位和导航,这使其易受欺骗攻击的影响,这些攻击可以秘密地重定向车辆或引发不安全的操作。在本文中,我们开发了首个基于视觉-语言模型(VLM)的GNSS欺骗检测框架,通过将前置摄像头的视觉数据与车载传感器读数(例如速度、加速度、偏航率)结合,针对GNSS派生的操作进行检测。我们的方法引入了一个三阶段的微调过程,首先将视觉线索与实际情况相结合,然后在共享的语义空间内校准传感器数据,以检测在三种攻击场景下预测的操作与GNSS派生操作之间的差异。我们还通过在阿拉巴马州塔斯卡卢萨的公共道路上驾驶一辆配备了时间同步的GNSS、IMU和摄像头日志的仪器化车辆,生成了一个独立的真实世界数据集,以验证我们微调模型在未见数据上的跨区域泛化能力。在该数据集上,我们生成了智能欺骗攻击,包括通过道路网络快照进行错误转弯攻击的轨迹镜像、用于超速场景的位置冻结,以及用于停车攻击的漂移生成。在该验证数据集上,零-shot VLM基线的F1分数范围为23%到32%,而我们的微调模型则达到了94%到95%的F1分数。结果表明,我们基于VLM的方法正确分类了每一次错误转弯和停车攻击,并在超速攻击中达到了88%-93%的准确率。此外,我们引入了一种自适应推理策略,将VLM调用减少到14%(约86%的计算减少),并在每4秒的窗口内实现了65毫秒到73毫秒的响应时间。这些结果表明,基于VLM的方法为信号级完整性检查提供了一个实用的、路面防御层。
cs.CV / 119 / 2607.23972

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

彩色眼底摄影分析:数据、预处理和建模的共同演进走向多模态人工智能
Li, Yu, He, Wengan, Xu, Wenhui, Jiang, Lihong, Xiao, Fan, Huang, Zhuohang, Liang, Yuanzhu, Liu, Jiayi, Chen, Yuxi, Luo, Yongsheng
Abstract
Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence. This review provides an integrated overview of CFP AI through the interplay of dataset evolution, preprocessing paradigms, and modeling frameworks. We show that CFP datasets have evolved from small single-center collections with task-specific labels to large multi-center resources featuring multimodal pairings and longitudinal clinical records. Preprocessing has progressed from conventional image enhancement to neural data-engineering pipelines, hardware-aware token optimization, and self-supervised imputation for incomplete electronic health records (EHRs). Meanwhile, modeling has advanced from convolutional neural networks (CNNs) to vision foundation models, state space models (SSMs), and multimodal expert architectures. At the multimodal frontier, CFP is increasingly integrated with EHRs and longitudinal patient information, enabling more comprehensive clinical reasoning beyond isolated image analysis. We conclude that future progress depends on the collaborative optimization of datasets, preprocessing, and multimodal modeling, providing a roadmap toward robust clinical deployment, improved cross-domain generalization, and resource-efficient edge intelligence.
Chinese Translation
彩色眼底摄影(Color Fundus Photography, CFP)是一种主要的非侵入性成像方式,用于大规模筛查眼科和系统性疾病。现有的调查主要独立总结了特定任务的算法、数据集或预处理技术,缺乏对它们与现代人工智能共同演进的统一视角。本综述通过数据集演变、预处理范式和建模框架之间的相互作用,提供了 CFP 人工智能的综合概述。我们展示了 CFP 数据集从小型单中心收集、具有特定任务标签,演变为大型多中心资源,具备多模态配对和纵向临床记录。预处理技术已从传统的图像增强进展到神经数据工程管道、硬件感知的令牌优化,以及针对不完整电子健康记录(EHRs)的自监督插补。同时,建模技术也从卷积神经网络(Convolutional Neural Networks, CNNs)发展到视觉基础模型、状态空间模型(State Space Models, SSMs)和多模态专家架构。在多模态前沿,CFP 正在与 EHRs 和纵向患者信息日益集成,使得临床推理超越孤立的图像分析变得更加全面。我们总结认为,未来的进展依赖于数据集、预处理和多模态建模的协同优化,为稳健的临床部署、改善跨域泛化能力和资源高效的边缘智能提供了路线图。
cs.CV / 120 / 2607.23979

Mutual Modality Trust with Lightweight Reconstruction Regularization for Fine-grained Tire Pattern Recognition

轻量重建正则化的互模态信任用于细粒度轮胎花纹识别
Yang, Jianning, Fang, Jie, Ma, Xinda, Song, Zirui, Wang, Dianwei, Wang, Nan
Abstract
Visual tire recognition serves as a core supporting technique for vehicle safety monitoring, autonomous driving perception and automated automotive maintenance. Existing fine-grained tire recognition techniques suffer from three prominent limitations. They tend to depend on only one visual source, lack the capacity to jointly model spatial and frequency cues for minute tread texture extraction, and suffer severe overfitting given limited annotated tire imagery. This paper proposes a lightweight fine-grained tire pattern recognition method incorporating dual-branch independent inference and enhanced feature fusion to boost recognition performance. The framework employs two task-specialized branches dedicated to tire surface and tread indentation, respectively, to extract modality-specific discriminative features. Each branch conducts independent prediction, while cross-branch feature fusion exploits Mutual Modality Trust (M$^2$T) to realize complementary feature enhancement across two modalities. Besides, a frequency-domain hierarchical guidance module is devised, which leverages bandpass filters to decompose feature maps into high- and low-frequency components and enables fine-grained cross-layer feature modulation. Furthermore, a Lightweight Reconstruction Regularization (LR$^2$) is introduced to retain abundant intrinsic information within feature embeddings, substantially improving feature stability and recognition robustness under limited labeled training data. In addition, we establish a surface-indentation multi-source dataset namely MTire299 for fine-grained tire tread recognition, which covers 299 categories with a total of 14795 paired image samples. Extensive experiments conducted on two public tire datasets validate the superiority and efficacy of the proposed algorithm.
Chinese Translation
视觉轮胎识别作为车辆安全监控、自动驾驶感知和自动化汽车维护的核心支持技术,面临三大显著限制。现有的细粒度轮胎识别技术往往仅依赖单一视觉来源,缺乏联合建模空间和频率线索以提取微小的胎面纹理,并且在有限的标注轮胎图像下严重过拟合。本文提出了一种轻量级的细粒度轮胎花纹识别方法,结合双分支独立推理和增强特征融合,以提升识别性能。该框架采用两个专门针对轮胎表面和胎面凹槽的任务分支,分别提取模态特定的判别特征。每个分支进行独立预测,而跨分支特征融合利用互模态信任(Mutual Modality Trust, M$^2$T)实现两种模态间的互补特征增强。此外,设计了一种频域层次引导模块,利用带通滤波器将特征图分解为高频和低频成分,从而实现细粒度的跨层特征调制。此外,引入了轻量重建正则化(Lightweight Reconstruction Regularization, LR$^2$),以保留特征嵌入中的丰富内在信息,显著提高在有限标注训练数据下的特征稳定性和识别鲁棒性。此外,我们建立了一个名为MTire299的表面-凹槽多源数据集,用于细粒度轮胎花纹识别,涵盖299个类别,共有14795对图像样本。在两个公共轮胎数据集上进行的广泛实验验证了所提算法的优越性和有效性。
cs.CV / 121 / 2607.23981

Multimodal Semantic-Probabilistic Objectness for Open World Object Detection

开放世界物体检测的多模态语义-概率物体性
Tian, Weijun, Liu, Rui
Abstract
Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space. However, visual objectness alone cannot determine whether an object-like query corresponds to a hard known instance, an unseen-category object, or background clutter, resulting in an ambiguous known-unknown decision boundary. We propose MSPO, a lightweight semantic calibration framework that augments PROB with task-aware known-category language priors while preserving its detector architecture and incremental learning protocol. For each currently known category, MSPO constructs an extended text description covering category attributes, visual appearance, typical scenes, and functional usage, and encodes it using a frozen CLIP text encoder. Decoder query features are projected into the same semantic space to estimate their support from the current known-category semantics. This semantic evidence is fused with PROB's visual objectness to calibrate known and unknown predictions without turning OWOD into open-vocabulary classification. Importantly, MSPO never uses future-category names, and all unseen categories remain unnamed during evaluation. Experiments on M-OWODB and S-OWODB show that MSPO improves the strong PROB baseline on the main aggregate metrics while retaining competitive unknown recall. It also improves early unknown-confusion metrics and raises PASCAL VOC final mAP by up to 2.7 points. These results demonstrate that known-category language semantics provide an effective calibration signal for probabilistic objectness under the standard OWOD setting.
Chinese Translation
开放世界物体检测(OWOD)要求检测器能够识别已知类别、发现来自未见类别的未命名物体,并逐步学习新标注的类别。PROB通过在解码器查询空间中建模与类别无关的概率物体性来改善未知物体的发现。然而,仅靠视觉物体性无法确定一个物体样的查询是对应于一个已知的困难实例、一个未见类别的物体,还是背景杂乱,从而导致已知与未知之间的决策边界模糊。我们提出了MSPO,一个轻量级的语义校准框架,它在保持PROB的检测器架构和增量学习协议的同时,增强了任务感知的已知类别语言先验。对于每个当前已知类别,MSPO构建了一个扩展的文本描述,涵盖类别属性、视觉外观、典型场景和功能使用,并使用冻结的CLIP文本编码器进行编码。解码器查询特征被投影到相同的语义空间,以估计它们在当前已知类别语义中的支持。该语义证据与PROB的视觉物体性融合,以校准已知和未知的预测,而不将OWOD转变为开放词汇分类。重要的是,MSPO从不使用未来类别名称,所有未见类别在评估过程中保持未命名。在M-OWODB和S-OWODB上的实验表明,MSPO在主要聚合指标上改善了强大的PROB基线,同时保持了竞争性的未知召回率。它还改善了早期未知混淆指标,并将PASCAL VOC最终mAP提高了最多2.7个百分点。这些结果表明,在标准OWOD设置下,已知类别语言语义为概率物体性提供了有效的校准信号。
cs.CV / 122 / 2607.23994

Effective Receptive Field Ordering Matters for Infrared Small Target Detection

有效感受野排序对红外小目标检测的重要性
Zhang, Guoyi, Du, Yanjin, Zhao, Zhengyao, Zhang, Tongsu, Xu, Guangsheng, Chen, Siyang, Xu, Xiangpeng, Wang, Han, Zhang, Xiaohu
Abstract
In this work, we investigate a previously unexplored architectural dimension for infrared small target detection: the organization of effective receptive fields (ERFs) during feature refinement. Unlike existing approaches that primarily improve individual feature operators, we argue that ERF organization constitutes an architectural dimension independent of receptive field design itself, and formulate deep feature transformation as a progressive residual correction process, from which a theoretical framework for ERF scheduling is established. Specifically, we reveal that ERF refinement is governed by two fundamental properties: scale-frequency correspondence, which aligns different ERF scales with distinct residual frequency characteristics, and nonlinear non-commutativity, which makes different ERF orderings produce fundamentally different refinement trajectories. Together, these properties show that ERF organization, rather than ERF scale alone, governs refinement dynamics. Guided by these principles, we propose Receptive Field Ordering Network (RFONet), which realizes hierarchical ERF scheduling through a multigrid-inspired V-cycle strategy using only standard $3\times3$ convolutions. RFONet achieves state-of-the-art performance on multiple benchmarks with only 1.16M parameters and over 157 FPS inference speed. Beyond empirical performance, our theoretical analysis provides theoretical guarantees for stable residual refinement under perturbations, frequency shifts, and partial occlusions, which are consistently reflected in superior noise robustness and cross-dataset generalization. Finally, our framework reformulates ERF organization as a task-dependent optimization objective, providing a principled foundation for future adaptive receptive field scheduling.
Chinese Translation
在本研究中,我们探讨了红外小目标检测中一个先前未被深入研究的架构维度:特征精炼过程中的有效感受野(ERFs)组织。与现有方法主要改进单个特征操作符不同,我们认为ERF组织构成了一个独立于感受野设计本身的架构维度,并将深度特征变换形式化为一个渐进的残差校正过程,从中建立了ERF调度的理论框架。具体而言,我们揭示了ERF精炼受两个基本属性的支配:尺度-频率对应关系,该关系将不同的ERF尺度与不同的残差频率特征对齐,以及非线性非交换性,使得不同的ERF排序产生根本不同的精炼轨迹。这些属性共同表明,ERF组织而非仅仅是ERF尺度主导了精炼动态。在这些原则的指导下,我们提出了感受野排序网络(Receptive Field Ordering Network, RFONet),该网络通过一种受多网格启发的V循环策略,仅使用标准的$3 imes3$卷积实现了分层ERF调度。RFONet在多个基准测试中达到了最先进的性能,仅需1.16M参数和超过157 FPS的推理速度。除了经验性能外,我们的理论分析为在扰动、频率偏移和部分遮挡下的稳定残差精炼提供了理论保证,这些保证在优越的噪声鲁棒性和跨数据集泛化中得到了持续体现。最后,我们的框架将ERF组织重新表述为一个任务依赖的优化目标,为未来自适应感受野调度提供了原则性的基础。
cs.CV / 123 / 2607.24002

Low-light Image Enhancement via Multi-scale Attention combined with Fourier Transform

基于多尺度注意力与傅里叶变换的低光照图像增强
Du, Wenbin, Long, Jian, Cao, Zhu
Abstract
Low-light image enhancement (LLIE) aims to improve image quality and clarity in diverse and demanding low-illumination environments. However, existing deep learning-based LLIE methods struggle to accurately capture real-world illumination and restore texture details, largely because their algorithmic strengths remain underutilized. To address these issues, we present a supervised frequency domain deep learning network for LLIE, named multi-scale attention combined with the Fourier transform (MSFT) which adopts a U-shaped, one-stage architecture that infuses guidance from low-light images into the network by channeling it through multi-scale attention. We further fuse the amplitude information from priori channels with that of the low-light image in MSFT's self-created module, and carry out multi-scale guidance along with the network. Subsequently, to better enhance the faint feature, such as fine content and textures, and to better fuse global context confidence in the decoding stage, we separately introduce a multi-shape synergistic attention and a lightweight network that effectively integrate information in high-dimensional space to embed into the superlative feature space channel containing rich texture information. Extensive experiments conducted on LOL, SID, SMID, and SDSD datasets demonstrate that MSFT significantly outperforms state-of-the-art competitors. For example, compared with Retinexformer, our method achieves a peak signal-to-noise ratio of up to 41.76 decibels on the SDSD-outdoor dataset with an increase of 11.92 decibels and a structural similarity index of 0.988 with a 13.80% improvement.
Chinese Translation
低光照图像增强(LLIE)旨在提高在多样化和苛刻的低照明环境中的图像质量和清晰度。然而,现有的基于深度学习的LLIE方法在准确捕捉现实世界的照明和恢复纹理细节方面存在困难,主要是因为它们的算法优势未得到充分利用。为了解决这些问题,我们提出了一种用于LLIE的监督频域深度学习网络,命名为多尺度注意力结合傅里叶变换(MSFT),该网络采用U型单阶段架构,通过多尺度注意力将低光照图像的引导信息注入网络。我们进一步将来自先验通道的幅度信息与MSFT自创模块中的低光照图像信息融合,并在网络中进行多尺度引导。随后,为了更好地增强微弱特征,如细腻内容和纹理,并在解码阶段更好地融合全局上下文信心,我们分别引入了多形状协同注意力和轻量级网络,有效地整合高维空间中的信息,以嵌入包含丰富纹理信息的优质特征空间通道。在LOL、SID、SMID和SDSD数据集上进行的大量实验表明,MSFT显著优于当前最先进的竞争对手。例如,与Retinexformer相比,我们的方法在SDSD户外数据集上实现了高达41.76分贝的峰值信噪比,提升了11.92分贝,并且结构相似性指数为0.988,改善了13.80%。
cs.CV / 124 / 2607.24009

Structural Loss Metrics for Tensor Approximation via Matrix Low-Rank Approximation

通过矩阵低秩近似的张量近似结构损失度量
Hasegawa, Hiroki
Abstract
Matricized low-rank approximation via SVD is a standard surrogate for tensor decompositions, but entry-wise reconstruction error fails to capture multiway geometric degradation. Under an orthogonal Tucker model, we characterize this degradation using two metrics: cross-mode Direction Loss, measuring geometric subspace deviation from rank truncation and noise rotation, and Interaction Loss, quantifying multilinear interaction distortion in the core tensor. We prove that squared relative reconstruction error orthogonally decomposes into interaction loss and out-of-subspace energy, and derive a Wedin-type bound establishing the stability of a plug-in Direction Loss estimator. Experiments on synthetic and hyperspectral datasets demonstrate that nearly identical reconstruction errors can yield markedly different structural-loss profiles; hyperspectral patches with comparable reconstruction errors exhibit up to a 4.6-fold difference in Direction Loss, correlating with severe visual blurring.
Chinese Translation
通过奇异值分解(SVD)的矩阵化低秩近似是张量分解的标准替代方法,但逐项重构误差无法捕捉多维几何退化。在正交Tucker模型下,我们使用两个度量来表征这种退化:交模式方向损失(cross-mode Direction Loss),用于测量几何子空间从秩截断和噪声旋转的偏差;交互损失(Interaction Loss),用于量化核心张量中的多线性交互失真。我们证明了平方相对重构误差正交分解为交互损失和超出子空间的能量,并推导出一种Wedin类型的界限,确立了插值方向损失估计器的稳定性。在合成数据集和高光谱数据集上的实验表明,几乎相同的重构误差可以产生显著不同的结构损失特征;具有可比重构误差的高光谱图块在方向损失上可展现出多达4.6倍的差异,这与严重的视觉模糊相关联。
cs.CV / 125 / 2607.24013

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

AptAvatar:用于生产就绪头像的快速生动长格式音频驱动视频生成
Zhang, Hengyuan, Sun, Jingna, Jin, Meiguang, Ma, Junfeng
Abstract
Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often compromise quality through restrictive architectural choices, such as causal attention and short temporal horizons, or by reducing model capacity and resolution. Without such compromises, we propose AptAvatar, a 14B-parameter long-form audio-driven avatar generation framework that delivers fast and expressive inference. For efficiency in production-level applications, AptAvatar addresses the extreme two-step generation challenge. To bridge the gap between the multi-step teacher model and the two-step student model, we introduce Endpoint-Anchored Distribution Distillation. It augments vanilla distribution matching with a dedicated Anchor Score Estimator trained on the trajectory-endpoint distribution defined from a frozen pretrained 4-step bridge generator. This provides an attainable endpoint-level anchor for the evolving two-step student. To improve long-horizon consistency, we further introduce Self-Generated History Replay, which reuses cached outputs from earlier generator checkpoints as history conditions during chunk-wise training. This approximates inference-time conditioning on self-generated histories without costly online rollouts, mitigating quality degradation from accumulated history errors. Extensive experiments demonstrate that AptAvatar generates vivid 720p long-form avatar videos with only 2 NFEs, achieving a 60x speedup while preserving visual fidelity and long-horizon identity. Code is available at https://github.com/TaoLiveAIGC/AptAvatar
Chinese Translation
生产就绪的音频驱动头像生成需要高效推理,同时不牺牲保真度或运动表现力。然而,现有的加速方法往往通过限制性的架构选择(如因果注意力和短时间范围)或通过降低模型容量和分辨率来妥协质量。为了解决这些妥协,我们提出了AptAvatar,一个拥有140亿参数的长格式音频驱动头像生成框架,能够提供快速且富有表现力的推理。为了在生产级应用中实现高效,AptAvatar解决了极端的两步生成挑战。为了弥合多步教师模型与两步学生模型之间的差距,我们引入了端点锚定分布蒸馏(Endpoint-Anchored Distribution Distillation)。它通过在冻结的预训练4步桥接生成器定义的轨迹端点分布上训练的专用锚得分估计器,增强了普通分布匹配。这为不断演变的两步学生提供了一个可实现的端点级锚点。为了改善长时间一致性,我们进一步引入了自生成历史重放(Self-Generated History Replay),在分块训练期间重用早期生成器检查点的缓存输出作为历史条件。这在推理时近似了对自生成历史的条件,而无需昂贵的在线展开,从而减轻了因历史错误累积而导致的质量下降。大量实验表明,AptAvatar能够生成生动的720p长格式头像视频,仅需2次NFEs,达成60倍的速度提升,同时保持视觉保真度和长时间身份一致性。代码可在 https://github.com/TaoLiveAIGC/AptAvatar 获取。
cs.CV / 126 / 2607.24016

DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models

DailyBench:一个统一的基准测试,用于评估现代生成模型生成和操纵的图像
Jiang, Xin, Tang, Hao, Gao, Junyao, Cao, Meiqi, Shen, Fei, Zhang, Dongming, Zhang, Yongdong
Abstract
Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identifying highly realistic content generated by both modern generation and manipulation pipelines. However, existing detection benchmarks are often built with outdated generative models and primarily emphasize full-image synthesis, creating a growing mismatch between benchmark data and the images encountered in real-world generation and editing scenarios. To bridge this gap, we introduce DailyBench, a high-quality unified benchmark for evaluating whether AI-generated image detectors can generalize across both modern full-image synthesis and object-level manipulation. DailyBench contains two complementary subsets: FakeBench, which includes high-quality images synthesized by recent open-source and commercial generative models, and ManipulationBench, which introduces challenging object-level edits applied to real images using advanced image-conditional models. This design makes DailyBench a realistic testbed for studying both generator-level generalization and manipulation-aware detection under subtle local edits. Experiments on DailyBench reveal substantial robustness gaps in current detectors: methods reporting 91-96% balanced accuracy on GenImage drop to 60-76% on FakeBench and 54-66% on ManipulationBench. These results show that existing detectors remain poorly generalized to realistic synthesis and manipulation, highlighting DailyBench as a rigorous testbed for developing robust and manipulation-aware AI-generated image detection methods. The project is available at https://dailybench.github.io/
Chinese Translation
最近在生成模型方面的进展使得人工智能生成图像的检测从识别易于区分的完全合成图像转变为识别由现代生成和操纵流程生成的高度逼真的内容。然而,现有的检测基准往往基于过时的生成模型构建,并主要强调全图像合成,这导致基准数据与现实世界生成和编辑场景中遇到的图像之间存在日益增长的不匹配。为了解决这一问题,我们引入了DailyBench,这是一个高质量的统一基准,用于评估人工智能生成图像检测器在现代全图像合成和对象级操纵之间的泛化能力。DailyBench包含两个互补的子集:FakeBench,包含由最近的开源和商业生成模型合成的高质量图像;ManipulationBench,介绍了使用先进的图像条件模型对真实图像进行的具有挑战性的对象级编辑。这种设计使DailyBench成为研究生成器级泛化和在微妙局部编辑下的操纵感知检测的现实测试平台。在DailyBench上的实验揭示了当前检测器存在显著的鲁棒性差距:在GenImage上报告91-96%平衡准确率的方法在FakeBench上下降到60-76%,在ManipulationBench上下降到54-66%。这些结果表明,现有检测器在现实合成和操纵方面的泛化能力仍然较差,突显了DailyBench作为开发鲁棒且具操纵感知的人工智能生成图像检测方法的严格测试平台的重要性。该项目可在https://dailybench.github.io/获取。
cs.CV / 127 / 2607.24017

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

从注意力流形中解构语义注意力与结构偏差
Jiao, Pengkun, Zhu, Bin, Chen, Jingjing, Jiang, Yu-gang
Abstract
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the semantic visual signal, precipitating multimodal hallucinations as the model prioritizes linguistic priors over valid visual evidence. To address this limitation, we introduce Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention. SPAR mitigates this generalized textual bias by purifying structural noise and subsequently redistributing the reclaimed attention budget to the most informative visual regions. Comprehensive evaluations across a diverse spectrum of hallucination benchmarks demonstrate that SPAR effectively restores authentic visual grounding with negligible computational overhead.
Chinese Translation
注意力机制在多模态大型语言模型(MLLMs)中的实证成功常常掩盖其固有的微妙缺陷。具体而言,MLLMs 一贯表现出对某些语义上无信息的视觉标记的不成比例的注意力,这一现象被称为“注册”或“视觉注意力沉没”。虽然现有的推理干预方法试图识别这些沉没标记并重新分配其注意力权重,但此类方法通常将这些标记孤立对待,且计算效率低下。相反,我们将这一现象重新框定为对视觉特征施加的广义文本偏差,这种偏差超出了孤立的沉没标记。从这个角度来看,普遍的结构偏差导致语义视觉信号的稀释,促使多模态幻觉的产生,因为模型优先考虑语言先验而非有效的视觉证据。为了解决这一局限性,我们引入了显著性引导的净化与自适应重分配(Saliency-guided Purification and Adaptive Redistribution, SPAR),这是一种无训练、即插即用的干预方法。SPAR 通过净化结构噪声并随后将回收的注意力预算重新分配到最具信息量的视觉区域,从而减轻这种广义文本偏差。对多种幻觉基准的全面评估表明,SPAR 能有效恢复真实的视觉基础,且计算开销微乎其微。
cs.CV / 128 / 2607.24024

A Unified Stereo Geometry Estimation Framework for Disparity and Surface Normal

一种统一的立体几何估计框架用于视差和表面法线
Wei, Qizhe, Guo, Xianda, Xu, Shaocong, Li, Hong, Yang, Runyi, Zhao, Hao
Abstract
Stereo matching and surface normal estimation are fundamental tasks in 3D vision. However, existing feed-forward stereo methods still struggle to produce reliable predictions in challenging regions, mainly due to the lack of strong geometric priors. In this paper, we propose $\textbf{GeoStereo}$, a unified stereo geometry estimation framework that leverages powerful diffusion priors to jointly predict disparity and surface normals. Specifically, GeoStereo couples a feed-forward stereo matching pipeline with a diffusion-based normal estimation branch. To enable effective interaction between the two tasks, we introduce a disparity to normal initialization strategy and construct a warp to left-view condition for the diffusion process. This coupled design allows the diffusion branch to provide strong structural priors that enhance disparity estimation in ill-posed regions, while the feed-forward branch offers reliable geometric guidance for accurate normal prediction. Extensive experiments show that GeoStereo performs reliably in challenging scenarios, including low-light environments, highly reflective surfaces, and transparent objects. Under zero-shot settings, it achieves Rank-1 disparity estimation on multiple benchmarks, including KITTI and NYUv2, and delivers the best normal estimation accuracy on many real indoor benchmarks, such as iBims-1 and ScanNet. Project page: https://qz-wei.github.io/GeoStereo.github.io/
Chinese Translation
立体匹配和表面法线估计是三维视觉中的基本任务。然而,现有的前馈立体方法在复杂区域仍然难以产生可靠的预测,主要是由于缺乏强大的几何先验。在本文中,我们提出了 $ extbf{GeoStereo}$,一种统一的立体几何估计框架,利用强大的扩散先验共同预测视差和表面法线。具体而言,GeoStereo 将前馈立体匹配管道与基于扩散的法线估计分支相结合。为了实现这两项任务之间的有效交互,我们引入了一种视差到法线的初始化策略,并为扩散过程构建了一个左视图条件的变换。这种耦合设计使得扩散分支能够提供强大的结构先验,从而增强在不适定区域的视差估计,而前馈分支则为准确的法线预测提供可靠的几何指导。大量实验表明,GeoStereo 在低光环境、高反射表面和透明物体等挑战性场景中表现可靠。在零-shot 设置下,它在多个基准测试(包括 KITTI 和 NYUv2)上实现了 Rank-1 视差估计,并在许多真实室内基准(如 iBims-1 和 ScanNet)上提供了最佳的法线估计精度。项目页面:https://qz-wei.github.io/GeoStereo.github.io/
cs.CV / 129 / 2607.24027

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

Sol-Attn:通过动态注意力稀疏化加速视频生成推理
Li, Haopeng, Li, Yitong, Chen, Junsong, Ye, Tian, Liu, Haozhe, Yu, Jincheng, Wang, Duomin, Zhang, Ruihua, Xie, Zeke, Xie, Enze, Han, Song
Abstract
Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing methods struggle to sparsify attention both efficiently and accurately for two reasons: (1) Rigid, unpredictable, and costly routing: selecting a fixed fraction of top-ranked blocks by proxy score imposes fixed budgets, whereas retaining blocks to reach a target cumulative proxy probability mass yields dynamic but potentially imbalanced budgets; both incur non-negligible overhead from computing and materializing proxy scores. (2) Lossy keep-or-drop sparsification: unselected blocks are discarded entirely, degrading accuracy under aggressive sparsity. These limitations motivate cheaper dynamic-budget routing while limiting accuracy degradation. In this paper, we introduce training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention. The core of Sol-Attn is on-the-fly block thresholding with proxy-score reuse, which selects critical blocks by comparing block proxy scores against a threshold during online softmax. This design enables dynamic yet controllable block budgets without materializing the proxy map, while directly reusing the proxy scores of unselected blocks to approximate their contribution. Experiments across image and video generation tasks show that Sol-Attn advances the quality-efficiency frontier of training-free sparse attention, delivering 2.1 times and 2.3 times end-to-end speedups for video generation and editing, respectively, while preserving visual quality.
Chinese Translation
扩散变换器对于高保真视频生成至关重要,但长序列的令牌使得注意力成为主要的推理瓶颈。无训练的动态稀疏注意力通过仅计算选定的键值块来缓解这一瓶颈,然而现有方法在高效且准确地稀疏化注意力方面面临两大挑战:(1)刚性、不可预测且成本高昂的路由:通过代理分数选择固定比例的高排名块会施加固定预算,而保留块以达到目标累积代理概率质量则会产生动态但可能不平衡的预算;这两者都导致计算和实现代理分数的开销不可忽视。(2)有损的保留或丢弃稀疏化:未选择的块被完全丢弃,在激进稀疏下降低了准确性。这些限制促使我们寻求更便宜的动态预算路由,同时限制准确性下降。在本文中,我们提出了无训练的Sol-Attn(在线稀疏注意力),它在单次在线-softmax传递中统一了动态路由、稀疏计算和近似校正,实现了稀疏注意力中更好的准确性与效率权衡。Sol-Attn的核心是动态块阈值设定与代理分数重用,它通过在在线softmax期间将块代理分数与阈值进行比较来选择关键块。该设计使得动态且可控的块预算成为可能,而无需实现代理图,同时直接重用未选择块的代理分数以近似其贡献。在图像和视频生成任务中的实验表明,Sol-Attn推动了无训练稀疏注意力的质量-效率前沿,分别为视频生成和编辑提供了2.1倍和2.3倍的端到端加速,同时保持了视觉质量。
cs.CV / 130 / 2607.24052

PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic Rectification

PointCHR:基于曲率感知的超曲面矫正的点云分析
Yu, Xinxing, Yang, Liying, Mo, Hao, Ma, Hui, Kai, Fang, Liu, Ajian, Liang, Yanyan
Abstract
High-curvature regions in 3D point clouds encapsulate critical fine-grained geometric semantics yet exhibit a distinct long-tail sparsity in their spatial distribution. The inherent limitations of polynomial volume growth in Euclidean space frequently render these intricate geometric features challenging to adequately resolve within a uniform-scale feature space. Consequently, these regions are frequently overshadowed by smooth global features dominated by low-curvature regions, thereby limiting the discriminative capacity of the network. To address this issue, we propose PointCHR, a curvature-aware hyperbolic rectification (CHR) for point cloud analysis. Utilising the property of exponential volume expansion in the vicinity of hyperbolic manifolds, CHR presents a learnable curvature-guided radial rectification mechanism. By adaptively projecting high-curvature points towards boundary regions endowed with larger effective embedding capacities, PointCHR effectively mitigates the representation crowding problem inherent in Euclidean settings. Extensive experimentation has demonstrated that PointCHR significantly enhances the ability of backbone to capture fine-grained geometric details, achieving state-of-the-art performance across multiple benchmarks.
Chinese Translation
三维点云中的高曲率区域 encapsulate 关键的细粒度几何语义,但在其空间分布中表现出明显的长尾稀疏性。欧几里得空间中多项式体积增长的固有限制常常使这些复杂的几何特征在均匀尺度特征空间中难以充分解析。因此,这些区域常常被低曲率区域主导的平滑全局特征所掩盖,从而限制了网络的区分能力。为了解决这个问题,我们提出了 PointCHR,一种用于点云分析的曲率感知超曲面矫正(CHR)。利用超曲面附近指数体积扩展的特性,CHR 提供了一种可学习的曲率引导径向矫正机制。通过自适应地将高曲率点投影到具有更大有效嵌入能力的边界区域,PointCHR 有效缓解了欧几里得环境中固有的表示拥挤问题。大量实验表明,PointCHR 显著增强了主干网络捕捉细粒度几何细节的能力,在多个基准测试中实现了最先进的性能。
cs.CV / 131 / 2607.24064

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

MarineEVT:通过视觉工具推理推进事件中心的海洋视频理解
To, Tuan-An, Wong, Yuk-Kwan, Vu, Tuan-Anh, Zheng, Ziqiang, Yeung, Sai-Kit
Abstract
Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs. However, the performance of VLMs often degrades in the video domain due to the essential need for temporal understanding and the scarcity of large-scale annotated video data. In this work, we focus on marine video understanding, which brings further challenges: first, it requires substantial domain expertise; and video VLMs usually struggle with localizing and interpreting critical information from marine videos, as the informative events are typically sparse, unpredictable, and unevenly distributed. To address these challenges, we carefully curate the first event-centric marine video understanding dataset called MarineEVT, which features 20K multi-task, video-level visual question-answering pairs spanning multiple dimensions of marine understanding and analysis. Meanwhile, based on MarineEVT, we decompose marine video understanding as an Event-centric Visual Tool-integrated Reasoning process EVT-R1 for short, where we leverage powerful visual tools to drive the model to localize and interpret critical information aligned with visual questions and human intent. To demonstrate its effectiveness, we compare EVT-R1 against 11 SOTA VLMs in different settings. EVT-R1 outperforms the top open-source and top commercial models by 5.22 and 11.09, respectively. MarineEVT and EVT-R1 lay the foundation for ecological discovery and marine education, fostering the development of VLMs capable of interpreting marine dynamics, reasoning about ecological interactions, and supporting sustainable ocean video understanding and analysis.
Chinese Translation
近期的视觉-语言模型(VLMs)在视觉理解方面取得了显著成功,这得益于高质量图像-文本对的日益丰富。然而,由于对时间理解的基本需求以及大规模标注视频数据的稀缺,VLMs在视频领域的性能往往下降。在本研究中,我们聚焦于海洋视频理解,这带来了进一步的挑战:首先,它需要大量的领域专业知识;而视频VLMs通常在从海洋视频中定位和解释关键信息方面面临困难,因为信息事件通常稀疏、不确定且分布不均。为了解决这些挑战,我们精心策划了第一个以事件为中心的海洋视频理解数据集,称为MarineEVT,该数据集包含2万对多任务、视频级的视觉问答对,涵盖了海洋理解和分析的多个维度。同时,基于MarineEVT,我们将海洋视频理解分解为事件中心的视觉工具集成推理过程(EVT-R1),在此过程中,我们利用强大的视觉工具推动模型定位和解释与视觉问题和人类意图相一致的关键信息。为了验证其有效性,我们将EVT-R1与11个不同设置下的最先进VLMs进行比较。EVT-R1分别超越了顶级开源模型和顶级商业模型5.22和11.09的性能。MarineEVT和EVT-R1为生态发现和海洋教育奠定了基础,促进了能够解释海洋动态、推理生态互动并支持可持续海洋视频理解和分析的VLMs的发展。
cs.CV / 132 / 2607.24077

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

当低字符错误率不足够时:对历史乌拉圭文档中的视觉-语言OCR系统幻觉的分析
Gardella, Marina, Mari{ñ}o, Camilo, Belzarena, Diego, Ram{í}rez, Ignacio, Randall, Gregory, Morel, Jean-Michel
Abstract
Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription remains insufficiently understood. In this work, we benchmark traditional OCR systems and VLM-based approaches on the Berrutti dataset, a challenging collection of Uruguayan dictatorship-era documents derived from microfilm scans. While VLMs consistently outperform traditional methods in terms of Character Error Rate (CER) and Word Error Rate (WER), we show that these improvements hide a more complex picture. Through a detailed qualitative analysis, we uncover systematic failure modes that are invisible to standard metrics, including orthographic normalization, spurious content generation, and semantic substitutions that preserve fluency while altering meaning. Errors affecting named entities are particularly critical, as they can introduce substantial semantic distortions with minimal impact on CER and WER. These findings reveal a critical gap between quantitative OCR performance and transcription fidelity in real-world archival settings, and highlight the need for evaluation frameworks that go beyond character-level accuracy to capture the semantic reliability of generated transcriptions.
Chinese Translation
光学字符识别(OCR)是历史档案数字化的关键组成部分。近年来,视觉-语言模型(VLMs)作为传统OCR系统的强有力替代方案出现,在标准基准测试中取得了最先进的性能。然而,它们在档案转录中的适用性仍然不够明确。在本研究中,我们在Berrutti数据集上对传统OCR系统和基于VLM的方法进行了基准测试,该数据集是一个具有挑战性的乌拉圭独裁时期文档的集合,来源于微缩胶卷扫描。尽管VLM在字符错误率(CER)和单词错误率(WER)方面始终优于传统方法,但我们表明这些改进掩盖了更复杂的情况。通过详细的定性分析,我们揭示了标准指标无法察觉的系统性失败模式,包括正字法规范化、虚假内容生成以及在保持流畅性的同时改变意义的语义替代。影响命名实体的错误尤其关键,因为它们可能在对CER和WER影响最小的情况下引入显著的语义扭曲。这些发现揭示了定量OCR性能与现实世界档案环境中的转录忠实度之间的关键差距,并强调了需要超越字符级准确性以捕捉生成转录的语义可靠性的评估框架。
cs.CV / 133 / 2607.24090

Cascade Forgery Mining Network for Fingerprint Presentation Attack Detection

级联伪造挖掘网络用于指纹呈现攻击检测
Fei, Hongyan, Huang, Chuanwei, Wang, Zheng, Luo, Pengcheng, Li, Jingwei, Feng, Jufu
Abstract
Fingerprint Presentation Attack Detection (PAD) is a critical component of fingerprint identification systems, serving as a protective measure against unauthorized access. In this paper, we observe that different regions of a fingerprint image can exhibit varying Artifact Extraction Difficulty (AED), with high-AED regions requiring more sophisticated extraction mechanisms to capture more subtle discriminative evidence. To address this issue, we propose to quantify AED using local Gabor feature certainty and partition fingerprint images into multiple regions based on their respective AED values. We then propose an AED guided Cascade Forgery Mining Network (CFM-Net) that employs an adaptive-depth feature extraction architecture to detect more precise and comprehensive artifact evidence across regions with heterogeneous AED values. Furthermore, we introduce an Orientation Guided Adversarial Training (OGAT) module to filter out identity information from PAD features while preserving the integrity of original artifact evidence. Experimental evaluations on LivDet datasets demonstrate the superior performance of our approach compared to state-of-the-art methods and achieve significant improvement in the classification ability of high AED fingerprints.
Chinese Translation
指纹呈现攻击检测(PAD)是指纹识别系统的关键组成部分,作为防止未经授权访问的保护措施。在本文中,我们观察到指纹图像的不同区域可能表现出不同的伪影提取难度(AED),高AED区域需要更复杂的提取机制来捕捉更微妙的区分证据。为了解决这个问题,我们提出通过局部Gabor特征确定性来量化AED,并根据各自的AED值将指纹图像划分为多个区域。然后,我们提出了一种AED引导的级联伪造挖掘网络(CFM-Net),该网络采用自适应深度特征提取架构,以便在具有异质AED值的区域中检测更精确和全面的伪影证据。此外,我们引入了一个方向引导的对抗训练(OGAT)模块,以在保留原始伪影证据完整性的同时过滤掉PAD特征中的身份信息。在LivDet数据集上的实验评估表明,我们的方法相比于最先进的技术具有更优越的性能,并在高AED指纹的分类能力上取得了显著提升。
cs.CV / 134 / 2607.24098

ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

ReflexTrack:一种无训练反馈驱动的引用视频目标分割代理
Li, Yuanjia, Xu, Tianyang, Zhou, Tao, Tang, Zhangyong, Wu, Xiao-Jun, Kittler, Josef
Abstract
Referring video object segmentation (RVOS) requires segmenting a target specified by natural language throughout a video. Recent agentic approaches combine multimodal large language models with promptable segmentation models to perform RVOS without task-specific training. However, most pipelines rely on one-shot spatial grounding followed by mask propagation, leaving both the initial prompts and temporal predictions largely unverified. We introduce ReflexTrack, a training-free, feedback-driven agent that closes this loop at both spatial and temporal levels. Mask-guided Spatial Refinement evaluates the mask induced by the current keyframe prompt and iteratively updates the bounding box together with positive and negative points, yielding a more reliable initialization. Video-level Mask Reflection assesses the complete mask sequence, localizes unreliable intervals, selects complementary repair keyframes, and generates candidate predictions through mask-guided re-propagation. Only candidates that provide a verified improvement are used to update the affected intervals, preserving reliable predictions elsewhere. All components remain frozen during inference. ReflexTrack achieves an overall $\mathcal{Q}$ score of $69.7$ on Ref-VPS and a $\mathcal{J}\&\mathcal{F}$ score of $67.2$ on ReasonVOS. These results demonstrate that prediction-level feedback substantially improves the reliability of training-free RVOS.
Chinese Translation
引用视频目标分割(RVOS)需要在整个视频中对由自然语言指定的目标进行分割。近期的代理方法将多模态大型语言模型与可提示的分割模型相结合,以在无需特定任务训练的情况下执行RVOS。然而,大多数流程依赖于一次性空间定位,随后进行掩码传播,这使得初始提示和时间预测在很大程度上未得到验证。我们提出了ReflexTrack,这是一种无训练的反馈驱动代理,能够在空间和时间层面上闭合这一循环。掩码引导的空间精细化评估当前关键帧提示所诱导的掩码,并通过正负点的迭代更新边界框,从而产生更可靠的初始化。视频级掩码反射评估完整的掩码序列,定位不可靠的区间,选择互补的修复关键帧,并通过掩码引导的重新传播生成候选预测。只有那些提供经过验证的改进的候选才会用于更新受影响的区间,从而在其他地方保留可靠的预测。所有组件在推理过程中保持冻结。ReflexTrack在Ref-VPS上实现了整体的$ ext{Q}$分数为$69.7$,在ReasonVOS上实现了$ ext{J} ext{&} ext{F}$分数为$67.2$。这些结果表明,预测级反馈显著提高了无训练RVOS的可靠性。
cs.CV / 135 / 2607.24101

LU-500: A Logo Benchmark for Concept Unlearning

LU-500:概念遗忘的标志基准
Li, Keyu, Gao, Jin, Zhang, Jialing, Wang, Dequan
Abstract
Concept unlearning is increasingly used to limit the reproduction of protected or unsafe visual concepts in text-to-image models. Existing evaluations, however, mostly study targets that dominate the whole image, such as styles, broad object categories, or portrait-like identities, leaving company logos comparatively underexamined. Logos create a different failure mode: a small localized mark can carry the entire protected concept, must be visually precise to remain recognizable, and can be triggered implicitly by products, storefronts, packaging, or advertisements even when the word ``logo'' is absent. We introduce LU-500, a logo-unlearning benchmark built from Fortune Global 500 companies to study this localized and semantically entangled setting. LU-500 contains nearly 10,000 curated text-query and logo-image pairs, with an explicit track (LUex-500) and an implicit contextual track (LUim-500). To avoid reducing the task to a binary detector score, we define a multi-grained protocol that evaluates both local logo removal and global image preservation in pixel and latent spaces. Experiments on representative inference-time methods, including NP, SLD, and SEGA, and compatible fine-tuning-based methods such as ESD and Forget-Me-Not, show that the evaluated methods struggle to remove logo evidence without changing non-target content. We further analyze ProLU, a prompt-space multi-agent baseline: it improves local erasure by removing logo-inducing semantics, but also illustrates why prompt filtering is not a substitute for weight-level disentanglement. Correlation analyses over logo area, location, and structural complexity suggest that future logo unlearning may need spatially aware controls, such as SSIM-guided constraints, rather than purely global concept suppression.
Chinese Translation
概念遗忘在限制文本到图像模型中受保护或不安全视觉概念的再现方面越来越多地被使用。然而,现有的评估主要研究主导整个图像的目标,如风格、广泛的物体类别或肖像般的身份,而对公司标志的研究相对较少。标志创造了一种不同的失败模式:一个小的局部标记可以承载整个受保护的概念,必须在视觉上精确以保持可识别性,并且即使在缺少“logo”一词的情况下,也可以通过产品、店面、包装或广告隐含触发。我们提出了LU-500,这是一个基于财富全球500强公司的标志遗忘基准,旨在研究这种局部和语义纠缠的设置。LU-500包含近10,000对精心策划的文本查询和标志图像,分为显式轨道(LUex-500)和隐式上下文轨道(LUim-500)。为了避免将任务简化为二元检测器评分,我们定义了一种多粒度协议,评估像素和潜在空间中的局部标志去除和全局图像保留。在代表性的推理时间方法(包括NP、SLD和SEGA)以及兼容的基于微调的方法(如ESD和Forget-Me-Not)上的实验表明,被评估的方法在不改变非目标内容的情况下难以去除标志证据。我们进一步分析了ProLU,一个提示空间的多代理基线:它通过去除标志诱导语义来改善局部消除,但也说明了为什么提示过滤不能替代权重级别的解耦。对标志区域、位置和结构复杂性的相关性分析表明,未来的标志遗忘可能需要空间感知控制,例如基于SSIM的约束,而不是单纯的全局概念抑制。
cs.CV / 136 / 2607.24110

BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

超越融合:自对齐潜在扩散用于无校准红外超分辨率和红外-可见融合
Chen, Minchong, Yuan, Xiaoyun, Cao, Minyu, Zhang, Jianing, Zhang, Jun, Liu, Shuyang, Yang, Xiaokang
Abstract
Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion, a unified latent diffusion framework for calibration-free visible-guided infrared super-resolution and infrared-visible fusion tasks. The proposed framework supports both task-specific training and joint training where two tasks are optimized and executed as two readouts of the same generative process. Instead of relying on explicit registration or geometric warping, BeyondFusion introduces a cross-modal self-aligning (CMSA) module into the denoising U-Net. CMSA reorganizes infrared and visible latent tokens into a shared attention space to learn content-adaptive cross-modal correspondence during the denoising process. Together with misalignment augmentation module, the model is facilitated to exploit visible structural and semantic cues while preserving thermal consistency, enabling high-frequency infrared reconstruction and informative fused-image generation under uncalibrated conditions. Extensive experiments on public benchmarks and a mobile infrared-visible imaging system show strong performance across aligned inputs, low-resolution infrared observations, synthetic misalignments, and real mobile captures with unsynchronized sensors. Ablation studies, unified training analysis, and downstream pedestrian detection further validate the effectiveness of BeyondFusion for calibration-free multimodal imaging.
Chinese Translation
移动红外-可见成像通常将紧凑的红外传感器与高分辨率的可见相机配对,以实现互补感知。然而,由于不同的光学系统、视角、视场和曝光时机导致的跨传感器不对齐,阻碍了其实际应用。在本文中,我们提出了BeyondFusion,一个统一的潜在扩散框架,用于无校准的可见引导红外超分辨率和红外-可见融合任务。该框架支持任务特定训练和联合训练,其中两个任务作为同一生成过程的两个输出进行优化和执行。BeyondFusion引入了一个跨模态自对齐(CMSA)模块到去噪U-Net中,而不是依赖于显式的配准或几何变形。CMSA在去噪过程中将红外和可见潜在标记重组到共享注意力空间,以学习内容自适应的跨模态对应关系。结合不对齐增强模块,该模型能够利用可见的结构和语义线索,同时保持热一致性,从而在无校准条件下实现高频红外重建和信息丰富的融合图像生成。在公共基准和移动红外-可见成像系统上的大量实验显示了该模型在对齐输入、低分辨率红外观测、合成不对齐和真实移动捕获(使用不同步传感器)下的强大性能。消融研究、统一训练分析和下游行人检测进一步验证了BeyondFusion在无校准多模态成像中的有效性。
cs.CV / 137 / 2607.24124

ViDS: Video Diffusion Shader using 3D Face Tracking

ViDS:基于3D人脸追踪的视频扩散着色器
Ji, Wenbo, Davoli, Davide, Chen, Zhe, Schoneveld, Liam, Nießner, Matthias, Tang, Jiapeng
Abstract
We introduce ViDS, a Video Diffusion Shader that leverages 3D face tracking for expressive and identity-preserving portrait animation. We first reconstruct the identity-specific 3DMM mesh from the reference image, and then animate it using expression and pose parameters from a driving video. Leveraging dense geometric cues from 3DMM normal maps, we employ a video diffusion model as a neural shader to synthesize lifelike portrait animations while preserving the appearance and identity of the reference image. We find that more accurate 3DMM tracking enables finer-grained expression control. We also introduce an autoregressive diffusion sampling process that extends generation beyond the model's native window while reducing discontinuities between adjacent clips. Compared with prior diffusion-based approaches for portrait animation that rely on landmark-based conditioning or implicit motion latents, our method achieves more detailed and consistent expression and pose control while faithfully preserving identity and appearance. Detailed ablation studies validate the effectiveness of our design choices. Project page: https://fusheng-ji.github.io/ViDS/
Chinese Translation
我们介绍了ViDS,一种利用3D人脸追踪技术进行富有表现力且保留身份特征的肖像动画的视频扩散着色器。我们首先从参考图像重建特定身份的3DMM网格,然后使用来自驱动视频的表情和姿态参数对其进行动画处理。通过利用来自3DMM法线图的密集几何线索,我们采用视频扩散模型作为神经着色器,合成逼真的肖像动画,同时保留参考图像的外观和身份。我们发现,更准确的3DMM追踪能够实现更细致的表情控制。我们还引入了一种自回归扩散采样过程,扩展生成超出模型的原生窗口,同时减少相邻剪辑之间的不连续性。与之前依赖于基于关键点的条件或隐式运动潜变量的肖像动画扩散方法相比,我们的方法在忠实保留身份和外观的同时,实现了更详细和一致的表情与姿态控制。详细的消融研究验证了我们设计选择的有效性。项目页面:https://fusheng-ji.github.io/ViDS/
cs.CV / 138 / 2607.24135

LoTA-N2N: Local Trace Adaptation for Zero-Shot Self-Supervised Image Denoising

LoTA-N2N:用于零-shot自监督图像去噪的局部轨迹适应
Hu, Jintong, Xia, Bin, Liu, Junlin, Liu, Jiayue, Yang, Wenming
Abstract
Single-image self-supervised denoising replaces unavailable clean targets with surrogate targets constructed from noisy observations. Its effectiveness therefore depends on how closely the surrogate objective remains aligned with supervised denoising, especially when noise is correlated, spatially nonstationary, or unknown. We express the discrepancy between a broad class of MSE-based self-supervised objectives and supervised MSE as a parameter-independent constant and a trace interaction between the surrogate-target residual and the prediction error. The corresponding gradient discrepancy is determined by the gradient of this interaction. This formulation provides a common view of paired-noise, blind-spot, weak-noise, re-corruption, and sub-image methods, while revealing that a small global interaction may conceal substantial positive and negative regional interactions through spatial cancellation. Building on these observations, we propose LoTA-N2N, a two-stage zero-shot adaptation framework. Stage 1 trains a denoiser on complementary sub-image pairs and freezes it to construct detached clean-sub-image proxies. Stage 2 estimates the residual--prediction interaction using these proxies and suppresses its patch-wise absolute magnitude. We show that the local construction prevents spatial cancellation and upper-bounds the magnitude of the corresponding global interaction. Experiments across natural, confocal, and X-ray images, complemented by iteration-matched controls, controlled noise shifts, and gradient diagnostics, show consistent gains over MSE-only adaptation under IID, spatially varying, and mixed noise. Overall, LoTA-N2N demonstrates that estimated local interaction and spatial cancellation control provide effective design principles for single-image self-supervised denoising without paired clean targets, repeated acquisitions, or a predefined re-corruption model.
Chinese Translation
单图像自监督去噪用来自噪声观测构建的替代目标来替代不可用的干净目标。因此,其有效性取决于替代目标与监督去噪之间的对齐程度,尤其是在噪声相关、空间非平稳或未知的情况下。我们将广泛类MSE(均方误差)自监督目标与监督MSE之间的差异表示为一个与参数无关的常数和替代目标残差与预测误差之间的轨迹交互。相应的梯度差异由这种交互的梯度决定。这种表述为配对噪声、盲点、弱噪声、再污染和子图像方法提供了一个共同的视角,同时揭示出小的全局交互可能通过空间抵消掩盖实质性的正负区域交互。基于这些观察,我们提出了LoTA-N2N,一个两阶段的零-shot适应框架。第一阶段在互补的子图像对上训练去噪器并将其冻结,以构建独立的干净子图像代理。第二阶段使用这些代理估计残差-预测交互,并抑制其块状绝对大小。我们展示了局部构造防止了空间抵消,并对相应的全局交互的大小进行了上界限制。在自然图像、共聚焦图像和X射线图像上的实验,辅以迭代匹配的对照、控制噪声转移和梯度诊断,显示出在IID、空间变化和混合噪声下,相较于仅使用MSE的适应方法具有一致的增益。总体而言,LoTA-N2N展示了估计的局部交互和空间抵消控制为没有配对干净目标、重复采集或预定义再污染模型的单图像自监督去噪提供了有效的设计原则。
cs.CV / 139 / 2607.24140

Image Inpainting via Stochastic Dynamics

通过随机动力学进行图像修复
Kuang, Jiaqi, Guo, Zihao, Qian, Zhongmin
Abstract
Image inpainting aims to recover missing regions while preserving structural consistency. We propose a non-parametric method without network training based on data-guided stochastic dynamics. Starting from a masked image, the missing pixels are evolved through a reverse-time stochastic differential equation with a kernel-weighted correction estimated directly from a reference dataset. This empirical correction guides the reconstruction toward high-density regions of the data distribution without training a neural network or fitting a parametric density model. Experiments on MNIST, Fashion-MNIST, and MVTec show that the proposed method outperforms Mean Fill, Telea, and Navier-Stokes inpainting in PSNR, SSIM, and visual quality. On CelebA, it remains competitive and produces plausible completions for structure-sensitive occlusions. These results demonstrate the effectiveness of empirical reference statistics as a non-parametric prior for image inpainting.
Chinese Translation
图像修复旨在恢复缺失区域,同时保持结构一致性。我们提出了一种基于数据引导的随机动力学的非参数方法,无需网络训练。从一个被遮挡的图像开始,缺失的像素通过反向时间随机微分方程演变,并通过直接从参考数据集中估计的核加权修正进行调整。这种经验修正引导重建朝向数据分布的高密度区域,而无需训练神经网络或拟合参数密度模型。在 MNIST、Fashion-MNIST 和 MVTec 上的实验表明,所提出的方法在 PSNR、SSIM 和视觉质量上优于均值填充(Mean Fill)、Telea 和 Navier-Stokes 修复。在 CelebA 上,该方法仍具竞争力,并为结构敏感的遮挡生成合理的补全。这些结果证明了经验参考统计作为图像修复的非参数先验的有效性。
cs.CV / 140 / 2607.24157

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

UniGen-AR:通过自回归建模统一视觉生成
Bao, Zhipeng, Zhu, Zhen, Kumari, Nupur, Bagchi, Anurag, Wang, Yu-Xiong, Tokmakov, Pavel, Hebert, Martial
Abstract
Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single model produces diverse image-valued outputs through a unified multimodal interface. While diffusion-based systems dominate UVG due to strong quality and controllability, their iterative sampling incurs substantial inference latency, limiting practical deployment. To address these limitations, we propose UniGen-AR, a framework that pairs a general-purpose multi-modal language model (MLLM) with an efficient next-scale visual auto-regressive (VAR) decoder. This design retains the flexibility of MLLM-based conditioning while leveraging the sampling efficiency and latent unification properties of VAR models. In our framework, the MLLM encodes free-form instructions and control signals into a unified sequence, which guides the VAR decoder to generate image-valued outputs for over 15 tasks spanning four families. Empirically, UniGen-AR achieves up to $19 \times$ lower inference latency than diffusion-based baselines while maintaining or improving output quality. Our ablations further reveal that VQ-VAE tokenizer design, particularly codebook size and hierarchy, is a critical factor for VAR scalability in UVG. These results establish visual auto-regressive modeling as a compelling and efficient backbone for unified visual generation. Our project page is at https://zpbao.github.io/projects/unigenar.
Chinese Translation
现代计算机视觉流程仍然存在碎片化现象,文本到图像生成、编辑、修复和经典感知等任务由不同的模型处理。我们研究统一视觉生成(Unified Visual Generation, UVG),在该框架下,单一模型通过统一的多模态接口生成多样的图像值输出。尽管基于扩散的系统因其强大的质量和可控性主导了UVG,但其迭代采样导致了显著的推理延迟,限制了实际应用。为了解决这些限制,我们提出了UniGen-AR,一个将通用多模态语言模型(Multi-Modal Language Model, MLLM)与高效的下一阶段视觉自回归(Visual Auto-Regressive, VAR)解码器相结合的框架。该设计保留了基于MLLM的条件化灵活性,同时利用VAR模型的采样效率和潜在统一特性。在我们的框架中,MLLM将自由形式的指令和控制信号编码为统一序列,指导VAR解码器为超过15个任务生成图像值输出,涵盖四个类别。实证结果表明,UniGen-AR的推理延迟比基于扩散的基线低达19倍,同时保持或提高输出质量。我们的消融实验进一步揭示,VQ-VAE标记器设计,特别是代码本大小和层次结构,是VAR在UVG中可扩展性的关键因素。这些结果确立了视觉自回归建模作为统一视觉生成的一个引人注目且高效的基础。我们的项目页面网址为 https://zpbao.github.io/projects/unigenar。
cs.CV / 141 / 2607.24165

Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval

当前检索器是否覆盖所有证据?关于联合跨页检索的对照研究
Cha, Sungguk, Kim, DongWook, Kim, Mintae, Han, Youngsub, Jeon, Byoung-Ki, Lee, Sangyeob
Abstract
Finding a long document relevant to a multi-part request is not the same as establishing that it contains every requested piece of evidence. We study this gap for conjunctive document retrieval, where two or three explicit conditions must be supported on different pages of one document. We use n-Clue as a controlled measurement instrument: 1{,}000 queries over 2{,}021 documents pair all-condition golds with naturally occurring documents that satisfy only a subset, and a complete-first success requires a top-10 gold to precede every released subset qrel. Across 70 configurations, condition-wise decomposition improves two dense backbones by 6.8--7.3 points and lexical--visual fusion adds 8.7, while four generic rerankers all reduce Gold-NDCG; these directions replicate on a four-source stress set. Scaling one dense family from 0.6B to 8B changes complete-first success by 0.0 points. The strongest displayed hybrid illustrates the resulting gap: it finds a gold for 81.1\% of queries but succeeds complete-first on only 35.8\%, and the gap persists across condition count, target length, candidate density, query rendering, and the four-source stress set. Finally, page-aware visual systems surface stored support for every condition on only 5.1--5.3\% of queries. These results identify condition coverage, rather than gold discovery alone, as the central bottleneck.
Chinese Translation
找到与多部分请求相关的长文档并不等同于确认它包含每一项请求的证据。我们研究了这一差距,聚焦于联合文档检索,其中两个或三个明确条件必须在同一文档的不同页面上得到支持。我们使用 n-Clue 作为控制测量工具:对 2,021 个文档进行 1,000 个查询,将所有条件的金标准与仅满足部分条件的自然文档配对,完整优先成功要求每个发布的子集 qrel 之前都有一个前十名的金标准。在 70 种配置中,条件-wise 分解使两个密集主干的性能提高了 6.8-7.3 分,而词汇-视觉融合增加了 8.7 分,四个通用重排序器则均降低了 Gold-NDCG;这些方向在一个四源压力集上得到了重复验证。将一个密集家族的规模从 0.6B 扩展到 8B 对完整优先成功的影响为 0.0 分。最强的混合模型展示了由此产生的差距:它为 81.1% 的查询找到金标准,但仅在 35.8% 的情况下成功实现完整优先,而这一差距在条件数量、目标长度、候选密度、查询呈现和四源压力集上持续存在。最后,基于页面的视觉系统仅在 5.1-5.3% 的查询中展示了对每个条件的存储支持。这些结果将条件覆盖而非仅仅是金标准发现确定为主要瓶颈。
cs.CV / 142 / 2607.24184

LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection

LCMamNet:一种轻量级跨尺度Mamba网络用于红外小目标检测
Fan, Yuhao, Hui, Le, Dai, Yuchao
Abstract
Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak targets in infrared imagery usually occupy only a few pixels and are easily submerged by cloud clutter, ground edges, and bright noise, making it difficult for lightweight segmentation-based methods to preserve local target structures while suppressing background interference. To address these challenges, we propose LCMamNet, a lightweight cross-scale Mamba network that progressively enhances local target structures, interacts cross-scale context in a latent space, and restores spatial details with background suppression. Specifically, a compact hierarchical encoder with cross-shaped directional bottleneck residual (CDBR) blocks strengthens direction-sensitive target structures under a small computation budget. A latent dense cross-scale fusion (LDCF) module then performs dense all-level interaction through bidirectional Mamba modeling and reorganizes the interacted features into stable hierarchical semantics. Finally, a progressive decoder selectively recovers shallow spatial details while suppressing irrelevant background textures. Extensive experiments on IRSTD-1k, NUAA-SIRST, and NUDT-SIRST show that the proposed network achieves mIoU scores of 71.25\%, 79.60\%, and 95.58\%, respectively, with only 1.175M parameters and 6.91 GFLOPs. It also runs with a mean inference latency of 6.62 ms, and deployment results on an NVIDIA Jetson Orin NX 16G SUPER further demonstrate its practical potential for real-time edge inference. The code and checkpoints are publicly available at https://github.com/Haoyu096/LCMamNet.
Chinese Translation
红外小目标检测(IRSTD)对于低空感知、无人系统预警和安全监控至关重要。然而,红外图像中的弱目标通常仅占用少量像素,容易被云杂波、地面边缘和明亮噪声淹没,这使得轻量级基于分割的方法在抑制背景干扰的同时难以保持局部目标结构。为了解决这些挑战,我们提出了LCMamNet,这是一种轻量级跨尺度Mamba网络,逐步增强局部目标结构,在潜在空间中进行跨尺度上下文交互,并通过背景抑制恢复空间细节。具体而言,一个紧凑的分层编码器结合交叉形状方向瓶颈残差(CDBR)模块在小计算预算下增强方向敏感的目标结构。随后,潜在密集跨尺度融合(LDCF)模块通过双向Mamba建模执行密集的全级别交互,并将交互特征重新组织为稳定的分层语义。最后,逐步解码器选择性地恢复浅层空间细节,同时抑制无关的背景纹理。在IRSTD-1k、NUAA-SIRST和NUDT-SIRST上的大量实验表明,所提网络分别实现了71.25%、79.60%和95.58%的mIoU分数,仅需1.175M参数和6.91 GFLOPs。它的平均推理延迟为6.62毫秒,并且在NVIDIA Jetson Orin NX 16G SUPER上的部署结果进一步证明了其在实时边缘推理中的实际潜力。代码和检查点已公开发布在https://github.com/Haoyu096/LCMamNet。
cs.CV / 143 / 2607.24194

Face Age Verification Vulnerabilities Under Simple Appearance Manipulations

简单外观操控下的面部年龄验证脆弱性
Sarridis, Ioannis, Kompatsiaris, Ioannis, Papadopoulos, Symeon
Abstract
Online platforms increasingly rely on automated age estimation systems to enforce minimum-age policies. Focusing on vision-based models designed for this task, concerns arise regarding their robustness to simple appearance changes that underage individuals may use to bypass such systems, such as drawing a mustache or applying lipstick. In this work, we present a systematic study of age verification robustness by simulating visual alterations that can be easily achieved by underage individuals. We evaluate seven models, including vision, vision-language, and multimodal large language models, across three datasets and four manipulation types. Interestingly, under drawn beard stubble, up to 61% of True Negatives are flipped into False Positives. Furthermore, we investigate how different demographics are affected by such manipulations, finding that Indians are more affected by beard stubble manipulations, while females are more affected than males across all manipulations. Finally, we explore how these biases can be mitigated using bias mitigation methodologies in lightweight linear probe settings.
Chinese Translation
在线平台越来越依赖自动化年龄估计系统来执行最低年龄政策。针对为此任务设计的基于视觉的模型,出现了对未成年人可能利用简单外观变化绕过这些系统的鲁棒性问题,例如画胡子或涂口红。在本研究中,我们通过模拟未成年人可以轻松实现的视觉变化,系统性地研究了年龄验证的鲁棒性。我们评估了七种模型,包括视觉模型、视觉-语言模型和多模态大型语言模型,涵盖了三个数据集和四种操控类型。有趣的是,在画胡须的情况下,最高可达61%的真实负例被转变为假阳性。此外,我们调查了不同人口统计特征如何受到这些操控的影响,发现印度人受到胡须操控的影响更大,而女性在所有操控中受到的影响均大于男性。最后,我们探讨了如何在轻量级线性探针设置中使用偏见缓解方法来减轻这些偏见。
cs.CV / 144 / 2607.24199

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

推理以规范:交通规则理解的思维链
Luo, Yueru, Yan, Xu, Zhou, Changqing, Yang, Yiming, Zhan, Chao, Mei, Shuqi, Zheng, Chao, Li, Zhen
Abstract
Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversity and context dependence of traffic signage. Importantly, regulation understanding is not a simple recognition task, but a reasoning problem: whether a rule applies depends on interpreting the sign in relation to the spatial layout of lanes and scene context. To support such reasoning, MapDR provide fine-grained annotations that link each traffic sign's regulatory rules to the specific lanes they govern. Existing methods, however, largely treat this as direct sequence prediction, ignoring the underlying reasoning that connects sign semantics and map structure. To address this limitation, we explicitly incorporate reasoning into this task and propose a framework that equips vision-language models (VLMs) with chain-of-thought (CoT) capabilities. We first design a scalable CoT curation pipeline that bootstraps rationales from a strong LLM through a two-round strategy and employs a VLM-based verifier to filter out incorrect cases, yielding a high-quality set of (CoT, answer) pairs. Building on this foundation, we adopt a two-stage training scheme: supervised fine-tuning (SFT) to teach rationale-to-answer generation, followed by GRPO reinforcement learning with answer-grounded, fine-grained rewards to further improve final answer accuracy. Extensive experiments on MapDR show that our approach significantly improves both interpretability and accuracy, establishing the first reasoning-based framework for regulation-aware autonomous driving.
Chinese Translation
理解和遵守交通法规是自动驾驶的安全关键要求,但由于交通标志的多样性和上下文依赖性,这一任务仍然具有挑战性。重要的是,法规理解并不是一个简单的识别任务,而是一个推理问题:规则的适用性取决于将标志与车道的空间布局和场景上下文进行解释。为了支持这种推理,MapDR提供了细粒度的注释,将每个交通标志的监管规则与其所管理的特定车道联系起来。然而,现有方法在很大程度上将其视为直接的序列预测,忽视了连接标志语义和地图结构的潜在推理。为了解决这一局限性,我们明确将推理纳入这一任务,并提出一个框架,使视觉-语言模型(VLMs)具备思维链(CoT)能力。我们首先设计了一个可扩展的CoT整理管道,通过两轮策略从强大的大型语言模型(LLM)中引导出推理,并采用基于VLM的验证器过滤掉不正确的案例,从而生成高质量的(CoT,答案)对。在此基础上,我们采用了两阶段训练方案:监督微调(SFT)用于教授推理到答案的生成,随后通过基于答案的细粒度奖励进行GRPO强化学习,以进一步提高最终答案的准确性。在MapDR上的大量实验表明,我们的方法显著提高了可解释性和准确性,建立了第一个基于推理的法规感知自动驾驶框架。
cs.CV / 145 / 2607.24210

Effect of User-Prompted Priors on Semi-Automated Cancer Lesion Segmentation in Whole-Body Computed Tomography

用户提示先验对全身计算机断层扫描中半自动癌症病灶分割的影响
Stark, Isac, Öfverstedt, Johan, Lundström, Elin, Ekström, Simon, Ahlström, Håkan, Kullberg, Joel
Abstract
In clinical oncology studies, metastatic cancer is commonly evaluated using "Response Evaluation Criteria in Solid Tumors" (RECIST), in which the diameter of up to five lesions is measured and followed over the course of treatment. However, RECIST shows limited correlation with overall survival. Total tumour volume (TTV) is a stronger predictor but typically relies on manual ground-truth segmentation of all lesions, which is time-consuming and requires expert domain knowledge. Semi-automated approaches leveraging user-prompted priors, such as bounding boxes and single-slice contours, as inputs to automated segmentation methods can facilitate the generation of ground-truth segmentations. This work investigates the impact of different user-prompted priors on semi-automated cancer lesion segmentation performance in whole-body computed tomography. Across 3-fold cross-validation and external testing, more complex spatial priors consistently improved performance, with contour priors from three orthogonal planes (axial, coronal and sagittal) achieving the best results. On the external test (n=3865 lesions), this approach achieved a mean Dice score of 0.882, compared to a mean Dice score of 0.671 for the baseline model with no spatial prior. These findings suggest that the use of multi-plane orthogonal user-prompted priors can improve semi-automated tumour lesion segmentation and support efficient generation of high-quality volumetric ground-truth data.
Chinese Translation
在临床肿瘤学研究中,转移性癌症通常使用“实体肿瘤反应评估标准”(Response Evaluation Criteria in Solid Tumors, RECIST)进行评估,其中测量最多五个病灶的直径并在治疗过程中进行随访。然而,RECIST与总体生存率的相关性有限。总体肿瘤体积(Total Tumour Volume, TTV)是更强的预测指标,但通常依赖于对所有病灶的手动真实分割,这既耗时又需要专家领域知识。利用用户提示先验(如边界框和单切片轮廓)作为输入的半自动方法可以促进真实分割的生成。本研究探讨了不同用户提示先验对全身计算机断层扫描中半自动癌症病灶分割性能的影响。在三折交叉验证和外部测试中,更复杂的空间先验始终提高了性能,来自三个正交平面(轴向、冠状和矢状)的轮廓先验取得了最佳结果。在外部测试中(n=3865个病灶),该方法的平均Dice分数为0.882,而没有空间先验的基线模型的平均Dice分数为0.671。这些发现表明,使用多平面正交用户提示先验可以改善半自动肿瘤病灶分割,并支持高质量体积真实数据的高效生成。
cs.CV / 146 / 2607.24215

TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation

TreeAdapter:基于层次分类法的适配器组合用于细粒度物种图像生成
Sun, Yuze, Duan, Zhongjie, Chen, Yingda
Abstract
Although general text-to-image models excel in open-domain generation, their performance degrades significantly in specialized downstream domains, particularly when generating images of rare biological species. Hindered by long-tailed distributions, general models struggle to capture subtle fine-grained details, while per-species fine-tuning methods over-isolate individual species and consequently ignore the shared visual features among closely related taxa. To address this, we propose TreeAdapter, a novel framework that explicitly leverages hierarchical taxonomic data. Rather than using a monolithic model or independent per-species modules, TreeAdapter attaches lightweight adapters to every node of the taxonomic tree. Specifically, leaf-node adapters capture species-specific visual traits, while internal-node adapters encapsulate shared semantics among descendant taxa. We introduce a two-stage training paradigm where ancestor adapters are optimized to model only the residual visual features unexplained by their descendants. This model architecture and training paradigm enable the model to fully leverage hierarchical information, ensuring the accurate generation of visual features for each species. Extensive experiments across three large-scale biodiversity benchmarks demonstrate that TreeAdapter achieves state-of-the-art fine-grained generation quality, outperforming both general-purpose and domain-specific baselines.
Chinese Translation
尽管通用的文本到图像模型在开放领域生成中表现出色,但在专门的下游领域,尤其是在生成稀有生物物种图像时,其性能显著下降。由于长尾分布的影响,通用模型难以捕捉细微的细节,而针对每个物种的微调方法则过于孤立个别物种,从而忽视了密切相关分类群之间的共享视觉特征。为了解决这一问题,我们提出了TreeAdapter,一个明确利用层次分类数据的新框架。TreeAdapter并不是使用单一的模型或独立的每物种模块,而是将轻量级适配器附加到分类树的每个节点上。具体而言,叶节点适配器捕捉物种特有的视觉特征,而内部节点适配器则封装了后代分类群之间的共享语义。我们引入了一个两阶段的训练范式,其中祖先适配器被优化以仅建模其后代未解释的残余视觉特征。这种模型架构和训练范式使模型能够充分利用层次信息,确保为每个物种准确生成视觉特征。在三个大规模生物多样性基准测试中的广泛实验表明,TreeAdapter在细粒度生成质量上达到了最先进的水平,超越了通用和领域特定的基准。
cs.CV / 147 / 2607.24224

MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving

MATS:一种用于自动驾驶3D感知的新型多模态多任务学习框架
Huo, Junchen, Hao, Wanming, Wang, Song, Chen, Enqing, Yang, Shouyi, Wang, Guanghui
Abstract
Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird's-eye-view (BEV) feature map for the joint learning of multiple perception tasks. However, such a single feature map hardly carries sufficient information to simultaneously meet the requirements of various perception tasks, leading to a very limited perception performance. To mitigate this limitation, this paper proposes MATS, a novel multi-modality multi-task learning approach with modality-adaptive BEV fusion and task-specific Mixture-of-Experts (MoE) for 3D perception. Specifically, a simple modality-adaptive BEV fusion module is designed to adaptively recalibrate the BEV features by modeling the global cross-modality dependencies, generating diverse BEV feature maps for various perception tasks. For joint multi-task learning, this paper proposes a task-specific MoE module to decouple the tasks and enable the network to automatically choose the appropriate BEV feature candidates for each specific task. To validate the effectiveness of the proposed approach, we conduct extensive experiments on the large-scale benchmark nuScenes. With the camera- and LiDAR-modality input data, the proposed approach outperforms the state-of-the-art (SOTA) by a significant margin. Furthermore, the experimental results on the single tasks show that the proposed approach significantly outperforms the baselines. The code and trained models will be available upon publication.
Chinese Translation
来自不同传感器的多模态数据提供了丰富的互补信息,对于3D感知而言,成为可靠自动驾驶系统的一个重要组成部分。目前的研究通常设计复杂的融合策略,将多模态数据的信息整合到统一的鸟瞰图(BEV)特征图上,以便共同学习多个感知任务。然而,这样的单一特征图往往难以同时满足各种感知任务的需求,导致感知性能非常有限。为了解决这一限制,本文提出了MATS,一种具有模态自适应BEV融合和任务特定专家混合(Mixture-of-Experts, MoE)的新型多模态多任务学习方法,旨在实现3D感知。具体而言,设计了一个简单的模态自适应BEV融合模块,通过建模全局跨模态依赖关系,自适应地重新校准BEV特征,为各种感知任务生成多样化的BEV特征图。为了实现联合多任务学习,本文提出了一个任务特定的MoE模块,以解耦任务并使网络能够自动选择每个特定任务的适当BEV特征候选。为了验证所提方法的有效性,我们在大规模基准nuScenes上进行了广泛的实验。使用相机和激光雷达模态输入数据,所提方法在性能上显著超越了现有的最先进技术(SOTA)。此外,单任务的实验结果表明,所提方法显著优于基线。代码和训练模型将在发表时提供。
cs.CV / 148 / 2607.24241

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

FilmBench:一个电影级基准用于电影视频生成
Wang, Shengyi, Li, Niantong, Hu, Guangzheng, Qi, Hong, Ding, Fei, Qiao, Weixu, Wang, Jinlin, Lv, Xiaotong, Han, Peng, Li, Zimeng, Ding, Fanshu, Wang, Yushu, Wu, Han, Chen, Jingjing, Wang, Chongxiao, Wu, Yanhao, Huang, Chenglong, Zhu, Xiaoqian, Tian, Jie, Li, Hua, Fan, Jingjing, Tang, Mingshuang, Li, Zhong, Qiang, Hengxia, Chen, Weibin, Zhen, Jinyang, Zhao, Bing, Qu, Lin, Li, Jing, Wei, Hu
Abstract
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
Chinese Translation
视频生成的进展不断缩小了人工智能生成的画面与专业制作的影片之间的视觉差距,但大多数基准仍然从网络来源或大型语言模型(LLM)模板中提取提示,并使用未经训练的通用多模态模型进行评分。更根本的是,它们的评估分类仍然是初步的(整体视觉质量、粗略文本对齐和时间平滑性),而非实际制作和评判电影的专业电影语言标准,因此它们评估的是基本视频的可信度,而非电影级的工艺。我们推出了FilmBench,一个基于电影学院传统的专业电影语言的文本到视频(T2V)和参考到视频(R2V)基准,并与北京电影学院和华景数字媒体与娱乐集团的导演和教师共同开发。该基准基于三个选择。首先,提示是从涵盖20种电影类型的获奖影片片段中反向工程而来,并由专业导演选择,因此每个提示都与经过验证的实拍参考相锚定;这些提示遵循真实的镜头列表,大多数剧本包含多个镜头(1,169个提示中有1,056个为多镜头),与之前的单片段基准不同。其次,评估遵循一个三层次的电影分类法,包括3个轴、12个组成部分和35个(T2V)+3个(仅R2V)子指标。第三,我们开发了一个内部专家级自动评估代理,并开源其核心电影语言操作符(FilmOps)。在对领先的视频生成模型进行基准测试(T2V 9个,R2V 7个)时,评估者在模型级别上重现了人类模型排名,Spearman {ho} = 0.95(T2V)和0.96(R2V)。得分远低于之前的网络风格基准,在动态美学方面存在两个一致的差距,并且在较弱模型中,从单镜头到多镜头的表现下降显著。
cs.CV / 149 / 2607.24249

SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation

SILICA:重新利用扩散先验进行联合玻璃分割和深度估计
R, Tarun, Verma, Anuj, Nanwani, Laksh, Garg, Sourav, Krishna, K. Madhava
Abstract
Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware sensors can detect glass, they lack modularity and have extensive hardware dependencies. Consequently, learning-based monocular depth estimation has emerged as a compelling alternative. However, domain-specific glass-aware monocular depth estimators struggle with unfamiliar indoor layouts; restricted by the severe scarcity of real-world glass depth annotations, they fail to generalize zero-shot to new settings. This motivates us to explore whether the extensive priors of text-to-image diffusion models can enable generalizable perception of transparent surfaces. We introduce SILICA, a unified pipeline leveraging these priors to jointly predict glass segmentation and glass-aware depth. This mutual information exchange establishes a robust spatial hierarchy, entirely eliminating the need for paired real-world glass depth annotations. Subsequently, we use the predicted segmentation mask to explicitly filter incorrect glass depth points from standard sensors, recovering accurate metric glass depth for downstream 3D mapping and autonomous collision avoidance. Supported by our novel Mirage 18k dataset, extensive experiments demonstrate that SILICA achieves remarkable zero-shot transfer across diverse, unseen environments, outperforming state-of-the-art models by almost 20% and setting a new benchmark for transparent surface perception.
Chinese Translation
标准深度传感器在透明表面上系统性失效,导致生成的3D地图受损并造成严重的导航风险。虽然专用硬件传感器可以检测玻璃,但它们缺乏模块化,且有广泛的硬件依赖。因此,基于学习的单目深度估计成为一种引人注目的替代方案。然而,特定领域的玻璃感知单目深度估计器在面对不熟悉的室内布局时表现不佳;由于真实世界玻璃深度注释的严重匮乏,它们无法在新环境中实现零-shot泛化。这促使我们探索文本到图像扩散模型的广泛先验是否能够实现对透明表面的可泛化感知。我们提出了SILICA,一个统一的管道,利用这些先验共同预测玻璃分割和玻璃感知深度。这种互信息交换建立了一个稳健的空间层次结构,完全消除了对配对真实世界玻璃深度注释的需求。随后,我们使用预测的分割掩码显式过滤标准传感器中的错误玻璃深度点,从而恢复准确的度量玻璃深度,以用于下游3D映射和自主避碰。得益于我们新颖的Mirage 18k数据集,广泛的实验表明,SILICA在多样化、未见环境中实现了显著的零-shot迁移,超越了最先进模型近20%,为透明表面感知设定了新的基准。
cs.CV / 150 / 2607.24288

Superpixel-Based QUBO for Scalable Quantum-Enhanced Medical Image Segmentation

基于超像素的 QUBO 用于可扩展的量子增强医学图像分割
Chalhoub, Mohammad, Chehimi, Mahdi, Domingo, Laia, Alhussein, Omar, Farouk, Ahmed, Al-Kuwari, Saif
Abstract
Quadratic unconstrained binary optimization (QUBO) has emerged as a powerful framework for medical computing problems. Binary decision variables naturally represent clinical choices, making QUBO formulations well-suited for quantum annealing hardware. However, a fundamental scalability challenge limits practical deployment: problem size grows rapidly with input dimensionality, creating computational bottlenecks that restrict applications to simplified scenarios. This paper addresses this challenge through hierarchical problem reduction, as demonstrated in medical image segmentation, where pixel-level QUBO formulations create over 65,000 variables for a 256x256 image, forcing existing approaches to downsample to 42x42 resolution and discard 97% of pixel information. A superpixel-based QUBO framework is proposed using simple linear iterative clustering (SLIC) to group pixels into perceptually meaningful regions, then formulate segmentation as QUBO over a region adjacency graph (RAG) combining min-cut and smoothness objectives. Validation on INbreast mammography breast cancer images demonstrates a 4.2% improvement in segmentation quality (mean IoU 0.76 vs 0.73) with 33 computational speedup (0.67s vs 21.97s) and a 97.3% reduction in problem size (1764 to 48 variables), all achieved while processing full-resolution images rather than downsampled versions. The reduced problem size also fits well within current quantum annealer connectivity limits, removing the embedding overhead that has historically blocked direct deployment of pixel-level QUBO segmentation on quantum hardware.
Chinese Translation
二次无约束二进制优化(QUBO)已成为医学计算问题的强大框架。二进制决策变量自然地表示临床选择,使得 QUBO 公式非常适合量子退火硬件。然而,一个根本的可扩展性挑战限制了其实际应用:随着输入维度的增加,问题规模迅速增长,造成计算瓶颈,限制了应用于简化场景。本文通过分层问题简化来解决这一挑战,以医学图像分割为例,像素级 QUBO 公式为 256x256 的图像创建了超过 65,000 个变量,迫使现有方法降采样到 42x42 的分辨率,并丢弃 97% 的像素信息。我们提出了一种基于超像素的 QUBO 框架,使用简单线性迭代聚类(SLIC)将像素分组为感知上有意义的区域,然后在区域邻接图(RAG)上将分割公式化为 QUBO,结合了最小割和光滑性目标。在 INbreast 乳腺癌乳腺摄影图像上的验证显示,分割质量提高了 4.2%(平均 IoU 0.76 对比 0.73),计算速度提升了 33 倍(0.67 秒对比 21.97 秒),问题规模减少了 97.3%(从 1764 个变量减少到 48 个变量),所有这些都是在处理全分辨率图像而非降采样版本的情况下实现的。减少的问题规模也很好地符合当前量子退火器的连接限制,消除了历史上阻碍像素级 QUBO 分割在量子硬件上直接部署的嵌入开销。
cs.CV / 151 / 2607.24298

UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing

UMI3D:通过同时聚焦交叉注意力路由在无约束多图像输入上实现稳健的3D生成
Qu, Zefan, Wang, Zhenwei, Hancke, Gerhard Petrus, Lau, Rynson W. H.
Abstract
Recent 3D foundation models can generate high-quality assets from a single image, but degrade markedly on unconstrained multi-image inputs, often producing distorted geometry, over-smoothed textures, and chaotic colors. We argue that this failure stems not from limited model capacity, but from a mismatch between single-image cross-attention and the multi-image setting: existing models lack a principled way to decide which image each 3D voxel should trust at each denoising step. Revisiting recent single-image 3D foundation models, we show that explicitly routing each voxel to its most informative image is sufficient to unlock strong performance on inconsistent multi-image inputs. Based on this observation, we propose UMI3D, a training-free and plug-and-play framework that restructures cross-attention for unconstrained multi-image 3D generation. Its core, Simultaneous Focus Cross-Attention (SFC-Attn), activates all conditioning images at each denoising step while allowing each voxel to focus on the single image that best explains it. To enable this routing, we derive the Voxel Reference Score (VRS), a model-intrinsic metric for voxel--image affinity that requires no external matching, segmentation, or correspondence models. Extensive experiments show that UMI3D unlocks the multi-image potential of single-image 3D generation frameworks across diverse tasks. Project Page: UMI3D-Project.github.io.
Chinese Translation
近期的3D基础模型能够从单幅图像生成高质量的资产,但在无约束的多图像输入上表现显著下降,常常产生扭曲的几何形状、过于平滑的纹理和混乱的颜色。我们认为,这一失败并非源于模型能力的限制,而是单图像交叉注意力与多图像设置之间的不匹配:现有模型缺乏一种原则性的方法来决定在每个去噪步骤中每个3D体素应信任哪幅图像。通过重新审视近期的单图像3D基础模型,我们展示了显式地将每个体素路由到其最具信息量的图像足以在不一致的多图像输入上解锁强大的性能。基于这一观察,我们提出了UMI3D,一个无训练且即插即用的框架,重构了无约束多图像3D生成的交叉注意力。其核心,称为同时聚焦交叉注意力(Simultaneous Focus Cross-Attention, SFC-Attn),在每个去噪步骤中激活所有条件图像,同时允许每个体素专注于最佳解释其内容的单幅图像。为了实现这一路由,我们推导了体素参考分数(Voxel Reference Score, VRS),这是一种无需外部匹配、分割或对应模型的体素-图像亲和度的模型内在度量。大量实验表明,UMI3D在多种任务中解锁了单图像3D生成框架的多图像潜力。项目页面:UMI3D-Project.github.io。
cs.CV / 152 / 2607.24302

Multiview Multi-Person Human Mesh Recovery Under Large Scenes with Occlusions

在大场景和遮挡条件下的多视角多人的人类网格恢复
Zhang, Qi, Yu, Tao, He, Jiechao, Chan, Antoni B., Huang, Hui
Abstract
Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruction from a single view or single-person reconstruction from multiple views, where the number of subjects and the scene scale are relatively limited. Such settings are insufficient for real-world applications with large scenes and severe inter-person occlusions. To address this limitation, we introduce a large-scale synthetic benchmark for multiview multi-person HMR, termed MVMP-HMR. The proposed dataset contains 15 complex scenes with up to 50 camera views and 30 interacting persons, featuring large spatial coverage and severe occlusions, which significantly increases the difficulty of human mesh recovery. Based on this benchmark, we further propose a multiview multi-person whole-body human mesh recovery model, referred to as MVMP-HMR model. The model first fuses multiview features into a scene-level 3D feature volume, and then leverages pelvis joints predicted by a 3D pose estimation network to extract person-specific queries from the 3D feature volume. These human queries are cross-attended with the 3D feature volume and integrated to decode each person's 3D mesh. Moreover, we introduce two novel losses--the orientation loss and the 3D joint density loss--to alleviate orientation and pose ambiguities under severe occlusions. Experiments demonstrate that existing state-of-the-art HMR methods struggle on the proposed MVMP-HMR benchmark, while our method consistently outperforms prior SOTAs in large-scale scenes with severe occlusions.
Chinese Translation
人类网格恢复(HMR)旨在从图像中恢复3D人类网格。现有的大多数HMR基准和方法要么关注于从单一视角进行多人的重建,要么关注于从多个视角进行单人的重建,其中受试者数量和场景规模相对有限。这种设置对于具有大场景和严重人物遮挡的真实世界应用来说是不够的。为了解决这一局限性,我们引入了一个大规模的合成基准,用于多视角多人的HMR,称为MVMP-HMR。该数据集包含15个复杂场景,最多可提供50个相机视角和30个互动人物,具有广泛的空间覆盖和严重的遮挡,这显著增加了人类网格恢复的难度。基于此基准,我们进一步提出了一种多视角多人的全身人类网格恢复模型,称为MVMP-HMR模型。该模型首先将多视角特征融合为场景级的3D特征体积,然后利用3D姿态估计网络预测的骨盆关节,从3D特征体积中提取特定于个体的查询。这些人类查询与3D特征体积进行交叉注意,并整合以解码每个人的3D网格。此外,我们引入了两种新颖的损失函数——方向损失和3D关节密度损失——以缓解在严重遮挡下的方向和姿态模糊。实验表明,现有的最先进HMR方法在所提出的MVMP-HMR基准上表现不佳,而我们的方法在具有严重遮挡的大规模场景中始终优于之前的最先进技术。
cs.CV / 153 / 2607.24353

PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation

PRISM:通过图像基础自奖励机制进行文本到图像生成的提示优化
Tang, Guo, Luo, HongJie, Wang, Tianxu, Zhang, Ying, Wang, Hao
Abstract
Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to prompt formulation. Existing prompt optimization methods mainly rely on text-side rewriting, prompt expansion, or external reward signals, offering limited image-grounded diagnosis and weak support for learning reusable optimisation policies. In this paper, we propose PRISM, a Prompt Refinement framework via Image-grounded Self-rewarding Mechanism. PRISM closes the prompt-image-feedback loop by interpreting generated images with structured visual diagnosis and scoring them along semantic consistency, aesthetic quality, and human preference alignment. It first initializes a unified VLM through multi-task supervised fine-tuning, and then improves the prompt policy via self-rewarding optimization with a hybrid ideal-point and Chebyshev reward. Extensive experiments show that PRISM improves holistic image quality and fine-grained semantic alignment, while providing interpretable feedback for targeted prompt refinement. The code is available at https://anonymous.4open.science/r/PRISM-FF81.
Chinese Translation
文本到图像生成模型能够根据自然语言描述合成高质量图像,但其性能对提示的构造仍然高度敏感。现有的提示优化方法主要依赖于文本端的重写、提示扩展或外部奖励信号,提供的图像基础诊断有限,并且对学习可重用的优化策略支持不足。本文提出了PRISM,一种通过图像基础自奖励机制进行提示优化的框架。PRISM通过结构化视觉诊断解释生成的图像,并根据语义一致性、美学质量和人类偏好对其进行评分,从而闭合提示-图像-反馈循环。它首先通过多任务监督微调初始化一个统一的视觉语言模型(VLM),然后通过混合理想点和切比雪夫奖励的自奖励优化来改善提示策略。大量实验表明,PRISM提高了整体图像质量和细粒度语义对齐,同时为有针对性的提示优化提供了可解释的反馈。代码可在 https://anonymous.4open.science/r/PRISM-FF81 获取。
cs.CV / 154 / 2607.24359

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

TaoMate:基于锚点的记忆桥接演变和参考状态的实时音视频数字人生成
Gan, Qijun, Zhang, Chenwei, Jin, Meiguang, Ma, Junfeng, Shen, Qiu
Abstract
Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across successive segments. A bounded cache retains local motion and phonetic context but discards older evidence, whereas attending to the complete generated history is computationally expensive and can propagate accumulated errors. We present \method, an anchor-guided persistent-memory framework for few-step joint audio-video generation. The framework preserves an immutable visual anchor, compresses completed video and audio blocks into fixed-capacity dynamic states, and retrieves those states through modality-specific residual attention without extending the active cache. A reference-aware modulation method additionally conditions video features on dynamic and anchor appearance statistics. Anchor-preserving causal-context distillation varies rollout horizon, prefix provenance, and cache-history reliability while keeping the immutable visual anchor unperturbed. By separating persistent memory from stage-local denoising dependencies, \method further admits stage-parallel execution across blocks, accelerating autoregressive inference without pipeline-specific retraining. We evaluate long-form video continuations with appearance, temporal, synchronization, facial, and speech diagnostics. Results show that \method preserves stable appearance across prompt-conditioned segments and strong audio-visual synchronization under autoregressive generation. Our project page is https://taoliveaigc.github.io/TaoMate.
Chinese Translation
实时长篇数字人生成依赖因果模型来扩展音视频内容,同时保持主体外观和音视频在连续片段中的同步。一个有界缓存保留局部运动和语音上下文,但会丢弃较旧的证据,而关注于完整的生成历史在计算上是昂贵的,并可能传播累积的错误。我们提出了 extit{TaoMate},一种基于锚点的持久记忆框架,用于少步联合音视频生成。该框架保持一个不可变的视觉锚点,将完成的视频和音频块压缩为固定容量的动态状态,并通过特定模态的残差注意力检索这些状态,而无需扩展活动缓存。参考感知调制方法进一步使视频特征依赖于动态和锚点外观统计。保持锚点的因果上下文蒸馏在保持不可变视觉锚点不变的同时,变化展开视野、前缀来源和缓存历史可靠性。通过将持久记忆与阶段局部去噪依赖分离, extit{TaoMate}进一步允许在块之间进行阶段并行执行,加速自回归推理而无需特定于管道的再训练。我们评估了具有外观、时间、同步、面部和语音诊断的长篇视频延续。结果表明, extit{TaoMate}在提示条件的片段中保持稳定的外观,并在自回归生成下实现强大的音视频同步。我们的项目页面是 https://taoliveaigc.github.io/TaoMate。
cs.CV / 155 / 2607.24364

HistoGPA: A Context-Conditioned Gene-Prior Attention Framework for Histology-Based Spatial Gene Expression Prediction

HistoGPA:一种基于上下文条件的基因优先注意框架,用于组织学基础的空间基因表达预测
Liu, Ziang, Chen, Xinhai, Feng, Yigui, Li, Shuai, Zhang, Qingyang, Liu, Jie
Abstract
Predicting spatial gene expression from routine hematoxylin and eosin (H&E) images provides a practical complement to experimental spatial transcriptomics. Existing approaches focus on local or multi-scale visual features and often treat pretrained gene representations as fixed priors, although the interpretation of local morphology and the relevance of gene priors depend on tissue context. We propose HistoGPA, a context-conditioned gene-prior attention framework that uses a shared slide-level representation in two parallel pathways: one modulates local morphological features, whereas the other conditions pretrained gene embeddings and retrieves gene-prior information through cross-attention. This design enables each spatial location to retrieve context-adapted gene-prior information using its local morphology, position, and slide context. Across ten cancer types in HEST-1k, HistoGPA achieves the highest macro-averaged gene-wise Pearson correlation coefficient among the compared methods under the same evaluation protocol for both the top-50 and top-1,500 highly variable gene sets. Additional analyses show that HistoGPA better recovers the spatial expression patterns of cancer-associated genes and yields greater agreement between clusters derived independently from predicted and ground-truth expression profiles. Together, these findings motivate a context-dependent view of histology-to-expression prediction, in which local morphological representations and gene priors are jointly adapted to the broader tissue context.
Chinese Translation
从常规的苏木精-伊红(H&E)图像中预测空间基因表达为实验空间转录组学提供了一个实用的补充。现有方法侧重于局部或多尺度视觉特征,并且通常将预训练的基因表征视为固定的先验,尽管局部形态的解释和基因先验的相关性依赖于组织上下文。我们提出了HistoGPA,这是一种基于上下文条件的基因优先注意框架,利用共享的幻灯片级表示在两个平行路径中:一个调节局部形态特征,而另一个则对预训练的基因嵌入进行条件处理,并通过交叉注意检索基因优先信息。这一设计使得每个空间位置能够利用其局部形态、位置和幻灯片上下文检索适应上下文的基因优先信息。在HEST-1k的十种癌症类型中,HistoGPA在相同评估协议下,对于前50个和前1500个高度变异基因集,均实现了比较方法中最高的宏平均基因级皮尔逊相关系数。额外分析表明,HistoGPA更好地恢复了与癌症相关基因的空间表达模式,并在独立于预测和真实表达谱的聚类之间产生了更大的一致性。这些发现共同激励了对组织学到表达预测的上下文依赖视角,其中局部形态表示和基因先验共同适应于更广泛的组织上下文。
cs.CV / 156 / 2607.24384

Ambient pressure compensation and robust position control of oil-filled electric joint systems for underwater manipulators

水下操纵器油填充电动关节系统的环境压力补偿与鲁棒位置控制
Wu, Hongrui, Wang, Songhui, Wang, Xin
Abstract
Electric joint systems are significant elements of an underwater manipulator for its actuation, drive, and control. Working in an underwater environment, the joints suffer huge ambient pressure. To withstand it, the pressure compensation method is usually deployed, whereas the pressurized oil introduces sealing problems as well as parametric uncertainties and unknown disturbances for the dynamic model of the joint. To tackle these issues, this study proposes a design framework for the underwater oil-filled electric joint. The dynamics of the pressure compensation module is analyzed and the structure of the joint is optimized to seal the internal hydraulic oil. An uncertainty dynamic model of the oil-filled joint is established and a robust position controller is designed based on the structured singular value synthesis (mu-synthesis). Experimental results validate the feasibility of the proposed methods.
Chinese Translation
电动关节系统是水下操纵器中用于驱动和控制的重要组成部分。在水下环境中,关节承受巨大的环境压力。为了抵御这种压力,通常采用压力补偿方法,但加压油会引入密封问题以及关节动态模型的参数不确定性和未知干扰。为了解决这些问题,本研究提出了一种水下油填充电动关节的设计框架。分析了压力补偿模块的动态特性,并优化了关节结构以密封内部液压油。建立了油填充关节的不确定性动态模型,并基于结构奇异值综合(mu-synthesis)设计了鲁棒位置控制器。实验结果验证了所提方法的可行性。
cs.CV / 157 / 2607.24403

GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion

GenSplatCodec:通过一步扩散实现前馈高斯点云压缩
Hu, Qiang, Wu, Zhenlong, Huang, Lei, Zheng, Zihan, Zhang, Xiaoyun, Zhang, Wenjun
Abstract
Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which becomes inadequate at low bitrates when high-frequency textures and view-dependent appearance are discarded. Although generative models offer a promising alternative, using them as standalone post-processing decouples generation from the transmitted scene structure, thereby compromising cross-view consistency. To address these limitations, we propose GenSplatCodec, a unified feed-forward Gaussian codec that reformulates low-bitrate Gaussian compression as geometry-guided generative decoding. We present a detail-aware feed-forward Gaussian coding scheme within a dual-stream formulation, where the resulting compact Gaussian structural stream is complemented by a lightweight reference appearance stream. We further introduce a geometry-guided one-step generative decoding approach that jointly exploits decoded structural and appearance cues through hierarchical geometry control to reconstruct high-fidelity and view-consistent novel views. Finally, we develop a three-stage optimization strategy that stabilizes the learning of the unified codec and adapts the generative decoder to codec-derived structural and appearance cues. Extensive experiments across multiple datasets demonstrate that GenSplatCodec consistently achieves superior rate-distortion (RD) performance over existing methods.
Chinese Translation
前馈三维高斯点云(3DGS)能够在无需针对每个场景进行优化的情况下实现可扩展的场景重建,但生成的密集高斯点云在存储和传输上成本较高。现有的前馈高斯压缩方法将解码形式化为确定性表示恢复,这在低比特率下对于高频纹理和视角依赖的外观信息的丢失显得不足。尽管生成模型提供了一个有前景的替代方案,但将其作为独立的后处理会使生成与传输的场景结构脱钩,从而影响跨视图的一致性。为了解决这些限制,我们提出了GenSplatCodec,这是一种统一的前馈高斯编解码器,将低比特率高斯压缩重新表述为几何引导的生成解码。我们在双流框架内提出了一种细节感知的前馈高斯编码方案,其中生成的紧凑高斯结构流由轻量级参考外观流补充。我们进一步引入了一种几何引导的一步生成解码方法,通过层次几何控制共同利用解码的结构和外观线索,以重建高保真且视图一致的新视图。最后,我们开发了一种三阶段优化策略,以稳定统一编解码器的学习,并使生成解码器适应基于编解码器生成的结构和外观线索。针对多个数据集的广泛实验表明,GenSplatCodec在比率失真(RD)性能上始终优于现有方法。
cs.CV / 158 / 2607.24407

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

思维令牌混合:统一自由形式多模态基础的感知与推理
Gao, Tianyi, Fang, Han, Ding, Tianyi, Li, Hao, Wei, Xin, Sun, Hongbo, Dong, Xiaodong, Yuan, Ye, Xu, Jinglin, Liang, Kongming, Sun, Hao, Xin, Jingmin
Abstract
Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing, text-based methods rely on coordinates or index prediction, severely limiting the perceptual capabilities of the model for dense visual objects. Meanwhile, latent token-based methods employ special tokens without inherent spatial references and use a decoding mechanism that lacks thinking steps, weakening high-level reasoning capabilities. Consequently, developing a unified framework that excels in both perception and reasoning remains challenging. To address this, we propose Mixture-of-Thought-Tokens (Motto), a new free-form multimodal grounding method that bridges the perception-reasoning gap, enabling MLLMs to empower diverse, arbitrary grounding queries. Specifically, we introduce Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability. We further design a Context-Adaptive Chain-of-Tokens that dynamically switch grounding modes within an interleaved reasoning chain, achieving robust grounding across tasks of varying complexity. In addition, we construct PR-Bench, a new referring expression comprehension benchmark to evaluate the perception-reasoning gap. Extensive experiments demonstrate that Motto achieves state-of-the-art performance across diverse free-form grounding tasks.
Chinese Translation
多模态大型语言模型在基础任务上取得了重大进展,但现有方法仍然难以统一精确定位与复杂推理。一方面,基于文本的方法依赖于坐标或索引预测,严重限制了模型对密集视觉对象的感知能力。与此同时,基于潜在令牌的方法使用没有固有空间参考的特殊令牌,并采用缺乏思考步骤的解码机制,削弱了高级推理能力。因此,开发一个在感知和推理方面都表现出色的统一框架仍然具有挑战性。为了解决这一问题,我们提出了思维令牌混合(Mixture-of-Thought-Tokens, Motto),这是一种新的自由形式多模态基础方法,旨在弥合感知与推理之间的差距,使多模态大型语言模型(MLLMs)能够支持多样化和任意的基础查询。具体而言,我们引入了空间基础思维令牌化(Spatially-Grounded Thought Tokenization),以明确将特殊令牌与空间位置对齐,从而实现清晰的空间对应关系和视觉可解释性。我们进一步设计了上下文自适应令牌链(Context-Adaptive Chain-of-Tokens),在交错的推理链中动态切换基础模式,实现跨不同复杂性任务的稳健基础。此外,我们构建了PR-Bench,这是一个新的指称表达理解基准,用于评估感知与推理之间的差距。大量实验表明,Motto在多样化的自由形式基础任务上实现了最先进的性能。
cs.CV / 159 / 2607.24409

Accuracy potential of visual localization exploiting high-end street-level imagery

利用高端街景影像的视觉定位精度潜力
Meyer, Jonas, Nebiker, Stephan, Theiler, Pascal, Haala, Norbert
Abstract
Accurate and reliable pose information with respect to a reference frame is increasingly demanded across applications such as autonomous navigation, surveying, robotics, and augmented and mixed reality. Visual localization can serve as a complementary positioning modality to GNSS, whose applicability and accuracy are often limited. Yet, the accuracy potential of visual localization has not been systematically investigated against survey-grade demands. This is mainly due to the lack of publicly available, large-scale outdoor datasets with ground-truth poses in the sub-centimeter range. In this work, we address both gaps. We introduce a scalable visual localization pipeline that employs precisely georeferenced, high-resolution street-level imagery directly as the scene representation. It combines prior-guided reference candidate selection with on-the-fly local Structure-from-Motion reconstruction and PnP-based pose estimation. We further present the FHNW Muttenz dataset, a real-world dataset covering a contiguous 10 km street network mapped in two mobile mapping campaigns approximately 1.5 years apart. It consists of high-resolution reference imagery and query sequences acquired by four different cameras across five representative scenes. All images are precisely co-registered, yielding 6-DoF ground-truth poses in the sub-centimeter range. Using this dataset, we evaluate the accuracy potential of visual localization. Our experiments demonstrate median pose accuracies in the range of 1-5 cm for translation and 0.05-0.1{\deg} for rotation, reaching as low as 1 cm and 0.03{\deg} under favorable conditions. These results show that visual localization can complement survey-grade GNSS positioning, paving the way for 3D geospatial data acquisition using consumer devices and fully automated georeferencing approaches. The dataset is publicly available at: https://fhnw-muttenz-vl-dataset.github.io/.
Chinese Translation
在自主导航、测绘、机器人技术以及增强现实和混合现实等应用中,对相对于参考框架的准确且可靠的位姿信息的需求日益增加。视觉定位可以作为全球导航卫星系统(GNSS)的补充定位方式,但其适用性和准确性往往受到限制。然而,视觉定位的精度潜力尚未针对测量级需求进行系统性的研究。这主要是由于缺乏公开可用的大规模户外数据集,其中包含亚厘米级的真实位姿。在本研究中,我们解决了这两个问题。我们引入了一种可扩展的视觉定位管道,直接利用精确地理参考的高分辨率街景影像作为场景表示。该管道结合了基于先验的参考候选选择、即时的运动重建(Structure-from-Motion)和基于PnP的位姿估计。我们进一步介绍了FHNW Muttenz数据集,这是一个真实世界的数据集,涵盖了一个连续的10公里街道网络,该网络在大约1.5年的两次移动测绘活动中进行了映射。该数据集由四个不同相机在五个代表性场景中获取的高分辨率参考影像和查询序列组成。所有影像都经过精确的配准,提供了亚厘米级的6自由度真实位姿。利用该数据集,我们评估了视觉定位的精度潜力。实验结果表明,平移的中位数位姿精度在1-5厘米范围内,旋转的中位数位姿精度在0.05-0.1°之间,在有利条件下可达到1厘米和0.03°。这些结果表明,视觉定位可以补充测量级GNSS定位,为使用消费设备和完全自动化的地理参考方法进行3D地理空间数据采集铺平道路。该数据集可在以下网址公开获取:https://fhnw-muttenz-vl-dataset.github.io/
cs.CV / 160 / 2607.24422

IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data

IJCB-AFMFR 2026:基于合成训练数据适应基础模型进行人脸识别的竞赛
Chettaoui, Tahar, Ozgur, Guray, Caldeira, Eduarda, Nakvosas, Arturas, Shahreza, Hatef Otroshi, Marcel, Sébastien, Shukla, Rishabh, Takkar, Aditya, Khullar, Rushil, Yadav, Lalak, Gupta, Gourav, Gupta, Anant, Yu, Shiqi, Struc, Vitomir, Damer, Naser, Boutros, Fadi
Abstract
This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2026 International Joint Conference on Biometrics (IJCB 2026). The competition received a total of eight valid submissions from four distinct teams across two complementary tracks: a Full Data Track, in which participants adapt the CLIP ViT-L/14 foundation model using large-scale synthetic identity data, and a Limited Data Track, designed to reflect more resource-constrained adaptation regimes. All training data was generated exclusively using IDPERTURB. Submitted solutions are ranked based on verification and identification performance across a diverse suite of benchmarks, including LFW, CFP-FP, AgeDB-30, CALFW, CPLFW, IJB-B, IJB-C, and TinyFace, using the Borda count method. Fairness evaluation is additionally conducted on the RFW dataset across four demographic groups. The results demonstrate that adaptation of the CLIP foundation model with synthetic training data substantially improves over the off-the-shelf model and, in several cases, surpasses the baseline. Notably, full fine-tuning with Sub-Center ArcFace (DMSTI-Neurotechnology) leads the Full Data Track, while rank-stabilized LoRA adaptation (Idiap-BSP) proves most effective under limited-data conditions.
Chinese Translation
本文总结了在2026年国际联合生物识别会议(IJCB 2026)上举行的基于合成训练数据适应基础模型进行人脸识别的竞赛(AFMFR)。该竞赛共收到来自四个不同团队的八份有效提交,分为两个互补的赛道:全数据赛道,参与者使用大规模合成身份数据对CLIP ViT-L/14基础模型进行适应;有限数据赛道,旨在反映更具资源限制的适应模式。所有训练数据均使用IDPERTURB独立生成。提交的解决方案根据在包括LFW、CFP-FP、AgeDB-30、CALFW、CPLFW、IJB-B、IJB-C和TinyFace等多样化基准上的验证和识别性能进行排名,采用Borda计数法。此外,还在RFW数据集上对四个不同人口群体进行了公平性评估。结果表明,使用合成训练数据对CLIP基础模型的适应显著优于现成模型,并且在多个案例中超越了基线。值得注意的是,使用Sub-Center ArcFace(DMSTI-Neurotechnology)进行的全面微调在全数据赛道中表现最佳,而在有限数据条件下,经过排名稳定的LoRA适应(Idiap-BSP)被证明是最有效的。
cs.CV / 161 / 2607.24424

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

MAViE:一种用于细粒度视觉感知和高效多模态推理的多尺度自适应视觉编码器
Lei, Shaofei
Abstract
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency. We introduce \method, a Multi-scale Adaptive Vision Encoder. \method uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure. It then performs question-conditioned token routing according to question relevance, local information content, global semantics, and spatial coverage, with a token budget that adapts to image complexity. To mitigate compression loss, we further introduce full-to-compressed representation distillation and a spatial diversity regularizer. In an illustrative simulation under a unified 7B language-model framework, \method reduces the average number of SigLIP-SO400M visual tokens from 729 to 146 (approximately 80.0\%) and improves the mean score on VQAv2, GQA, TextVQA, ScienceQA-IMG, and MMBench by 2.2 percentage points, while reducing single-image time to first token from 228\,ms to 129\,ms. We provide the full model design and evaluation protocol. All reported numbers currently serve only as placeholders for paper organization and experimental design; formal claims require real training runs, independent replications, and official benchmark evaluation.
Chinese Translation
视觉-语言模型通常将预训练视觉编码器生成的所有标记投影到大型语言模型中。然而,最终层特征可能会丢弃文本、局部属性和空间关系,同时高分辨率输入显著增加了上下文长度和推理延迟。我们提出了 extit{MAViE},一种多尺度自适应视觉编码器。MAViE使用位置依赖的门控机制融合来自视觉Transformer的浅层、中间层和深层特征,保留全局语义,同时增强边缘、文本和局部结构。然后,它根据问题相关性、局部信息内容、全局语义和空间覆盖进行问题条件的标记路由,标记预算根据图像复杂性进行调整。为了减轻压缩损失,我们进一步引入了全到压缩表示蒸馏和空间多样性正则化器。在一个统一的7B语言模型框架下的示例模拟中,MAViE将SigLIP-SO400M视觉标记的平均数量从729减少到146(约80.0%),并在VQAv2、GQA、TextVQA、ScienceQA-IMG和MMBench上提高了平均得分2.2个百分点,同时将单图像首次标记的时间从228毫秒减少到129毫秒。我们提供了完整的模型设计和评估协议。所有报告的数字目前仅作为论文组织和实验设计的占位符;正式声明需要真实的训练运行、独立的重复实验和官方基准评估。
cs.CV / 162 / 2607.24431

InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting

InterOCF:用于仅依靠摄像头的4D占用预测的时空2D-3D交互
Zhang, Qi, Yu, Xinquan, Zhang, Kaiyi, Huang, Hui
Abstract
Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety. Even though current methods have achieved good performance, the strong spatial-temporal modeling between the input multi-view frames is still underexplored, which limits the performance of those methods in future 4D forecasting. To address this gap, we introduce a novel framework, InterOCF, for 4D occupancy forecasting that jointly models temporal dynamics in both 3D voxel-based representations and multi-view segmentation sequences, while explicitly incorporating feature interaction between the 2D and 3D branches. Our framework incorporates three core components: 1) A 3D Spatio-Temporal (3DST) module that learns volumetric dynamics from historical voxel states to predict future voxel states; 2) A 2D Spatio-Temporal (2DST) module employing an auxiliary multi-view temporal segmentation forecasting task to enhance temporal semantic dynamics; 3) A Spatio-Temporal Interaction Modeling (STIM) module that enables feature interaction between 2D and 3D representations. Experiments on the nuScenes, Lyft-Level5, and nuScenes-Occupancy datasets show that InterOCF consistently outperforms existing baseline approaches.
Chinese Translation
仅依靠摄像头的4D占用预测使得自动驾驶车辆能够仅通过历史多视角图像预测未来的3D语义场景,这对驾驶安全至关重要。尽管当前的方法已取得良好性能,但输入多视角帧之间的强时空建模仍然未得到充分探索,这限制了这些方法在未来4D预测中的表现。为了解决这一问题,我们提出了一种新颖的框架InterOCF,用于4D占用预测,该框架联合建模3D体素表示和多视角分割序列中的时间动态,同时明确地结合2D和3D分支之间的特征交互。我们的框架包含三个核心组件:1)一个3D时空(3DST)模块,从历史体素状态中学习体积动态,以预测未来的体素状态;2)一个2D时空(2DST)模块,采用辅助的多视角时间分割预测任务来增强时间语义动态;3)一个时空交互建模(STIM)模块,使得2D和3D表示之间的特征交互成为可能。在nuScenes、Lyft-Level5和nuScenes-Occupancy数据集上的实验表明,InterOCF始终优于现有的基线方法。
cs.CV / 163 / 2607.24436

MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

MSVS-VAE:用于高保真3D重建的多尺度锚定VecSet
Hao, Dehao, Zhang, Kaiyi, Jia, Tanghui, Gao, Xiangjun, Yan, Dongyu, Chen, Weikai, Hu, Zeyu, Zhu, Lingting, Yin, Yingda, Zhang, Runze, Yuan, Li, Wang, Xin, Quan, Long
Abstract
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. Existing approaches largely follow two paradigms: sparse voxel-based representations achieve strong reconstruction quality but incur significant memory and computational overhead, while set-based representations are compact and continuous yet typically lag in fidelity due to latent sparsity and excessive global smoothness. We propose MSVS-VAE, a hierarchical set-based VAE that closes this fidelity gap without sacrificing compactness. Our key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling. To efficiently decode from the densified hierarchy, we replace global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the exhaustive latent set. We further introduce multi-scale query decoding to fuse coarse-to-fine latent features, where coarse scales provide stable global context, and fine scales refine localized geometry, reducing artifacts from overly local receptive fields. Extensive experiments on Objaverse, ABO, and in-the-wild benchmarks demonstrate that MSVS-VAE consistently outperforms prior set-based and voxel-based VAEs, delivering approximately 10x faster decoding than prior set-based methods and approximately 10x higher compactness than voxel-based baselines.
Chinese Translation
高保真3D生成建模越来越依赖于潜在扩散范式,其中基础3D变分自编码器(VAE)的重建质量成为主要瓶颈。现有方法主要遵循两种范式:稀疏体素基础表示实现了强大的重建质量,但带来了显著的内存和计算开销;而基于集合的表示则紧凑且连续,但通常由于潜在稀疏性和过度的全局平滑性而在保真度上滞后。我们提出了MSVS-VAE,一种分层的基于集合的VAE,旨在不牺牲紧凑性的情况下弥补这一保真度差距。我们的关键思想是通过分层点洗牌上采样逐步密集化锚定的VecSet潜变量,从而增加细粒度几何建模的空间容量。为了高效地从密集化的层次中解码,我们用AVS-Conv替换了全局交叉注意力,这是一种几何感知的局部聚合算子,在局部邻域内操作,而不是在耗时的潜在集合上。我们进一步引入多尺度查询解码,以融合粗到细的潜在特征,其中粗尺度提供稳定的全局上下文,而细尺度则细化局部几何,减少来自过于局部感受野的伪影。在Objaverse、ABO和野外基准上的广泛实验表明,MSVS-VAE始终优于先前的基于集合和体素的VAE,解码速度比先前的基于集合的方法快约10倍,且比体素基线的紧凑性高出约10倍。
cs.CV / 164 / 2607.24440

Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

更大还是更便宜?图像降质下视觉语言模型中不确定性信号的规模与量化效应
Ferdous, M M Asif
Abstract
Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness. A practitioner with a fixed memory budget faces a choice between a small model at full precision, the same small model quantized, and a larger model quantized into the same footprint -- three configurations that push the confidence signal in opposing directions. We measure, on identical inputs, how model scale and 4-bit quantization affect two confidence signals in the Qwen2-VL family: the confidence a model states in natural language, and its own mean token probability over the answer it generates. Across 5,700 predictions spanning six realistic photographic degradations at three severities, we find that scale sharply improves the model's internal uncertainty signal (mean error-detection AUROC 0.80 to 0.98 from 2B to 7B) while its verbalized confidence stays weak and often at chance (mean 0.61 to 0.69): the gap between what the model knows and what it says widens rather than closes with size. We find that 4-bit quantization is nearly free for accuracy (-1.6 points) but expensive for the confidence signal (internal AUROC 0.95 to 0.80, and the verbalized-confidence parse rate collapses from 99% to 64%). For a fixed memory budget the recommendation is therefore to prefer a larger quantized model over a smaller full-precision one: 7B-4bit gives both the best accuracy and the best uncertainty signal (internal AUROC 0.98) of the three configurations that fit. We frame the results as selective-prediction operating points so they translate directly into a deployment recommendation, and we argue that error-detection AUROC, not calibration error, is the metric that exposes the difference between the two signals.
Chinese Translation
部署在消费硬件上的视觉语言模型(VLMs)必须决定何时回答和何时推迟,而这一决策依赖于能够跟踪正确性的信心信号。具有固定内存预算的从业者面临选择:是使用全精度的小模型,还是量化的小模型,或是量化为相同占用空间的更大模型——这三种配置在信心信号上产生相反的影响。我们在相同输入下测量模型规模和4位量化如何影响Qwen2-VL系列中的两个信心信号:模型在自然语言中表述的信心,以及其生成答案的平均标记概率。在涵盖六种现实摄影降质和三种严重程度的5,700个预测中,我们发现规模显著提升了模型的内部不确定性信号(平均错误检测AUROC从2B的0.80提升至7B的0.98),而其口头表述的信心则保持较弱,且常常接近随机(平均值从0.61提升至0.69):模型所知与所言之间的差距随着规模的增大而扩大。我们发现4位量化对准确性几乎没有影响(-1.6分),但对信心信号却代价高昂(内部AUROC从0.95降至0.80,口头信心解析率从99%降至64%)。因此,在固定内存预算下,建议优先选择更大的量化模型,而非较小的全精度模型:7B-4位模型在三种适配配置中提供了最佳的准确性和最佳的不确定性信号(内部AUROC为0.98)。我们将结果框架化为选择性预测操作点,使其可以直接转化为部署建议,并且我们认为错误检测AUROC,而非校准误差,是揭示这两种信号差异的指标。
cs.CV / 165 / 2607.24447

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models

RP-OPSD:针对多模态大语言模型的分辨率特权在线自蒸馏
Zhu, Qihui, Wang, Yuchen, Wen, Zijian, Zhang, Tao, Zhang, Mengjie, Liu, Yang, Chen, Shuangwu, Wu, Siying, Yang, Jian, Jiang, Xiaofeng
Abstract
On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. However, existing methods often rely on verified solution traces, explanations generated by external models, or manually localized visual evidence, which limits their scalable application to multimodal large language models. To address this issue, we exploit the information gap between high- and low-resolution views of the same image and propose RP-OPSD (Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models). During training, the student policy generates on-policy trajectories from images at one-quarter of the original resolution, while the teacher policy provides supervision using the original-resolution images. By minimizing the divergence between their output distributions along the student trajectories, the student learns the predictive behavior of the teacher under high-resolution inputs, thereby strengthening its low-resolution capability and transferring the learned improvement to original-resolution inference. RP-OPSD requires neither additional human annotations nor external models to generate solution traces but only image--question pairs. Experiments on Qwen3.5-9B show that RP-OPSD achieves a 5.45\% relative improvement in average performance at the original resolution and a $1.78\times$ training speedup over OPSD. These results demonstrate that resolution differences can serve as a simple and scalable source of privileged information, providing an effective and efficient approach to on-policy self-distillation for multimodal large language models.
Chinese Translation
在线自蒸馏(On-Policy Self-Distillation, OPSD)利用仅对教师可用的特权信息,为学生生成的轨迹提供密集的令牌级监督。然而,现有方法通常依赖于经过验证的解决方案轨迹、由外部模型生成的解释或手动定位的视觉证据,这限制了它们在多模态大语言模型中的可扩展应用。为了解决这一问题,我们利用同一图像的高分辨率和低分辨率视图之间的信息差距,提出了RP-OPSD(针对多模态大语言模型的分辨率特权在线自蒸馏)。在训练过程中,学生策略从原始分辨率的四分之一的图像生成在线轨迹,而教师策略则使用原始分辨率的图像提供监督。通过最小化学生轨迹上输出分布之间的差异,学生学习在高分辨率输入下教师的预测行为,从而增强其低分辨率能力,并将学习到的改进转移到原始分辨率推理中。RP-OPSD不需要额外的人类标注或外部模型来生成解决方案轨迹,只需图像-问题对。对Qwen3.5-9B的实验表明,RP-OPSD在原始分辨率下实现了5.45%的平均性能相对提升,并且在训练速度上比OPSD快了1.78倍。这些结果表明,分辨率差异可以作为一种简单且可扩展的特权信息来源,为多模态大语言模型的在线自蒸馏提供了一种有效且高效的方法。
cs.CV / 166 / 2607.24453

ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image

ESRVS:基于单个标注图像的极端半监督视网膜血管分割
Xu, Mingzhi, Zhang, Yizhe
Abstract
Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly. We study retinal vessel segmentation in an extreme semi-supervised setting with one annotated image and a pool of unlabeled images. We propose ESRVS, which selects a representative reference image for manual annotation and transfers vessel cues using target-domain-adapted DINOv3 features. ESRVS constructs a multi granular vessel prototype, combines prototype-similarity maps with a physics-inspired prior to generate initial pseudo-labels, and refines the transferred supervision through weighted pseudo-label training and adversarial refinement. Across eight public datasets, ESRVS achieves the best Dice and clDice on six datasets, and the best HD95 on all eight datasets among the compared semi-supervised methods, although those methods use 10 to 20% labeled data. With Mask2Former, ESRVS retains on average 93.7% of fully supervised Dice and 95.1% of fully supervised clDice. These results demonstrate the potential of foundation-model label propagation for highly label-efficient retinal vessel segmentation. Code is available at https://github.com/IAANNH/ESRVS.
Chinese Translation
在医学图像分析中,从最少的人类监督中学习一直是一个长期目标,因为密集的专家标注成本高昂。我们研究在极端半监督环境下的视网膜血管分割,使用一个标注图像和一组未标注图像。我们提出了ESRVS,该方法选择一个代表性的参考图像进行手动标注,并利用目标领域适应的DINOv3特征传递血管线索。ESRVS构建了一个多粒度的血管原型,将原型相似性图与物理启发的先验结合,以生成初始伪标签,并通过加权伪标签训练和对抗性精炼来优化传递的监督。在八个公共数据集上,ESRVS在六个数据集上达到了最佳的Dice和clDice,在所有八个数据集上达到了最佳的HD95,尽管这些对比的半监督方法使用了10%到20%的标注数据。结合Mask2Former,ESRVS平均保留了93.7%的完全监督Dice和95.1%的完全监督clDice。这些结果展示了基础模型标签传播在高效标注的视网膜血管分割中的潜力。代码可在 https://github.com/IAANNH/ESRVS 获取。
cs.CV / 167 / 2607.24465

Rethinking Expert Training for Model Merging with Prompt Learning

重新思考基于提示学习的模型合并专家训练
Georgakilas, Christos, Panariello, Aniello, Moreno, Samir El Karrat, Calderara, Simone, Karatzas, Dimosthenis, van de Weijer, Joost
Abstract
Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approaches largely focus on improving the merging procedure itself and typically assume experts obtained through full-parameter fine-tuning. In this work, we revisit expert training for model merging. We first show that prompt-based adaptation provides a strong baseline: independently learned prompts can be exploited across tasks while keeping the backbone fixed, avoiding the interference introduced by weight merging. Building on this observation, we introduce Dual-Tuned Experts (DTEs), a two-stage training strategy that first learns prompts and then fine-tunes the vision encoder. This reduces the magnitude of task-specific parameter updates and produces experts with higher merge compatibility. Experiments across multiple CLIP architectures, full fine-tuning, and LoRA experts show that DTEs consistently improve merged performance of standard merging approaches and remain effective even when combining heterogeneous sets of experts.
Chinese Translation
模型合并旨在将多个从共享基础模型训练而来的领域专业专家组合成一个单一的多任务模型。现有的方法主要集中在改善合并过程本身,并通常假设专家是通过全参数微调获得的。在本研究中,我们重新审视了模型合并的专家训练。我们首先展示了基于提示的适应提供了一个强有力的基线:独立学习的提示可以在任务之间进行利用,同时保持主干网络不变,从而避免权重合并带来的干扰。在此观察的基础上,我们引入了双调优专家(Dual-Tuned Experts, DTEs),这是一种两阶段的训练策略,首先学习提示,然后微调视觉编码器。这减少了任务特定参数更新的幅度,并产生了具有更高合并兼容性的专家。在多个 CLIP 架构、全微调和 LoRA 专家的实验中,DTEs 一致提高了标准合并方法的合并性能,并且即使在组合异构专家集时仍然有效。
cs.CV / 168 / 2607.24495

NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction

NSL-SLAM:用于实用SLAM和重建的高保真神经结构光深度
Li, Jiaheng, Zhang, Binsheng, Chang, Xinhai, Chen, Wenzheng
Abstract
Structured-light (SL) cameras power depth sensing in millions of devices, and recent neural SL decoding methods have substantially improved their depth quality. SLAM systems can benefit greatly from such strong depth sensing, where reliable geometry enables stable tracking and faithful reconstruction. In this work, we present NSL-SLAM, a practical SLAM system tailored for high-fidelity structured-light depth. We first strengthen SL depth sensing: inspired by the neural structured-light (NSL) method, we further incorporate strong monocular depth priors into the SL stereo decoding, reducing depth RMSE by 35% on Replica-SL compared to NSL. We then build a depth-centric SLAM pipeline with this stronger depth: because structured-light geometry is dense and metrically accurate, we keep it as the primary tracking signal, and add only sparse visual correspondences for geometrically degenerate cases and lightweight bundle adjustment for long-range drift. Our depth estimator and SLAM design reinforce each other: stronger depth makes a simple SLAM pipeline effective, and the depth-centric pipeline ensures this advantage transfers to downstream reconstruction. Experimentally, on the synthetic Replica-SL benchmark, NSL-SLAM achieves the best tracking accuracy and improves reconstruction F-score by 1.6 points over the SOTA baseline under a shared-depth protocol. On a real benchmark of 8 challenging scenes, it is the only method that avoids catastrophic failure on all sequences while achieving 43.3% lower trajectory deviation than selected baselines. The SLAM system runs online at 20.9 FPS, demonstrating that stronger structured-light depth and depth-centric system design together enable practical, robust SLAM.
Chinese Translation
结构光(SL)相机在数百万设备中驱动深度感知,近期的神经结构光解码方法显著提升了其深度质量。SLAM系统可以从这种强大的深度感知中获益匪浅,可靠的几何信息能够实现稳定的跟踪和真实的重建。在本研究中,我们提出了NSL-SLAM,一个为高保真结构光深度量身定制的实用SLAM系统。我们首先增强了SL深度感知:受到神经结构光(NSL)方法的启发,我们进一步将强大的单目深度先验融入SL立体解码中,与NSL相比,在Replica-SL上将深度均方根误差(RMSE)降低了35%。然后,我们基于这种更强的深度构建了一个以深度为中心的SLAM管道:由于结构光几何信息密集且度量准确,我们将其作为主要跟踪信号,仅在几何退化情况下添加稀疏视觉对应关系,并进行轻量级束调整以应对长距离漂移。我们的深度估计器和SLAM设计相辅相成:更强的深度使得简单的SLAM管道有效,而以深度为中心的管道确保这一优势能够传递到下游重建中。实验表明,在合成的Replica-SL基准测试中,NSL-SLAM实现了最佳的跟踪精度,并在共享深度协议下将重建F-score提高了1.6分,相较于当前最优基线(SOTA)。在8个具有挑战性的真实场景基准测试中,它是唯一在所有序列中避免灾难性失败的方法,同时实现了比选定基线低43.3%的轨迹偏差。该SLAM系统以20.9 FPS的速度在线运行,证明了更强的结构光深度和以深度为中心的系统设计共同实现了实用且稳健的SLAM。
cs.CV / 169 / 2607.24516

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

DecoupleMix:可扩展 VLM 数据配方的解耦比例搜索与凸分配
Xie, Jiahao, Guo, Zhongbin, Wang, Qianle, Lu, Ruiqi, Xiao, Dongling, Sun, Wanxuan, Yang, Cheng
Abstract
While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion for admitting new data, while frontier recipes remain undisclosed. We formulate data construction as a systematic mixture-optimization problem and turn it into a reproducible engineering discipline by decoupling the mixture into two orthogonal sub-problems: inter-class ratios across capabilities and intra-class ratios within a category. For inter-class allocation, we use a single-variable iterative search; for intra-class composition, we apply a multidimensional, dataset-level assessment scoring Quality and Difficulty, and formulate selection as a constrained convex optimization with a diversity objective. The DecoupleMix framework delivers two critical capabilities: guiding what data to collect next and rendering dataset validation a controlled, attributable experiment. Experiments show our approach consistently surpasses heuristic baselines. Moreover, optimal ratios discovered on small-scale proxies transfer seamlessly to larger scales without retuning. Using 80B additional multimodal continue-pretraining tokens, our VLM is competitive with strong open-source models trained with substantially larger multimodal budgets.
Chinese Translation
尽管视觉语言模型(VLM)的数据策划活动日益频繁,但构建预训练混合数据的公共实践仍然主要依赖于经验法则:从业者堆叠通过质量筛选的数据集,凭直觉设定跨领域比例,并缺乏一个有原则、可归因的标准来接纳新数据,而前沿配方仍然未公开。我们将数据构建形式化为一个系统的混合优化问题,并通过将混合解耦为两个正交子问题:跨能力的类间比例和类内类别的比例,将其转化为可重复的工程学科。对于类间分配,我们使用单变量迭代搜索;对于类内组成,我们应用多维数据集级别的评估,评分质量和难度,并将选择形式化为具有多样性目标的约束凸优化。DecoupleMix 框架提供了两个关键能力:指导下一步应收集何种数据,并将数据集验证转变为一个可控的、可归因的实验。实验表明,我们的方法始终超越经验基线。此外,在小规模代理上发现的最佳比例能够无缝转移到更大规模,而无需重新调优。使用 80B 个额外的多模态继续预训练标记,我们的 VLM 在与训练预算显著更大的强开源模型竞争时表现出色。
cs.CV / 170 / 2607.24560

EgoPlay: Event-Triggered Video Editing for Egocentric Streams

EgoPlay:基于事件触发的自我中心视频编辑
Mai, Jinjie, Qian, Gordon Guocheng, Menapace, Willi, Sahni, Arpit, Wang, Chaoyang, Mirzaei, Ashkan, Li, Runjia, Tulyakov, Sergey, Ghanem, Bernard, Wonka, Peter, Abdal, Rameen
Abstract
We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form "when X happens, do Y," EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the post-event continuation. Rather than cascading a separate event detector with an editor, EgoPlay learns event recognition, temporal restraint, and pixel-level editing jointly in a single end-to-end model, while also handling negative and multi-event prompts. To support this, we construct a large-scale dataset of 106K event-triggered clip-prompt pairs spanning positive triggers, fabricated-trigger negatives, and multi-event prompts. We then train a bidirectional video diffusion editor with event-triggered supervision and derive a causal variant for chunk-by-chunk streamable inference. We further introduce an event-aware evaluation protocol that separately measures post-trigger editing quality, pre-trigger preservation, and false-trigger robustness. On the Ego4D benchmark, EgoPlay substantially outperforms EgoEdit, the state-of-the-art instruction-based egocentric video editing baseline, with relative gains of 17.7%, 16.9%, and 16.4% in editing quality, visual quality, and background consistency. It also surpasses a VLM-guided detector-editor baseline by 15.7%, 14.5%, and 13.5% on the same metrics, while using less than half the GPU memory.
Chinese Translation
我们介绍了EgoPlay,一种用于自我中心流的事件触发视频到视频编辑器,该编辑器通过在主要基于Ego4D的数据上对预训练的V2V扩散变换器进行微调而获得。给定单目视频和形式为“当X发生时,执行Y”的事件触发提示,EgoPlay推断事件X是否发生以及何时发生,保留事件前的帧,并仅对事件后的延续应用编辑Y。EgoPlay不是将单独的事件检测器与编辑器级联,而是在一个端到端模型中联合学习事件识别、时间约束和像素级编辑,同时处理负面和多事件提示。为此,我们构建了一个大规模数据集,包含106K个事件触发的剪辑-提示对,涵盖正触发、虚构触发的负例和多事件提示。然后,我们训练了一个具有事件触发监督的双向视频扩散编辑器,并推导出一种因果变体以支持逐块流式推理。我们进一步引入了一种事件感知评估协议,分别测量触发后的编辑质量、触发前的保留情况和虚假触发的鲁棒性。在Ego4D基准测试中,EgoPlay在编辑质量、视觉质量和背景一致性方面相较于最先进的基于指令的自我中心视频编辑基线EgoEdit,分别提升了17.7%、16.9%和16.4%。它在相同指标上也超越了VLM引导的检测器-编辑器基线,提升幅度为15.7%、14.5%和13.5%,同时使用的GPU内存不到一半。
cs.CV / 171 / 2607.24570

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

视觉瓶颈:多模态大语言模型在联合时空视频定位中的稀疏帧适应
Zhang, Jiameng, Madikeri, Srikanth
Abstract
Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to sparse inputs of 8 to 16 frames per video. Yet state-of-the-art multimodal large language models (MLLMs) are pretrained on dense sequences of hundreds of frames, creating a fundamental mismatch between training and deployment conditions. This mismatch causes severe performance collapse: the Qwen3-VL 8B model drops from 56.0% to 22.3% temporal mIoU when frames are reduced to 16, a 60.2% relative degradation. We present a systematic empirical study of training strategies to close this gap for spatial-temporal video grounding. Our results suggest that visual feature extraction is the dominant bottleneck under sparse-frame inputs. Adapting only the final three ViT layers, 4% of total parameters, achieves 68.8% temporal mIoU and surpasses a zero-shot 8B model using dense inputs by 12.8 points. Language model fine-tuning, by contrast, offers negligible or negative returns. A boundary-aware sampling strategy, Hybrid16, further improves temporal mIoU by 26 points over uniform sampling when temporal boundaries are available. We conclude that for sparse-frame video grounding, training strategy dominates model scale: a fine-tuned 2B model consistently outperforms a zero-shot 8B model, with or without dense frame access.
Chinese Translation
大型视频平台每小时处理数百万个上传视频,要求有能够定位每个视频中政策违规发生的时间和地点的审核系统。在规模上处理每一帧是不可行的,因此系统被限制为每个视频的稀疏输入,通常为8到16帧。然而,最先进的多模态大语言模型(MLLMs)是在数百帧的密集序列上进行预训练的,这在训练和部署条件之间造成了根本的不匹配。这种不匹配导致了严重的性能崩溃:当帧数减少到16时,Qwen3-VL 8B模型的时间mIoU从56.0%下降到22.3%,相对降幅达到60.2%。我们对缩小这一差距的训练策略进行了系统的实证研究,以实现时空视频定位。我们的结果表明,在稀疏帧输入下,视觉特征提取是主要瓶颈。仅适应最后三层ViT(视觉变换器),占总参数的4%,便实现了68.8%的时间mIoU,并超越了使用密集输入的零-shot 8B模型12.8个百分点。相比之下,语言模型的微调几乎没有或产生负收益。当时间边界可用时,一种边界感知的采样策略Hybrid16进一步使时间mIoU比均匀采样提高了26个百分点。我们得出结论,对于稀疏帧视频定位,训练策略主导模型规模:一个经过微调的2B模型在有或没有密集帧访问的情况下,始终优于一个零-shot 8B模型。
cs.CV / 172 / 2607.24582

CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

CADER:基于信心的动态证据推理用于长视频理解
Yang, Jinlong, Zhang, Wenhao, Lin, Kuanwei, Cheng, Sijie
Abstract
Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example regardless of difficulty. This uniform strategy invokes unnecessary tool-assisted processing for easy questions and provides limited control when difficult questions require fine-grained temporal evidence. We propose CADER (Confidence-Aware Dynamic Evidence Reasoning), a training-free framework for adaptive and reliable long-video reasoning. CADER first performs global reasoning over uniformly sampled frames and estimates answer confidence with a logit-margin signal, allowing high-confidence examples to exit early. For uncertain examples, CADER activates a second-stage tool-augmented loop that combines temporal cropping, lightweight semantic verification, and Relevance-Guided Resampling to progressively localize question-relevant evidence. This design treats tool use as a sample-level decision: a single global pass handles easy cases, while additional reasoning is reserved for examples where uncertainty suggests that more evidence is needed. Experiments on multiple VideoQA benchmarks show that CADER improves long-video reasoning while bypassing Stage~2 for high-confidence samples. Moreover, when applied to a backbone trained only with tool-free chain-of-thought supervision, CADER achieves competitive performance against specialized tool-augmented frameworks, suggesting a practical inference-time route for adaptive long-video reasoning.
Chinese Translation
长视频理解越来越依赖于大型视觉-语言模型和工具增强推理,但大多数系统对每个示例应用相同的推理过程,而不考虑其难度。这种统一策略在简单问题上引发了不必要的工具辅助处理,而在困难问题需要细粒度时间证据时提供了有限的控制。我们提出了CADER(基于信心的动态证据推理),这是一个用于自适应和可靠的长视频推理的无训练框架。CADER首先对均匀采样的帧进行全局推理,并通过logit-margin信号估计答案的信心,从而允许高信心示例提前退出。对于不确定的示例,CADER激活第二阶段的工具增强循环,结合时间裁剪、轻量级语义验证和相关性引导重采样,逐步定位与问题相关的证据。这一设计将工具使用视为样本级决策:单个全局传递处理简单案例,而额外的推理则保留给那些不确定性表明需要更多证据的示例。在多个VideoQA基准上的实验表明,CADER在提高长视频推理的同时,能够为高信心样本跳过第二阶段。此外,当应用于仅通过无工具链式思维监督训练的主干网络时,CADER在与专门的工具增强框架的竞争中表现出色,这表明了一种适用于自适应长视频推理的实用推理时间路径。
cs.CV / 173 / 2607.24591

CameraAnything: Refilming Videos with Arbitrary Camera Control

CameraAnything:具有任意相机控制的视频重新拍摄
Li, Yixuan, Zeng, Yanhong, Cheng, Ka Leong, Zhu, Jiayi, Wang, Hanlin, Wang, Wen, Meng, Yihao, Ouyang, Hao, Wang, Qiuyu, Yu, Yue, Wang, ZiDong, Zhang, Yiyuan, Shen, Yujun, Lin, Dahua
Abstract
We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters. Existing approaches either rely on expensive 3D reconstruction to achieve full camera functionality or restrict editing to extrinsic parameter manipulation. Moreover, the coupled influence of intrinsic and extrinsic parameters on video appearance makes disentangled modeling particularly challenging. To address this, we adopt per-pixel Pl\"ucker ray injection alongside resolution-aware 3D RoPE in self-attention, building both camera conditioning and spatial positional encoding on the target latent to jointly control camera position, focal length, and native resolution editing without cropping or outpainting. To overcome the scarcity of paired training data, we further develop a scalable synthetic pipeline that constructs diverse dynamic scenes through structured multi-camera recording and generates synchronized videos with varied camera configurations. With a tailored orthogonal training strategy, CameraAnything enables expressive video reshooting with arbitrary viewpoint control, focal length adjustment, resolution adaptation, and multi-shot transitions within a single generation process, offering strong practical value for cinematic video editing and cross-platform content adaptation in video production.
Chinese Translation
我们介绍了CameraAnything,这是第一个统一的相机控制视频编辑框架,能够同时控制内在和外在相机参数。现有的方法要么依赖昂贵的3D重建来实现完整的相机功能,要么将编辑限制在外在参数的操控上。此外,内在和外在参数对视频外观的耦合影响使得解耦建模尤其具有挑战性。为了解决这个问题,我们采用了逐像素的Plücker光线注入,并结合了在自注意力中考虑分辨率的3D RoPE,在目标潜变量上构建相机条件和空间位置编码,以共同控制相机位置、焦距和原生分辨率编辑,而无需裁剪或外绘。为了克服配对训练数据的稀缺性,我们进一步开发了一个可扩展的合成管道,通过结构化的多相机录制构建多样的动态场景,并生成具有不同相机配置的同步视频。通过量身定制的正交训练策略,CameraAnything能够在单一生成过程中实现具有表现力的视频重新拍摄,支持任意视角控制、焦距调整、分辨率适应和多镜头过渡,为电影视频编辑和视频制作中的跨平台内容适配提供了强大的实用价值。
cs.CV / 174 / 2607.24598

QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

QueenVIS:通过查询增强重新思考仅图像训练的视频实例分割
Kheirandish, Arian, Ayar, Fardin, Javanmardi, Ehsan, Tsukada, Manabu, Javanmardi, Mahdi
Abstract
Video instance segmentation (VIS) requires models to detect, segment, and track object identities across frames, and most methods enforce temporal consistency through video-level supervision. Image-only training approaches, with MinVIS as one prominent example, have challenged this assumption, reaching competitive VIS without video training by treating frames as independent images and associating instances only at inference. The field has nonetheless moved toward ever more elaborate video-trained trackers, which depend on costly identity-consistent annotations, leaving the image-only direction under-explored. A diagnostic analysis identifies object query quality as the bottleneck: queries trained only to localize objects within a frame drift apart across frames and destabilize association. QueenVIS introduces a query-centric framework for strengthening image-trained VIS. During single-frame training, we enrich Mask2Former queries with two auxiliary heads: a feature-prediction loss that aligns each query with the pooled backbone descriptor of its instance, and a center-prediction loss that injects spatial structure. Both heads are discarded at inference, adding zero parameters, and temporal identity is maintained by a training-free query-propagation and memory-bank scheme. On YouTube-VIS and OVIS with a ResNet-50 backbone, QueenVIS improves over MinVIS, up to +6.7 AP on YouTube-VIS, +4.8 AP on OVIS, and +10.3 AP on the long-sequence YouTube-VIS split. QueenVIS achieves 50.9 AP on YouTube-VIS and remains competitive with recent video-supervised state-of-the-art, without processing a single video clip during training. Our findings suggest that strengthening the discriminative power and temporal stability of object queries is an important, underexplored axis for VIS. Code and models: https://github.com/ArianKheir/QueenVIS
Chinese Translation
视频实例分割(VIS)要求模型在帧之间检测、分割和跟踪对象身份,大多数方法通过视频级监督来强制执行时间一致性。仅图像训练的方法,以 MinVIS 为一个显著例子,挑战了这一假设,通过将帧视为独立图像并仅在推理时关联实例,从而在没有视频训练的情况下实现了具有竞争力的 VIS。尽管如此,该领域仍然朝着越来越复杂的视频训练跟踪器发展,这些跟踪器依赖于昂贵的一致性身份注释,使得仅图像方向的研究仍未得到充分探索。诊断分析表明,目标查询质量是瓶颈:仅训练以定位帧内对象的查询在帧之间漂移,导致关联不稳定。QueenVIS 引入了一种以查询为中心的框架,以增强仅图像训练的 VIS。在单帧训练期间,我们通过两个辅助头增强 Mask2Former 查询:一个特征预测损失,使每个查询与其实例的池化主干描述符对齐,另一个中心预测损失则注入空间结构。这两个头在推理时被丢弃,增加的参数为零,而通过无训练的查询传播和记忆库方案保持时间一致性。在 YouTube-VIS 和 OVIS 上,使用 ResNet-50 主干,QueenVIS 相比 MinVIS 提升了性能,在 YouTube-VIS 上提高了 +6.7 AP,在 OVIS 上提高了 +4.8 AP,在长序列 YouTube-VIS 拆分上提高了 +10.3 AP。QueenVIS 在 YouTube-VIS 上达到了 50.9 AP,并在没有处理任何视频片段的情况下,与最近的视频监督最先进技术保持竞争力。我们的研究结果表明,增强对象查询的区分能力和时间稳定性是 VIS 领域一个重要且未被充分探索的方向。代码和模型可在:https://github.com/ArianKheir/QueenVIS
cs.CV / 175 / 2607.24611

Test-Time Adaptation via Dual Distillation for Videos Under Severe Distribution Shifts

在严重分布变化下通过双重蒸馏进行视频测试时适应
Sacilotti, André, Santos, Samuel Felipe dos, Almeida, Jurandy
Abstract
Deep learning models have achieved state-of-the-art performance in several computer vision tasks. However, they experience severe performance degradation when applied to real-world scenarios due to unanticipated distribution shifts. Test-Time Adaptation (TTA) attempts to solve this problem by using unlabeled data from the target domain to dynamically adapt to the test distribution at inference time, without access to the source data. However, TTA remains a challenging problem when adapting to continuous, temporally correlated data, such as videos, and in scenarios where the target domain contains severe domain shifts. For this reason, few works in the literature explore TTA for videos under such extreme conditions. To overcome these limitations, we propose Test-time Adaptation via Dual Distillation (TADD), an online adaptation framework that relies on a lightweight projection adapter to bridge the domain gap. The adapter module is pre-trained on the source domain and then adapted to the target using our proposed complementary losses: (i) zero-shot distillation, which encourages alignment with the domain-agnostic features from a pre-trained vision-language model (VLM); and (ii) target distillation, which retains the source domain discriminative knowledge encoded in the pre-trained adapter. Built upon a frozen CLIP backbone, our method introduces this lightweight projection adapter as the sole updatable component during inference. We conducted extensive evaluations on three well-known video action recognition benchmarks: UCF-HMDB, Daily-DA, and Sports-DA. Our experiments in the closed-set scenario demonstrate that our method consistently outperforms state-of-the-art TTA baselines. Notably, our TTA approach improves upon previous methods by up to +3.81% on UCF-HMDB, +2.63% on Daily-DA, and +3.03% on Sports-DA.
Chinese Translation
深度学习模型在多个计算机视觉任务中已实现最先进的性能。然而,当应用于现实场景时,由于不可预见的分布变化,它们会经历严重的性能下降。测试时适应(Test-Time Adaptation, TTA)试图通过使用来自目标领域的未标记数据,在推理时动态适应测试分布,而无需访问源数据。然而,当适应于连续的、时间相关的数据(如视频)以及目标领域存在严重领域变化的场景时,TTA仍然是一个具有挑战性的问题。因此,文献中很少有研究探讨在如此极端条件下的视频TTA。为克服这些限制,我们提出了通过双重蒸馏进行测试时适应(Test-time Adaptation via Dual Distillation, TADD),这是一种在线适应框架,依赖于轻量级投影适配器来弥合领域差距。适配器模块在源领域上进行预训练,然后使用我们提出的互补损失进行适应:(i)零样本蒸馏,鼓励与来自预训练视觉-语言模型(Vision-Language Model, VLM)的领域无关特征对齐;(ii)目标蒸馏,保留编码在预训练适配器中的源领域判别知识。在一个冻结的CLIP骨干网络基础上,我们的方法引入了这个轻量级投影适配器作为推理过程中唯一可更新的组件。我们在三个著名的视频动作识别基准上进行了广泛评估:UCF-HMDB、Daily-DA和Sports-DA。我们在闭集场景中的实验表明,我们的方法始终优于最先进的TTA基线。值得注意的是,我们的TTA方法在UCF-HMDB上提高了最多+3.81%,在Daily-DA上提高了+2.63%,在Sports-DA上提高了+3.03%。
cs.CV / 176 / 2607.24651

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

无需坐标或区域标签的视觉文档理解中的证据归因
Liu, Zhuchenyang, Zhang, Yao, Xiao, Yu
Abstract
Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investigates whether this failure is partially limited by what the model can express through coordinates. On a verified bilingual CiteVQA subset, we compare the coordinate interface with a language interface in which the model outputs only text, quoting its evidence verbatim, and a multimodal retriever returns the location of each quote as a page region proposed by a layout parser (tables and figures are quoted through their captions or notes); the comparison is repeated over six open vision-language models. Compared with the coordinate interface, evidence recall rises from at most 8 points to between 26 and 47 and the hallucination rate roughly halves, with little change in answer quality. Building on this comparison, we use the same quote-and-retrieve pipeline as a training scaffold: because region-level evidence labels are expensive to collect for long documents, we introduce a GRPO recipe whose reward is a judge's reading of the gold answer and crops of the retrieved regions, training the model to quote better evidence without any region labels and raising an 8B backbone's strict attributed accuracy from 22.4 to 33.8. These findings indicate a practical path to improve attribution"without a coordinate interface and without costly region-level supervision.
Chinese Translation
可靠的视觉文档理解要求模型将每个答案归因于支持该答案的证据区域。最近的基准测试和系统通过坐标接口表达这一步骤:模型输出标记文档中证据区域的边界框坐标。在这一接口下,视觉-语言模型即使在答案正确的情况下,往往也无法识别正确的区域,这种失败被称为归因幻觉(Attribution Hallucination)。我们进行了一项研究,探讨这种失败是否部分受限于模型通过坐标所能表达的内容。在经过验证的双语CiteVQA子集上,我们比较了坐标接口与一种语言接口,在该接口中,模型仅输出文本,逐字引用其证据,而多模态检索器返回每个引用的页面区域位置,由布局解析器提出(表格和图形通过其标题或注释进行引用);这一比较在六个开放的视觉-语言模型上重复进行。与坐标接口相比,证据召回率从最多8个百分点提高到26到47之间,幻觉率大致减半,而答案质量变化不大。在此比较的基础上,我们使用相同的引用-检索管道作为训练支架:由于长文档的区域级证据标签收集成本高,我们引入了一种GRPO(Gradient Reward Policy Optimization)方法,其奖励是评审对黄金答案的阅读和检索区域的裁剪,训练模型在没有任何区域标签的情况下更好地引用证据,并将一个8B主干的严格归因准确率从22.4提高到33.8。这些发现表明了一条在没有坐标接口和昂贵的区域级监督的情况下改善归因的实用路径。
cs.CV / 177 / 2607.24665

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

MMOE:通过高效专家设计现代化扩散变换器
Jia, Yanhao, Wang, Jiepeng, Huang, Haibin, Zhang, Chi, Cambria, Erik, Li, Xuelong
Abstract
Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly enlarge total parameter counts and sparsity ratios without importing the efficiency mechanisms that made LLM scaling practical, so generation quality is seldom balanced against training and deployment cost. This raises a natural question: can the architectural principles behind efficient LLM scaling be adapted to AFMs in a more balanced way? We introduce ModernMOE (MMOE), a modernization of SiT-style diffusion transformers that systematically adapts routed experts, shared and lightweight experts, gate-residual routing, and attention-residual information reuse to AIGC generation. Rather than treating MoE as a single plug-in replacement, MMOE studies how different modern expert components affect convergence, efficiency, and generation quality when composed inside a diffusion transformer. Every experiment in this paper is trained on a single eight-GPU H100 node with batch size 256 for 400k steps, an accessible single-machine budget. Under matched training and sampling protocols and at this budget, MMOE reaches lower FID at every recorded checkpoint, that is, it converges faster per training step, than dense and intermediate sparse-expert baselines, and among the sparse variants it attains the best quality-cost balance. Routing analysis further shows stable expert specialization across depth, substantial use of lightweight routes, and modest step-to-step routing changes during denoising. These results suggest that AFMs can follow the balanced scaling path of LLMs by importing proven efficiency designs, rather than by simply increasing total parameters and sparsity ratios.
Chinese Translation
现代大型语言模型通过将容量增长与效率相结合,成功实现了规模化,在容量增长的同时控制每个标记和部署成本。AIGC基础模型(AFMs),尤其是扩散变换器骨干,已经开始采用稀疏专家,但最近的努力大多是增加总参数数量和稀疏比,而没有引入使大型语言模型(LLM)规模化变得可行的效率机制,因此生成质量往往无法与训练和部署成本相平衡。这引发了一个自然的问题:高效LLM规模化背后的架构原则能否以更平衡的方式适应AFMs?我们介绍了ModernMOE(MMOE),这是对SiT风格扩散变换器的现代化,系统性地将路由专家、共享和轻量级专家、门控残差路由以及注意力残差信息重用适应于AIGC生成。MMOE并不是将MoE视为单一的插件替代品,而是研究不同现代专家组件在扩散变换器内部组合时如何影响收敛、效率和生成质量。本文中的每个实验均在单个八GPU H100节点上训练,批量大小为256,训练400k步,符合可访问的单机预算。在匹配的训练和采样协议下,在此预算下,MMOE在每个记录的检查点都达到了更低的FID,即在每个训练步骤上收敛速度快于稠密和中间稀疏专家基线,并且在稀疏变体中实现了最佳的质量-成本平衡。路由分析进一步显示了在深度上的稳定专家专业化、轻量级路由的显著使用,以及在去噪过程中适度的逐步路由变化。这些结果表明,AFMs可以通过引入经过验证的效率设计,遵循LLMs的平衡规模化路径,而不仅仅是简单地增加总参数和稀疏比。
cs.CV / 178 / 2607.24683

Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification

多模态分类中缺失任意模态的协同学习
Mena, Francisco, Ienco, Dino, Interdonato, Roberto, Dantas, Cassio F., Besnard, Simon
Abstract
Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to inconsistent modality availability between training and inference times. To handle missing modalities, prior studies have mainly covered bimodal data setups and focused on designing robust fusion processes. Instead, we adopt a multi-modal co-learning framework that prioritizes inter-modal collaboration rather than multi-modal fusion. Specifically, we consider that any subset of modalities may be absent, without assuming predefined missing-modality patterns, an inference scenario we refer to as missing arbitrary modalities. To address this challenge, we introduce two alternative approaches that leverage information at both feature- and decision-level. Experiments on two multi-modal classification benchmarks demonstrate significant robustness gains in various missing modality conditions. The first method shows more robust behavior under minimal missing conditions, where a single modality is absent, whereas the second performs better under extreme missing conditions, where all-but-one modalities are missing. Our code is available at https://github.com/fmenat/Co4Miss.
Chinese Translation
多模态分类利用来自多种数据源的互补信息来增强预测性能。然而,现实场景中受到操作限制(如传感器故障或隐私限制)导致训练和推理时模态可用性不一致。为了处理缺失模态,以往的研究主要集中于双模态数据设置,并专注于设计稳健的融合过程。相反,我们采用了一种多模态协同学习框架,优先考虑模态间的协作而非多模态的融合。具体而言,我们考虑到任何模态子集可能缺失,而不假设预定义的缺失模态模式,这种推理场景我们称之为缺失任意模态。为了解决这一挑战,我们提出了两种替代方法,利用特征级和决策级的信息。对两个多模态分类基准的实验表明,在各种缺失模态条件下显著提高了稳健性。第一种方法在缺失条件较轻的情况下表现出更强的稳健性,即仅缺失单一模态,而第二种方法在极端缺失条件下表现更佳,即所有模态中仅缺失一个。我们的代码可在 https://github.com/fmenat/Co4Miss 获取。
cs.CV / 179 / 2607.24701

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

用于模态缺失的 RGBT 跟踪的时空条件去噪变换器
Lu, Andong, Zha, Ziyi, Jin, Jiandong, Li, Shihao, Li, Chenglong, Tang, Jin, Luo, Bin
Abstract
Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be unsatisfactory. In addition, current approaches exhibit limited flexibility in processing both missing and complete data. To overcome these limitations, we propose a Spatio-temporal Conditional Denoising Transformer (SCDT), which integrates the spatial cues and the temporal context to adaptively perform information reconstruction of missing modalities and feature enhancement of weak modalities in a unified framework, for robust modality-missing RGBT tracking. In particular, SCDT leverages the short-term temporal cues from recent historical frames to capture the fine-grained temporal correlations and the long-term temporal cues encoding modality evolution to capture the global context. By jointly exploiting long short-term temporal contexts as the conditions, SCDT progressively guides noisy features of available modalities to learn reliable and temporally consistent multimodal representations. Furthermore, SCDT introduces a noisemodulated adaptation mechanism that dynamically adjusts its behavior according to the modal availability, enabling a single framework to unify feature learning under both modality-missing and complete scenarios without changing the architecture or parameters. Extensive experiments on three public benchmark datasets demonstrate that our method consistently outperforms state-of-the-art methods. The code is available here.
Chinese Translation
在 RGBT 跟踪中,模态缺失通常导致不完整和不稳定的多模态特征表示,从而严重降低性能。现有方法通常尝试从可用模态中恢复缺失模态,但在具有挑战性的场景中生成的数据质量可能不尽如人意。此外,目前的方法在处理缺失数据和完整数据方面灵活性有限。为克服这些局限性,我们提出了一种时空条件去噪变换器(SCDT),该方法整合了空间线索和时间上下文,以自适应地在统一框架中执行缺失模态的信息重建和弱模态的特征增强,从而实现稳健的模态缺失 RGBT 跟踪。具体而言,SCDT 利用来自最近历史帧的短期时间线索捕捉细粒度的时间相关性,并利用编码模态演变的长期时间线索捕捉全局上下文。通过共同利用长期和短期时间上下文作为条件,SCDT 逐步引导可用模态的噪声特征学习可靠且时间一致的多模态表示。此外,SCDT 引入了一种噪声调制适应机制,根据模态可用性动态调整其行为,使得单一框架能够在模态缺失和完整场景下统一特征学习,而无需改变架构或参数。在三个公共基准数据集上的大量实验表明,我们的方法始终优于最先进的方法。代码可在此处获取。
cs.CV / 180 / 2607.24703

Panda: Unsupervised Pelvic Anomaly Detection for Real-Time MR Imaging

Panda:用于实时磁共振成像的无监督盆腔异常检测
Knupfer, Anika, Lindholz, Maximilian, Müller, Johanna Paula, Verdera, Jordina Aviles, Tripathy, Smiti, Schulz-Heise, Susanne, Hutter, Jana
Abstract
Female pelvic diseases remain an under researched area characterized by often delayed diagnosis. While pelvic MRI offers superior soft-tissue contrast for diagnosis and image-guided procedures, real-time anomaly detection remains challenging due to physiological motion, tissue deformation, and instrument artifacts. Existing supervised approaches are impractical, as adverse events are rare, heterogeneous, and difficult to annotate. We present a Dinomaly-based unsupervised anomaly detection framework adapted for pelvic MRI that learns normative representations from healthy cases and flags deviations without requiring labels. Our approach leverages a frozen DINOv3 Vision Transformer encoder combined with a noisy MLP bottleneck and Linear Attention decoder to prevent identity mapping while maintaining computational efficiency. Anomalies are localized via per-token cosine distance between encoder and decoder representations, yielding spatial anomaly maps that provide immediate feedback at the scanner to support radiologist decision-making and adaptive protocol adjustment. Evaluated on a curated subset of the Uterine Myoma Dataset, the framework achieves a pixel-level AUROC of 88.06% and high specificity (95.45%) at frame level at 40.5 slices/s, meeting real-time clinical deployment requirements. The spatial anomaly maps and frame-level scores provide immediate, localized feedback at the scanner to support radiologist decision-making and adaptive protocol adjustment during active procedures.
Chinese Translation
女性盆腔疾病仍然是一个研究不足的领域,常常面临诊断延迟的问题。尽管盆腔磁共振成像(MRI)在诊断和图像引导程序中提供了优越的软组织对比度,但由于生理运动、组织变形和仪器伪影,实时异常检测仍然具有挑战性。现有的监督方法不切实际,因为不良事件稀少、异质且难以标注。我们提出了一种基于Dinomaly的无监督异常检测框架,适用于盆腔MRI,该框架从健康病例中学习规范表示,并在不需要标签的情况下标记偏差。我们的方法利用了一个冻结的DINOv3视觉变换器编码器,结合了一个噪声多层感知器(MLP)瓶颈和线性注意力解码器,以防止身份映射,同时保持计算效率。通过编码器和解码器表示之间的每个标记余弦距离来定位异常,生成空间异常图,为放射科医生提供即时反馈,以支持决策和自适应协议调整。在经过精心挑选的子集Uterine Myoma Dataset上进行评估,该框架在40.5切片/秒的帧级别上实现了88.06%的像素级AUROC和95.45%的高特异性,满足实时临床部署的要求。空间异常图和帧级分数在扫描仪上提供即时、局部的反馈,以支持放射科医生在主动程序中的决策和自适应协议调整。
cs.CV / 181 / 2607.24706

SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations

SADe:用于弱支持注释的少样本分割的稀疏原子支持去污染
Xing, Hang, Liu, Guangjun, Xia, Yan, Ding, Xueming
Abstract
Few-shot segmentation (FSS) commonly assumes clean pixel-level support masks, yet practical support supervision often uses boxes, scribbles, coarse masks, or pseudo-masks. These weak annotations may include texture-similar distractors and background context alongside the target, contaminating class prototypes or visual prompts before query prediction. We introduce SADe, a predictor-agnostic support decontamination layer that estimates the reliability of selected support patches without query information. Central to SADe is sparse autoencoder (SAE) atom evidence: dense similarity may respond to both target and texture-similar context, whereas contrasting atom activations inside and outside the weak-support region provides factor-level reliability cues. A lightweight router combines atom evidence with dense similarity and episode statistics to predict patch reliability and generate a cleaned support mask. Trained once on synthetic weak-support episodes from FSS-1000, the router is frozen for all target evaluations. The resulting mask supports standalone prediction or can be supplied to heterogeneous FSS models through native support interfaces without altering query-side inference. Under a matched weak-support protocol, SADe achieves the highest query mIoU in six of nine standalone prompt-shot combinations. With the same ProMi query head, it is within 0.03 mIoU of SAM3-derived masks under tight boxes and surpasses them by 11.17 and 19.49 points under box-r2 and box-r4, respectively. As a plug-in, SADe improves over raw support in 70 of 72 matched box-family comparisons across four frozen downstream models and two datasets. On point and scribble prompts, its average performance remains close to the corresponding raw-support baseline. Ablations and atom-removal controls show that atom evidence contributes reliability information beyond dense similarity.
Chinese Translation
少样本分割(FSS)通常假设干净的像素级支持掩码,然而实际的支持监督往往使用框、涂鸦、粗略掩码或伪掩码。这些弱注释可能包含与目标相似的纹理干扰物和背景上下文,从而在查询预测之前污染类原型或视觉提示。我们提出了SADe,一种与预测器无关的支持去污染层,它在没有查询信息的情况下评估所选支持补丁的可靠性。SADe的核心是稀疏自编码器(SAE)原子证据:密集相似性可能对目标和与纹理相似的上下文都有反应,而弱支持区域内外的对比原子激活提供了因子级的可靠性线索。一个轻量级路由器将原子证据与密集相似性和情节统计结合,以预测补丁的可靠性并生成清洁的支持掩码。该路由器在FSS-1000的合成弱支持情节上训练一次后被冻结,用于所有目标评估。生成的掩码支持独立预测,或可以通过本地支持接口提供给异构FSS模型,而无需改变查询侧推理。在匹配的弱支持协议下,SADe在九种独立提示-样本组合中的六种中实现了最高的查询mIoU。在相同的ProMi查询头下,它在紧框下与SAM3派生掩码相差仅0.03 mIoU,而在box-r2和box-r4下分别超越它们11.17和19.49点。作为一个插件,SADe在72个匹配框族比较中的70个上优于原始支持,跨越四个冻结的下游模型和两个数据集。在点和涂鸦提示下,其平均性能仍接近相应的原始支持基线。消融实验和原子移除控制显示,原子证据提供了超越密集相似性的可靠性信息。
cs.CV / 182 / 2607.24721

DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement

DreamStyle3D:通过双重注意力解耦实现高效的3D风格化资产生成
Wang, Kai, Ouyang, Ziheng, Zhang, Xuying, Cheng, Ming-Ming, Hou, Qibin
Abstract
With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. However, existing approaches still struggle to jointly preserve style fidelity, geometric consistency, and generation efficiency, as most of them still rely on indirect 2D-to-3D stylization pipelines. This motivates a native 3D stylization framework that can explicitly disentangle style from geometry while remaining efficient. To this end, we propose DreamStyle3D, an efficient framework for stylized 3D asset generation built on a Decoupled Dual Cross-Attention mechanism. Our method explicitly separates geometric and stylistic features to enable efficient style injection while preserving structural consistency, and further adopts a lightweight training strategy to enhance style consistency and model generalization. In addition, we build an automated data pipeline and construct a dataset of about 15K content-style-stylized triplets for training and evaluation. Extensive experiments demonstrate that our DreamStyle3D can generate high-fidelity, geometrically consistent stylized 3D assets within 10 seconds, substantially improving efficiency while maintaining superior style quality and offering a new solution for 3D content creation. The code and data are available at https://github.com/HVision-NKU/DreamStyle3D.
Chinese Translation
随着游戏、动画和虚拟现实产业的快速发展,对高效生成风格化3D资产的需求也在迅速增加。然而,现有方法在共同保持风格保真度、几何一致性和生成效率方面仍然面临挑战,因为大多数方法仍依赖于间接的2D到3D风格化管道。这促使我们开发一种本地3D风格化框架,能够明确地将风格与几何解耦,同时保持高效。为此,我们提出了DreamStyle3D,这是一个基于解耦双重交叉注意力机制的高效风格化3D资产生成框架。我们的方法明确分离几何特征和风格特征,以实现高效的风格注入,同时保持结构一致性,并进一步采用轻量级训练策略以增强风格一致性和模型泛化能力。此外,我们构建了一个自动化数据管道,并构建了一个包含约15K内容-风格-风格化三元组的数据集用于训练和评估。大量实验表明,我们的DreamStyle3D能够在10秒内生成高保真、几何一致的风格化3D资产,显著提高了效率,同时保持了优越的风格质量,为3D内容创作提供了一种新的解决方案。代码和数据可在 https://github.com/HVision-NKU/DreamStyle3D 获取。
cs.CV / 183 / 2607.24727

Infrared Imaging Empowered by Artificial Intelligence for Pediatric Skeletal Triage: A Narrative Review and Future Perspectives

人工智能赋能的红外成像在儿童骨骼分诊中的应用:叙述性综述与未来展望
Amiri, Sajad, Afshar, Pardis, Anjomshoa, Elham
Abstract
Background. Pediatric musculoskeletal trauma represents up to 18% of pediatric ED visits, yet diagnosis still depends on ionizing radiography. Cumulative low-dose radiation in early life raises lifetime leukemia and brain malignancy risk, motivating radiation-free triage alternatives. Objective. To synthesize evidence for a hybrid framework coupling broad-spectrum infrared (IR) imaging with deep-learning cross-modal translation to generate clinically interpretable synthetic-radiograph reconstructions from non-ionizing data. Approach. We review five IR spectral windows spanning 650 nm to 1 mm - NIR-I, NIR-II, SWIR, MIR/LWIR, and THz - and how dual-geometry (transmission/reflection) acquisition exploits wavelength-specific tissue depth and biochemical sensitivity. We summarize image-to-image translation networks (Pix2Pix, CycleGAN, Swin-Unet) and feature-matching algorithms (SuperPoint, SuperGlue, ALIKED, LightGlue) used to align and fuse IR data into radiograph-equivalent reconstructions. Implications. Pediatric anatomy - smaller cross-sections, thinner cortical bone - favors IR penetration, enabling compact, portable, non-ionizing triage hardware. Feasibility is grounded in fNIRS and transcranial photobiomodulation evidence: near-infrared light passes through skin, skull, and cortex with sufficient signal for hemodynamic monitoring - a longer, more attenuating path than through a pediatric forearm or distal leg. Key barriers: paired IR/X-ray dataset construction, AI-as-medical-device regulatory pathways, generalization across body habitus and skin pigmentation, and acquisition-protocol standardization. Conclusions. Integrated multi-spectral IR+AI imaging is a promising radiation-free complement to pediatric skeletal radiography. Progress requires multi-center paired datasets, externally validated models, and IR source safety qualification under IEC 60825-1.
Chinese Translation
背景:儿童肌肉骨骼创伤占儿童急诊就诊的18%,但诊断仍依赖于电离辐射成像。早期生活中累积的低剂量辐射增加了终生白血病和脑恶性肿瘤的风险,因此迫切需要无辐射的分诊替代方案。目标:综合证据,构建一个将宽谱红外(IR)成像与深度学习跨模态转换相结合的混合框架,从非电离数据生成临床可解释的合成放射影像重建。方法:我们回顾了五个红外光谱窗口,涵盖650纳米至1毫米的范围——近红外I(NIR-I)、近红外II(NIR-II)、短波红外(SWIR)、中波/长波红外(MIR/LWIR)和太赫兹(THz),以及双几何(透射/反射)采集如何利用特定波长的组织深度和生化敏感性。我们总结了用于将红外数据对齐和融合为放射影像等效重建的图像到图像转换网络(Pix2Pix、CycleGAN、Swin-Unet)和特征匹配算法(SuperPoint、SuperGlue、ALIKED、LightGlue)。影响:儿童解剖结构——较小的横截面、较薄的皮质骨——有利于红外穿透,使得紧凑、便携的非电离分诊硬件成为可能。可行性基于功能性近红外光谱(fNIRS)和经颅光生物调制的证据:近红外光能够穿透皮肤、颅骨和皮层,提供足够的信号进行血流动力学监测——这一路径比通过儿童前臂或远端腿部的路径更长且衰减更大。主要障碍:配对的红外/X射线数据集构建、作为医疗设备的人工智能的监管路径、跨体型和皮肤色素的泛化,以及采集协议的标准化。结论:集成多光谱红外+人工智能成像是儿童骨骼放射成像的有前景的无辐射补充。进展需要多中心配对数据集、外部验证模型以及根据IEC 60825-1的红外源安全资格认证。
cs.CV / 184 / 2607.24729

MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale

MicroZoom:极大尺度下的结构保留细节合成
Huynh, Huy, Ma, Jingwei, Curless, Brian, Kemelmacher-Shlizerman, Ira, Seitz, Steven M.
Abstract
We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of consumer-grade microscope close-ups, MicroZoom synthesizes a seamless, gigapixel-resolution image grounded in the material character of the real references, enabling exploratory visualization of microscopic texture across the full spatial extent of an object. Our goal is plausible synthesis, not exact reconstruction. We focus on full-image, reference-based, extreme-scale super-resolution at magnification levels of up to 350x, a setting that introduces two major challenges: (1) recovering texture-specific detail from highly lossy inputs near ambiguous material boundaries, and (2) preserving correct large-scale pattern structure, such as the repeating geometry of a fabric weave, across millions of local predictions. We address these with a two-stage cascaded design, where the first stage recovers global pattern coherence and the second refines local texture detail, supplemented by a segmentation mask to guide synthesis at ambiguous boundaries. We verify our approach on a collection of self-captured everyday objects and demonstrate globally coherent, materially grounded gigapixel imagery.
Chinese Translation
我们提出了MicroZoom,一个用于微观尺度下千兆像素图像合成的生成框架。给定一张标准照片和一组稀疏的消费级显微镜特写,MicroZoom合成了一幅无缝的千兆像素分辨率图像,基于真实参考的材料特性,使得能够在物体的整个空间范围内探索微观纹理的可视化。我们的目标是实现可信的合成,而非精确重建。我们专注于基于参考的全图极大尺度超分辨率,放大倍数高达350倍,这一设置带来了两个主要挑战:(1)从模糊材料边界附近的高度损失输入中恢复特定纹理细节,以及(2)在数百万个局部预测中保持正确的大尺度模式结构,例如织物编织的重复几何形状。我们通过一个两阶段级联设计来解决这些问题,其中第一阶段恢复全局模式一致性,第二阶段细化局部纹理细节,并辅以分割掩膜以指导在模糊边界处的合成。我们在一组自捕获的日常物体上验证了我们的方法,并展示了全球一致、基于材料的千兆像素图像。
cs.CV / 185 / 2607.24730

KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

KANEx:将Kolmogorov-Arnold网络的可解释性转化为医学可解释性
Shailya, Krithi, Ravi, Ananya Lakshmi, V., Venkatanathan K., Sundaram, Sowmya S., Krishnan, Gokul S., Anand, Aditi, Ravindran, Balaraman
Abstract
Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In clinical workflows, chest X-ray classifiers are increasingly paired with Vision-Language Models (VLMs) to generate natural-language explanations. However, these systems add linguistic fluency without addressing the underlying opacity of the visual model. With the emergence of Kolmogorov-Arnold Networks (KANs), whose spline-based components provide inherently interpretable functional units, we investigate whether this architectural transparency can be leveraged to produce more trustworthy textual explanations. We introduce KANEx, the first ever framework that leverages the symbolic transparency of KANs to ground VLM reasoning. This interpretability also made it possible to design KAN-Map, a novel heatmap generation method derived directly from KAN models rather than gradient approximations. We feed these grounded contexts into downstream VLMs for enhanced explainability. Benchmarked on the MIMIC-CXR dataset, we demonstrate that KAN-based architectures with ResNet/ViT baselines demonstrate improved semantic similarity while producing significantly more faithful saliency maps. KAN architectures improve visual localization and downstream reasoning quality by 10%. Our findings suggest that grounding linguistic explanations and visual attributions in mathematically interpretable units is a necessary step toward trustworthy medical AI.
Chinese Translation
计算机视觉模型在医学应用中已变得极为有效,但其黑箱特性仍然削弱了临床医生的信任。在临床工作流程中,胸部X光分类器越来越多地与视觉-语言模型(VLMs)配对,以生成自然语言解释。然而,这些系统虽然增加了语言流畅性,却未能解决视觉模型的内在不透明性。随着Kolmogorov-Arnold网络(KANs)的出现,其基于样条的组件提供了固有的可解释功能单元,我们探讨这种架构透明性是否可以被利用,以生成更可信的文本解释。我们提出了KANEx,这是第一个利用KAN的符号透明性来支撑VLM推理的框架。这种可解释性还使得设计KAN-Map成为可能,这是一种直接从KAN模型而非梯度近似生成的新型热图生成方法。我们将这些有依据的上下文输入到下游VLM中,以增强可解释性。在MIMIC-CXR数据集上的基准测试中,我们证明了基于KAN的架构与ResNet/ViT基线相比,表现出更高的语义相似性,同时生成了显著更真实的显著性图。KAN架构在视觉定位和下游推理质量上提高了10%。我们的研究结果表明,将语言解释和视觉归因基于数学可解释单元是迈向可信医疗人工智能的必要步骤。
cs.CV / 186 / 2607.24731

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

重新思考在策略内扩散蒸馏中的无分类器引导
Li, Bingnan, Wang, Haozhe, Xiong, Haozhong, Wu, Fangtai, Yu, Jinpeng, Shi, Yang, Liu, Jiaming, Huang, Ruihua
Abstract
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases, we find that naive matching remains effective under shared negative conditioning, where both branch errors decrease jointly. When the model's native CFG schema retains privileged information in the teacher's negative branch that is unavailable to the student, however, this joint reduction breaks down and the composed objective induces antagonistic branch-error dynamics, reducing the positive-branch error while increasing the negative-branch error. We term this failure mode Negative Branch Asymmetry (NBA). To address NBA, we introduce Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction. We apply PDM to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.
Chinese Translation
策略内蒸馏(On-policy distillation, OPD)通过在当前学生生成的轨迹上查询教师来适应扩散模型,但在无分类器引导(classifier-free guidance, CFG)这一现代扩散系统的默认组件下,其应如何表现仍然不甚明了。现有的 OPD 方法自然地将速度匹配扩展到由 CFG 组成的预测,直接匹配教师和学生的引导速度。我们展示了这一目标在分支层面上是欠识别的:正分支和负分支的误差可以在引导预测中相互补偿。通过两个对比案例,我们发现,在共享负条件下,简单的匹配仍然有效,此时两个分支的误差共同减少。然而,当模型的原生 CFG 结构在教师的负分支中保留了学生无法获得的特权信息时,这种联合减少就会失效,组合目标会引发对抗性的分支误差动态,减少正分支误差的同时增加负分支误差。我们将这种失效模式称为负分支不对称(Negative Branch Asymmetry, NBA)。为了解决 NBA 问题,我们引入了正向匹配(Positive--Direction Matching, PDM),这是一种分支感知的 OPD 目标,分别约束正预测和 CFG 条件方向。我们将 PDM 应用于密集到稀疏的视频控制,其中简单的引导匹配对推理引导尺度高度敏感,而分支感知的监督则使知识转移更加稳健和有效。
cs.CV / 187 / 2607.24743

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

ClinFusion:一种以视觉为中心的多模态大语言模型系统,用于全面的医学理解
Yuan, Hangjie, Qian, Yichen, Tang, Zhiwei, Xu, Xianzhe, Wu, Lirong, Yang, Sicheng, Wang, Jinwang, Wang, Pengju, Zeng, Zhitao, Han, Yizeng, Xing, Yan, Luo, Shengxuan, Feng, Tao, Xie, Qing, Yao, Weigen, Yang, Yi, Liu, Zuozhu, Tang, Jiasheng, Wang, Shaocheng, Wang, Jitao, Dong, Jiahong, Chen, Weihua, Xu, Feng, Wang, Fan
Abstract
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (\textit{e.g.}, Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.
Chinese Translation
多模态大语言模型(MLLMs)具有革命性改变临床实践的巨大潜力,但在医学领域的应用本质上是一个以视觉为中心的挑战:模型必须从异构的二维和三维医学图像中吸收知识,评估协议必须与放射科医生的临床实践对齐,并提供准确、细致且以事实为驱动的评估。在本文中,我们介绍了ClinFusion,一种旨在全面医学理解的以视觉为中心的MLLM,系统性地解决了这些局限性。我们提出了一种组合和级联的视觉编码器架构,采用级联空间感知局部融合操作符,将多样的二维和原生三维医学图像理解统一于一个融合编码器中。我们进一步引入了一个以视觉为基础的评估框架,包括用于指令跟随评估的MedIF-Bench和一种以兴趣区域为基础的方法,用于临床对齐和以事实为驱动的报告生成评估。我们展示了ClinFusion在一系列全面的二维和三维多模态医学基准测试中设立了新的最先进水平——涵盖视觉问答、报告生成和指令跟随,以及文本医学任务,在24个基准测试中超越了领先的开源医学MLLM(例如,Hulu-Med,Lingshu)的20个,并在16个基准测试中表现出比强大的专有模型如GPT-5.2和Gemini-3-Flash更好的多模态能力。ClinFusion还可以通过代理工具使用进一步增强,以支持检索增强和工具辅助的临床工作流程。经过认证的放射科医生的盲评确认ClinFusion生成了排名最高的报告,并验证了我们的兴趣区域基础指标在所有检查的自动评估指标中与专家判断的相关性最强。
人工智能 (Artificial Intelligence)
185
cs.AI / 1 / 2607.22544

Concept-based Visual Counterfactual Explanations with Diffusion Models

基于概念的视觉反事实解释与扩散模型
Oueslati, Yassine, Kirilenko, Daniil, Gjoreski, Martin, Langheinrich, Marc
Abstract
Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?", and are increasingly important as vision models are deployed in safety-critical domains (e.g., medicine). Existing diffusion-based methods can produce realistic edits, but they rely on external classifiers that must work reliably on noisy images, which makes them fragile and hard to deploy for robust explanations. We introduce C-VCE, a new diffusion framework that builds the classifier directly into the generative model via a concept bottleneck layer, so that counterfactuals are guided by human-interpretable features (concepts) instead of a separate noise robust classifier that works with pixel-level edits. Our model lets users to toggle on/off semantic concepts during sampling, then minimally adjusts relevant image regions, while preserving the rest of the image, respecting feature correlations. To keep edits small and controlled, we add a simple probabilistic regularizer that balances "change the prediction" against "stay close to the original", plus a gradient-based mask that confines modifications to the most relevant regions. On benchmarks such as CelebA, C-VCE matches or improves flip rates while producing counterfactuals that are visually closer to the input and less distorted than baselines that depend on separate noisy-image classifiers. These properties make C-VCE a practical tool for vision systems where users need concrete "what-if" images without having to trust an additional, noise-robust classifier. More broadly, our results suggest that exposing and controlling an internal concept layer is a promising way to make powerful generative models easier to understand and safer to use.
Chinese Translation
视觉反事实解释旨在回答“对这张图像进行什么最小的改变会改变模型的预测?”随着视觉模型在安全关键领域(例如医学)中的应用,这一问题变得越来越重要。现有的基于扩散的方法能够生成逼真的编辑,但它们依赖于必须在噪声图像上可靠工作的外部分类器,这使得它们脆弱且难以用于稳健的解释。我们提出了C-VCE,一种新的扩散框架,通过概念瓶颈层将分类器直接构建到生成模型中,从而使反事实由人类可解释的特征(概念)引导,而不是依赖于与像素级编辑配合使用的单独噪声鲁棒分类器。我们的模型允许用户在采样过程中开启/关闭语义概念,然后对相关图像区域进行最小调整,同时保留图像的其余部分,尊重特征相关性。为了保持编辑的小规模和可控性,我们添加了一个简单的概率正则化项,平衡“改变预测”和“保持接近原始”的目标,以及一个基于梯度的掩码,将修改限制在最相关的区域。在CelebA等基准测试中,C-VCE的翻转率与基线相匹配或有所提高,同时生成的反事实在视觉上更接近输入且失真程度低于依赖于单独噪声图像分类器的基线。这些特性使得C-VCE成为视觉系统中的一种实用工具,用户无需信任额外的噪声鲁棒分类器即可获得具体的“如果……会怎样”的图像。更广泛地说,我们的结果表明,暴露和控制内部概念层是一种有前景的方法,可以使强大的生成模型更易于理解和使用更安全。
cs.AI / 2 / 2607.22548

SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series

SeT-Diff:面向高性能计算遥测和时间序列的语义基础模型
Esposito, Giovanni B., Antici, Francesco, Cesarini, Daniele, Bartolini, Andrea
Abstract
Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics. Current machine learning approaches for HPC and its telemetry typically rely on a static subset of anonymous, fixed-position sensor variables tailored to single tasks. Consequently, these models become obsolete when target tasks change or sensor metrics vary. We propose SeT-Diff, the first foundational model for compute node telemetry and time-series. Unlike rigid architectures, our diffusion-based approach conditions the generative process on each sensor's semantic description, decoupling the system dynamics from the structure of the dataset. Experiments on a real-world supercomputer dataset demonstrate a Mean Absolute Error (MAE) of 0.0470 on reconstruction tasks. SeT-Diff exhibits zero-shot permutation stability, maintaining accuracy with negligible degradation even when sensors are shuffled. A single pre-trained model effectively performs data imputation, forecasting, and virtual sensing - achieving a 0.033 MAE in thermal inference - making SeT-Diff an effective data-driven digital twin for HPC systems.
Chinese Translation
数据中心及其计算节点需要准确且灵活的数字双胞胎,以能够建模工作负载、环境参数和物理指标之间的复杂相互作用。目前针对高性能计算(HPC)及其遥测的机器学习方法通常依赖于一组静态的匿名固定位置传感器变量,这些变量是为单一任务量身定制的。因此,当目标任务发生变化或传感器指标变化时,这些模型便会失效。我们提出了SeT-Diff,这是首个用于计算节点遥测和时间序列的基础模型。与刚性架构不同,我们的基于扩散的方法将生成过程条件化于每个传感器的语义描述,从而将系统动态与数据集的结构解耦。对真实世界超级计算机数据集的实验表明,在重建任务中,平均绝对误差(MAE)为0.0470。SeT-Diff展现出零-shot排列稳定性,即使在传感器被打乱时,仍能保持准确性,且降级微乎其微。一个单一的预训练模型有效地执行数据插补、预测和虚拟传感,热推断的MAE达到0.033,使SeT-Diff成为高性能计算系统中有效的数据驱动数字双胞胎。
cs.AI / 3 / 2607.22549

QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction

QFoldAgent:一种用于蛋白质结构预测的自主量子优化多智能体系统
Chen, Winson, Zhang, Yuqi, Chen, Sixu, Xu, Nuo, Guan, Qiang, Ding, Caiwen
Abstract
Hybrid quantum-classical protein structure prediction depends strongly on Hamiltonian penalty weights, yet existing lattice-based workflows typically fix these coefficients by hand and evaluate only very short fragments in simulation. We present QFoldAgent, a closed-loop multi-agent framework for 5-residue tetrahedral-lattice folding in which a design agent proposes sequence-conditioned penalties, a VQE-based quantum-classical pipeline optimizes the resulting Hamiltonian under Qiskit Aer noise, and a feedback agent uses energy-landscape diagnostics and MolProbity validation signals to refine penalties across cycles. Ground-truth metrics such as RMSD are never exposed to the agents and are used only for evaluation. We study the framework on two complementary datasets: 55 QDockBank-derived fragments with known structures and 100 coverage-optimized unseen sequences. On the QDockBank benchmark, QFoldAgent reduces median RMSD from 3.64 \AA{} to 3.20 \AA{}, with the largest gains on the hardest targets. On unseen sequences, the closed loop raises structural validity from 87.5% to 98.7%, recovers 87% of initially invalid cases, and the strongest controller improves cycle-3 energy on 87% of sequences while maintaining 96% Ramachandran-favored geometry. These results show that iterative agent control can systematically improve optimization behavior and reduce failure cases in a 5-residue quantum setting.
Chinese Translation
混合量子-经典蛋白质结构预测在很大程度上依赖于哈密顿惩罚权重,然而现有的基于晶格的工作流程通常手动固定这些系数,并且仅在模拟中评估非常短的片段。我们提出了QFoldAgent,这是一个闭环多智能体框架,用于5个残基的四面体晶格折叠,其中设计智能体提出基于序列的惩罚,基于变分量子特征求解器(VQE)的量子-经典管道在Qiskit Aer噪声下优化生成的哈密顿量,而反馈智能体利用能量景观诊断和MolProbity验证信号在多个周期中细化惩罚。真实度量(如均方根偏差RMSD)从未暴露给智能体,仅用于评估。我们在两个互补的数据集上研究该框架:55个具有已知结构的QDockBank衍生片段和100个覆盖优化的未见序列。在QDockBank基准测试中,QFoldAgent将中位数RMSD从3.64 Å降低到3.20 Å,最大的增益出现在最难的目标上。在未见序列中,闭环将结构有效性从87.5%提高到98.7%,恢复了87%的最初无效案例,而最强的控制器在87%的序列上提高了第3轮的能量,同时保持96%的Ramachandran偏好几何形状。这些结果表明,迭代智能体控制可以系统性地改善优化行为并减少5个残基量子设置中的失败案例。
cs.AI / 4 / 2607.22554

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

相同问题,不同答案:超越准确性的LLM可靠性评估
Faghih, Kazem, Cheng, Yize, Saha, Shoumik, Pournemat, Mobina, Gerami, Armin, Feizi, Soheil
Abstract
Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact wording of the prompt. While overall accuracy typically changes only modestly across paraphrases, instance-level behavior is far less stable: for many questions, models alternate between correct and incorrect answers depending on phrasing, with mismatch rates reaching more than 23%. Conditioning on questions that are answered correctly in their original form reveals even larger failures measured by answer flip rates, showing that single-prompt correctness is often a poor indicator of reliability. At the same time, we find that models often produce a correct answer for at least one paraphrase of a question, suggesting that the underlying knowledge is present but inconsistently retrieved. Building on this observation, we show that a simple self-paraphrasing strategy can partially recover this latent knowledge and improve performance at inference time. Together, these findings suggest that standard accuracy metrics can mask substantial instability, and that evaluating consistency across equivalent inputs provides a clearer picture of LLM reliability.
Chinese Translation
大型语言模型(LLMs)在基准测试中通常能够实现较高的准确性,但当同一问题以不同但等效的方式表述时,它们在应用这些知识时的可靠性仍然不清楚。在本研究中,我们探讨了在事实问答和数学推理任务中,模型答案在保留意义的释义下如何变化。在四个基准和13个模型的研究中,我们发现模型输出往往依赖于提示的确切措辞。尽管整体准确性在释义之间通常变化不大,但实例级别的行为则远不稳定:对于许多问题,模型根据措辞在正确和错误答案之间交替,错配率超过23%。对以原始形式正确回答的问题进行条件处理显示,答案翻转率的测量揭示了更大的失败,表明单一提示的正确性往往是可靠性的差劣指标。同时,我们发现模型通常会为至少一个问题的释义产生正确答案,这表明潜在知识存在但检索不一致。基于这一观察,我们展示了一种简单的自我释义策略可以部分恢复这一潜在知识,并在推理时提高性能。这些发现共同表明,标准准确性指标可能掩盖了显著的不稳定性,而评估等效输入之间的一致性则提供了更清晰的LLM可靠性图景。
cs.AI / 5 / 2607.22555

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs

DeepLens 诊断代理:代理工作流程设计使小型推理模型能够与前沿大型语言模型竞争
Bayeshi, Mahmood, Kocaman, Veysel, Naqvi, Muhammed Ali, Gul, Yigit, Talby, David
Abstract
Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagnostic reasoning. We present the DeepLens Diagnosis Agent, a five-stage harnessing pipeline (combining model capabilities with disciplined process constraints) centered on a small medical reasoning model (JSL Medical Small 7B v2) and retrieval-augmented generation (RAG). The pipeline enforces structured clinical extraction, disciplined retrieval, constrained candidate generation, explicit evidence triangulation, and an auditable final decision. On the 915-case DiagnosisArena benchmark, the agent achieved 60.14% top-1 diagnostic accuracy, the highest among small and medium-sized models. The same model without the agent workflow achieved 23.99%, a +36-point gain from workflow design alone, despite 88.2% on standard medical benchmarks, showing that diagnostic reasoning under uncertainty requires more than knowledge recall. The agent costs USD 0.0072 per case (24K tokens on A100) with 24-second latency, 35-45% cheaper than Claude Sonnet 4.5 (USD 0.0110) and Gemini 3.1 Pro (USD 0.0128) while outperforming them by +9.70pp and +9.17pp. Harnessing can also correct frontier model failures; workflow constraints can outweigh parameter count or API cost. Beyond aggregate accuracy, the pipeline produces structured intermediate artifacts that make each stage inspectable and support error localization. These properties support high-stakes settings where traceability, reproducibility, and auditable evidence matter alongside benchmark performance.
Chinese Translation
医学诊断是一个多阶段的过程:提取事实、咨询知识、生成鉴别分析,并选择最佳诊断及其解释。前沿大型语言模型(LLMs)是强大的通用模型,但单次提示往往导致脆弱的诊断推理。我们提出了 DeepLens 诊断代理,一个五阶段的整合管道(结合模型能力与严格的过程约束),以一个小型医学推理模型(JSL Medical Small 7B v2)和检索增强生成(RAG)为中心。该管道强制执行结构化的临床提取、严格的检索、受限的候选生成、明确的证据三角验证以及可审计的最终决策。在 915 个案例的 DiagnosisArena 基准测试中,该代理实现了 60.14% 的顶级诊断准确率,在小型和中型模型中最高。没有代理工作流程的同一模型达到了 23.99%,仅凭工作流程设计就获得了 +36 个百分点的提升,尽管在标准医学基准上达到了 88.2%,这表明在不确定性下的诊断推理需要的不仅仅是知识回忆。该代理每个案例的成本为 0.0072 美元(在 A100 上为 24K 令牌),延迟为 24 秒,比 Claude Sonnet 4.5(0.0110 美元)和 Gemini 3.1 Pro(0.0128 美元)便宜 35-45%,同时在准确率上超越它们 +9.70 个百分点和 +9.17 个百分点。整合还可以纠正前沿模型的失败;工作流程约束可以超越参数数量或 API 成本。除了整体准确率外,该管道还生成结构化的中间文档,使每个阶段可检查并支持错误定位。这些特性支持高风险环境,在这些环境中,可追溯性、可重复性和可审计证据与基准性能同样重要。
cs.AI / 6 / 2607.22556

MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models

MIITA:用于小型语言模型的持续学习的记忆诱导推理时适应
Li, Dong, Liu, Yanchi, Zhao, Xujiang, Cheng, Wei, Chen, Zhengzhang, Wu, Xintao, Chen, Zhong, Zhao, Chen, Chen, Haifeng
Abstract
Continual learning (CL) is essential for small language models (SLMs) to adapt to evolving real-world needs in resource-constrained deployments. However, directly updating their limited parameter space causes catastrophic forgetting. While memory-based methods naturally address this by decoupling knowledge retention from parameters, existing approaches designed for large language models (LLMs) rely on abundant storage and strong in-context reasoning that SLMs lack. To address these challenges, we propose MIITA, a Memory-Induced Inference-Time Adaptation framework for supervised CL under constrained storage. MIITA stores supervised experiences as compact correction-direction prototypes with semantic anchors, and retrieves them at inference time using semantic and uncertainty-based cues. The retrieved directions are applied through gated temporary hidden-state adaptation, enabling non-destructive reuse of past supervision without backbone updates, prompt extensions, or test-time backpropagation. A local theoretical analysis links this design to first-order loss reduction, uncertainty-guided retrieval, and directional coverage for retaining old-stage knowledge. Extensive experiments across diverse supervised CL settings show that MIITA consistently improves final performance and mitigates forgetting under fixed memory budgets.
Chinese Translation
持续学习(CL)对于小型语言模型(SLMs)在资源受限的部署中适应不断变化的现实需求至关重要。然而,直接更新其有限的参数空间会导致灾难性遗忘。虽然基于记忆的方法通过将知识保留与参数解耦,自然地解决了这一问题,但现有针对大型语言模型(LLMs)设计的方法依赖于丰富的存储和强大的上下文推理能力,而这些是SLMs所缺乏的。为了解决这些挑战,我们提出了MIITA,一个用于受限存储下监督持续学习的记忆诱导推理时适应框架。MIITA将监督经验存储为具有语义锚点的紧凑修正方向原型,并在推理时通过语义和不确定性线索进行检索。检索到的方向通过门控临时隐藏状态适应应用,使得在不更新主干、扩展提示或测试时反向传播的情况下,能够非破坏性地重用过去的监督。局部理论分析将这一设计与一阶损失减少、不确定性引导的检索以及保留旧阶段知识的方向覆盖联系起来。在多种监督持续学习设置下的广泛实验表明,MIITA始终提高最终性能,并在固定内存预算下减轻遗忘。
cs.AI / 7 / 2607.22561

Codifying the Judge: Scalable Evaluation via Program Distillation

法官的编码:通过程序蒸馏实现可扩展评估
Huang, Tzu-Heng, Qiu, Shengqi, Sala, Frederic
Abstract
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.
Chinese Translation
大型语言模型(LLM)作为法官已成为自动评估的标准,但其高成本、显著延迟和不透明的决策等局限性削弱了其可扩展性和可靠性。我们提出了一种简单而高效的替代方案:程序蒸馏。我们不再在评估时提示LLM,而是将其决策逻辑蒸馏成一个程序委员会,直接对候选者进行评分。这些程序法官提供透明性,易于检查或编辑,并消除了每个样本的API成本。在此基础上,我们引入了PAJAMA,一个合成程序作为法官的系统,汇总它们的决策形成联合裁决,并包含一个后备机制,以选择性地将低信心案例升级到LLM。在五个数据集和四个模型系列中,我们展示了程序法官能够匹配一个13B规模的LLM法官的性能。当使用程序输出作为路由信号时,PAJAMA提高了准确性和吞吐量,并推动了帕累托前沿。除了评估,程序法官还产生廉价且有效的奖励信号:在RewardBench上,从程序裁决中蒸馏出的奖励模型在API成本低两个数量级的情况下,优于基于专有LLM标签训练的模型。
cs.AI / 8 / 2607.22562

SF-AMS: Strategic Forgetting for Structured Memory in LLM Agent

SF-AMS:大规模语言模型代理中的结构化记忆的战略遗忘
Yang, Ning, Li, Siqi, Shen, Miaoxin, Zhou, Yuan, Zhang, Meng, Li, Tong, Zhang, Haijun
Abstract
Managing long-context dependencies remains a primary bottleneck in LLM agents, as redundant and irrelevant information can degrade multi-step reasoning. Strategic Forgetting for Agent Memory Systems (SF-AMS) is proposed as a framework for maintaining compact high-utility memory by modeling the long-term importance of memory units. SF-AMS replaces static retrieval and heuristic decay with a utility-driven survival mechanism that updates memory importance from usage redundancy and temporal signals, inducing a hierarchical memory structure that prioritizes stable entity-consistent information while filtering noise. On top of this, Composite Importance Scoring integrates semantic and entity level signals to improve retrieval robustness. Experiments on LoCoMo and LongMemEval-s show consistent gains over strong state of the art baselines including LightMem MemO and A-Mem. The largest improvement appears in multi-hop reasoning under Qwen2.5-7B where SF-AMS achieves plus 9.65 F1 over the strongest baseline followed by temporal reasoning under GPT-4o-mini plus 6.91 F1 and open-domain tasks plus 6.53 F1 demonstrating strong cross backbone generalization. These results show that modeling memory importance as a dynamic utility signal is critical for reliable long-context reasoning.
Chinese Translation
管理长上下文依赖性仍然是大规模语言模型(LLM)代理的主要瓶颈,因为冗余和无关的信息会降低多步推理的效果。战略遗忘代理记忆系统(SF-AMS)被提出作为一个框架,通过对记忆单元的长期重要性建模,来维持紧凑的高效用记忆。SF-AMS用一种基于效用的生存机制替代了静态检索和启发式衰减,该机制根据使用冗余和时间信号更新记忆的重要性,从而诱导出一种层次化的记忆结构,优先考虑稳定的实体一致信息,同时过滤噪声。在此基础上,复合重要性评分(Composite Importance Scoring)整合了语义和实体级别的信号,以提高检索的鲁棒性。在LoCoMo和LongMemEval-s上的实验显示,SF-AMS在包括LightMem MemO和A-Mem在内的强基线之上取得了一致的提升。在Qwen2.5-7B下,多跳推理的最大提升为9.65 F1,紧随其后的是在GPT-4o-mini下的时间推理提升6.91 F1,以及开放域任务提升6.53 F1,显示出强大的跨骨干网络泛化能力。这些结果表明,将记忆重要性建模为动态效用信号对于可靠的长上下文推理至关重要。
cs.AI / 9 / 2607.22563

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

用于评估工业4.0代理的合成场景生成
Kumar, Sagar Chethan, Kanathur, Rohith, Patel, Dhaval, Maghraoui, Kaoutar El
Abstract
Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks such as AssetOpsBench rely on manually authored scenarios and cover a limited set of asset classes. We extend AssetOpsBench with a Smart Grid Transformer asset class and four IEC-grounded diagnostic tools for health-index prediction, dissolved-gas analysis, winding-temperature assessment, and load-profile assessment. We further introduce ScenarioGeneratorAgent, a pipeline for synthetic industrial-agent scenario generation. The pipeline constructs evidence-grounded asset profiles, allocates coverage-aware scenario budgets across operational domains, and generates candidates through a hybrid validation-and-repair loop that enforces schema validity, tool reachability, physical plausibility, standards alignment, and deduplication. To improve scalability, we apply two-level caching, parallel focus-group generation, thread-pool offloading, batched LLM calls, and early rejection filtering. On Smart Grid Transformer scenario generation, these optimizations reduce end-to-end runtime by $8\times$ for 50 scenarios while preserving quality, achieving a composite quality score of $74.2 \pm 1.9$ compared with $73.8 \pm 3.0$ for the unoptimized baseline. These results show that standards-grounded synthetic scenario generation can efficiently expand industrial-agent benchmarks without sacrificing scenario quality.
Chinese Translation
工业代理基准测试需要现实的评估场景,这些场景整合了遥测、故障模式、维护记录和领域标准。然而,现有的基准测试如AssetOpsBench依赖于手动编写的场景,并且覆盖的资产类别有限。我们通过引入智能电网变压器资产类别和四种基于IEC的诊断工具(用于健康指数预测、溶解气体分析、绕组温度评估和负载特征评估)来扩展AssetOpsBench。我们进一步引入了ScenarioGeneratorAgent,这是一个用于合成工业代理场景生成的管道。该管道构建基于证据的资产档案,在操作领域之间分配覆盖感知的场景预算,并通过一个混合验证与修复循环生成候选场景,该循环强制执行模式有效性、工具可达性、物理合理性、标准一致性和去重。为了提高可扩展性,我们应用了两级缓存、并行焦点小组生成、线程池卸载、批量LLM调用和早期拒绝过滤。在智能电网变压器场景生成中,这些优化将50个场景的端到端运行时间减少了8倍,同时保持了质量,获得了74.2 ± 1.9的综合质量评分,而未优化基线的评分为73.8 ± 3.0。这些结果表明,基于标准的合成场景生成可以有效扩展工业代理基准,而不牺牲场景质量。
cs.AI / 10 / 2607.22564

Loss-Aware Feature-Map Pruning in Convolutional Neural Networks Using Multi-Armed Bandits

基于多臂赌博机的卷积神经网络损失感知特征图剪枝
Ameen, Salem, Vadera, Sunil
Abstract
Convolutional neural networks often contain redundant feature maps that increase storage and inference cost. This paper presents a loss-aware feature-map pruning framework using multi-armed bandits. Feature-map pruning is structured because it removes complete convolutional output channels and their producing filters rather than isolated scalar weights. Each candidate feature map is treated as an arm. At each play time, one map is temporarily masked and evaluated on a sampled mini-batch; the map is then restored and the observed loss change is converted into a safe-removal reward. After a fixed play budget, candidate maps are ranked by learned scores and the top-k maps are permanently removed with their filters, biases and corresponding next-layer input-channel kernels. The study evaluates UCB1 and Thompson Sampling, compares them with direct/oracle-style evaluation on LeNet/MNIST, and extends the evaluation to MNIST, CIFAR-10, CIFAR-100, SVHN, CUB-200-2011 and Oxford Flowers 102. Results show that UCB1 and Thompson Sampling preserve accuracy close to unpruned models while removing feature maps and reducing convolutional computation. Friedman and Nemenyi tests show that UCB1 obtains the highest mean rank, followed by Thompson Sampling; both significantly outperform greedy and magnitude-based pruning while remaining statistically comparable to the original unpruned model.
Chinese Translation
卷积神经网络通常包含冗余的特征图,这增加了存储和推理成本。本文提出了一种基于多臂赌博机的损失感知特征图剪枝框架。特征图剪枝是结构化的,因为它移除完整的卷积输出通道及其生成的滤波器,而不是孤立的标量权重。每个候选特征图被视为一个臂。在每次游戏时,一个特征图被暂时屏蔽,并在一个采样的小批量上进行评估;然后该特征图被恢复,观察到的损失变化被转换为安全移除奖励。在固定的游戏预算之后,候选特征图根据学习到的分数进行排名,前k个特征图及其滤波器、偏置和对应的下一层输入通道内核被永久移除。研究评估了UCB1和汤普森采样,并将其与在LeNet/MNIST上的直接/Oracle风格评估进行比较,并将评估扩展到MNIST、CIFAR-10、CIFAR-100、SVHN、CUB-200-2011和Oxford Flowers 102。结果表明,UCB1和汤普森采样在移除特征图并减少卷积计算的同时,保持了接近未剪枝模型的准确性。Friedman和Nemenyi测试显示,UCB1获得了最高的平均排名,其次是汤普森采样;两者在统计上显著优于贪婪和基于幅度的剪枝,同时与原始未剪枝模型在统计上可比。
cs.AI / 11 / 2607.22565

DSTFView: Multi-View Cloud-Edge Workload Forecasting with Dual-Input Spatio-Temporal-Frequency Modeling

DSTFView:基于双输入时空频率建模的多视角云边工作负载预测
Li, Qingzhong, Ma, Hui, Zhang, Yajun, Ma, Qingchang, Long, Zhou
Abstract
With the widespread deployment of edge-side AI inference, edge platforms are increasingly required to support latency-sensitive, highly concurrent, and reliability-critical applications. However, existing methods often struggle to balance multidimensional feature modeling and forecasting efficiency in collaborative cloud-edge environments. To address this issue, we propose DSTFView, a dual-input spatio-temporal-frequency multi-view workload forecasting framework for collaborative cloud-edge environments. It jointly models closeness and period dependencies and extracts spatial, temporal, and frequency-domain dependencies. Besides, it designs an adaptive fusion mechanism and adjusts the contribution of each view to capture abrupt changes. Experimental results on the CPU and TP datasets demonstrate that DSTFView consistently outperforms representative baselines across multiple forecasting horizons and evaluation metrics.
Chinese Translation
随着边缘侧人工智能推理的广泛部署,边缘平台越来越需要支持对延迟敏感、高并发和可靠性要求高的应用。然而,现有方法往往难以在协同云边环境中平衡多维特征建模和预测效率。为了解决这一问题,我们提出了DSTFView,一个针对协同云边环境的双输入时空频率多视角工作负载预测框架。该框架共同建模了相似性和周期依赖性,并提取了空间、时间和频域依赖性。此外,它设计了一种自适应融合机制,并调整每个视角的贡献,以捕捉突发变化。在CPU和TP数据集上的实验结果表明,DSTFView在多个预测时域和评估指标上始终优于代表性的基线方法。
cs.AI / 12 / 2607.22566

MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models

MedLoCoMo:针对大型语言模型的长上下文多会话医学对话基准
Zhang, Zeyu, Wang, Ziqing, Ding, Kaize
Abstract
MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existing medical QA benchmarks largely test short context knowledge or single document grounding, leaving open whether LLMs can use, connect, and abstain over longitudinal patient histories. We build MedLoCoMo from deidentified MIMIC-IV and MIMIC-IV-Note records by constructing admission-level clinical packets, synthesizing grounded doctor-patient conversations, and generating evidence linked QA items over single-admission, cross-admission, and adversarial unanswerable settings. The benchmark contains 100 patient timelines averaging 1,669.8 turns, 29.7 sessions, and 74,512.2 tokens per conversation. Across the evaluated baselines, cross-admission reasoning is consistently harder than localized evidence use, even when models have long context windows or use external memory or retrieval methods. The code and MedLoCoMo benchmark release is available at https://github.com/leozzy13/MedLoCoMo for use and reproducibility.
Chinese Translation
MedLoCoMo 是一个针对多次入院医学对话的患者特定临床推理的医学长上下文记忆基准。现有的医学问答基准主要测试短上下文知识或单一文档的基础,尚未探讨大型语言模型(LLMs)是否能够利用、连接和抑制纵向患者历史。我们通过构建入院级临床数据包、合成基于证据的医患对话,并在单次入院、跨入院和对抗性不可回答的设置中生成与证据相关的问答项目,从去标识化的 MIMIC-IV 和 MIMIC-IV-Note 记录中构建了 MedLoCoMo。该基准包含 100 个患者时间线,平均每个对话 1,669.8 次交互、29.7 次会话和 74,512.2 个标记。在评估的基线中,跨入院推理始终比局部证据使用更具挑战性,即使模型具有长上下文窗口或使用外部记忆或检索方法。代码和 MedLoCoMo 基准的发布可在 https://github.com/leozzy13/MedLoCoMo 获取,以供使用和复现。
cs.AI / 13 / 2607.22568

Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting

关键词的重要性:揭示设备端大语言模型提示的能量敏感性
Tao, Ruiyi, Tu, Xiaolong, Wang, Haoxin
Abstract
Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency. Yet on-device inference faces a fundamental constraint: high energy consumption on battery-powered, resource-limited hardware. While model compression and runtime acceleration have been widely studied, the effect of \emph{prompt design} on energy efficiency remains underexplored. This paper presents an empirical study of the relationship between prompt wording and energy consumption for on-device LLMs. Using real power measurements collected on a smartphone, we quantify how linguistic features, particularly imperative keywords and instruction structure, affect decoding length and total energy. Our results show consistent energy differences across verbs and tasks, indicating that prompt engineering is a lightweight lever for improving energy efficiency.
Chinese Translation
大型语言模型(LLMs)越来越多地部署在移动和嵌入式设备上,以提高隐私性并减少网络延迟。然而,设备端推理面临一个基本限制:在电池供电、资源有限的硬件上高能耗。尽管模型压缩和运行时加速已被广泛研究,但 extit{提示设计}对能效的影响仍然未得到充分探讨。本文呈现了一项关于提示措辞与设备端LLMs能耗之间关系的实证研究。通过在智能手机上收集的实际功率测量,我们量化了语言特征,特别是命令性关键词和指令结构,如何影响解码长度和总能耗。我们的结果显示,在动词和任务之间存在一致的能量差异,表明提示工程是改善能效的轻量级杠杆。
cs.AI / 14 / 2607.22569

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines

基于执行的安全测试在软件工程管道中的编码代理
Ge, Yifei, Sun, Weisong, Xiao, Jinkun, Chen, Yuchen, Feng, Yebo, Lv, Peizhuo, Feng, Xia, Fang, Chunrong, Zhao, Zhihong, Chen, Zhenyu, Liu, Yang
Abstract
Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, and the underlying system. For example, if a coding agent inserts a hook into a system startup or configuration script, that change can persist after the interaction, be triggered later, and abuse delegated user or system privileges to modify the system. This makes security testing a system problem: the key question is not only what the agent says, but what it actually does to the surrounding environment. We present an execution-grounded red-team testing framework for probing this execution-layer security boundary using observable sandbox evidence, including tool invocations, runtime traces, and file-system diffs. Our framework embeds target unsafe operations into routine software engineering workloads, including unit testing, regression testing, crash reproduction, and validation, and uses an execution oracle to guide refinement when an initial probe is rejected or fails. Across multiple agent frameworks and model backbones, our red-team workload reformulation substantially increases verified unsafe execution, reaching 73.61% on code carriers and 53.93% on text carriers. These results show that coding agents in system operations remain insecure under task disguise: once risky intent is hidden inside plausible engineering tasks, the agent can be induced to carry out unsafe actions on the surrounding system. More broadly, coding agents in system operations still demand stronger security testing and safeguards.
Chinese Translation
编码代理越来越多地集成到系统操作中,它们的工具使用可以直接修改项目工件、执行环境和底层系统。例如,如果一个编码代理在系统启动或配置脚本中插入一个钩子,这一改变在交互后可能会持续存在,并在之后被触发,滥用委托的用户或系统权限来修改系统。这使得安全测试成为一个系统性问题:关键问题不仅在于代理所说的内容,还在于它实际上对周围环境所做的事情。我们提出了一种基于执行的红队测试框架,利用可观察的沙箱证据(包括工具调用、运行时跟踪和文件系统差异)来探测这一执行层安全边界。我们的框架将目标不安全操作嵌入常规软件工程工作负载中,包括单元测试、回归测试、崩溃重现和验证,并使用执行预言机在初始探测被拒绝或失败时引导细化。在多个代理框架和模型骨干上,我们的红队工作负载重构显著增加了验证的不安全执行,代码载体的验证率达到73.61%,文本载体的验证率达到53.93%。这些结果表明,系统操作中的编码代理在任务伪装下仍然不安全:一旦风险意图隐藏在合理的工程任务中,代理就可能被诱导在周围系统上执行不安全的操作。更广泛地说,系统操作中的编码代理仍然需要更强的安全测试和保障措施。
cs.AI / 15 / 2607.22570

Reference Feature Atlases for Mechanistic Auditing of Language Models

用于语言模型机制审计的参考特征图谱
Wu, Rui, Che, Tong
Abstract
Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by fitting only a linear decoder. This yields two complementary views. The atlas channel reads the target on already interpreted panel features, providing a stable coordinate system across models. The residual channel learns features only from what the atlas fails to reconstruct, making "outside the reference panel" an explicit audit signal. We train leave-one-out atlases over five 7-9B instruction-tuned models and audit held-out Mistral and Qwen targets. On three controlled LoRA hidden objectives injected into both targets, the residual channel makes the planted mechanism perfectly controllable at runtime while matched controls stay unaffected and recovers the planted objective as the top-ranked latent across both targets; on Mistral, where the per-target SAE and pairwise crosscoder baselines are retrained for a head-to-head benchmark, both baselines fail to do so. On Qwen-2.5, the same channel additionally reveals a panel-relative political-framing cluster; steering it shifts the audited framing metrics while out-of-domain controls remain unchanged.
Chinese Translation
审计一个新的语言模型通常意味着从头开始重新学习和重新解释其内部特征。我们提出了一种参考特征图谱:一个稀疏特征库,首次在参考面板上训练,并在新目标上重复使用,只需通过拟合线性解码器即可附加。这提供了两种互补的视角。图谱通道在已经解释的面板特征上读取目标,提供了一个跨模型的稳定坐标系统。残差通道仅从图谱未能重构的部分学习特征,使得“超出参考面板”成为一个明确的审计信号。我们在五个7-9B指令调优模型上训练了留一法图谱,并审计了保留的Mistral和Qwen目标。在注入到两个目标中的三个受控LoRA隐藏目标上,残差通道使得植入机制在运行时完全可控,而匹配的控制保持不变,并在两个目标中恢复植入目标作为排名最高的潜在目标;在Mistral上,针对每个目标的SAE和成对交叉编码器基线重新训练以进行头对头基准测试,但这两个基线都未能做到这一点。在Qwen-2.5上,同一通道还揭示了一个面板相关的政治框架聚类;对其进行引导会改变审计框架指标,而域外控制保持不变。
cs.AI / 16 / 2607.22571

SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs

SCAIR:面向企业知识图谱的模式条件代理迭代推理
Chaturvedi, Prateek, Zhu, Yuqicheng, Zhou, Hongkuan, Zhou, Dongzhuoran, He, Yunjie, Staab, Steffen, Du, Fei, Tang, Jie, Kharlamov, Evgeny
Abstract
Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing agentic approaches that perform well on public benchmarks often fail to generalize to real-world enterprise Knowledge Graphs (KGs), which are dense, schema-driven, and operationally constrained. To address these limitations, we propose SCAIR (Schema-Conditioned Agentic Iterative Reasoning), a training-free framework that integrates structured planning with controlled iterative reasoning by injecting schema-conditioned structural priors and enforcing schema-aware traversal during multi-hop reasoning. Experiments on an enterprise-oriented benchmark constructed from a real-world Configuration Management DataBase (CMDB) demonstrate that SCAIR substantially improves performance over existing KG-RAG methods. Crucially, our study highlights that reliable enterprise graph reasoning cannot rely on generic agentic designs; instead, it must explicitly incorporate the target domain's structural and operational constraints into the reasoning process. We demonstrate that by aligning agent design with business logic, substantial performance gains can be achieved without the need for costly model retraining.
Chinese Translation
基于知识图谱的检索增强生成(KG-RAG)使得与结构化企业知识的自然语言交互成为可能,但现有在公共基准上表现良好的代理方法往往无法推广到真实世界的企业知识图谱(KG),这些图谱通常是密集的、以模式驱动的,并且在操作上受到限制。为了解决这些局限性,我们提出了SCAIR(模式条件代理迭代推理),这是一个无需训练的框架,通过注入模式条件的结构先验并在多跳推理过程中强制执行模式感知的遍历,整合了结构化规划与受控的迭代推理。我们在一个基于真实世界配置管理数据库(CMDB)构建的企业导向基准上的实验表明,SCAIR在性能上显著优于现有的KG-RAG方法。至关重要的是,我们的研究强调,可靠的企业图推理不能依赖于通用的代理设计;相反,它必须明确将目标领域的结构和操作约束纳入推理过程。我们证明,通过将代理设计与业务逻辑对齐,可以在无需昂贵模型重训练的情况下实现显著的性能提升。
cs.AI / 17 / 2607.22572

Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL

模式感知本地化 (SAL):用于 Oracle NL2SQL 的实时模式基础和幻觉验证
Mishra, Sanjay, Chukkapalli, Divya, Naik, Ganesh R.
Abstract
Large language models can generate fluent SQL from natural language, but on real enterprise Oracle databases they frequently fail at execution time: columns and aliases are hallucinated and dialect-specific syntax is missed, leading to ORA-00904 invalid-identifier errors. In this setting, failures are primarily due to missing schema grounding: the model cannot know which tables and columns actually exist. This paper introduces Schema-Aware Localisation (SAL), a lightweight middleware layer for Oracle NL2SQL that requires no model retraining. SAL queries Oracle's USER_TAB_COLUMNS catalog to build a live schema map, selects a relevant table subset for each question (falling back to the full schema for multi-table queries), and injects this ground-truth context into the LLM prompt. Generated SQL is then checked by the Hallucination Index (Hidx), which validates every alias.column reference against the live catalog, automatically rewrites predictable prefix errors, and otherwise triggers a structured retry with itemised corrections. We evaluate SAL on 500 TPC-H natural language questions executed against a live Oracle Autonomous Database 23c instance using GPT-4o-mini. Without any schema grounding, execution-grounded truth (EGT; executes and matches the reference result set) is 2.2% (12/500). A hand-written static schema hint brings EGT to 62.0%. SAL, with no manual schema curation, achieves 62.6% EGT (96% simple, 95% medium, 40.7% complex) while reducing execution failures from 97.6% to 2.6%.
Chinese Translation
大型语言模型可以从自然语言生成流畅的 SQL,但在真实的企业 Oracle 数据库上,它们在执行时常常失败:列和别名被虚构,特定方言的语法被遗漏,导致 ORA-00904 无效标识符错误。在这种情况下,失败主要是由于缺乏模式基础:模型无法知道哪些表和列实际上存在。本文介绍了模式感知本地化 (SAL),这是一个轻量级的 Oracle NL2SQL 中间件层,无需重新训练模型。SAL 查询 Oracle 的 USER_TAB_COLUMNS 目录以构建实时模式映射,为每个问题选择相关的表子集(对于多表查询则回退到完整模式),并将这一真实上下文注入 LLM 提示中。生成的 SQL 然后通过幻觉指数 (Hidx) 进行检查,该指数验证每个 alias.column 引用与实时目录的匹配,自动重写可预测的前缀错误,并在其他情况下触发结构化重试并进行逐项修正。我们在 500 个 TPC-H 自然语言问题上评估 SAL,这些问题在使用 GPT-4o-mini 的实时 Oracle Autonomous Database 23c 实例上执行。在没有任何模式基础的情况下,执行基础真实值 (EGT;执行并匹配参考结果集) 为 2.2% (12/500)。手动编写的静态模式提示将 EGT 提高到 62.0%。SAL 在没有手动模式整理的情况下,实现了 62.6% 的 EGT(96% 简单,95% 中等,40.7% 复杂),同时将执行失败率从 97.6% 降低到 2.6%。
cs.AI / 18 / 2607.22573

PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability

PhononBench-MP40:一个用于声子稳定性的光谱分辨基准数据集
Li, Wen-Kao, Gao, Ze-Feng, Lu, Zhong-Yi
Abstract
Imaginary phonon modes remain a practical bottleneck in computational materials screening because otherwise plausible structures can be locally dynamically unstable under a chosen workflow. Here we present PhononBench-MP40, a spectrum-resolved benchmark dataset of Materials Project-derived crystals for workflow-defined phonon stability. The dataset starts from 47,969 MP40 workflow tasks and provides 46,899 completed records with paired stability labels and local phonopy YAML spectra, including 16,683 Stable records and 30,216 completed-phonon unstable records. A further 1,067 relaxation failures are reported separately rather than merged into the completed phonon denominator. The release centers on the local YAML spectrum: the stability label, the lowest sampled frequency and any threshold-dependent relabeling are derived from that spectrum. The dataset is openly available through Science Data Bank at https://doi.org/10.57760/sciencedb.38735. A companion GitHub repository provides the calculation code and lightweight access utilities. PhononBench-MP40 provides an auditable reference for workflow-defined stability classification, minimum-frequency analysis, threshold studies and failure-aware triage, while keeping the reference workflow, data schema and interpretation boundaries explicit.
Chinese Translation
虚声子模式在计算材料筛选中仍然是一个实际瓶颈,因为在所选择的工作流程下,其他看似合理的结构可能在局部上动态不稳定。本文介绍了PhononBench-MP40,这是一个基于材料项目(Materials Project)衍生晶体的光谱分辨基准数据集,用于工作流程定义的声子稳定性。该数据集源自47,969个MP40工作流程任务,提供了46,899条完成记录,包含配对的稳定性标签和局部phonopy YAML光谱,其中包括16,683条稳定记录和30,216条完成声子不稳定记录。另有1,067条松弛失败的记录被单独报告,而不是合并到完成声子的分母中。该发布的核心是局部YAML光谱:稳定性标签、最低采样频率及任何阈值依赖的重新标记均源自该光谱。该数据集通过科学数据银行(Science Data Bank)公开提供,链接为https://doi.org/10.57760/sciencedb.38735。一个配套的GitHub仓库提供计算代码和轻量级访问工具。PhononBench-MP40为工作流程定义的稳定性分类、最低频率分析、阈值研究和失败感知分流提供了可审计的参考,同时保持了参考工作流程、数据模式和解释边界的明确性。
cs.AI / 19 / 2607.22574

Too much evidence, too little time: From text to actionable recommendations through multi-objective evidence reasoning

证据过多,时间过少:通过多目标证据推理从文本到可操作建议
Bara, Adela, Oprea, Simona-Vasilica
Abstract
Evidence-based clinical decision making requires specialists to identify, evaluate and synthesize relevant scientific literature. However, PubMed searches for complex clinical cases often return hundreds of publications that cannot be reviewed manually under time constraints. This study proposes SCEPTER (Single-Case Evidence-driven PubMed-To-rEcommendation Reasoner), a framework for transforming clinical case descriptions into evidence-based recommendations. SCEPTER combines PubMed retrieval, PubMedBERT semantic ranking, large language model (LLM)-based claim extraction, evidence-level weighting, contradiction detection, consensus analysis and multi-objective Pareto claim selection. The framework generates structured evidence syntheses and grounded actionable recommendations. A Paper Q&A module further enables interactive exploration of selected publications. The proposed framework introduces multi-objective reasoning model that integrates literature support, contradiction analysis and interactive literature interrogation into a unified clinical decision-support pipeline. Evaluation on 150 case studies demonstrated that the framework reduced an average search space of 576 papers to 53 retained papers, 7 Pareto-optimal claims and 3 final recommendations, corresponding to an overall compression ratio of 192:1. Despite this reduction, the retained evidence maintained high diversity (entropy=0.901). The ablation study showed that Pareto-based selection increased evidence diversity and recommendation utility compared with conventional ranking approaches.
Chinese Translation
基于证据的临床决策要求专家识别、评估和综合相关的科学文献。然而,针对复杂临床案例的PubMed搜索通常会返回数百篇文献,这在时间限制下无法手动审阅。本研究提出了SCEPTER(单案例证据驱动的PubMed到推荐推理器),这是一个将临床案例描述转化为基于证据的建议的框架。SCEPTER结合了PubMed检索、PubMedBERT语义排名、基于大型语言模型(LLM)的主张提取、证据水平加权、矛盾检测、共识分析和多目标帕累托主张选择。该框架生成结构化的证据综合和基于事实的可操作建议。一个论文问答模块进一步支持对所选文献的互动探索。所提出的框架引入了一个多目标推理模型,将文献支持、矛盾分析和互动文献审查整合到一个统一的临床决策支持流程中。在150个案例研究上的评估表明,该框架将平均搜索空间从576篇论文减少到53篇保留论文、7个帕累托最优主张和3个最终建议,对应的总体压缩比为192:1。尽管有这种减少,保留的证据仍然保持了较高的多样性(熵=0.901)。消融研究表明,与传统排名方法相比,基于帕累托的选择增加了证据的多样性和建议的实用性。
cs.AI / 20 / 2607.22575

Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models

时间上下文重建驱动长上下文语言模型中的类情节顺序记忆
Pink, Mathis, Vo, Vy Ai, Wu, Qinyuan, Mu, Jianing, Turek, Javier, Hasson, Uri, Norman, Kenneth A., Michelmann, Sebastian, Huth, Alexander, Toneva, Mariya
Abstract
Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlying this ability remain debated due to the limited mechanistic accessibility in long-term memory experiments in humans. Long-context LLMs may offer promising ways to reveal plausible computational mechanisms that drive this type of retrieval. Here, we investigate whether and how LLMs capture the core behavioral signatures of episodic memory via a temporal order memory task. Using a new dataset of human behavior based on memory of a full-length novel, we show that models exhibit the same characteristic distance effect observed in humans on this task. We next apply long-context mechanistic interpretability analyses to uncover how models solve this task, and find that model performance relies on a one-dimensional temporal code that is reinstated during retrieval by a single time-reinstatement attention head. These findings support temporal context reinstatement as an important mechanism for episodic-like temporal-order memory in LLMs, offering new insights into how temporal aspects of long-term episodic memory may be instantiated in both artificial and biological systems.
Chinese Translation
人类的情节记忆支持对在较长时间尺度上展开的经历的检索,然而,由于人类长期记忆实验中机制的有限可及性,这种能力背后的计算机制仍然存在争议。长上下文语言模型(LLMs)可能提供了揭示驱动这种检索类型的合理计算机制的有希望的方法。在此,我们研究了LLMs是否以及如何通过时间顺序记忆任务捕捉情节记忆的核心行为特征。基于对完整小说记忆的人类行为的新数据集,我们展示了模型在该任务中表现出与人类相同的特征性距离效应。接下来,我们应用长上下文机制可解释性分析来揭示模型如何解决该任务,并发现模型的表现依赖于一个在检索过程中由单个时间重建注意力头重建的一维时间编码。这些发现支持时间上下文重建作为LLMs中类情节时间顺序记忆的重要机制,为理解长期情节记忆的时间方面如何在人工和生物系统中体现提供了新的见解。
cs.AI / 21 / 2607.22577

cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs

大规模 cMoLLM:混合大语言模型的横向扩展法则
Yang, Xin, Wang, Yemin, Liu, Mingda, Li, Letian, Cao, Shuaishuai, He, Zhengxiao, Dong, Ryan
Abstract
Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes. We aim to scale capacity through MoE-style mixture throughout the LLM pipeline rather than only the FFN. Prior pipeline-level approaches include ParaScale, which introduces virtual tokens and parallel streams but incurs substantial overhead and suffers from homogenized routing and gradient collapse, and AltUp, which uses an auxiliary prediction branch but offers limited adaptivity and slow convergence. We establish that MoE-style mixture layers can be reformulated as variable-kernel dynamic convolutions, where each expert corresponds to a $1{\times}1$ convolutional kernel and routing implements input-conditioned kernel aggregation. Building on this equivalence, we introduce cMoLLM: a convolutionally gated mixture-of-LLMs that routes over end-to-end streams through fully differentiable dynamic convolution. In GPT-2-style models trained on FineWeb, cMoLLM improves language modeling perplexity and downstream GLUE and SQuAD accuracy under matched compute, with better stream utilization, more stable optimization, and favorable scaling compared to ParaScale- and AltUp-style baselines.
Chinese Translation
大规模语言模型(LLMs)的扩展推动了其成功,然而,密集型变换器将容量与计算紧密结合:每个参数在每个标记上都被激活,这使得训练和推理成本随着模型规模线性增长——当模型接近万亿参数时,这成为一个关键瓶颈。我们的目标是通过在 LLM 流水线中采用 MoE 风格的混合来扩展容量,而不仅仅是在前馈网络(FFN)中。先前的流水线级方法包括 ParaScale,该方法引入了虚拟标记和并行流,但带来了可观的开销,并且遭受了同质化路由和梯度崩溃的问题;还有 AltUp,该方法使用辅助预测分支,但适应性有限且收敛速度较慢。我们证明 MoE 风格的混合层可以重新表述为可变核动态卷积,其中每个专家对应于一个 $1{ imes}1$ 卷积核,路由实现了基于输入的核聚合。基于这一等价性,我们引入了 cMoLLM:一种卷积门控的混合大语言模型,通过完全可微的动态卷积在端到端流中进行路由。在在 FineWeb 上训练的 GPT-2 风格模型中,cMoLLM 在匹配计算下提高了语言建模的困惑度以及下游 GLUE 和 SQuAD 的准确性,相比于 ParaScale 和 AltUp 风格的基线,具有更好的流利用率、更稳定的优化和更有利的扩展性。
cs.AI / 22 / 2607.22578

HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization

HeraSys:通过细粒度端到端优化协同服务多个大语言模型工作流
Li, Size, Tang, Zhiqing, Liang, Hongrui, Guo, Jianxiong, Lou, Jiong, Wang, Tian, Jia, Weijia
Abstract
The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. However, existing solutions typically prioritize intra-workflow optimization, largely neglecting the significant potential for inter-workflow optimization. In this paper, we propose HeraSys, an LLM serving system designed to optimize the end-to-end performance of concurrent workflows. Through fine-grained orchestration, HeraSys eliminates cross-workflow computational redundancy via structural node merging and reuse. Furthermore, HeraSys introduces a load-aware joint scheduling policy that dynamically manages execution order by evaluating both inter- and intra-query priorities. By integrating a resource skewing mechanism with adaptive batching and pipeline decomposition, HeraSys effectively mitigates tail latency while maintaining low average latency, thereby substantially improving system throughput. Extensive experiments demonstrate that HeraSys reduces P99 latency by up to 2.17$\times$ and increases serving throughput by up to 1.85$\times$ under strict latency guarantees.
Chinese Translation
大语言模型(LLMs)的快速发展使得服务系统从处理孤立请求转变为协调高并发、多租户的自主工作流。然而,现有解决方案通常优先考虑工作流内部的优化,忽视了工作流之间优化的巨大潜力。本文提出了HeraSys,一个旨在优化并发工作流端到端性能的LLM服务系统。通过细粒度的协调,HeraSys通过结构节点合并和重用消除了跨工作流的计算冗余。此外,HeraSys引入了一种负载感知的联合调度策略,通过评估查询之间和查询内部的优先级动态管理执行顺序。通过将资源偏斜机制与自适应批处理和管道分解相结合,HeraSys有效减轻了尾延迟,同时保持低平均延迟,从而显著提高了系统吞吐量。大量实验表明,在严格的延迟保证下,HeraSys将P99延迟降低了最多2.17倍,将服务吞吐量提高了最多1.85倍。
cs.AI / 23 / 2607.22583

Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization

针对延迟和模型大小优化的多目标结构化剪枝方法
Ali, Muhammad Junaid, Niar, Smail, Talbi, El-Ghazali
Abstract
Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, memory, and energy constraints. Their large parameter counts and computational demands hinder efficient execution on resource-constrained platforms. Although model pruning has emerged as a viable solution for reducing scale while preserving performance, jointly optimizing layers, attention heads, and Multi-Layer Perceptron (MLP) dimensions remains highly complex. Exhaustively exploring this combined design space is computationally expensive and often leads to local optima or unstable configurations. To address these limitations, we propose a hardware-aware, multi-objective structured pruning framework. The proposed two-stage method explicitly targets latency and model size for efficient deployment on edge devices. In the coarse-grained stage, multi-objective depth pruning removes entire attention and MLP blocks to reduce computational load and memory usage. In the subsequent fine-grained stage, Parallel Bayesian Optimization (PBO) searches for the optimal layer-wise pruning ratios for pruning under latency constraints, while importance-based strategies rank the specific components to be pruned within each layer's allocated budget. Experimental results show that our approach reduces model complexity with minimal impact on commonsense reasoning tasks and zero-shot performance. Our method achieves a favorable trade-off among accuracy, latency, and model size, making it suitable for edge deployment. Across multiple LLMs at 37.5% and 50% pruning ratios, the proposed approach achieves better performance on commonsense reasoning tasks than existing methods while significantly reducing inference cost.
Chinese Translation
大型语言模型(LLMs)因其强大的推理和问答能力而得到广泛应用。然而,在嵌入式和边缘计算环境中部署这些模型仍然面临挑战,因为需要满足严格的延迟、内存和能耗限制。它们庞大的参数数量和计算需求阻碍了在资源受限平台上的高效执行。尽管模型剪枝已成为降低规模而保持性能的可行解决方案,但联合优化层、注意力头和多层感知器(MLP)维度仍然非常复杂。全面探索这一组合设计空间计算成本高昂,且常常导致局部最优或不稳定配置。为了解决这些限制,我们提出了一种硬件感知的多目标结构化剪枝框架。该方法的两阶段设计明确针对延迟和模型大小,以实现高效的边缘设备部署。在粗粒度阶段,多目标深度剪枝移除整个注意力和MLP模块,以减少计算负载和内存使用。在随后的细粒度阶段,采用并行贝叶斯优化(PBO)搜索在延迟约束下的最佳层级剪枝比例,同时基于重要性的方法对每层分配预算内的具体剪枝组件进行排序。实验结果表明,我们的方法在对常识推理任务和零-shot性能影响最小的情况下,降低了模型复杂性。我们的方法在准确性、延迟和模型大小之间实现了良好的权衡,适合边缘部署。在37.5%和50%的剪枝比率下,所提方法在常识推理任务上的表现优于现有方法,同时显著降低了推理成本。
cs.AI / 24 / 2607.22584

Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach

源感知重排序用于检索增强生成:一种可靠性先验方法
Koganti, Yuktha Tata, Belinchon, Hugo Garrido-Lestache
Abstract
Standard Retrieval-Augmented Generation pipelines rank retrieved documents by semantic similarity alone, without accounting for source provenance or credibility. This work evaluates a simple and interpretable modification to RAG retrieval ranking that incorporates domain-informed source reliability priors. Each document is assigned a prior lambda(s) based on its source type, and retrieval scores are reweighted using score(q, d) = sim(q, d) * lambda(s). The framework is evaluated against a similarity-only baseline on a 120-document health-domain corpus. In this controlled setting, source-aware reranking improves Precision@5 from 0.48 to 0.72 and reduces average adversarial document retrieval under the evaluated threat model, where low-credibility sources are identifiable via metadata. All experiments were executed on Rosie, the high-performance computing cluster at the Milwaukee School of Engineering, which provided the GPU-accelerated infrastructure necessary to run the full experimental pipeline reliably and reproducibly. These results suggest a potential mitigation strategy for source quality degradation in RAG pipelines, within the limits of the experimental setup described.
Chinese Translation
标准的检索增强生成(RAG)流程仅通过语义相似性对检索到的文档进行排序,而未考虑源的来源或可信度。本文评估了一种简单且可解释的RAG检索排名修改,该修改结合了基于领域的源可靠性先验。每个文档根据其源类型被分配一个先验λ(s),并使用公式score(q, d) = sim(q, d) * λ(s)重新加权检索分数。在一个包含120个文档的健康领域语料库中,该框架与仅基于相似性的基线进行了评估。在这一受控环境中,源感知重排序将Precision@5从0.48提高到0.72,并在评估的威胁模型下减少了平均对抗性文档检索,其中低可信度源可以通过元数据识别。所有实验均在密尔沃基工程学院的高性能计算集群Rosie上执行,该集群提供了运行完整实验流程所需的GPU加速基础设施,确保了实验的可靠性和可重复性。这些结果表明,在所描述的实验设置范围内,源质量下降的潜在缓解策略在RAG流程中是可行的。
cs.AI / 25 / 2607.22585

The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation

编码代理中的支架效应:将选择作为编码代理评估中的隐变量
Vats, Naman, Golev, Oleg
Abstract
Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified. Model-to-model comparison is valid when the harness is fixed; when it varies, performance and efficiency conflate model and scaffold effects. We evaluate Qwen 3.6 Plus and MiniMax M2.5 across three open-source harnesses (Goose, OpenCode, OpenHands-SDK) on a stratified 50-task subset of Terminal-Bench Pro. Harness choice induces up to a 40x difference in tokens per solved task, while paired within-model pass-rate differences remain 0-8 percentage points (95% paired-task bootstrap CIs include zero except for the largest gap). Failure fingerprints replicate across models (REASON for Goose, VERIFY/MAX_TURNS for OpenHands-SDK, idle-loop/TIME for OpenCode), indicating harness-level biases that are largely model-independent. For human-centered coding-agent evaluation, model name alone is an incomplete comparison unit: harness-model pairs determine real-world cost, latency, and oversight burden; no-action turns are a per-task wait tax, not just a token tax. We therefore recommend selecting harness-model pairs by pass rate under token/latency budgets, and reporting token usage, latency, and full harness specifications alongside any model comparison. We release anonymized configs, raw trial logs, aggregated snapshots, and analysis scripts.
Chinese Translation
公共编码代理排行榜通常根据模型名称和通过率对系统进行排名,而周围的支架(提供工具、管理上下文并决定何时停止的支架)往往描述不够明确。当支架固定时,模型间比较是有效的;当支架变化时,性能和效率混淆了模型和支架的效应。我们在 Terminal-Bench Pro 的一个分层 50 任务子集上,评估了 Qwen 3.6 Plus 和 MiniMax M2.5 在三个开源支架(Goose、OpenCode、OpenHands-SDK)下的表现。支架选择导致每个解决任务的标记数差异高达 40 倍,而同模型间的通过率差异保持在 0-8 个百分点(95% 配对任务自助法置信区间包含零,除了最大差距)。失败指纹在模型间重复出现(Goose 的 REASON,OpenHands-SDK 的 VERIFY/MAX_TURNS,OpenCode 的 idle-loop/TIME),表明存在大部分与模型无关的支架级偏差。对于以人为中心的编码代理评估,仅凭模型名称是不完整的比较单位:支架-模型对决定了现实世界的成本、延迟和监督负担;无操作轮次是每个任务的等待税,而不仅仅是标记税。因此,我们建议在标记/延迟预算下,根据通过率选择支架-模型对,并在任何模型比较中报告标记使用、延迟和完整的支架规格。我们发布了匿名配置、原始试验日志、汇总快照和分析脚本。
cs.AI / 26 / 2607.22586

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

MM-ShiftKV:面向多模态大语言模型的解码感知预填充阶段KV选择
Shu, Jinsong, Wu, Chenyang, Xie, Zhongle, Wang, Baokun, Shou, Lidan
Abstract
Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual tokens. Recent prefill-stage KV selection methods estimate KV importance from prefilling statistics, implicitly assuming that prefilling-time queries are representative of those encountered during decoding. We show that this assumption breaks down in multimodal inference, where decoding-time queries exhibit substantially larger variance than prefilling-stage representations, leading to unstable KV importance estimation under tight cache budgets. As a result, small ranking errors can disproportionately discard semantically critical visual tokens and degrade grounding and reasoning performance. We propose MM-ShiftKV, a training-free, decode-aware and strictly prefill-only KV selection method. MM-ShiftKV approximates decoding-time query behavior during prefilling by constructing variance-expanded query proxies and estimates prompt KV importance based on their aggregated attention mass. Experiments on multimodal benchmarks demonstrate that MM-ShiftKV consistently outperforms existing methods under strict KV-cache budgets. Our code is available at https://github.com/zjuDBxAI/MM-ShiftKV.
Chinese Translation
键值(KV)缓存对于多模态大语言模型(MLLMs)中的高效推理至关重要,但其内存占用随着上下文长度线性增长,并因大量视觉标记而成为主要瓶颈。近期的预填充阶段KV选择方法通过预填充统计来估计KV的重要性,隐含地假设预填充时的查询能够代表解码过程中遇到的查询。我们表明,这一假设在多模态推理中失效,因为解码时的查询表现出显著更大的方差,导致在严格的缓存预算下KV重要性估计不稳定。因此,小的排名误差可能会不成比例地丢弃语义上关键的视觉标记,从而降低基础和推理性能。我们提出了MM-ShiftKV,这是一种无训练、解码感知且严格限于预填充阶段的KV选择方法。MM-ShiftKV通过构造方差扩展的查询代理来近似预填充阶段的解码查询行为,并基于其聚合的注意力质量来估计提示KV的重要性。在多模态基准测试中的实验表明,MM-ShiftKV在严格的KV缓存预算下始终优于现有方法。我们的代码可在 https://github.com/zjuDBxAI/MM-ShiftKV 获取。
cs.AI / 27 / 2607.22587

TriSP: Tri-Signal Structured Pruning for Large Language Models

TriSP:大语言模型的三信号结构化剪枝
laoua, Manel Kara, Bouyahiaoui, Soumia, Boutorh, Aicha
Abstract
Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their parameters. Structured pruning addresses this by removing entire structures such as attention heads and Multi-Layer Perceptron (MLP) neurons to produce smaller dense models that run efficiently on standard hardware. However, existing methods rely on either gradient-based importance estimation, which is memory-prohibitive, or activation-based statistical proxies, which do not directly measure the effect of removal on the loss. Furthermore, the interaction between the importance criterion and the post-pruning recovery strategy has not been systematically studied. We propose TriSP (Tri-Signal Structured Pruning), an importance metric that combines weight magnitude scaled by activation norm with first-order gradient sensitivity via a geometric mean, producing a channel-level score that captures both structural and loss-sensitivity signals. Combined with adaptive per-layer budget allocation and low-rank adaptation (LoRA) recovery, TriSP achieves the lowest perplexity and highest zero-shot accuracy across all tested configurations, reaching 6.80 WikiText-2 perplexity at 20% pruning on LLaMA-7B. Inference throughput improves by 82% at 50% pruning, while still maintaining competitive performance.
Chinese Translation
大语言模型(LLMs)在各种任务中表现出色,但其部署受到参数的内存和计算成本的限制。结构化剪枝通过移除整个结构,如注意力头和多层感知器(MLP)神经元,来解决这一问题,从而生成在标准硬件上高效运行的更小的稠密模型。然而,现有方法依赖于基于梯度的重要性估计,这在内存上是不可行的,或者依赖于基于激活的统计代理,这并不能直接衡量移除对损失的影响。此外,重要性标准与剪枝后恢复策略之间的相互作用尚未得到系统研究。我们提出了TriSP(三信号结构化剪枝),这是一种重要性度量,它结合了通过激活范数缩放的权重大小与通过几何平均计算的一阶梯度敏感性,生成一个通道级别的得分,捕捉结构和损失敏感性信号。结合自适应每层预算分配和低秩适应(LoRA)恢复,TriSP在所有测试配置中实现了最低的困惑度和最高的零-shot准确率,在LLaMA-7B上以20%的剪枝达到了6.80的WikiText-2困惑度。在50%的剪枝下,推理吞吐量提高了82%,同时仍保持竞争力的性能。
cs.AI / 28 / 2607.22588

ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation

ParBench:用于可靠评估大规模语言模型并行代码翻译的基准测试
Jhaveri, Samyak, Kaplan, Erel, Yotam, Tom, Chen, Le, Bitan, Tomer, Hasabnis, Niranjan, Oren, Gal
Abstract
Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers, including CUDA, OpenMP, OpenCL, and OpenMP target offload. Large language models and autonomous coding agents are increasingly proposed for such migration, but the field lacks reliable ways to measure whether they preserve the low-level parallel semantics that make translations behaviorally valid, including thread indexing, synchronization, memory management, host-device coordination, and API-specific execution structure. We present ParBench, a kernel-centric benchmark framework for evaluating LLM-based parallel API translation under executable, reproducible conditions. ParBench fixes the surrounding build, run, and verification infrastructure through declarative benchmark specifications and asks models to translate only the computational kernels. It draws on multiple open-source HPC suites and covers representative cross-API translation directions among CUDA, OpenMP, OpenCL, and OpenMP target offload. To test whether success reflects robust translation rather than surface-form memorization, ParBench includes AST-driven, intended behavior-preserving, baseline-validated source augmentation. Evaluations on state-of-the-art open and proprietary LLMs show persistent barriers to reliable parallel code translation, including direction asymmetry, multi-file coordination, incomplete API adaptation, and uneven robustness to source-level perturbations. Code is available at https://github.com/Scientific-Computing-Lab/ParBench.
Chinese Translation
现代计算密集型软件必须在不断变化的加速器、编程API、编译器栈和可移植性层之间迁移,包括CUDA、OpenMP、OpenCL和OpenMP目标卸载。越来越多的研究提出使用大型语言模型和自主编码代理进行此类迁移,但该领域缺乏可靠的方法来衡量它们是否保留了使翻译在行为上有效的低级并行语义,包括线程索引、同步、内存管理、主机-设备协调和特定API的执行结构。我们提出了ParBench,一个以内核为中心的基准测试框架,用于在可执行、可重现的条件下评估基于LLM的并行API翻译。ParBench通过声明性基准测试规范固定了周围的构建、运行和验证基础设施,并要求模型仅翻译计算内核。它借鉴了多个开源高性能计算套件,并涵盖了CUDA、OpenMP、OpenCL和OpenMP目标卸载之间的代表性跨API翻译方向。为了测试成功是否反映了稳健的翻译而非表面形式的记忆,ParBench包括基于抽象语法树(AST)驱动的、旨在保留行为的、基线验证的源代码增强。对最先进的开放和专有LLM的评估显示,可靠的并行代码翻译面临持续的障碍,包括方向不对称、多文件协调、不完整的API适配以及对源级扰动的不均匀鲁棒性。代码可在 https://github.com/Scientific-Computing-Lab/ParBench 获取。
cs.AI / 29 / 2607.22591

Lexical discovery in unknown environments orchestrated by Large Language Models

大型语言模型主导的未知环境中的词汇发现
Sendra-Arranz, Rafael, Varela, Iñaki Dellibarda, Rocon, Eduardo, Gutiérrez, Álvaro, Cebrian, Manuel
Abstract
Populations of autonomous agents deployed in unknown environments (e.g. planetary or deep-sea exploration) must develop shared vocabularies to refer to entities that have no name in any human language. We propose the Neuro-Symbolic Lexical Discovery (NSLD) framework, in which a population of LLM-based agents plays a referential game over out-of-distribution visual referents, autonomously self-organising a shared alien lexicon. Each agent combines a frozen CLIP vision encoder with a private FAISS vector index and a text-only LLM. Crucially, discovered alien words are anchored to natural language via semantic proximity in the embedding space, enlarging the human vocabulary with new perceptually grounded words. Consensus is reached in simulations with populations of up to twenty agents and ten visual referents. Convergence dynamics are characterised through three analytical models achieving R^2 > 0.95, representing a first step towards pre-deployment planning in autonomous exploration missions.
Chinese Translation
部署在未知环境(例如行星或深海探索)中的自主代理群体必须发展共享词汇,以指代在任何人类语言中没有名称的实体。我们提出了神经符号词汇发现(Neuro-Symbolic Lexical Discovery, NSLD)框架,其中基于大型语言模型(LLM)的代理群体在分布外的视觉参照物上进行指称游戏,自主组织共享的外星词汇。每个代理结合了一个冻结的CLIP视觉编码器、一个私有的FAISS向量索引和一个仅文本的LLM。重要的是,发现的外星词汇通过嵌入空间中的语义接近性与自然语言相锚定,从而用新的感知基础词汇扩展人类词汇。在多达二十个代理和十个视觉参照物的模拟中达成共识。收敛动态通过三个分析模型进行表征,达到R^2 > 0.95,代表了在自主探索任务中进行预部署规划的第一步。
cs.AI / 30 / 2607.22592

Structure Over Scale: Schema-Constrained Causal Graphs for RAG

结构与规模:用于检索增强生成的模式约束因果图
Saouda, Marc, Bale, Rajprakash, Aldis, Eren, Almeida, Cloves
Abstract
Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships exhaustively, producing graphs whose size and construction cost scale with corpus length rather than with the reasoning a query requires. We introduce HCG-RAG (Hierarchical Causal Graph RAG), which replaces open-ended extraction with schema-constrained causal graphs: an automated pipeline distills a corpus into a fixed, typed vocabulary of causal variables and materializes a compact two-tier graph over it. Our schema-constrained graphs match entity-relation baselines on answer quality at a fraction of the cost: 3-20x fewer nodes, 8x-135x fewer build-time LLM calls than the most LLM-intensive baseline (MS-GraphRAG), and graphs compact enough for a domain expert to audit, correct, and extend. On medical and clinical benchmarks, including a neurologist-validated epilepsy dataset, HCG-RAG matches or exceeds the best entity-relation systems. An ablation isolates the causal graph as a structured retrieval filter, contributing +6 percentage points (pp) over embedding-only retrieval. Across all domains with discoverable hierarchical causal structure, only methods imposing higher-level organization outperform flat entity-relation retrieval, indicating that what is placed in the graph matters more than how many nodes it contains.
Chinese Translation
基于图的检索增强生成(GraphRAG)将答案建立在结构化知识之上,但当前系统对实体和关系的提取过于全面,导致生成的图的大小和构建成本与语料库长度成正比,而非与查询所需的推理相匹配。我们提出了HCG-RAG(层次因果图检索增强生成),它用模式约束的因果图替代了开放式提取:一个自动化管道将语料库提炼为固定的、类型化的因果变量词汇,并在其上物化一个紧凑的双层图。我们的模式约束图在答案质量上与实体-关系基线相匹配,但成本仅为其一小部分:节点数量减少3-20倍,构建时间的LLM调用减少8倍至135倍,相较于最依赖LLM的基线(MS-GraphRAG),而且图的紧凑程度足以让领域专家进行审计、修正和扩展。在医学和临床基准测试中,包括经过神经科医生验证的癫痫数据集,HCG-RAG的表现与最佳实体-关系系统相当或更优。消融实验将因果图作为结构化检索过滤器进行隔离,相较于仅使用嵌入的检索贡献了6个百分点(pp)。在所有具有可发现层次因果结构的领域中,只有施加更高层次组织的方法优于平面实体-关系检索,这表明图中所放置的内容比节点数量更为重要。
cs.AI / 31 / 2607.22595

xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps

xMIx:高性能机制可解释性应用的服务时间平台
Blum, Michael, Silberstein, Mark, David, Yaniv
Abstract
Mechanistic interpretability (MI) has emerged as a powerful approach for analyzing and intervening in inference computations, with a growing number of applications such as jailbreak attempt detection, truthfulness evaluation, and hallucination detection. Unfortunately, MI deployment in production model-serving systems is currently not practical, as most existing MI frameworks introduce prohibitively high runtime overheads. The fundamental problem is that MI functions do not compose cleanly with served models: they fragment deployment, often force draining requests and rebuilding serving state, and conflict with critical performance optimizations such as continuous batching and CUDA-graph execution, essential for production deployments. We present xMIx, a serving-native framework for deploying MI applications in production inference serving environments. xMIx enables attaching MI functions to a predefined set of locations in the model runtime, interposing on activations within the layers and residual streams. xMIx supports conditional invocation of MI functions depending on the outputs in preceding model layers. Multiple MI applications can be deployed in a single model instance. xMIx compiles them all into the serving path but activates them dynamically at runtime only when necessary, with negligible performance cost, and without requiring a separate model instance or alternative execution stack. We integrate xMIx with the vLLM serving system and evaluate it across three major models and seven diverse MI applications. xMIx achieves performance comparable to native vLLM execution, incurring a slowdown of 1.3% mean inter-token latency (ITL), 1.2% for tail P99 ITL, 2.6% for mean time to first token (TTFT), and 1.6% for mean total token throughput (TTT).
Chinese Translation
机制可解释性(MI)作为一种强大的方法,已逐渐成为分析和干预推理计算的重要手段,应用范围不断扩大,包括越狱尝试检测、真实性评估和幻觉检测等。然而,目前在生产模型服务系统中部署MI并不实际,因为大多数现有的MI框架引入了过高的运行时开销。根本问题在于MI函数与服务模型的组合不够清晰:它们会导致部署碎片化,常常迫使请求被排空并重建服务状态,并与关键的性能优化(如连续批处理和CUDA图执行)发生冲突,而这些对于生产部署至关重要。我们提出了xMIx,这是一个用于在生产推理服务环境中部署MI应用的服务原生框架。xMIx允许将MI函数附加到模型运行时的预定义位置,插入到层内和残差流中的激活上。xMIx支持根据前面模型层的输出有条件地调用MI函数。多个MI应用可以在单个模型实例中部署。xMIx将它们全部编译到服务路径中,但仅在运行时必要时动态激活,性能损失微乎其微,并且不需要单独的模型实例或替代执行栈。我们将xMIx与vLLM服务系统集成,并在三个主要模型和七个不同的MI应用上进行了评估。xMIx的性能与原生vLLM执行相当,平均每个令牌延迟(ITL)仅增加1.3%,尾部P99 ITL增加1.2%,首次令牌平均时间(TTFT)增加2.6%,总令牌吞吐量(TTT)平均增加1.6%。
cs.AI / 32 / 2607.22596

An Agentic Orchestration of Atomistic Simulations

原子级模拟的自主编排
Somasundaram, Rahul, Habib, Adela, Dang, Khanh, Shivakumar, Sachin, Hill, Ryley G., Wimmer, Golo, Mishra, Avanish, Pachalieva, Aleksandra, Lui, Arthur, Viswanathan, Hari, Grosskopf, Michael, Fensin, Saryu, Bent, Russell, DeBardeleben, Nathan, Lawrence, Earl
Abstract
Atomistic simulations are central to materials design, but their execution involves complex, multi-step workflows that require significant human expertise. Here, we present an agent-based system embedded within the URSA (Universal Research and Scientific Agent) framework that automates the design, execution, and validation of atomistic simulations, demonstrated using the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) tool. Our system autonomously selects interatomic potentials, constructs and runs simulations, and performs iterative error recovery within a closed-loop workflow. We evaluate the scientific reliability of the agent by benchmarking its outputs against LAVA, a high-throughput toolkit for LAMMPS and the Vienna Ab initio Simulation Package (VASP) calculations. Our framework reduces manual intervention and trial-and-error, thereby improving the rigor, reproducibility, and scalability of atomistic modeling.
Chinese Translation
原子级模拟在材料设计中至关重要,但其执行涉及复杂的多步骤工作流程,需要显著的人类专业知识。在此,我们提出了一种嵌入在URSA(通用研究与科学代理)框架中的基于代理的系统,该系统自动化了原子级模拟的设计、执行和验证,使用大型原子/分子大规模并行模拟器(LAMMPS)工具进行演示。我们的系统自主选择原子间势,构建并运行模拟,并在闭环工作流程中执行迭代错误恢复。我们通过将其输出与LAVA(LAMMPS的高通量工具包)和维也纳第一性原理模拟包(VASP)计算进行基准测试,评估该代理的科学可靠性。我们的框架减少了人工干预和试错,从而提高了原子级建模的严谨性、可重复性和可扩展性。
cs.AI / 33 / 2607.22597

HyCE-RAG: Hypergraph Chain-of-Evidence Retrieval-Augmented Generation for Explainable Multi-hop Question Answering

HyCE-RAG:用于可解释的多跳问答的超图证据链检索增强生成
An, Hong-Yu, Zhang, Yun-Jian, Liang, Chen-Wei, Zhang, Tian-Yi, Ding, Jian, Wu, Yi-Lun, Li, Ao-Bo, Su, Wei-Cong, Saifullah, Wang, Mujiangshan
Abstract
Multi-hop question answering requires systems to retrieve evidence from multiple documents and connect scattered facts into a coherent reasoning process. Standard retrieval-augmented generation (RAG) mainly relies on semantic similarity between a query and text chunks, and therefore often fails to model structural relations among entities, facts, and evidence units. Graph-based RAG improves this by introducing graph-structured knowledge, but pairwise edges are still limited in representing higher-order associations involving multiple entities and contexts. We propose HyCE-RAG, a Hypergraph Chain-of-Evidence Retrieval-Augmented Generation framework for explainable multi-hop question answering. HyCE-RAG organizes entities, relations, and contextual evidence into hyperedges, builds a query-aware evidence hypergraph, and performs confidence propagation over entity--hyperedge incidence structures. It then uses confidence-guided evidence assembly to select, connect, and rank evidence paths before answer generation. The scoring process jointly considers semantic relevance, entity connectivity, evidence coverage, relation reliability, extraction confidence, and propagated confidence. By providing the language model with structured evidence chains rather than flat retrieved passages, HyCE-RAG supports more faithful and interpretable reasoning. Experiments on HotpotQA, 2WikiMultihopQA, MuSiQue, and two GraphRAG-Bench subsets show that HyCE-RAG consistently outperforms standard RAG and graph-based RAG baselines in answer accuracy, context relevance, and faithfulness. These results suggest that hypergraph-based evidence organization is a promising direction for post-retrieval reasoning in complex question answering.
Chinese Translation
多跳问答要求系统从多个文档中检索证据,并将分散的事实连接成一个连贯的推理过程。标准的检索增强生成(RAG)主要依赖于查询与文本片段之间的语义相似性,因此往往无法建模实体、事实和证据单元之间的结构关系。基于图的RAG通过引入图结构知识来改善这一点,但成对边仍然在表示涉及多个实体和上下文的高阶关联方面有限。我们提出了HyCE-RAG,一种用于可解释的多跳问答的超图证据链检索增强生成框架。HyCE-RAG将实体、关系和上下文证据组织成超边,构建一个查询感知的证据超图,并在实体-超边发生结构上执行置信度传播。然后,它使用置信度引导的证据组装来选择、连接和排序证据路径,随后生成答案。评分过程共同考虑语义相关性、实体连通性、证据覆盖率、关系可靠性、提取置信度和传播置信度。通过向语言模型提供结构化的证据链,而不是平面的检索段落,HyCE-RAG支持更真实和可解释的推理。在HotpotQA、2WikiMultihopQA、MuSiQue和两个GraphRAG-Bench子集上的实验表明,HyCE-RAG在答案准确性、上下文相关性和可信度方面始终优于标准RAG和基于图的RAG基线。这些结果表明,基于超图的证据组织是复杂问答中检索后推理的一个有前景的方向。
cs.AI / 34 / 2607.22599

Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting

针对不确定成分的扩散轨迹差分法在时间序列预测中的应用
Su, Chen, Tian, Yuanhe, Song, Yan
Abstract
Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values given an observed history. In time series forecasting, however, the future continues the observed history, creating an asymmetry the standard diffusion process leaves unaddressed, with slowly-varying content largely determined by the observed continuity while higher-frequency dynamics carry most of the residual uncertainty. Existing diffusion-based forecasters decouple this asymmetry through an external rule before generation, leaving the corruption trajectory blind to which parts of the target the history can already anchor. We propose DiffDiff, a diffusion framework that embeds this predictability asymmetry into the diffusion trajectory itself, so that a single end-to-end diffusion process becomes aware of which parts of the target the history can already anchor. DiffDiff makes the forward operator step-dependent so that the noisy intermediate state progressively shifts from the target itself toward its second-order differenced structure, while a conditioning pathway supplies the denoiser with both value-domain and differential history information balanced by a stage-adaptive gate at each diffusion step. The terminal distribution approaches a standard Gaussian, preserving compatibility with existing samplers. On seven benchmarks across four prediction horizons, DiffDiff outperforms six diffusion baselines, and our analysis confirms that DiffDiff concentrates the diffusion's generative effort on the most uncertain components of the target while relieving it from rebuilding the history-anchored content.
Chinese Translation
扩散模型已成为概率时间序列预测的广泛使用框架,通过建模给定观察历史的未来值分布。然而,在时间序列预测中,未来是对观察历史的延续,这种不对称性是标准扩散过程未能解决的,缓慢变化的内容在很大程度上由观察到的连续性决定,而高频动态则承载了大部分剩余的不确定性。现有的基于扩散的预测方法在生成之前通过外部规则解耦这种不对称性,使得腐蚀轨迹无法识别历史可以锚定目标的哪些部分。我们提出了DiffDiff,这是一种将这种可预测性不对称性嵌入扩散轨迹本身的扩散框架,使得单一的端到端扩散过程能够意识到历史可以锚定目标的哪些部分。DiffDiff使得前向算子依赖于步骤,从而使得噪声中间状态逐步从目标本身转向其二阶差分结构,同时一个条件路径为去噪器提供了价值域和差分历史信息,并在每个扩散步骤通过阶段自适应门进行平衡。在终端分布上,DiffDiff接近标准高斯分布,保持与现有采样器的兼容性。在四个预测视野的七个基准测试中,DiffDiff的表现优于六个扩散基线,我们的分析确认DiffDiff将扩散的生成努力集中在目标中最不确定的成分上,同时减轻了其重建历史锚定内容的负担。
cs.AI / 35 / 2607.22600

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

视觉-语言模型中的图表欺骗:从脆弱性到缓解
Mahbub, Ridwan, Islam, Mohammed Saidul, Laskar, Md Tahmid Rahman, Rahman, Mizanur, Nayeem, Mir Tafseer, Hoque, Enamul
Abstract
Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematically alter interpretation while preserving the underlying data. As Vision-Language Models (VLMs) are increasingly used for chart understanding and analytical reasoning, assessing their robustness to such deceptive visualizations has become critical for trustworthy data analysis. We introduce VisDeception, the first controlled paired benchmark for evaluating the robustness of VLMs to misleading chart designs. The benchmark contains 1,600 paired faithful and misleading charts spanning eight major categories of deceptive visualization tactics, where each misleading chart is paired with a faithful counterpart generated from the same underlying data. To isolate deception-induced reasoning errors from baseline chart-understanding errors, we introduce the Deception Score, a paired evaluation metric that quantifies how misleading visualizations shift model responses away from the faithful interpretation of the data. Across 32,000 responses from 10 state-of-the-art VLMs, we find that even advanced models remain highly vulnerable to deceptive visual manipulations. To improve robustness, we further propose an inference-time multi-agent mitigation framework that grounds reasoning in structured chart metadata extracted from the visualization before answer generation, enabling models to reduce the influence of deceptive visual cues without requiring explicit user instructions. Together, our findings reveal important reliability gaps in current chart-understanding systems and establish benchmark-driven evaluation, deception-aware metrics, and structured reasoning as promising directions for developing more trustworthy VLMs for visual analytics.
Chinese Translation
信息可视化广泛用于传达模式、趋势和异常值,然而,欺骗性的设计选择——如截断或反转的坐标轴、扭曲的纵横比、不当的编码和误导性的颜色映射——可以系统性地改变解读,同时保留基础数据。随着视觉-语言模型(VLMs)在图表理解和分析推理中的日益应用,评估它们对这些欺骗性可视化的鲁棒性已成为可信数据分析的关键。我们引入了VisDeception,这是第一个用于评估VLMs对误导性图表设计鲁棒性的受控配对基准。该基准包含1,600对忠实和误导性的图表,涵盖八大类欺骗性可视化策略,其中每个误导性图表都与从相同基础数据生成的忠实对应图表配对。为了将因欺骗引起的推理错误与基线图表理解错误隔离,我们引入了欺骗评分(Deception Score),这是一种配对评估指标,用于量化误导性可视化如何使模型的响应偏离数据的忠实解读。在来自10个最先进VLM的32,000个响应中,我们发现即使是先进的模型也仍然对欺骗性视觉操控高度脆弱。为了提高鲁棒性,我们进一步提出了一种推理时多智能体缓解框架,该框架在生成答案之前,将推理基于从可视化中提取的结构化图表元数据,使模型能够减少欺骗性视觉线索的影响,而无需明确的用户指令。综上所述,我们的研究揭示了当前图表理解系统中的重要可靠性缺口,并确立了基准驱动的评估、关注欺骗的指标和结构化推理作为开发更可信的视觉分析VLM的有前景的方向。
cs.AI / 36 / 2607.22602

DeepLook: Deeper Thinking with Lookahead

DeepLook:通过前瞻性思维深化推理
Yang, Tingxin, Wang, Zefeng, Wang, Mengyue, Zhou, Xingcheng, Ma, Yunpu
Abstract
Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reasoning tasks than parameter scaling alone. However, existing approaches remain inefficient in how compute is allocated within a reasoning trace. Motivated by the observation that reasoning failures often exhibit an early onset of uncertainty before a wrong answer become explicit, we introduce DeepLook, a training-free monitor-and-intervene decoding framework that concentrates lookahead compute at uncertainty bottlenecks. DeepLook aggregates token-level confidence into segment-level signals, triggers when confidence drops relative to recent history, and explores candidate continuations with fixed-horizon lookahead. Branches are ranked by Average Lookahead Confidence (ALC), the average segment-level confidence over rollout continuations, then pruned and aggregated through voting. On four competition-style mathematics benchmarks across DeepSeek-R1-8B, Qwen3-32B, GPT-OSS-20B, and GPT-OSS-120B, DeepLook shifts the accuracy--token-cost Pareto frontier: it improves accuracy over DeepConf-low in 11 of 16 settings while reducing dataset-level token generation by 87.3% on average, including gains of +3.1 on AIME25 with Qwen3-32B and +8.8 on BRUMO25 with GPT-OSS-20B. These results show that selective, future-aware intervention yields substantially stronger accuracy--cost trade-offs than uniformly scaling complete reasoning trajectories. Code is available here.
Chinese Translation
推理时的扩展已成为改善大型语言模型推理的强大范式,通常在困难的推理任务上比单纯的参数扩展带来更大的收益。然而,现有方法在推理轨迹中计算资源的分配效率仍然较低。基于观察到的推理失败通常在错误答案显现之前就会表现出不确定性的早期征兆,我们提出了DeepLook,这是一种无训练的监控与干预解码框架,专注于不确定性瓶颈处的前瞻性计算。DeepLook将令牌级别的置信度聚合为段级信号,当置信度相对于最近的历史下降时触发,并通过固定范围的前瞻性探索候选延续。分支通过平均前瞻置信度(Average Lookahead Confidence, ALC)进行排名,即在展开延续中的平均段级置信度,然后通过投票进行修剪和聚合。在DeepSeek-R1-8B、Qwen3-32B、GPT-OSS-20B和GPT-OSS-120B的四个竞赛风格数学基准上,DeepLook改变了准确率与令牌成本的帕累托前沿:在16个设置中,DeepLook在11个设置上提高了准确率,同时平均减少了数据集级别的令牌生成87.3%,包括在Qwen3-32B上AIME25提高了+3.1,在GPT-OSS-20B上BRUMO25提高了+8.8。这些结果表明,选择性、面向未来的干预在准确率与成本的权衡上显著优于均匀扩展完整的推理轨迹。代码可在此处获取。
cs.AI / 37 / 2607.22603

Group Preference Collapse in Personalized Multimodal Large Language Models

个性化多模态大型语言模型中的群体偏好崩溃
Lyu, Fan, Zhang, Wenqi, van de Weijer, Joost
Abstract
Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse user preferences. We identify group preference collapse, where multi-user personalized MLLMs become insensitive to individual preferences and drift toward dominant population-level choices due to suppressed preference signals and unreliable preference use during generation. We propose PrefMoE, a preference-centric framework that separates stable profile information from preference-related representations. PrefMoE decomposes preferences into shared prototypes and personalized residuals, preserves individualized residuals with imbalance-aware learning, counterfactual pseudo-user augmentation, and residual decorrelation, and routes profile and preference factors through separate LoRA adaptation paths. Experiments across multiple MLLM backbones show that PrefMoE improves preference-sensitive personalization while substantially reducing preference collapse. Project page: https://prefmoe.github.io/.
Chinese Translation
个性化多模态大型语言模型(MLLMs)旨在生成用户特定的响应,但现有方法主要依赖于档案级信息,忽视了多样化的用户偏好。我们识别出群体偏好崩溃现象,即多用户个性化 MLLMs 对个体偏好变得不敏感,并因偏好信号被压制和生成过程中偏好使用不可靠而向主导的人口级选择漂移。我们提出了 PrefMoE,一个以偏好为中心的框架,旨在将稳定的档案信息与偏好相关的表示分离。PrefMoE 将偏好分解为共享原型和个性化残差,通过关注不平衡的学习、反事实伪用户增强和残差去相关,保留个性化的残差,并通过独立的 LoRA 适应路径引导档案和偏好因素。多个 MLLM 骨干网络的实验表明,PrefMoE 在显著减少偏好崩溃的同时,提高了对偏好的敏感个性化。项目页面:https://prefmoe.github.io/
cs.AI / 38 / 2607.22609

Evaluating LLMs as Interpretable Controllers for Dynamical Systems

评估大型语言模型作为动态系统可解释控制器的能力
Østensen, Aleksander, Calero, Alberto Mino, Lekkas, Anastasios M., Rasheed, Adil
Abstract
Large Language Models (LLMs) are increasingly used for decision-making and reasoning tasks, yet their potential as controllers for physical systems remains largely unexplored. This work investigates whether LLMs can function as interpretable controllers for a dynamic thermal environment, examining their ability to follow setpoints, interpret natural-language commands, reason about actuator effects, and incorporate prior model-based knowledge. Five LLMs of varying scales are evaluated under multiple scenarios, including settings with penalties on heater or fan usage and cases where the models have access to a physics-based prediction tool. The results show that control performance depends on model complexity: while low- and mid-scale models frequently misinterpret actuator dynamics or generate inconsistent reasoning, high-complexity models such as Qwen-3~14B and GPT-4o achieve accurate temperature tracking, stable actuator usage, and coherent explanations aligned with physical principles. Incorporating a physics-based model significantly improves control smoothness and energy efficiency by enabling anticipatory decision-making. A detailed reasoning taxonomy further reveals a clear progression from causal misinterpretation in smaller models to cohesive and temporally aware reasoning in larger ones. The findings demonstrate that LLMs can act as interpretable controllers when sufficiently capable and appropriately grounded in domain knowledge, highlighting promising opportunities for hybrid model-based and language-driven control strategies that can provide plausible explanations.
Chinese Translation
大型语言模型(LLMs)在决策和推理任务中的应用日益增多,但它们作为物理系统控制器的潜力仍然未被充分探索。本研究探讨了LLMs是否可以作为动态热环境的可解释控制器,考察它们跟踪设定点、理解自然语言命令、推理执行器效应以及结合先前基于模型的知识的能力。我们在多种场景下评估了五种不同规模的LLMs,包括对加热器或风扇使用施加惩罚的设置,以及模型访问基于物理的预测工具的情况。结果表明,控制性能依赖于模型的复杂性:低中规模模型经常误解执行器动态或产生不一致的推理,而高复杂性模型如Qwen-3~14B和GPT-4o则实现了准确的温度跟踪、稳定的执行器使用以及与物理原理一致的连贯解释。结合基于物理的模型显著提高了控制的平滑性和能效,使得预期决策成为可能。详细的推理分类法进一步揭示了较小模型在因果误解方面的明显进展,以及较大模型在连贯和时间感知推理方面的提升。研究结果表明,当LLMs具备足够的能力并适当地扎根于领域知识时,可以作为可解释的控制器,这突显了基于混合模型和语言驱动的控制策略的有希望的机会,这些策略能够提供合理的解释。
cs.AI / 39 / 2607.22610

Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations

Tokengeist:代理对话中的多轮归因追踪
Tang, Jessica, Barke, Shraddha, Agarwal, Sharad
Abstract
When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies propagate across prior turns? Existing context attribution methods process the full context in a single pass, recovering surface-level dependencies but missing the layered, non-linear structure of real-world dialogues and multi-step reasoning tasks. We introduce multi-turn context attribution (MTCA): given a target span in a model response, the task of tracing attribution backward across turns to identify not only which prior turns were directly relevant, but also how those turns themselves depended on earlier context. We propose Tokengeist, an attribution-method-agnostic and scalable framework that recovers full dependency paths by casting attribution as a recursive traversal of a directed acyclic graph (DAG) over conversation turns. We will release MTCABench, a benchmark of 3,845 target spans across 665 multi-turn conversations, annotated with gold provenance graphs reaching depths of up to 14, across four dependency types. Across four open-weight models, flat attribution methods fail to recover multi-hop dependencies, achieving under 20% source recall, while Tokengeist reaches 90%. Our results reveal systematic failure modes of single-pass attribution -- which we term provenance collapse -- and motivate attribution methods that reason recursively across turns.
Chinese Translation
当语言模型在多轮对话中生成响应时,之前的哪些词元影响了该答案,以及这些依赖关系是如何在之前的轮次中传播的?现有的上下文归因方法在单次处理完整上下文时,能够恢复表层依赖关系,但却忽视了现实对话和多步推理任务中层次化的非线性结构。我们提出了多轮上下文归因(MTCA):给定模型响应中的目标范围,追踪归因的任务是向后遍历轮次,以识别哪些之前的轮次是直接相关的,以及这些轮次本身是如何依赖于更早的上下文。我们提出了Tokengeist,一个与归因方法无关且可扩展的框架,通过将归因视为在对话轮次上对有向无环图(DAG)的递归遍历,来恢复完整的依赖路径。我们将发布MTCABench,这是一个包含665个多轮对话中3,845个目标范围的基准数据集,附有深度可达14的金标准来源图,涵盖四种依赖类型。在四个开放权重模型中,平面归因方法未能恢复多跳依赖,源召回率低于20%,而Tokengeist达到了90%。我们的结果揭示了单次归因的系统性失败模式——我们称之为来源崩溃——并激励了在轮次间递归推理的归因方法。
cs.AI / 40 / 2607.22611

Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure

关键基础设施中代理人工智能系统的去中心化粒度访问控制
Malik, Arun, Jayasinghe, Deepal, Klemick, Bradley, Shah, Prachi, Talasu, Nitish, Trivedi, Vineet Tushar
Abstract
The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access control (RBAC) models cannot address. Unlike deterministic automation, AI agents exhibit stochastic behavior, making conventional trust models insufficient for governing their access to critical systems. This paper presents a decentralized, multi-layered access control architecture designed specifically for agentic AI systems operating in critical cloud infrastructure. Our framework introduces four key innovations: (1) a compound identity model that binds agent actions to delegated human authority, (2) a hierarchical permission system spanning five granularity levels from global platform access to per-parameter constraints, (3) a decentralized policy ownership model where tool teams independently govern their authorization boundaries, and (4) progressive trust escalation with safety interlocks that prevent autonomous agents from executing high-risk operations. We ground our design in the OWASP Top 10 for LLM Applications (2025) threat taxonomy and demonstrate how each architectural decision mitigates specific attack vectors. Deployed in production at a major cloud provider managing network infrastructure across hundreds of datacenters, the system enforces granular access control for 20+ specialized AI agents and 60+ deterministic playbooks processing thousands of operations daily while maintaining zero unauthorized write operations over eight months of production deployment. We present empirical data on access pattern distributions, denial rates, and the effectiveness of layered authorization in preventing privilege escalation by non-deterministic actors.
Chinese Translation
在生产基础设施中部署自主人工智能代理引入了传统基于角色的访问控制(RBAC)模型无法解决的基本安全挑战。与确定性自动化不同,人工智能代理表现出随机行为,这使得传统的信任模型不足以管理它们对关键系统的访问。本文提出了一种专为在关键云基础设施中运行的代理人工智能系统设计的去中心化多层次访问控制架构。我们的框架引入了四项关键创新:(1)一个复合身份模型,将代理行为与委托的人类权威绑定,(2)一个层级权限系统,涵盖从全球平台访问到每个参数约束的五个粒度级别,(3)一个去中心化的政策所有权模型,工具团队独立管理其授权边界,以及(4)逐步信任升级与安全联锁,防止自主代理执行高风险操作。我们的设计基于OWASP Top 10 for LLM Applications(2025)威胁分类法,并展示了每个架构决策如何减轻特定攻击向量。在一家主要云服务提供商的生产环境中部署,该提供商管理着数百个数据中心的网络基础设施,该系统为20多个专业人工智能代理和60多个确定性剧本实施了粒度访问控制,每天处理数千个操作,并在八个月的生产部署中保持零未授权写操作。我们提供了关于访问模式分布、拒绝率以及分层授权在防止非确定性行为者特权升级方面有效性的实证数据。
cs.AI / 41 / 2607.22614

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

DynaResize:用于分散式大语言模型后训练的运行时 GPU 重新分配
Du, Hanlin, Yan, Zhiyuan, Chen, Haiquan, Fang, Jiarui, Bao, Yungang, wang, Sa
Abstract
RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.
Chinese Translation
基于强化学习的大语言模型后训练日益将 Rollout 和 Training 分散到不同的 GPU 资源上,但静态 GPU 分区在长尾 Rollout 延迟下遭遇严重的管道气泡。我们提出了 DynaResize,一种运行时 GPU 重新分配系统,它动态地在 Rollout 和 Training 之间切换 GPU,以平衡阶段执行时间而不改变强化学习的语义。DynaResize 将调整大小分解为细粒度操作,并通过通信器重用、有限状态分级和基于滞后的调整大小,从关键路径中移除非启动关键工作。实验结果表明,DynaResize 能够将端到端吞吐量提高 66.5%,并在最优静态配置基础上将总执行时间减少 33%,同时隐藏 27% 的角色切换开销。
cs.AI / 42 / 2607.22621

Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning

Opti-Q:一种基于约束的多LLM问题规划优化框架
Hamid, Aamir, Barot, Bharg, Racharla, Satvik, Finin, Tim, Pappachan, Primal, Yus, Roberto
Abstract
While large language models (LLMs) enable strong question answering (QA), budgeted deployment is complicated by nondeterminism and heterogeneous resource profiles (cost, latency, and energy). We present OPTI-Q, a database-inspired, cost-based optimizer that implements a plan-before-execute paradigm for multi-LLM orchestration. OPTI-Q models LLM invocations as physical operators in an execution DAG and, for each question, searches for plans that optimize answer quality (QoA) while trading off financial cost, latency, and energy under user-specified resource constraints. Plans can include sequential operators that pass intermediate answers as context and parallel/blend operators that run models concurrently and merge their outputs. To search this space without executing each candidate plan, OPTI-Q uses PERFDB, a statistics catalog populated and refreshed from benchmarks and execution traces, to estimate the QoA and resource costs of both individual operators and composed subplans. Using these estimates, OPTI-Q performs Pareto-frontier search and selects a final plan based on user preferences. On MMLU-Pro and SimpleQA under user-specified budgets, OPTI-Q improves average QoA by ~58% and ~41% over baselines at comparable cost, demonstrating that database-style planning yields better quality-resource trade-offs for multi-LLM QA.
Chinese Translation
虽然大型语言模型(LLMs)能够实现强大的问答(QA)能力,但由于非确定性和异构资源配置(成本、延迟和能耗),预算部署变得复杂。我们提出了OPTI-Q,这是一种受数据库启发的基于成本的优化器,采用计划优先执行的范式来进行多LLM的协调。OPTI-Q将LLM调用建模为执行有向无环图(DAG)中的物理操作符,并为每个问题搜索优化答案质量(QoA)的计划,同时在用户指定的资源约束下权衡财务成本、延迟和能耗。计划可以包括将中间答案作为上下文传递的顺序操作符,以及并行/混合操作符,这些操作符同时运行模型并合并其输出。为了在不执行每个候选计划的情况下搜索这个空间,OPTI-Q使用PERFDB,这是一个从基准测试和执行轨迹中填充和更新的统计目录,用于估计单个操作符和组合子计划的QoA和资源成本。利用这些估计,OPTI-Q执行帕累托前沿搜索,并根据用户偏好选择最终计划。在用户指定预算下的MMLU-Pro和SimpleQA上,OPTI-Q在可比成本下将平均QoA提高了约58%和41%,证明了数据库风格的规划为多LLM问答提供了更好的质量与资源权衡。
cs.AI / 43 / 2607.22624

CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process

CHS-SQL:基于信心引导启发式搜索的文本到SQL方案链接方法
Yang, Minghao, Xu, Yanjun
Abstract
Recently, there have been several works in the Text-to-SQL domain that utilize Small Language Models (SLMs) for training. These approaches achieve performance close to that of large models in generating SQL, using only the computational power of a single NVIDIA RTX 4090 GPU, while also ensuring data security. Most existing methods filter out redundant tables and columns during Schema Linking to improve Text-to-SQL accuracy. However, they do not consider the precision-recall trade-off when selecting the candidate schema subset. Our research found that both the precision and recall of Schema Linking directly affect the final SQL accuracy. Therefore, we propose a novel framework for efficiently fine-tuning SLMs on Text-to-SQL tasks, CHS-SQL, that not only balances precision and recall but also improves overall performance on Text-to-SQL tasks. Its main innovation lies in the Schema Linking phase, where a heuristic search combined with model internal confidence is employed to achieve an optimal precision-recall trade-off. This elaborated mechanism maximizes the precision of relevant schema candidates for the generated SQL queries while suppressing irrelevant noise. The same strategy is further applied during SQL generation to refine candidate queries while helping the SLM to avoid trapping in a local optimum. Our method achieves state-of-the-art (SOTA) results on Text-to-SQL tasks via SLMs.
Chinese Translation
近年来,在文本到SQL领域中,有多项研究利用小型语言模型(Small Language Models, SLMs)进行训练。这些方法在生成SQL时的性能接近大型模型,仅使用单个NVIDIA RTX 4090 GPU的计算能力,同时确保数据安全。现有大多数方法在方案链接过程中过滤掉冗余的表和列,以提高文本到SQL的准确性。然而,它们在选择候选方案子集时并未考虑精确率与召回率之间的权衡。我们的研究发现,方案链接的精确率和召回率直接影响最终的SQL准确性。因此,我们提出了一种新颖的框架CHS-SQL,用于高效微调SLMs在文本到SQL任务上的表现,该框架不仅平衡了精确率和召回率,还提高了文本到SQL任务的整体性能。其主要创新在于方案链接阶段,采用结合模型内部信心的启发式搜索,以实现最佳的精确率-召回率权衡。这一精细化机制最大化了生成SQL查询的相关方案候选的精确率,同时抑制了无关噪声。在SQL生成过程中同样应用这一策略,以精炼候选查询,并帮助SLM避免陷入局部最优。我们的方法在SLMs的文本到SQL任务上实现了最先进的(SOTA)结果。
cs.AI / 44 / 2607.22625

TokenMem: Faithful Knowledge Injection for Frozen LLMs

TokenMem:为冻结的语言模型注入可靠知识
Yu, Chengzhang, Zheng, Chenyang, Lu, Zening, He, Yingru, Huang, Yutong, Zhang, Yiming, Xu, Yue, Jin, Zhanpeng
Abstract
Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved information contradicts parametric memory, the shared self-attention pathway produces unpredictable outputs. We present TokenMem, a lightweight memory system that injects knowledge into frozen LLMs through a dedicated cross-attention channel, bypassing competition with parametric memory in the residual stream. TokenMem trains only a thin gating adapter ($\sim$3-7M parameters) via a two-phase curriculum: first learning general knowledge utilization, then strengthening faithful compliance under counterfactual knowledge. In controlled experiments on five models spanning three families (Qwen3-4B/8B/14B, LLaMA-3.1-8B, OLMo-3-7B), TokenMem achieves 69-70% Knowledge Compliance (KC) on counterfactual benchmarks, compared to 20-52% for vanilla RAG, a gap of up to 49 percentage points. Ablation studies show that the two-phase curriculum is critical: removing Phase 2 collapses KC to near-zero. Mechanistic analysis reveals that the gate adapter learns a conflict-aware, layer-specific injection strategy without explicit supervision.
Chinese Translation
检索增强生成(RAG)通过外部知识增强大型语言模型(LLMs),但面临知识冲突的问题:当检索到的信息与参数记忆相矛盾时,共享的自注意力通道会产生不可预测的输出。我们提出了TokenMem,这是一种轻量级的记忆系统,通过专用的交叉注意力通道将知识注入冻结的LLMs,从而避免在残差流中与参数记忆的竞争。TokenMem仅通过一个两阶段的课程训练一个薄的门控适配器(约3-7M参数):首先学习一般知识的利用,然后在反事实知识下加强可靠的遵从性。在对五个模型(Qwen3-4B/8B/14B, LLaMA-3.1-8B, OLMo-3-7B)进行的控制实验中,TokenMem在反事实基准上实现了69-70%的知识遵从性(KC),而普通RAG仅为20-52%,差距高达49个百分点。消融研究表明,两阶段课程至关重要:去除第二阶段会使KC降至接近零。机制分析表明,门控适配器学习了一种冲突感知的、层特定的注入策略,而无需显式监督。
cs.AI / 45 / 2607.22629

Masked Distillation: Internalizing the Chain-of-Thought in Language Models

掩蔽蒸馏:在语言模型中内化思维链
Kalwar, Durgesh, Palod, Vardhan, Kambhampati, Subbarao
Abstract
Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though the final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises a natural question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers directly (or with much shorter intermediate traces)? We introduce \textit{masked distillation}, a knowledge-distillation framework in which a student LLM is trained to predict only the solution tokens conditioned on the question, while a reasoning teacher provides feedback on the student's responses after conditioning on the question and its own CoT trace. We instantiate this framework in two settings: (i) a \textit{self-distillation} setting, in which the same model serves as the teacher in thinking mode and as the student in non-thinking mode, and (ii) a \textit{dual-model} setting, in which a larger reasoning teacher supervises a separate smaller non-thinking student over the solution tokens. By treating intermediate tokens as a scaffold which reasoning models use to fit over the solution tokens, We additionally vary the length of intermediate-token scaffolding the student is supervised on, interpolating between full internalization (the student emits only the solution) and no internalization (the student emits the full trace before the answer). We evaluate the framework through controlled experiments on two reasoning domains: GSM8K (grade-school arithmetic) and Countdown (a number-puzzle search task).
Chinese Translation
大型推理模型(LRMs)在推理时生成长且明确的中间步骤链,最终得出答案。这些中间痕迹主导了延迟、内存使用和服务成本,尽管最终答案的正确性与痕迹的正确性并无因果关系,且痕迹长度并不是问题复杂性的可靠指示。这引发了一个自然的问题:能否将这些中间标记中表达的计算内化到语言模型的参数中,使其能够直接生成答案(或使用更短的中间痕迹)?我们提出了 extit{掩蔽蒸馏},一种知识蒸馏框架,其中学生LLM在条件为问题的情况下,仅预测解决方案标记,而推理教师在条件为问题及其自身的思维链(CoT)痕迹后,对学生的回答提供反馈。我们在两种设置中实例化该框架:(i) extit{自我蒸馏}设置,其中同一模型在思考模式下充当教师,在非思考模式下充当学生;(ii) extit{双模型}设置,其中一个更大的推理教师监督一个单独的较小的非思考学生,针对解决方案标记。通过将中间标记视为推理模型用来拟合解决方案标记的支架,我们还改变了学生所监督的中间标记支架的长度,在完全内化(学生仅输出解决方案)和无内化(学生在答案之前输出完整痕迹)之间进行插值。我们通过在两个推理领域(GSM8K(小学算术)和Countdown(数字谜题搜索任务))进行控制实验来评估该框架。
cs.AI / 46 / 2607.22632

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

VlogReward:学习多维度评估以优化Vlog编辑
Liu, Yexiang, Zhong, Wen, Zhu, Sijie, Gu, Xin, Chen, Fan, Duan, Junxian, Cao, Jie, Wen, Longyin, Chen, Zhenfang
Abstract
The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions, i.e., Creativity, Consistency, Concept Design, Cinematography, Narration, and Pacing. Subsequently, we curate a large-scale dataset of 100k vlog edits and a dedicated benchmark, VRMBench, to evaluate the vlog rewarding capabilities of Multimodal Large Language Models (MLLMs). Finally, we present VlogReward, a robust vlog reward model that can provide both fine-grained multi-dimensional scores and actionable feedback for iterative refinement. Technically, we enhance the Group Relative Policy Optimization (GRPO) framework by introducing an adjustable inter-group comparison reward, which mitigates the "direction blindness" issue of standard GRPO and enables the model to better distinguish varied-quality edits. VlogReward achieves state-of-the-art results that significantly outperform existing MLLMs, including GPT-5 and Gemini-3-Pro. We hope that our study can help vlog creators and foster automated vlog evaluation and refinement systems.
Chinese Translation
Vlog作为一种个性化叙事媒介的快速崛起,催生了对自动化系统以评估和优化Vlog编辑计划的需求。然而,Vlog评估高度主观,因缺乏标准化的评估标准、数据集和基准,以及有效的奖励模型而面临挑战。为了解决这些问题,我们定义了一个全面的Vlog评估框架,该框架由专业Vlog创作者和产品经理指导,建立了六个关键维度的分类法,即创造力(Creativity)、一致性(Consistency)、概念设计(Concept Design)、摄影(Cinematography)、叙述(Narration)和节奏(Pacing)。随后,我们整理了一个包含10万条Vlog编辑的大规模数据集和一个专门的基准VRMBench,以评估多模态大型语言模型(Multimodal Large Language Models, MLLMs)在Vlog奖励能力上的表现。最后,我们提出了VlogReward,一个强大的Vlog奖励模型,能够提供细粒度的多维度评分和可操作的反馈,以便进行迭代优化。从技术上讲,我们通过引入可调的组间比较奖励来增强组相对策略优化(Group Relative Policy Optimization, GRPO)框架,这缓解了标准GRPO的“方向盲”问题,使模型能够更好地区分不同质量的编辑。VlogReward在性能上达到了最先进的结果,显著超越了现有的MLLM,包括GPT-5和Gemini-3-Pro。我们希望我们的研究能够帮助Vlog创作者,并促进自动化Vlog评估和优化系统的发展。
cs.AI / 47 / 2607.22633

Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering

从经验中演变:基于技能增强的操作级表格推理用于表格问答
Su, Guixin, Pi, Qiankun, Xu, Mayi, Li, Wenli, Zhong, Ming, Zhu, Yuanyuan, Jiang, Jiawei, Qian, Tieyun
Abstract
Table Question Answering (TableQA) aims to reason over tables to answer user queries. Existing research treats all questions uniformly and evaluates solely through overall accuracy, obscuring a critical reality that LLMs excel at simple lookups yet struggle with complex operations like aggregation and arithmetic. To reveal this disparity, we introduce a novel \emph{Operation-wise TableQA} task with a fine-grained question taxonomy and release two datasets named WikiTQ-ow and TabFact-ow for evaluation. As for modeling bottlenecks, existing methods flatten tables into linearized texts, disrupting inherent structures and inducing the ``lost-in-the-middle'' issue, which poses a primary barrier to complex cross-row reasoning. Moreover, they typically reason from scratch, neglecting reusable patterns shared across similar operations. To address these limitations, we propose a Skill-augmented Table Graph Reasoning (SkillTGR) framework for self-evolving structured reasoning. Specifically, SkillTGR represents tables as attributed graphs with explicit row-column-cell structures, where LLMs plan and execute dynamic chains to retrieve evidence subgraphs for graph traversal reasoning. Based on this, SkillTGR builds a hierarchical SkillBank to distill reason trajectories into abstract skills under cognitive heuristics, then hybrid retrieves both successful and failed skills for contrastive augmented table graph reasoning, thereby enabling the continual self-evolution. Extensive experiments demonstrate that SkillTGR achieves superior performance with an average of 5.91\% overall and 6.03\% operation-wise improvement, also reducing 19.76\% token consumption and 27.64\% inference latency. Our codes and data will be released upon publication.
Chinese Translation
表格问答(TableQA)旨在对表格进行推理以回答用户查询。现有研究将所有问题视为统一,并仅通过整体准确率进行评估,这掩盖了一个关键现实:大型语言模型(LLMs)在简单查找方面表现出色,但在聚合和算术等复杂操作上却面临挑战。为了揭示这种差异,我们引入了一种新颖的操作级表格问答(Operation-wise TableQA)任务,采用细粒度的问题分类法,并发布了两个用于评估的数据集,分别命名为WikiTQ-ow和TabFact-ow。针对建模瓶颈,现有方法将表格展平为线性文本,破坏了固有结构,并引发了“迷失在中间”的问题,这成为复杂跨行推理的主要障碍。此外,它们通常从头推理,忽视了在类似操作中共享的可重用模式。为了解决这些局限性,我们提出了一种技能增强的表格图推理框架(Skill-augmented Table Graph Reasoning, SkillTGR),用于自我演变的结构化推理。具体而言,SkillTGR将表格表示为具有明确行-列-单元格结构的属性图,其中LLMs规划并执行动态链以检索证据子图进行图遍历推理。基于此,SkillTGR构建了一个层次化的技能库(SkillBank),将推理轨迹提炼为认知启发下的抽象技能,然后混合检索成功和失败的技能以进行对比增强的表格图推理,从而实现持续的自我演变。大量实验表明,SkillTGR在整体性能上平均提高了5.91%,操作级性能提高了6.03%,同时减少了19.76%的令牌消耗和27.64%的推理延迟。我们的代码和数据将在发表时发布。
cs.AI / 48 / 2607.22634

PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding

PRESTO:用于扩散推测解码的前缀对齐树草拟
Wang, Zheng, Ye, Zhifan, Cheng, Qi, Fu, Yonggan, Wang, Ziyan, Zhu, Feng, Zhao, Haozhe, Kautz, Jan, Molchanov, Pavlo, Shi, Humphrey, Zhang, Minjia
Abstract
Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, generating tokens in parallel. This makes them effective draft models for speculative decoding (SD), producing an entire block of draft tokens in a single forward pass. Yet existing diffusion-based drafting methods rely on linear drafting, even though dLLMs emit multiple candidate tokens across positions, inducing a large combinatorial space of decoding paths. Consequently, they limit acceptance length and decoding efficiency. To exploit this multi-candidate structure, we apply tree-based drafting to diffusion drafters, enabling exploration of diverse candidate paths. However, we find that naive tree drafting is suboptimal: diffusion marginals are prefix-blind, mismatching the prefix-based AR verification and yielding unreliable path ranking. We propose PRESTO, a principled framework that extends tree-based drafting to diffusion drafters while resolving the fundamental mismatch between diffusion draft confidence and prefix-based AR verification through PREfix-aligned Scoring and priority-based Tree search for diffusion speculative decOding. The key principles behind PRESTO are that (1) candidate ranking should align with the prefix-based nature of AR verification, and (2) tree construction should prioritize candidate paths with high verification potential to maximize acceptance length. Extensive experiments show that PRESTO achieves up to an average of $1.5\times$ end-to-end throughput speedup on the state-of-the-art dedicated diffusion drafter SD and an average of $1.12\times$ on self-speculative diffusion LLMs across diverse benchmarks.
Chinese Translation
扩散大型语言模型(dLLMs)作为自回归(AR)LLMs的有希望的替代方案,能够并行生成标记。这使得它们成为推测解码(SD)的有效草拟模型,能够在一次前向传递中生成整个草拟标记块。然而,现有的基于扩散的草拟方法依赖于线性草拟,尽管dLLMs在各个位置发出多个候选标记,导致解码路径的组合空间巨大。因此,它们限制了接受长度和解码效率。为了利用这种多候选结构,我们将基于树的草拟应用于扩散草拟器,从而能够探索多样的候选路径。然而,我们发现简单的树草拟并不是最优的:扩散边际对前缀是盲目的,与基于前缀的AR验证不匹配,导致路径排名不可靠。我们提出了PRESTO,一个原则性框架,扩展了基于树的草拟到扩散草拟器,同时通过前缀对齐评分(PREfix-aligned Scoring)和基于优先级的树搜索(priority-based Tree search)解决了扩散草拟置信度与基于前缀的AR验证之间的根本不匹配。PRESTO的关键原则是(1)候选排名应与基于前缀的AR验证的特性对齐,以及(2)树构建应优先考虑具有高验证潜力的候选路径,以最大化接受长度。大量实验表明,PRESTO在最先进的专用扩散草拟器SD上实现了高达$1.5 imes$的端到端吞吐量加速,在自我推测扩散LLMs的多样基准上实现了平均$1.12 imes$的加速。
cs.AI / 49 / 2607.22635

CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

CallBench:一个用于电话助手双目标协调的基准测试
Geng, Xuzhao, Wang, Haozhao, Li, Xuelian, Yang, Zhenyu, Lu, Haonan, Zhang, Rui, Li, Ruixuan
Abstract
Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations. However, existing studies are primarily designed for single, explicit goal completion, while phone call assistants face a proxy setting that requires coordinating the device owner's explicit preset goal with the caller's implicit and dynamic goal. We introduce \textsc{CallBench}, a Chinese benchmark for evaluating dual-goal coordination in phone call assistants. \textsc{CallBench} contains 50,000 complete multi-turn phone call dialogues across six scenarios: takeout, delivery, taxi, work, life, and harassment. It covers regular presets, emergent presets, and no-preset cases, and includes diverse relations between owner-side and caller-side goals, such as alignment, complementarity, irrelevance, and conflict. We further design a preset-aware turn-level evaluation protocol covering semantic understanding, context use, active guidance, response quality, preset compliance, dialogue rhythm, and safety. Experiments on representative dialogue methods show that existing approaches still struggle with this task, highlighting the need for phone call assistants that can make reliable turn-level decisions between two independent goals under proxy constraints.
Chinese Translation
以目标为导向的对话系统在通过互动对话完成用户目标方面展现了强大的能力。然而,现有研究主要针对单一、明确的目标完成,而电话助手面临一种代理设置,需要协调设备所有者的明确预设目标与来电者的隐含和动态目标。我们介绍了 extsc{CallBench},一个用于评估电话助手双目标协调的中文基准测试。 extsc{CallBench} 包含了在六种场景下的50,000个完整的多轮电话对话:外卖、配送、出租车、工作、生活和骚扰。它涵盖了常规预设、突发预设和无预设情况,并包括所有者侧与来电者侧目标之间的多种关系,如一致性、互补性、无关性和冲突。我们进一步设计了一种考虑预设的轮次级评估协议,涵盖语义理解、上下文使用、主动引导、响应质量、预设遵循、对话节奏和安全性。对代表性对话方法的实验表明,现有方法在此任务上仍然面临挑战,突显了需要能够在代理约束下在两个独立目标之间做出可靠轮次级决策的电话助手。
cs.AI / 50 / 2607.22636

Answering Path Queries under Linear and Guarded Existential Rules

在线性和受限存在规则下回答路径查询
Baget, Jean-François, Bienvenu, Meghyn, Mugnier, Marie-Laure, Thomazo, Michaël
Abstract
Ontology-mediated query answering is concerned with the problem of answering queries over knowledge bases consisting of a database instance and an ontology. While most work in the area focuses on conjunctive queries (CQs), navigational queries have gained increasing attention. In this paper, we investigate the complexity of answering two-way (conjunctive) regular path queries ((C)RPQs) over knowledge bases whose ontology is given by a set of guarded existential rules. We first consider the subclass of linear existential rules and show that (C)RPQ answering is NL-complete in data complexity, which matches the data complexity of answering RPQs over plain graph databases (i.e., without an ontology). In combined complexity, both tasks are ExpTime-complete in the general case, but RPQ and CRPQ answering drop to PTime-complete and PSpace-complete respectively if there is a bound on predicate arity. For guarded rules, we provide a non-trivial reduction to the linear case, which allows us to show that the complexity of (C)RPQ answering is the same as for CQs, namely 2ExpTime-complete in combined complexity (ExpTime-complete in the bounded-arity case) and PTime-complete in data complexity.
Chinese Translation
本体介导的查询回答关注于在由数据库实例和本体组成的知识库上回答查询的问题。尽管该领域的大多数研究集中在合取查询(CQs)上,但导航查询却越来越受到关注。本文研究了在由一组受限存在规则给定的本体下,回答双向(合取)正则路径查询((C)RPQs)的复杂性。我们首先考虑线性存在规则的子类,并证明在数据复杂性方面,(C)RPQ 的回答是 NL 完全的,这与在普通图数据库(即没有本体的情况下)上回答 RPQs 的数据复杂性相匹配。在综合复杂性方面,两个任务在一般情况下都是 ExpTime 完全的,但如果对谓词的元数有界,RPQ 和 CRPQ 的回答分别降为 PTime 完全和 PSpace 完全。对于受限规则,我们提供了一个非平凡的归约到线性情况,这使我们能够证明 (C)RPQ 的回答复杂性与 CQs 相同,即在综合复杂性中为 2ExpTime 完全(在有界元数情况下为 ExpTime 完全),在数据复杂性中为 PTime 完全。
cs.AI / 51 / 2607.22637

Fast Cross-Scenario Adaptation of CSI Models via Channel Conditional Parameter Generation

通过信道条件参数生成实现CSI模型的快速跨场景适应
Zou, Xudong, Wu, Siyu, Feng, Zunlei, Song, Jie, Wan, Yuanyu, Song, Mingli, Hu, Jiacong
Abstract
Deep learning has shown strong potential for massive multiple-input multiple-output (Massive MIMO) physical-layer tasks, including channel state information (CSI) feedback and channel estimation. However, environmental heterogeneity can severely degrade CSI models in unseen scenarios, while conventional adaptation requires target-domain data and substantial computation. This paper proposes Channel Conditional Parameter Generation (CCPG), an end-to-end pipeline for rapid deployment of CSI models in dynamic wireless environments. CCPG identifies scene-sensitive adaptation bottlenecks through component-freezing experiments and generates only lightweight LoRA weights instead of full model parameters. It compresses high-dimensional channel features into compact latent conditions using cascaded SVD and a Perceiver Resampler. An energy-based canonicalization mechanism mitigates permutation and sign ambiguities in LoRA weights, while a diffusion-based generator incorporates structural information and an asymmetric size-aware loss for topology-aware parameter generation. Experiments on DeepMIMO and WAIR-D for CSI feedback and channel estimation show that CCPG adapts to new scenarios in about 3 seconds with a single forward pass, without target-scenario training or fine-tuning, and achieves cross-domain recovery performance comparable to costly online adaptation. These results demonstrate that CCPG enables efficient deployment of CSI models in large-scale dynamic wireless scenarios for intelligent 6G communications.
Chinese Translation
深度学习在大规模多输入多输出(Massive MIMO)物理层任务中展现了强大的潜力,包括信道状态信息(CSI)反馈和信道估计。然而,环境异质性可能严重降低在未见场景中的CSI模型性能,而传统的适应方法需要目标领域数据和大量计算。本文提出了信道条件参数生成(Channel Conditional Parameter Generation,CCPG),这是一个用于在动态无线环境中快速部署CSI模型的端到端管道。CCPG通过组件冻结实验识别场景敏感的适应瓶颈,并仅生成轻量级的LoRA权重,而不是完整的模型参数。它利用级联奇异值分解(SVD)和感知重采样器将高维信道特征压缩为紧凑的潜在条件。基于能量的标准化机制缓解了LoRA权重中的排列和符号歧义,而基于扩散的生成器则结合了结构信息和不对称的尺寸感知损失,以实现拓扑感知的参数生成。在DeepMIMO和WAIR-D上进行的CSI反馈和信道估计实验表明,CCPG在大约3秒内通过单次前向传播适应新场景,无需目标场景的训练或微调,并且实现了与昂贵的在线适应相当的跨领域恢复性能。这些结果表明,CCPG能够高效地在大规模动态无线场景中部署CSI模型,为智能6G通信提供支持。
cs.AI / 52 / 2607.22639

TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs

TRACE:基于业务规则的推理课程用于企业大型语言模型中的知识保留参数化工具检索
Sistla, Sai Shruthi, Hathidara, Ashutosh, Toukmaji, Christopher, Shrivastava, Mayank, Asokkumar, Karthikeyan
Abstract
Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search. Toolsense shows that this regime has two critical drawbacks: it destroys parametric tool knowledge during training, and its beam-search decoding is too slow for real-time deployment. We introduce TRACE (Tool Retrieval via Augmented Chain-of-thought and Enterprise rules), a two-stage curriculum that resolves this dissociation. Stage 1 reuses the multi-format memorization SFT from ToolSense to seed tool knowledge with LoRA. Stage 2 is our core contribution: the model is trained to emit a thinking trace before producing a JSON list of tool tokens, using two data sources -- RRB pairs from ToolSense and queries synthesized to target business rules curated by domain experts -- both augmented with reasoning traces. This training objective preserves Stage 1 MCQ and QA probing accuracy while enabling single-beam greedy decoding at production latency. Evaluated on a combined enterprise catalog of 8,300+ tools across two enterprise product lines, TRACE training for Stage 2 not only preserves but improves tool understanding: MCQ accuracy gains +3.2 pp and QA probing gains +9 pp over Stage 1. On retrieval, TRACE achieves ~86% recall on Domain A and ~60% on Domain B -- compared to embedding baseline performance of ~27% & ~52% -- both with single-beam greedy decoding, making it directly deployable at production latency.
Chinese Translation
参数化检索使大型语言模型(LLMs)能够通过为每个API分配一个唯一的虚拟令牌并训练模型通过受限束搜索生成该令牌来隐式检索工具。Toolsense显示,这种机制存在两个关键缺陷:在训练过程中破坏了参数化工具知识,并且其束搜索解码速度对于实时部署来说过于缓慢。我们引入了TRACE(通过增强的思维链和企业规则进行工具检索),这是一个两阶段的课程,解决了这种脱离。第一阶段重用来自ToolSense的多格式记忆SFT,通过LoRA为工具知识提供种子。第二阶段是我们的核心贡献:模型在生成工具令牌的JSON列表之前被训练发出思维轨迹,使用两个数据源——来自ToolSense的RRB对和由领域专家策划的目标业务规则合成的查询——这两者都增强了推理轨迹。这个训练目标在保留第一阶段多项选择题(MCQ)和问答(QA)探测准确性的同时,使得在生产延迟下能够实现单束贪婪解码。在对跨两个企业产品线的8300多个工具的综合企业目录进行评估时,TRACE第二阶段的训练不仅保留了工具理解,还提高了工具理解:MCQ准确性提高了3.2个百分点,QA探测提高了9个百分点。检索方面,TRACE在领域A上实现了约86%的召回率,在领域B上实现了约60%——相比之下,嵌入基线性能约为27%和52%——两者均采用单束贪婪解码,使其能够直接在生产延迟下部署。
cs.AI / 53 / 2607.22642

CRAFT: Learn the Schema, Execute the Plan

CRAFT:学习模式,执行计划
Kolekar, Aakash, Genc, Sahika, Shariat, Shahriar, Sisman, Bunyamin, Mezi, Tibor, Poblete, Barbara, Kachroo, Shree Vandana, Chi, Calvin, Parmar, Parth, Singer, Ari, Jain, Prayaas, Barker, Cindy, Dumoulin, Benoit
Abstract
Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet the prevailing deployment pattern injecting exhaustive schema and tool documentation into each prompt increases inference overhead, complicates schema evolution, and undermines reliability in multi-turn analysis. We investigate whether stable schema knowledge and tool-use behavior can instead be acquired through post-training while preserving the consistency required for production-facing analytics. We present CRAFT, a two-stage post-training recipe for schema-grounded coding agents. First, schema-stripped PLAN supervised fine-tuning learns domain-structured plans and executable behaviors from validated trajectories without exhaustive prompt-time schema injection. Second, execution-shaped reinforcement learning aligns the policy for tool selection, code quality, plan-code consistency, and recovery from failed executions. Training trajectories are curated through a Tri-Gate filter combining execution validation, data-integrity checks, and LLM-judge reasoning audit. We evaluate CRAFT for planned rollout in advertising analytics, covering campaign performance analysis, metric drill-downs, entity-level performance analysis, and multi-turn analytical refinement. The enterprise evaluation environment incorporates beta APIs as the agent-facing tool surface and spans 25 schema-linked core entities and 30 agentic workflows. Relative to a schema-stuffed baseline, CRAFT improves composite Agent Score by +9.6 pp, consistency by +4.1 pp, and multi-turn coherence by +4.2 pp, while reducing input-token burden by approximately 9x and schema-discovery loops by up to 5x. We further report deployment tradeoffs, reward-shaping limitations, and training-infrastructure extensions required for multi-turn tool-use reinforcement learning in enterprise settings.
Chinese Translation
企业编码代理将自然语言分析请求转换为可执行代码,涉及专有API、模式和指标定义。然而,当前的部署模式在每个提示中注入详尽的模式和工具文档,增加了推理开销,复杂化了模式演变,并削弱了多轮分析中的可靠性。我们研究了是否可以通过后训练获得稳定的模式知识和工具使用行为,同时保持生产分析所需的一致性。我们提出了CRAFT,一种针对基于模式的编码代理的两阶段后训练方案。首先,去除模式的PLAN监督微调从经过验证的轨迹中学习领域结构化的计划和可执行行为,而无需在提示时注入详尽的模式。其次,执行形状的强化学习对工具选择、代码质量、计划与代码的一致性以及从失败执行中恢复的策略进行对齐。训练轨迹通过结合执行验证、数据完整性检查和LLM-judge推理审计的Tri-Gate过滤器进行策划。我们评估了CRAFT在广告分析中的计划推广,涵盖了活动绩效分析、指标深入分析、实体级绩效分析和多轮分析精细化。企业评估环境结合了作为代理工具界面的beta API,涵盖25个与模式相关的核心实体和30个代理工作流。与模式填充的基线相比,CRAFT将复合代理评分提高了9.6个百分点,一致性提高了4.1个百分点,多轮连贯性提高了4.2个百分点,同时将输入令牌负担减少了约9倍,将模式发现循环减少了最多5倍。我们还报告了部署权衡、奖励塑造的局限性以及在企业环境中进行多轮工具使用强化学习所需的训练基础设施扩展。
cs.AI / 54 / 2607.22643

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

检索之前的推理:多模态 RAG 的自主规划
Yang, Tianyu, Simon, Shir, Li, Zhenzhen, Cheng, Minhao, Zhang, Xiangliang
Abstract
Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space. This design often struggles with two key challenges: the retrieval target is under-specified because the question intent must be grounded to the correct visual referent, and the search space is weakly structured, forcing semantically distinct evidence to compete in a single global ranking step. We propose MM-R2, a multimodal agentic retrieval framework that reasons before retrieval by explicitly modeling both what to retrieve and where to search. MM-R2 first constructs an intent-grounded retrieval state from the image-question pair, capturing the information need, grounded referent, and retrieval constraints. It then performs retrieval over a structured KnowledgeMap, where the agent selects relevant retrieval units before issuing grounded queries within them. To enable this capability, we build MM-R2-Traj, a large-scale trajectory dataset of multi-step retrieval processes, and adopt a two-stage post-training strategy with supervised fine-tuning and GRPO. Experiments on Infoseek and Encyclopedic VQA datasets show that MM-R2 substantially outperforms strong baselines on answer accuracy while also yielding more interpretable and verifiable retrieval trajectories.
Chinese Translation
多模态检索增强生成 (mRAG) 旨在利用外部知识回答图像-文本查询,但现有大多数系统仍直接从原始多模态输入中在平坦的证据空间中进行检索。这种设计通常面临两个关键挑战:检索目标不明确,因为问题意图必须与正确的视觉参照物相结合,并且搜索空间结构较弱,迫使语义上不同的证据在单一的全局排名步骤中竞争。我们提出了 MM-R2,一种多模态自主检索框架,通过明确建模检索内容和搜索位置,在检索之前进行推理。MM-R2 首先从图像-问题对构建一个意图驱动的检索状态,捕捉信息需求、基础参照物和检索约束。然后,它在结构化的知识图谱 (KnowledgeMap) 上进行检索,在此过程中,代理选择相关的检索单元,然后在这些单元内发出有依据的查询。为了实现这一能力,我们构建了 MM-R2-Traj,一个大规模的多步骤检索过程轨迹数据集,并采用了一个包含监督微调和 GRPO 的两阶段后训练策略。在 Infoseek 和百科全书 VQA 数据集上的实验表明,MM-R2 在答案准确性上显著优于强基线,同时也产生了更具可解释性和可验证性的检索轨迹。
cs.AI / 55 / 2607.22644

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

DocHRL:一种用于成本优化文档分类的层次强化学习框架
Yousif, Mohammed, Singh, Prabhjot, Pankajakshan, Arjun, Reddiboina, Madhu
Abstract
Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type. This leads to inefficient use of compute and human resources: simple documents are over-processed while difficult ones may not receive enough scrutiny. We introduce DocHRL, a hierarchical reinforcement learning framework that learns to adaptively and dynamically select the most cost-effective classification policy on a per-document basis. DocHRL formulates document classification as a sequential decision problem with a two-level policy hierarchy: a top-level policy selects among broad options (vision classifiers, LLMs, OCR, and human-in-the-loop review), while option-specific sub-policies choose the concrete model or tool to invoke. The reward signal is the negative total expected cost, which captures inference cost, cost of misclassification, and cost of human labelling. Trained with Proximal Policy Optimisation on the RVL-CDIP benchmark, DocHRL achieves a macro F1 of 0.973 across 16 document classes while reducing average per-document cost to 2.74 normalised units compared to substantially higher costs incurred by fixed standalone classifiers. Our results demonstrate that cost-aware reinforcement learning can simultaneously improve classification performance and operational efficiency in document understanding systems.
Chinese Translation
现实世界中的文档分类流程通常对每个传入文档应用相同的模型序列,而不考虑其复杂性或类型。这导致计算和人力资源的低效使用:简单文档被过度处理,而复杂文档可能得不到足够的审查。我们提出了DocHRL,一种层次强化学习框架,能够在每个文档的基础上自适应和动态地选择最具成本效益的分类策略。DocHRL将文档分类公式化为一个具有两级策略层次的序列决策问题:顶层策略在广泛选项(视觉分类器、LLMs、OCR和人类审查)中进行选择,而特定选项的子策略则选择具体的模型或工具进行调用。奖励信号是负的总预期成本,涵盖推理成本、误分类成本和人工标注成本。在RVL-CDIP基准上使用近端策略优化(Proximal Policy Optimisation)进行训练,DocHRL在16个文档类别中实现了0.973的宏观F1分数,同时将每个文档的平均成本降低到2.74个标准化单位,相比之下,固定独立分类器的成本显著更高。我们的结果表明,关注成本的强化学习可以同时提高文档理解系统的分类性能和操作效率。
cs.AI / 56 / 2607.22646

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models

预训练大语言模型中的算法提取:以隐马尔可夫模型为例
Dai, Yijia, Gao, Zhaolin, Sattar, Yahya, Sun, Jennifer J., Dean, Sarah
Abstract
Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several candidates without consensus, and none has been grounded in the model's internal activations. We close this gap with a three-stage pipeline. First, we empirically compare LLM behavior against a suite of candidate algorithms and narrow the space to three classes -- though no single class explains LLM behavior across all HMM settings and sequence lengths. Second, we derive theoretical connections between the three classes and show how each can be implemented in-context by a Transformer, validating the construction in a small trained Transformer. Third, returning to pre-trained LLMs, we introduce the Principal Activations Probe (PAP), a layer-wise probing and intervention method that isolates algorithmic signals in model activations. PAP reveals low-dimensional linear representations that causally drive model predictions and track empirical ICL performance. PAP further reveals how these representations shift with properties of the underlying HMM regime; distinct computational stages are localized to different layers. Together, our results connect the in-context behavior of pre-trained LLMs to the underlying internal mechanisms and advance our understanding of how LLMs perform ICL on HMMs.
Chinese Translation
大型语言模型(LLMs)通过上下文学习(ICL)展现出从隐马尔可夫模型(HMMs)中预测下一观察值的显著能力,但支撑这一能力的算法仍未确定:先前的研究提出了几种候选算法,但并未达成共识,且没有一种算法基于模型的内部激活进行验证。我们通过三阶段的流程填补了这一空白。首先,我们对LLM的行为与一系列候选算法进行了实证比较,并将候选空间缩小到三类——尽管没有单一类别能够解释LLM在所有HMM设置和序列长度下的行为。其次,我们推导出这三类之间的理论联系,并展示每一类如何通过Transformer在上下文中实现,在一个小型训练的Transformer中验证了这种构建。第三,回到预训练的LLM,我们引入了主激活探测器(Principal Activations Probe,PAP),这是一种逐层探测和干预的方法,能够在模型激活中隔离算法信号。PAP揭示了低维线性表示,这些表示因果驱动模型预测并跟踪实证ICL性能。PAP进一步揭示了这些表示如何随基础HMM状态的属性而变化;不同的计算阶段被定位于不同的层。综合来看,我们的结果将预训练LLM的上下文行为与其内部机制联系起来,推进了我们对LLM在HMM上执行ICL的理解。
cs.AI / 57 / 2607.22648

PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving

PTStore(前缀张量存储):用于高吞吐量推理服务的分布式前缀缓存和复制
Maghyastha, Meghana, Underwood, Robert, Burns, Randal, Nicolae, Bogdan
Abstract
Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the main technique used by state of art approaches to accelerate inferences. This reduces the latency of accessing the KV cache and alleviates load imbalance caused by a disproportionately large number of requests on servers containing popular tensors. Furthermore, thanks to decentralization, PTStore allows the expansion of the size of the KV cache for LLM inference by orders of magnitude. As a result, PTStore can execute inferences on long passage Q\&A datasets 5-6 times more efficiently than current baselines, which do not aggregate memory across different nodes and GPUs and therefore require regenerating the KV cache.
Chinese Translation
受到内容分发网络(CDN)中客户端缓存设计的启发,PTStore 分布和复制形成可重用 KV 缓存前缀的热门张量,这些前缀是加速推理的先进方法所采用的主要技术。这减少了访问 KV 缓存的延迟,并缓解了由于对包含热门张量的服务器请求数量不成比例而导致的负载不平衡。此外,得益于去中心化,PTStore 允许将 LLM 推理的 KV 缓存大小扩展几个数量级。因此,PTStore 在长段问答数据集上的推理效率比当前基线高出 5-6 倍,而后者并未在不同节点和 GPU 之间聚合内存,因此需要重新生成 KV 缓存。
cs.AI / 58 / 2607.22649

STAIF: A Stage-wise Optimization for Complex Instruction Following

STAIF:一种针对复杂指令遵循的分阶段优化方法
Hong, Jian, Cheng, Chen, Liu, Quan, Chen, Yuhao, Chen, Enhong
Abstract
Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment methods, such as DPO, optimize holistic reward signals that often underemphasize strict satisfaction of individual constraints, particularly under out-of-distribution or multi-constraint settings. In this paper, we propose STAIF, a stage-wise optimization framework that decouples the alignment of subjective (soft) constraints from the optimization of objectively verifiable (hard) constraints. Stage 1 applies preference optimization with multiple negative samples to sharpen sensitivity to soft constraints, while Stage 2 applies Reinforcement Learning with Verifiable Rewards (RLVR) to enforce strict compliance with hard constraints. To support this method, we construct STAINSTRUCT, a high-quality bilingual (English, Chinese) dataset of approximately 31,000 complex multi-constraint instructions. Extensive analyses validate the design of STAIF and show state-of-the-art performance on representative benchmarks against strong baselines, as well as genuine generalization.
Chinese Translation
遵循具有多个明确约束的复杂指令仍然是大型语言模型(LLMs)面临的一个基本挑战。现有的对齐方法,如DPO,优化整体奖励信号,这往往低估了对个别约束的严格满足,特别是在分布外或多约束设置下。在本文中,我们提出了STAIF,一种分阶段优化框架,它将主观(软)约束的对齐与客观可验证(硬)约束的优化解耦。第一阶段应用带有多个负样本的偏好优化,以提高对软约束的敏感性,而第二阶段则应用可验证奖励的强化学习(RLVR),以强制严格遵守硬约束。为了支持该方法,我们构建了STAINSTRUCT,一个高质量的双语(英语、中文)数据集,包含约31,000个复杂的多约束指令。大量分析验证了STAIF的设计,并在强基线下的代表性基准测试中显示出最先进的性能,以及真正的泛化能力。
cs.AI / 59 / 2607.22651

ARdena: Scenario-driven control of real-time LLM agents

ARdena:基于场景驱动的实时大语言模型代理控制
Borozan, Luka, Matijević, Domagoj
Abstract
Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive environments remains a significant challenge. Existing approaches often rely on model fine-tuning or alignment procedures that are difficult to adapt to changing interaction requirements. This paper introduces layered scenario-driven LLM control, a framework that enables runtime behavior control through structured prompting. By combining persistent context with scenario-specific constraints, the approach allows agent behavior to be modified during interaction without changing the underlying model. The framework is implemented in ARDena, a real-time multimodal embodied agent that integrates speech interaction, visual perception, tool use, and avatar-based response generation. The proposed approach is evaluated with respect to control effectiveness, response latency, and operational stability. The results demonstrate that scenario definitions alone can produce substantially different interaction behaviors while maintaining stable real-time operation, highlighting the effectiveness of scenario-driven prompting for controlling LLM agents.
Chinese Translation
大型语言模型(LLMs)使得对话代理的能力不断增强,但在实时交互环境中可靠地控制其行为仍然是一个重大挑战。现有方法通常依赖于模型微调或对齐程序,这些方法难以适应不断变化的交互需求。本文介绍了一种分层场景驱动的LLM控制框架,通过结构化提示实现运行时行为控制。通过将持久上下文与场景特定约束相结合,该方法允许在交互过程中修改代理行为,而无需更改底层模型。该框架在ARDena中实现,ARDena是一个实时多模态具身代理,集成了语音交互、视觉感知、工具使用和基于虚拟形象的响应生成。所提出的方法在控制有效性、响应延迟和操作稳定性方面进行了评估。结果表明,仅场景定义就可以产生显著不同的交互行为,同时保持稳定的实时操作,突显了场景驱动提示在控制LLM代理方面的有效性。
cs.AI / 60 / 2607.22652

KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering

KG2Code:通过可执行代码连接知识图谱与大型语言模型以实现问答
Wu, Yike, Hu, Nan, Qi, Guilin, Xiao, Guohui, Jiang, Chen, Zou, Xinchun, Lu, Yuchen, Zhai, Songlin, Chen, Yongrui, Zhang, Yuyang, Li, Xiaoguang, Shang, Lifeng, Chen, Jiaoyan, Pan, Jeff Z.
Abstract
Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA). Existing approaches primarily combine LLMs with KGs through retrieval-augmented generation (RAG)-based, agent-based, and SPARQL-based methods. Although these methods have achieved notable success, they still suffer from several limitations, including structural information loss, unfaithful reasoning, and limited flexibility and generalization. To address these challenges, this paper proposes KG2Code, a novel approach that transforms knowledge graphs into a code-based representation, preserving structural semantics while naturally aligning with the code-aware pretraining of modern LLMs. Based on KG2Code, KG2Code-QA is further introduced as a KGQA framework that formulates KGQA as a code generation task. This formulation enables the generation of verifiable reasoning traces and executable code, thereby substantially mitigating the impact of hallucinations. In addition, an automated pipeline is developed to construct a large-scale, high-quality code corpus for effectively training open-source LLMs on KG2Code-QA. After training, LLMs are able to perform KGQA in zero-shot scenarios. Extensive experiments demonstrate that the proposed approach significantly outperforms existing KG-enhanced LLM methods for KGQA, while exhibiting strong generalization to unseen KGs. The code and data are available at Github.
Chinese Translation
近期研究探讨了将知识图谱(KGs)与大型语言模型(LLMs)相结合,以提升其在知识密集型下游任务中的表现,特别是知识图谱问答(KGQA)。现有的方法主要通过基于检索增强生成(RAG)、基于代理和基于SPARQL的方法将LLMs与KGs结合。尽管这些方法取得了显著成功,但仍然存在一些局限性,包括结构信息丢失、不可靠推理以及灵活性和泛化能力有限。为了解决这些挑战,本文提出了KG2Code,这是一种将知识图谱转化为基于代码的表示的新方法,保留结构语义,同时自然地与现代LLMs的代码感知预训练对齐。在KG2Code的基础上,进一步提出了KG2Code-QA,作为一个将KGQA表述为代码生成任务的KGQA框架。这种表述使得生成可验证的推理轨迹和可执行代码成为可能,从而显著减轻幻觉的影响。此外,开发了一个自动化流程,以构建一个大规模、高质量的代码语料库,以有效训练开源LLMs在KG2Code-QA上的表现。经过训练后,LLMs能够在零样本场景中执行KGQA。大量实验表明,所提出的方法在KGQA上显著优于现有的KG增强LLM方法,同时对未见过的KGs表现出强大的泛化能力。代码和数据可在Github上获取。
cs.AI / 61 / 2607.22653

Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual Relaxation

语言模型是否会收敛于自身?递归自我精炼作为文本放松
Wu, Xuening, Xu, Qianya, Kang, Yanlan, Chen, Zeping, Liu, Yubin, Yin, Shenqin
Abstract
Large language models are increasingly used in recursive refinement workflows, where an initial draft is repeatedly revised by the same model. Despite their growing use, the long-term dynamics of such workflows remain poorly understood. Does repeated refinement continue to improve outputs indefinitely, or does it converge toward a stable textual form? We study recursive self-refinement as a dynamical process in which repeated LLM revision drives text toward a model-preferred soft fixed-point region. Using GPT-5.5, we generate 10-step refinement trajectories for 50 ICML 2025 abstracts under both default-temperature and deterministic decoding, and additionally evaluate 15 ICML 2020 abstracts. We analyze normalized edit distance, exact and approximate fixed points, word-count stability, exponential relaxation, and external LLM-as-a-judge evaluation. Across all settings, refinement trajectories rapidly saturate. Most edits occur within the first few iterations, after which trajectories enter a soft fixed-point region with only minor surface-level changes. Deterministic decoding reaches exact fixed points earlier and exhibits smaller residual fluctuations than default-temperature decoding, while both achieve universal approximate convergence. The average edit magnitude follows a consistent exponential relaxation pattern, suggesting convergence toward a model-preferred textual equilibrium rather than open-ended optimization. External evaluation indicates that converged abstracts improve clarity, conciseness, and scientific style while preserving technical meaning. These findings support a dynamical-systems view of LLM self-refinement and motivate practical stopping criteria based on edit-magnitude saturation.
Chinese Translation
大型语言模型在递归精炼工作流程中被越来越多地使用,其中初始草稿由同一模型反复修订。尽管其使用日益增长,但此类工作流程的长期动态仍然不甚了解。反复精炼是否会无限期地改善输出,还是会收敛到一个稳定的文本形式?我们将递归自我精炼视为一个动态过程,其中反复的LLM(大型语言模型)修订将文本驱动向模型偏好的软固定点区域。使用GPT-5.5,我们为50个ICML 2025摘要生成了10步精炼轨迹,采用默认温度和确定性解码,并额外评估了15个ICML 2020摘要。我们分析了归一化编辑距离、精确和近似固定点、字数稳定性、指数放松以及外部LLM作为评判的评估。在所有设置中,精炼轨迹迅速饱和。大多数编辑发生在前几次迭代中,之后轨迹进入一个软固定点区域,仅有微小的表面变化。确定性解码比默认温度解码更早达到精确固定点,并表现出较小的残余波动,而两者都实现了普遍的近似收敛。平均编辑幅度遵循一致的指数放松模式,表明收敛于模型偏好的文本平衡,而非开放式优化。外部评估表明,收敛的摘要在清晰度、简洁性和科学风格上有所改善,同时保留了技术意义。这些发现支持了LLM自我精炼的动态系统视角,并激励基于编辑幅度饱和的实际停止标准。
cs.AI / 62 / 2607.22654

MINT-V2X: A Mobility-Integrated Network Trajectory Dataset for Predictive Resource Management

MINT-V2X:一种用于预测资源管理的移动集成网络轨迹数据集
Anjum, Abdullah, Rezaei, Abdolazim, Sookhak, Mehdi
Abstract
Vehicle-to-Everything (V2X) communication systems are based on datasets that not only contain vehicle trajectory data but also wireless network parameters with a realistic level of fidelity, enabling the creation of prediction and optimization models. There is a very critical research infrastructure gap today, and publicly available datasets are likely to be limited to one of the two: mobility or network parameters, and rarely provide a single, integrated view that combines both. This paper introduces MINT-V2X, a comprehensive dataset generated by coupling SUMO traffic dynamics with OMNeT++/Simu5G network simulation. The validation framework is composed of 14 standardized tests based on 3GPP Release 14 (C-V2X), ETSI standards and Shannon capacity theory. The resulting dataset contains 9.87 million synchronized data points from 1,386 vehicles from 15 roadside units (RSUs) during 3 hours of urban traffic simulation. We demonstrate strict algorithmic consistency through network metric correlations (CQI-SINR: 0.993; SINR-PDR: 0.946). Finally, we demonstrate the value of the dataset by conducting an RSU load prediction case study, showing that using trajectory data yields better predictive performance than network-history-only baselines. The dataset, experiments, and complete SUMO configuration files are available in the GitHub repository to facilitate reproduction on alternative simulation stacks.
Chinese Translation
车对一切(V2X)通信系统基于的数据集不仅包含车辆轨迹数据,还包含具有现实保真度的无线网络参数,从而能够创建预测和优化模型。目前存在一个非常关键的研究基础设施缺口,公开可用的数据集往往限于两者之一:移动性或网络参数,且很少提供一个综合视角来结合两者。本文介绍了MINT-V2X,这是一个通过将SUMO交通动态与OMNeT++/Simu5G网络仿真相结合生成的综合数据集。验证框架由基于3GPP Release 14(C-V2X)、ETSI标准和香农容量理论的14个标准化测试组成。生成的数据集包含来自15个路边单元(RSUs)中1,386辆车辆的9.87百万个同步数据点,涵盖3小时的城市交通仿真。我们通过网络指标相关性(CQI-SINR: 0.993; SINR-PDR: 0.946)展示了严格的算法一致性。最后,我们通过进行RSU负载预测案例研究展示了数据集的价值,结果表明使用轨迹数据的预测性能优于仅基于网络历史的基线。数据集、实验和完整的SUMO配置文件已在GitHub仓库中提供,以便在替代仿真堆栈上进行复现。
cs.AI / 63 / 2607.22655

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation

EventOD:通过大语言模型引导的语义调制实现事件感知的起讫点流量生成
Zhao, Jie, Feng, Jie, Rong, Can, Hou, Zhihan, Lu, Peng, Li, Yong
Abstract
Estimating origin-destination (OD) flows under disruptive events is important for disaster response and urban resilience. Existing deep OD models trained on routine mobility often degrade when extreme events abruptly alter regional functions and population activities, while retraining a new generator for each event is impractical under limited event-time supervision. We propose EventOD, an event-adaptive OD generation framework that steers a pretrained OD generator using structured event semantics. EventOD first uses a large language model to infer region-level functional and demographic control vectors from coarse event observations. It then learns two lightweight adaptation modules, AlphaNet and BetaNet, to calibrate the magnitude of these semantic shifts, and further introduces a retrieval-augmented fallback pathway for scenarios with sparse supervision. The resulting event-conditioned features are injected into a pretrained graph diffusion OD model through input-level modulation, enabling event-aware adaptation without updating generator parameters. Experiments on hurricane- and pandemic-induced mobility across U.S. counties show that EventOD consistently improves both reconstruction accuracy and distributional fidelity over strong baselines. Source code is available at https://anonymous.4open.science/r/EventOD-5C11/.
Chinese Translation
在突发事件下估计起讫点(OD)流量对于灾害响应和城市韧性至关重要。现有的基于常规出行模式训练的深度OD模型在极端事件突然改变区域功能和人口活动时往往表现不佳,而在有限的事件时间监督下为每个事件重新训练新的生成器是不切实际的。我们提出了EventOD,一个事件自适应的OD生成框架,通过结构化事件语义引导预训练的OD生成器。EventOD首先利用大语言模型从粗略的事件观察中推断区域级功能和人口控制向量。然后,它学习两个轻量级适应模块,AlphaNet和BetaNet,以校准这些语义变化的幅度,并进一步引入增强检索的后备路径,以应对监督稀疏的场景。最终生成的事件条件特征通过输入级调制注入到预训练的图扩散OD模型中,实现事件感知的适应,而无需更新生成器参数。在美国各县的飓风和疫情引发的出行实验中,EventOD在重建精度和分布保真度上均显著优于强基线。源代码可在 https://anonymous.4open.science/r/EventOD-5C11/ 获取。
cs.AI / 64 / 2607.22658

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

StanceBench:基于音频大语言模型的人际立场评估基准
Wang, Yuzhe, Thebaud, Thomas, Hu, Jennifer, Villalba-Lopez, Jesús, Ravichandran, Venkatesh, Tinchev, Georgi, Dehak, Najim, Moro-Velázquez, Laureano
Abstract
Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceBench, a benchmark for measuring interpersonal stance in conversational speech and evaluating audio-capable LLMs as automated judges. Using the Seamless Interaction corpus, StanceBench (1) specifies 9 stance dimensions via role-prompt poles, (2) standardizes single-speaker and interaction-based evaluations, and (3) reports LLM-as-a-judge robustness, bias, and stance inference. Across evaluated stances, empathy and politeness are the easiest. Warmth and assertiveness are moderately separable with positivity skew/asymmetry. Honesty is the hardest and shows high prompt order bias, consistent with needing cross-turn evidence. Attentiveness is separable but aligns weakly with humans. Interaction stances are more context-sensitive, with threshold gaps and high variance, especially conflict regulation.
Chinese Translation
语音对话模型越来越依赖韵律和互动细微差别来传达社会意图,但对于这些线索的基准仍然有限。我们引入了StanceBench,这是一个用于测量对话语音中人际立场并评估音频能力的大语言模型(LLM)作为自动评判者的基准。通过无缝互动语料库,StanceBench (1) 通过角色提示极指定了9个立场维度,(2) 标准化了单一发言者和基于互动的评估,(3) 报告了LLM作为评判者的稳健性、偏见和立场推断。在评估的立场中,同理心和礼貌是最容易的。温暖和自信在积极性偏斜/不对称的情况下适度可分离。诚实是最难的,并显示出较高的提示顺序偏见,这与需要跨轮次证据一致。专注性是可分离的,但与人类的对齐较弱。互动立场更具上下文敏感性,存在阈值差距和高方差,尤其是在冲突调节方面。
cs.AI / 65 / 2607.22661

TRE: Training-Free Hallucination Detection for Diffusion Language Models

TRE:无训练的扩散语言模型幻觉检测
Weng, Pengcheng, Qian, Yanyu, Tan, Yue, Liu, Yixin
Abstract
Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination problem. Existing hallucination detection approaches for D-LLMs mainly follow a training-based paradigm, relying on data-driven training to optimize the detector. Such reliance not only limits their generalizability across domains models but also incurs additional training cost and deployment overhead. To address these limitations, we propose TRE, a training-free hallucination detection metric for D-LLMs. TRE is a parameter-free and single-run metric that estimates hallucination risk directly from the entropy signals of a single generation, without requiring any detector training or repeated sampling. TRE extracts entropy signals within the D-LLM decoding process along both the spatial and temporal dimensions. From a token-level spatial perspective, we focus on revealing tokens as the most informative carriers of uncertainty, capturing where uncertainty is actively committed. From a diffusion step-level temporal perspective, we empirically identify the dominance of late-step entropy and hence aggregate these signals with a simple linear weighting scheme to obtain TRE. Extensive experiments on multiple D-LLMs and QA datasets demonstrate that TRE achieves competitive performance, while enjoying strong generalizability, efficiency, and robustness.
Chinese Translation
扩散大型语言模型(D-LLMs)近年来受到越来越多的关注,但其可靠性受到幻觉问题的显著影响。现有的D-LLMs幻觉检测方法主要遵循基于训练的范式,依赖数据驱动的训练来优化检测器。这种依赖不仅限制了它们在不同领域模型中的泛化能力,还增加了额外的训练成本和部署开销。为了解决这些限制,我们提出了TRE,一种针对D-LLMs的无训练幻觉检测指标。TRE是一个无参数且单次运行的指标,能够直接从单次生成的熵信号中估计幻觉风险,而无需任何检测器训练或重复采样。TRE在D-LLM解码过程中提取熵信号,考虑空间和时间两个维度。从token级的空间角度来看,我们专注于揭示token作为不确定性最具信息性的载体,捕捉不确定性积极发生的地方。从扩散步骤级的时间角度来看,我们实证识别出后期熵的主导性,因此通过简单的线性加权方案聚合这些信号以获得TRE。在多个D-LLMs和问答数据集上的广泛实验表明,TRE在性能上具有竞争力,同时享有强大的泛化能力、效率和鲁棒性。
cs.AI / 66 / 2607.22662

CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data

CuraWeb:针对网络规模预训练数据的质量、冗余和多样性的联合优化
Li, Peiguang, Zhou, Yongwei, Diao, Juncheng, Fan, Yuchun, Yang, Jian, Yang, Jianxiao, Su, Zhongda, Jiao, Shuguang, Wei, Xiao, Zou, Zhiye, Dong, Gan, Zeng, Zhizhao, Weng, Rongxiang, Wang, Jingang, Cai, Xunliang
Abstract
Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objectives, which inevitably narrows distributional diversity and marginalizes long-tail knowledge, thereby restricting data coverage and underutilizing the vast potential of the open web. To address this limitation, we propose a novel curation paradigm that shifts from linear pruning to the joint optimization of quality, redundancy, and diversity. This framework synergizes dual-track cleaning (rule-based and model-driven) with hybrid deduplication (n-gram and semantic), while employing a multi-objective sampler to balance informational quality with distributional breadth. Applying this framework to Common Crawl, we construct CuraWeb, a 2T-token English corpus. Unlike existing resources, CuraWeb establishes an industrial-grade standard for data curation by recovering a more holistic data distribution with enhanced diversity and minimal redundancy, achieving broader coverage of long-tail knowledge across diverse domains. Experimental evaluations at the 3B scale demonstrate that CuraWeb significantly outperforms state-of-the-art baselines, yielding an average performance gain of 1.8\% across a wide range of benchmarks, particularly in knowledge-intensive and reasoning tasks.
Chinese Translation
通过高度选择性过滤器(如 FineWeb-Edu 和 DCLM)策划的开放网络语料库构成了大型语言模型(LLM)预训练数据的核心,并显著提升了 LLM 的性能。然而,这些流程通常依赖于单一的优化目标,这不可避免地缩小了分布多样性并边缘化了长尾知识,从而限制了数据覆盖面并未充分利用开放网络的巨大潜力。为了解决这一局限性,我们提出了一种新的策划范式,转变为质量、冗余和多样性的联合优化。该框架将双轨清理(基于规则和基于模型)与混合去重(n-gram 和语义)相结合,同时采用多目标采样器来平衡信息质量与分布广度。将该框架应用于 Common Crawl,我们构建了 CuraWeb,一个包含 2T 令牌的英语语料库。与现有资源不同,CuraWeb 通过恢复更全面的数据分布,增强多样性和最小冗余,确立了数据策划的工业级标准,实现了对各个领域长尾知识的更广泛覆盖。在 3B 规模的实验评估中,CuraWeb 显著超越了最先进的基准,平均性能提升达 1.8\%,尤其在知识密集型和推理任务中表现突出。
cs.AI / 67 / 2607.22663

Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

超越块边界:扩散大语言模型的多块编辑
Mou, Xingyu, Huang, Zijin, Zhang, Tianze, Ma, Yuxin, Wei, Lanning, Huang, Zengfeng, Zheng, Da, Du, Lun
Abstract
Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks preserves parallel generation within each block while keeping the quadratic attention cost tractable. However, this efficiency comes with a structural limitation: tokens near the end of a block are generated without access to future cross-block context, and once a block is finalized, its uncertain predictions become irreversible context for all subsequent blocks. This creates a block boundary problem, in which uncertainty accumulates toward block boundaries and early mistakes propagate throughout later generation. To address this issue, we propose Multi-Block Editing (MBE), to mitigate this problem by editing decoded tokens based on cross-block context. Following this principle, MBE first proposes a training-free decoding algorithm to edit the decoded tokens in previous blocks, which is achieved by re-opening a full-attention window over selected blocks. Given the mismatched attention mechanism between block diffusion training and MBE inference, MBE further introduces a supervised Fine-tuning strategy, which equips the model with bidirectional attention masks that progressively expands the editing span. Furthermore, it also extends SGLang with a multi-shape CUDA Graph pool and fine-grained KV cache control to make these variable-length editing passes efficient in practice. Experiments on LLaDA2.1-Mini across 13 benchmarks show that training-free MBE outperforms all existing decoding baselines while maintaining comparable throughput, and MBE SFT further brings a performance gain of 2.7. The largest improvements appear on tasks requiring strong long-range consistency, including +13.3 on AIME 2025 and +5.9 on ZebraLogic, demonstrating the effectiveness of MBE.
Chinese Translation
块扩散已成为扩展离散扩散语言模型(dLLMs)的主导范式,因为在固定大小块中解码文本能够在每个块内保持并行生成,同时使得二次注意力成本可控。然而,这种效率带来了结构上的限制:靠近块末尾的标记在生成时无法访问未来的跨块上下文,并且一旦一个块被最终确定,其不确定的预测就成为所有后续块的不可逆上下文。这造成了块边界问题,其中不确定性在块边界处累积,早期的错误在后续生成中传播。为了解决这个问题,我们提出了多块编辑(Multi-Block Editing,MBE),通过基于跨块上下文编辑解码的标记来缓解这一问题。根据这一原则,MBE 首先提出了一种无训练的解码算法,通过在选定块上重新开启全注意力窗口来编辑先前块中解码的标记。鉴于块扩散训练与 MBE 推理之间的注意力机制不匹配,MBE 进一步引入了一种监督微调策略,为模型提供双向注意力掩码,逐步扩展编辑范围。此外,它还通过多形状 CUDA 图池和细粒度 KV 缓存控制扩展了 SGLang,使这些可变长度的编辑过程在实践中高效。对 LLaDA2.1-Mini 在 13 个基准上的实验表明,无训练的 MBE 在保持可比吞吐量的同时超越了所有现有的解码基线,而 MBE 的监督微调进一步带来了 2.7 的性能提升。最大的改进出现在需要强长程一致性的任务上,包括 AIME 2025 上的 +13.3 和 ZebraLogic 上的 +5.9,证明了 MBE 的有效性。
cs.AI / 68 / 2607.22665

Obliviate: Efficient Unlearning in Recommender Systems

Obliviate:推荐系统中的高效遗忘
Prakash, Tushar, Singh, Brijraj, Pedanekar, Niranjan, Chaturvedi, Narayan
Abstract
Machine unlearning is becoming increasingly critical in the context of data privacy regulations, particularly for recommendation systems that are directly trained on user interaction data. The goal of this work is to remove requested interaction data and their downstream influence from trained model while preserving recommendation quality, and to do so without incurring the substantial computational cost of full retraining. Existing approaches exhibit several limitations, including limited unlearning completeness and degradation in recommendation performance, while having substantial computational overhead. In this paper, we propose Obliviate, an efficient two-stage unlearning framework for recommender systems that achieves high unlearning completeness while maintaining good utility. In the first stage, we introduce a Low-Rank Unlearning Adapter (LUA), which employs a lightweight Hessian proxy to enable curvature-aware and efficient unlearning through localized low-rank adapters rather than full parameters. In the second stage, we propose Locality-Aware Calibration (LAC), a lightweight refinement stage that updates only the adapter parameters to improve the performance by enforcing unlearning via ranking-based objectives while preserving utility through knowledge distillation. Extensive empirical evaluations demonstrate that Obliviate achieves high level of forgetting with minimal loss in recommendation quality and at significantly reduced computational cost, offering a practical and scalable solution for large-scale recommender systems.
Chinese Translation
在数据隐私法规的背景下,机器遗忘变得越来越重要,尤其是对于直接基于用户交互数据训练的推荐系统。本研究的目标是移除请求的交互数据及其对训练模型的下游影响,同时保持推荐质量,并且在不承担全面重新训练的巨大计算成本的情况下实现这一目标。现有方法存在若干局限性,包括有限的遗忘完整性和推荐性能的下降,同时伴随显著的计算开销。本文提出了Obliviate,一个高效的两阶段推荐系统遗忘框架,能够在保持良好效用的同时实现高遗忘完整性。在第一阶段,我们引入了低秩遗忘适配器(Low-Rank Unlearning Adapter, LUA),该适配器利用轻量级的Hessian代理,通过局部低秩适配器而非全参数实现曲率感知和高效遗忘。在第二阶段,我们提出了局部感知校准(Locality-Aware Calibration, LAC),这是一个轻量级的精炼阶段,仅更新适配器参数,通过基于排名的目标强制遗忘,同时通过知识蒸馏保持效用。大量实证评估表明,Obliviate在推荐质量损失最小的情况下实现了高水平的遗忘,并显著降低了计算成本,为大规模推荐系统提供了一个实用且可扩展的解决方案。
cs.AI / 69 / 2607.22667

Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance

用于海洋监视中异构传感器选择的强化学习
Starodubov, Andrei, Prabowo, Yaqub Aris, Hadjipieris, Andreas, Galeazzi, Roberto, Kyriakides, Ioannis
Abstract
This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: instead of activating all sensors or repeatedly performing computationally expensive online expected-information-gain evaluation, a learned policy selects one tracking-relevant sensor at each decision epoch. A Bayesian sequential Monte Carlo tracker estimates the vessel state from noisy measurements and provides a belief representation for scheduling under nonlinear and non-Gaussian conditions. A Proximal Policy Optimization agent selects one of five sensors deployed in a georeferenced simulation of the CMMI Smart Marina testbed at Ayia Napa Marina, Cyprus. The agent observes belief-state, detection-history, coverage, sensor-geometry, and realized-information-gain features. The reward is defined as a realized-information-gain term gated by an observability mask. Final-test simulations compare the proposed framework with random single-sensor selection, always-on sensing using all sensors simultaneously, and the expected-information-gain sensor-selection baseline proposed in our previous work. Results show that the learned policy achieves tracking performance close to always-on sensing while activating only one sensor per decision time step and avoiding the computationally expensive online entropy search required by expected-information-gain selection.
Chinese Translation
本文提出了一种基于信息增益引导的强化学习传感器选择框架,用于在异构海洋传感器网络中进行单船舶跟踪。该方法的提出受到信息论传感器管理的启发:与其激活所有传感器或反复执行计算成本高昂的在线期望信息增益评估,不如通过学习的策略在每个决策时刻选择一个与跟踪相关的传感器。贝叶斯序贯蒙特卡洛跟踪器从噪声测量中估计船舶状态,并在非线性和非高斯条件下提供信念表示以进行调度。一个近端策略优化(Proximal Policy Optimization)代理在塞浦路斯阿依亚纳帕港的CMMI智能码头测试床的地理参考模拟中选择五个传感器之一。该代理观察信念状态、检测历史、覆盖范围、传感器几何形状和实现的信息增益特征。奖励定义为由可观测性掩码限制的实现信息增益项。最终测试模拟将所提出的框架与随机单传感器选择、同时使用所有传感器的始终开启感知,以及我们之前工作中提出的期望信息增益传感器选择基线进行了比较。结果表明,学习的策略在每个决策时间步仅激活一个传感器的情况下,跟踪性能接近始终开启感知,同时避免了期望信息增益选择所需的计算成本高昂的在线熵搜索。
cs.AI / 70 / 2607.22671

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

AIR-BENCH Live:一个不断演进的基础模型安全基准
Naphade, Rohan, Pan, Minzhou, Li, Bo
Abstract
Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective. We present AIR-BENCH Live, a self-evolving successor to AIR-BENCH 2024. An automated update pipeline monitors government regulation and classifies new policies against the current four-tier risk taxonomy, either matching them to existing categories or proposing new granular categories. Then, a multi-agent, persona-driven prompt generation algorithm generates realistic, multilingual prompts with minimal human review, leaving room for improvement with modern jail breaking techniques. This algorithm is used to overhaul legacy prompts and generate prompts for new categories. In our current version, the pipeline has expanded the benchmark from 314 to 335 granular risks, with the 21 new categories drawing from 31 truly novel policy clauses across seven jurisdictions. Evaluating 14 recent models, we find a wide safety spread (from 0.17 to 0.89 among the models judged on their own behavior), that the modernized prompts are on average 0.06 points harder than the 2024 set, with the largest drops concentrated among the most compliant models, and that most models are modestly less safe on non-English prompts. By continuously absorbing new regulation and regenerating prompts, AIR-BENCH Live is designed to evolve alongside a fast-moving field.
Chinese Translation
基础模型安全基准捕捉了其发布时的人工智能风险:随着模型的改进和政府通过新的人工智能安全立法,其风险分类变得不够全面,攻击提示也变得无效。我们提出了AIR-BENCH Live,这是AIR-BENCH 2024的自我演进继任者。一个自动更新管道监测政府法规,并根据当前的四级风险分类法对新政策进行分类,既可以将其匹配到现有类别,也可以提出新的细分类别。然后,一个多代理、基于角色的提示生成算法生成现实的多语言提示,几乎不需要人工审核,为现代越狱技术的改进留出空间。该算法用于改造旧的提示并为新类别生成提示。在我们当前的版本中,该管道已将基准从314个细分风险扩展到335个,21个新类别源自七个司法管辖区的31个真正新颖的政策条款。评估14个近期模型时,我们发现安全性差异较大(在根据模型自身行为评判时,范围从0.17到0.89),现代化的提示平均比2024年版本难度高0.06分,最大降幅集中在最合规的模型上,并且大多数模型在非英语提示上安全性略有下降。通过不断吸收新法规并重新生成提示,AIR-BENCH Live旨在与快速发展的领域共同演进。
cs.AI / 71 / 2607.22676

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

大语言模型任务适应如何重塑对齐:行为和表征漂移的多维研究
Elcock, James, Shen, William F., Qiu, Xinchi, Lane, Nicholas D.
Abstract
Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
Chinese Translation
后训练是将大型语言模型适应于下游任务的关键机制。尽管先前的研究表明,任务适应可以改变模型的预先对齐,特别是其安全行为,但其在对齐领域的更广泛影响仍然不甚了解。我们通过对代表性任务适应方法的系统评估来填补这一空白,包括监督微调(SFT)、KL正则化SFT和具有可验证奖励的强化学习(RLVR),涵盖了六个关键领域的15个对齐方面:安全性、事实性、立场稳定性、社会危害、可控性和可指令性。我们的结果表明,后训练并不会均匀地重塑对齐。RLVR在提高任务性能的同时,导致相对较小但非零的特定指标漂移,而SFT在各个领域导致了显著更大的对齐漂移。KL正则化减轻了这一影响:更强的参考模型锚定减少了与基线的对齐漂移,尽管KL-SFT在保持对齐方面仍不及RLVR。表征层面的分析进一步支持了这一模式,对齐相关表征的变化与行为漂移相一致。综合来看,这些结果表明,任务适应不仅仅是提升能力的步骤,而是一个独立的对齐干预,促使多维对齐评估成为后训练流程的标准组成部分。
cs.AI / 72 / 2607.22679

DOSA: A Tree-Guided, Self-Regressive Framework for Long Document Structure Analysis

DOSA:一种树导向的自回归框架用于长文档结构分析
Li, Bohou, Sowell, Benjamin, Shah, Mehul, Lindblad, Mark, Lindeman, Henry
Abstract
In visually-rich documents, information is encoded not only in individual page objects such as tables, headers, and text blocks, but also in the structural relations among them, making document structure analysis fundamental to information retrieval and document understanding. However, accurately inferring such relations remains challenging in multi-page documents with long-range dependencies and heterogeneous layouts. To address this, we propose a tree-guided and self-regressive framework, termed DOcument Structure Analyzer (DOSA), for inferring relations among page objects and reconstructing document-level semantic trees. DOSA processes documents chunk-by-chunk, fusing visual, textual, and layout features for each page object and predicting hierarchical and ordering relations. The predicted relations are used to incrementally construct a semantic tree, which is then leveraged as structural context to guide inference on subsequent chunks. Experimental results on five benchmarks demonstrate the effectiveness of DOSA, with improvements of up to 4 F1 points and 19 TEDS points on DocHieNet, the most challenging multi-page hierarchy benchmark.
Chinese Translation
在视觉丰富的文档中,信息不仅编码在各个页面对象中,例如表格、标题和文本块,还编码在它们之间的结构关系中,这使得文档结构分析对于信息检索和文档理解至关重要。然而,在具有长距离依赖和异构布局的多页文档中,准确推断这些关系仍然具有挑战性。为此,我们提出了一种树导向的自回归框架,称为文档结构分析器(DOSA),用于推断页面对象之间的关系并重建文档级语义树。DOSA 逐块处理文档,融合每个页面对象的视觉、文本和布局特征,并预测层次和顺序关系。预测的关系用于逐步构建语义树,随后作为结构上下文指导对后续块的推断。在五个基准测试上的实验结果表明,DOSA 的有效性,在最具挑战性的多页层次基准 DocHieNet 上,F1 分数提高了最多 4 分,TEDS 分数提高了最多 19 分。
cs.AI / 73 / 2607.22682

A Vocabulary for Multi-Agent Automated Research Systems

多智能体自动化研究系统的词汇
Akhbari, Bardiya
Abstract
We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies 1) who the agents are, 2) what operations are available in the system, 3) who may invoke them, 4) how agents communicate, 5) what information is visible within and across runs, 6) how the next action is chosen, 7) how a run begins, and 8) how outputs are evaluated. A trajectory records one run from the input task to the returned artifact. Because agents, operations, and initialization may be stochastic, repeated runs on the same task induce a distribution over trajectories rather than a single behavior. Our vocabulary turns structural design questions, such as when agents should communicate, gain or lose a capability, or carry information across runs, into testable choices. It also makes the evaluator a component of the system, since reported gains depend on how closely the proxy score matches true quality. That separation also splits the vague complaint that these systems lack taste into two failures with different solutions. Generative taste is the rate at which a system proposes novel trajectories before any score is observed, and evaluative taste is the gap between the proxy score and the quality it should match. We instantiate the vocabulary on recent autoresearch systems to illustrate that it covers designs that differ widely in structure.
Chinese Translation
我们提出了一种用于自动化研究系统的词汇,该系统由一个或多个智能体构成,以便更容易描述和比较其设计选择。该词汇指定了1)智能体是谁,2)系统中可用的操作是什么,3)谁可以调用这些操作,4)智能体如何进行通信,5)在不同运行中可见的信息是什么,6)如何选择下一个动作,7)一次运行如何开始,以及8)如何评估输出。轨迹记录了从输入任务到返回工件的一个运行。由于智能体、操作和初始化可能是随机的,因此在同一任务上重复运行会产生一个轨迹的分布,而不是单一行为。我们的词汇将结构设计问题,例如智能体何时应进行通信、获得或失去能力,或在运行之间传递信息,转化为可测试的选择。它还使评估者成为系统的一个组成部分,因为报告的收益取决于代理评分与真实质量的匹配程度。这种分离还将对这些系统缺乏品味的模糊抱怨分解为两个具有不同解决方案的失败。生成品味是系统在观察到任何评分之前提出新轨迹的速率,而评估品味是代理评分与其应匹配的质量之间的差距。我们在最近的自研究系统上实例化该词汇,以说明它涵盖了在结构上差异很大的设计。
cs.AI / 74 / 2607.22683

Imprompt: A Language Framework for Prompt Programming

Imprompt:一种用于提示编程的语言框架
Wu, Chentian, Yang, Shengyuan, Murali, Adithya
Abstract
With the unprecedented success of Language Models (LMs), the science of Prompt Engineering has evolved the powerful idea of Prompt Programming, where prompts are treated as a programmable control surface for describing complex tasks and leveraging LM capabilities. However, existing prompt programming frameworks suffer from various complexities and inelegances, which make them hard to utilize in practice for effectively describing tasks. We propose Imprompt, a new language framework for the study and practice of prompt programming. We undertake a foundational investigation of prompt programming, and contend that prompt programs must contain only the task descriptions and must be decoupled from lower-level 'execution' details. We further develop this position by illustrating structured prompting as a combination of prompt programming and prompt program 'compilation'. We exemplify this view by formally defining two compilers for Imprompt programs. We then explore the idea of typing for prompt programs and draw a correspondence between type checking and constrained decoding. Finally, we implement our compilers and type checkers and evaluate them on a variety of case studies. We believe our work contributes programming-language foundations toward the emerging area of prompt programming.
Chinese Translation
随着语言模型(Language Models, LMs)前所未有的成功,提示工程的科学发展出了强大的提示编程(Prompt Programming)理念,其中提示被视为描述复杂任务和利用语言模型能力的可编程控制界面。然而,现有的提示编程框架存在各种复杂性和不优雅之处,使得它们在实践中难以有效描述任务。我们提出了Imprompt,一种用于提示编程研究和实践的新语言框架。我们对提示编程进行了基础性研究,并主张提示程序必须仅包含任务描述,并且必须与较低层次的“执行”细节解耦。我们通过将结构化提示(structured prompting)阐述为提示编程与提示程序“编译”的结合,进一步发展了这一观点。我们通过正式定义两个Imprompt程序的编译器来举例说明这一观点。随后,我们探讨了提示程序的类型系统,并在类型检查与受限解码之间建立了对应关系。最后,我们实现了我们的编译器和类型检查器,并在多种案例研究中进行了评估。我们相信我们的工作为新兴的提示编程领域贡献了编程语言基础。
cs.AI / 75 / 2607.22688

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

共适应:为大型语言模型代理共同演化的适配器和模型权重
Chen, Zhengyu, Xiao, Teng, Zhu, Huaisheng, Yuan, Yige, Zhang, Luan, Wang, Jingang
Abstract
Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typically train models under a fixed harness, including prompts, tools, skills, middleware, and memory, while leaving the data-generating process outside the optimization objective. This creates a mismatch between model updates and the static scaffolding that determines trajectory quality. We introduce Co-Harness, a framework that jointly optimizes the agent harness and model parameters during post-training. Co-Harness alternates between harness optimization and model optimization. An LLM-based HarnessCritic analyzes failed trajectories, identifies harness-level failure modes, and proposes validated local updates. The model is then fine-tuned on high-quality trajectories generated by the improved harness, distilling effective scaffolding into model parameters. A 200+ hour autonomous case study further shows that Co-Harness can recover from system crashes, improve inference efficiency, and discover ensemble strategies without human intervention. These results suggest that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.
Chinese Translation
后训练代理用于自动化人工智能研究,不仅需要优化模型参数,还需要优化运行时适配器,这影响研究轨迹的生成、评估和学习方式。现有的流程通常在固定的适配器下训练模型,包括提示、工具、技能、中间件和内存,同时将数据生成过程排除在优化目标之外。这导致模型更新与决定轨迹质量的静态支架之间存在不匹配。我们提出了共适应(Co-Harness),一个在后训练期间共同优化代理适配器和模型参数的框架。共适应在适配器优化和模型优化之间交替进行。基于大型语言模型的HarnessCritic分析失败的轨迹,识别适配器级别的失败模式,并提出经过验证的局部更新。然后,模型在改进的适配器生成的高质量轨迹上进行微调,将有效的支架提炼为模型参数。一项超过200小时的自主案例研究进一步表明,共适应能够从系统崩溃中恢复,提高推理效率,并在没有人工干预的情况下发现集成策略。这些结果表明,共同优化适配器和模型是一种超越固定适配器后训练的有效方法,以提升代理的性能。
cs.AI / 76 / 2607.22689

Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents

超越顺序交互:图形用户界面代理的并行执行与协调基准测试
Yu, Zedong, Li, Qianxing, Gao, Zhi, Xiang, Liuyu, Shi, Chenrui, Liu, Yang, Wu, Huiming, Wei, Yujie, Fei, Yuhao, Fu, Yubo, He, Zhaofeng
Abstract
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as clicking, typing, and scrolling on desktops and mobile devices. However, current agents scale poorly to long-horizon tasks: actions incur costly LMM inferences, and performance degrades as context grows. Humans divide such workloads among collaborators who complete sub-tasks in parallel. Yet parallel coordination among GUI agents has received little attention. To close this gap, we introduce ParaGUIBench, to our knowledge, the first benchmark dedicated to parallel execution and coordination of multiple GUI agents on separate desktop instances. It consists of three components: a multi-device Docker infrastructure with a shared file system; a dataset of 233 tasks spanning six task categories; and an evaluation system with efficiency metrics, including step reduction ratio and token cost. We further introduce ParaGUI, a planner-worker agent that decomposes GUI tasks and dispatches sub-tasks to concurrent workers on separate desktop instances. On ParaGUIBench, ParaGUI reaches a 46.4% success rate, outperforming the strongest serial baseline (Claude Sonnet 4.6) by 12.9 points while using roughly half the steps and less than half the tokens. These results show that parallel execution can improve both success rate and efficiency on decomposable, long-horizon GUI tasks, pointing to a direction worth further study.
Chinese Translation
图形用户界面(GUI)代理是由大型多模态模型(LMMs)驱动的系统。它们通过在桌面和移动设备上进行点击、输入和滚动等GUI操作来感知屏幕状态并执行用户指令。然而,当前的代理在处理长时间任务时表现不佳:操作会产生高昂的LMM推理成本,且随着上下文的增加,性能会下降。人类通常将此类工作负载分配给协作者,由他们并行完成子任务。然而,GUI代理之间的并行协调却鲜有关注。为了解决这一问题,我们推出了ParaGUIBench,这是我们所知的第一个专门用于多个GUI代理在不同桌面实例上进行并行执行和协调的基准测试。它由三个部分组成:一个具有共享文件系统的多设备Docker基础设施;一个涵盖六个任务类别的233个任务的数据集;以及一个包含效率指标的评估系统,包括步骤减少比率和令牌成本。我们进一步介绍了ParaGUI,一个规划者-工作者代理,它将GUI任务分解并将子任务分派给在不同桌面实例上并行工作的工人。在ParaGUIBench上,ParaGUI的成功率达到了46.4%,比最强的串行基线(Claude Sonnet 4.6)高出12.9个百分点,同时使用的步骤大约只有一半,令牌数量也不到一半。这些结果表明,并行执行可以提高可分解的长时间GUI任务的成功率和效率,指向一个值得进一步研究的方向。
cs.AI / 77 / 2607.22690

LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory

LazyMem:广泛检索,选择性构建以实现高效的长期智能体记忆
Yu, Jing, Zhao, Yibo, Zhang, Jiaming, Li, Xiang
Abstract
Long-term memory lets LLM agents reuse past interactions, but raw dialogue histories are verbose and information-sparse. Retrieving broadly improves evidence coverage yet overwhelms downstream reasoning with noise; compressing at write time reduces noise but irreversibly discards details the future query may need. We introduce LazyMem, which sidesteps this dilemma by deferring all memory construction to query time. A lightweight 4B model processes the retrieved candidate pool in overlapping parallel windows, selectively retaining and compressing only query-relevant content. The model is trained through supervised fine-tuning followed by group-based reinforcement learning with a format-gated composite reward that combines a rule-based action signal measuring selection accuracy with an LLM-judged quality signal measuring source faithfulness and query utility. On the LongMemEval benchmark, LazyMem-4B achieves an LLM-judge accuracy of 0.85 with only 213 memory tokens, 68.7$\times$ fewer than retrieval-only, and generalizes to LoCoMo (0.68) without target-domain training, while reducing mean latency over the prior query-time baseline. The 32B variant reaches 0.93, surpassing oracle-context references on aggregation-heavy question types. The code associated with this work is publicly available at https://github.com/allacnobug/LazyMem.
Chinese Translation
长期记忆使得大型语言模型(LLM)智能体能够重用过去的交互,但原始对话历史冗长且信息稀疏。广泛检索提高了证据覆盖率,但会用噪声淹没下游推理;在写入时进行压缩虽然减少了噪声,但不可逆地丢弃了未来查询可能需要的细节。我们提出了LazyMem,通过将所有记忆构建推迟到查询时来规避这一困境。一个轻量级的4B模型在重叠的并行窗口中处理检索到的候选池,仅选择性地保留和压缩与查询相关的内容。该模型通过监督微调训练,随后进行基于组的强化学习,采用格式门控的复合奖励,该奖励结合了基于规则的动作信号(衡量选择准确性)与LLM评估的质量信号(衡量来源的可信度和查询的实用性)。在LongMemEval基准上,LazyMem-4B在仅使用213个记忆标记的情况下,达到了0.85的LLM评估准确率,比仅检索的情况减少了68.7倍,并且在没有目标领域训练的情况下,能够推广到LoCoMo(0.68),同时减少了相较于之前查询时基线的平均延迟。32B变体达到了0.93,超越了在聚合重的问题类型上的oracle上下文参考。与本工作相关的代码已公开可用,网址为https://github.com/allacnobug/LazyMem。
cs.AI / 78 / 2607.22691

HiLLTS: Zero-Shot Hierarchical LLM-Guided Traffic Signal Control for Sustainable Transportation

HiLLTS:零样本层次化LLM引导的交通信号控制以实现可持续交通
Ding, Yue, Mukande, Tendai, Liu, Mingming
Abstract
Urban traffic congestion significantly increases fuel consumption, greenhouse gas emissions, and commuter delays, resulting in substantial economic losses and environmental harm in modern cities. Traditional traffic signal control strategies such as fixed-time scheduling, actuated control, and reinforcement learning (RL)-based methods, offer different degrees of adaptability; however, RL-based methods can require extensive retraining, careful reward design, and substantial simulation data when transferred across networks or demand regimes. To address these challenges, we propose HiLLTS, an LLM-guided traffic signal control framework that employs a hierarchical three-layer architecture consisting of a central coordination agent, a district layer and multiple cluster-level intersection agents. Experimental results demonstrate consistent improvements in both congestion and environmental performance. Compared with the strongest non-LLM baseline in each scenario, HiLLTS reduces average waiting time by 36.73% under the low-congestion scenario and 14.71% under the high-congestion scenario, while reducing average CO2 emissions by 7.87% and 8.57%, respectively. Larger gains are observed against weaker baselines: under low congestion, HiLLTS achieves reductions of up to 18.00% in emissions and 62.07% in waiting time relative to Fixed-Time control; under high congestion, reductions of up to 28.89% in emissions and 40.36% in waiting time are observed relative to Max Pressure. The ablation study further validates the contribution of LLM-guided coordination over rule-based control
Chinese Translation
城市交通拥堵显著增加了燃料消耗、温室气体排放和通勤延误,导致现代城市中巨大的经济损失和环境危害。传统的交通信号控制策略,如固定时间调度、感应控制和基于强化学习(RL)的方法,提供了不同程度的适应性;然而,基于RL的方法在跨网络或需求模式转移时,可能需要大量的再训练、精心设计的奖励机制和大量的仿真数据。为了解决这些挑战,我们提出了HiLLTS,一个LLM引导的交通信号控制框架,采用由中央协调代理、区域层和多个集群级交叉口代理组成的三层层次结构。实验结果表明,在拥堵和环境性能方面均有持续改善。与每种场景中最强的非LLM基线相比,HiLLTS在低拥堵场景下将平均等待时间减少了36.73%,在高拥堵场景下减少了14.71%,同时在平均CO2排放方面分别减少了7.87%和8.57%。在较弱的基线下观察到更大的收益:在低拥堵情况下,HiLLTS在排放和等待时间上相对于固定时间控制分别实现了高达18.00%和62.07%的减少;在高拥堵情况下,相对于最大压力控制观察到排放和等待时间分别高达28.89%和40.36%的减少。消融研究进一步验证了LLM引导的协调相较于基于规则控制的贡献。
cs.AI / 79 / 2607.22692

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

生成性人工智能心理健康支持的风险治理:多轮安全架构
Areias, Anabela C., Botelho, Catarina, Farinhas, António, Vassilopoulos, Areti, Janela, Dora, Tong, Xin, Guerreiro, Nuno M., D'Eon, Maya, Costa, Fabíola, Rei, Ricardo
Abstract
Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational risk unfolds. We developed a model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions. Synthetic conversations grounded in real-world mental health narratives were used to evaluate the architecture's performance, tested with GPT-5-chat and Qwen3.5-27B, achieving high risk detection performance (specificity: 0.85 (95\%CI: 0.78;0.91), sensitivity: 0.92 (95\%CI: 0.88;0.95)) and increasing clinician-preferred escalation responses by 25.6--59.2pp while preserving rapport and connection. Performance remained stable across conversation length and generalized across both proprietary and open-source models. These findings demonstrate that clinically-grounded safety governance can extend beyond risk detection to improve how LLMs manage evolving mental health risk, providing a scalable framework for safer deployment across models.
Chinese Translation
大型语言模型(LLMs)在情感支持方面的应用日益增多,尽管它们缺乏安全治理不断演变的心理健康风险的机制。现有的安全方法主要检测风险,但很少影响模型在对话风险展开时的响应方式。我们开发了一种模型无关的安全治理架构,该架构结合了上下文风险检测、基于推理的验证和协议指导的响应生成,适用于多轮心理健康互动。我们使用基于现实世界心理健康叙事的合成对话来评估该架构的性能,测试对象为GPT-5-chat和Qwen3.5-27B,取得了高风险检测性能(特异性:0.85(95%CI:0.78;0.91),敏感性:0.92(95%CI:0.88;0.95)),同时在保持良好关系和连接的情况下,增加了临床医生偏好的升级响应25.6--59.2个百分点。性能在对话长度上保持稳定,并在专有和开源模型之间具有良好的泛化能力。这些发现表明,基于临床的安全治理可以超越风险检测,改善LLMs管理不断演变的心理健康风险的方式,为模型的更安全部署提供了可扩展的框架。
cs.AI / 80 / 2607.22694

Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models

贝叶斯重复惩罚:一种原则性相邻条件框架,用于逆转自回归语言模型中的注意力崩溃
Fan, Wenjie, Ma, Bin, Li, Dong
Abstract
Attention collapse in autoregressive language models -- manifested as repetitive token loops where the model becomes trapped in self-reinforcing attractors -- is a persistent pathology that existing decoding-time heuristics fail to address at its root cause. We present a principled framework that penalises or compensates anomalous confidence arising from collapsed generation patterns, by comparing a token's observed frequency against its corpus prior through an adjacent-conditional probability construction. The resulting self-normalising penalty ratio $R=f(m,n,p)/f(np,n,p)$ requires no ad hoc standardisation and admits a closed-form logit offset with zero approximation error. The correction is isolated from the loss gradient and accumulated into a frozen output-layer bias via exponential moving average, enabling deployment as a repair mechanism for models that have already collapsed without requiring intrusive modifications to standard training pipelines. Experimental validation on a 1.5B-parameter model demonstrates that the frozen-bias mechanism can rescue a model already trapped in a collapsed attractor, reducing 2-gram repetition from 0.073 to near 0 while preserving generation quality.
Chinese Translation
自回归语言模型中的注意力崩溃表现为模型陷入自我强化吸引子的重复标记循环,这是一种持续存在的病理现象,现有的解码时间启发式方法未能从根本上解决这一问题。我们提出了一种原则性框架,通过相邻条件概率构造,将标记的观察频率与其语料库先验进行比较,从而惩罚或补偿因崩溃生成模式而产生的异常置信度。所得到的自归一化惩罚比率 $R=f(m,n,p)/f(np,n,p)$ 不需要临时标准化,并且允许以闭合形式的logit偏移,具有零近似误差。该修正与损失梯度相隔离,并通过指数移动平均累积到冻结的输出层偏置中,使其能够作为已经崩溃模型的修复机制进行部署,而无需对标准训练流程进行侵入性修改。在一个15亿参数的模型上的实验验证表明,冻结偏置机制能够拯救已经陷入崩溃吸引子的模型,将2-gram重复率从0.073降低到接近0,同时保持生成质量。
cs.AI / 81 / 2607.22695

PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window

PANOPTICON:基于个人可识别信息(PII)的自然输出令牌集合,用于研究大型语言模型(LLM)上下文窗口中的隐私泄露
Thornton, Ryan, Pritom, Mir Mehedi Ahsan, Gupta, Maanak
Abstract
Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment. While this automation provides clear utility, completing these tasks often requires the insertion of Personally Identifiable Information (PII), strings of information that uniquely identify some individual, raising privacy concerns. However, ethics has prevented the curation of a public, authentic dataset of PII. Without an appropriate dataset, it is difficult to quantify privacy risks. Thus, we introduce the PANOPTICON pipeline and dataset. The dataset, generated by Meta's Llama-3.1-8B-Instruct model, contains 67, 718 prompts, intended for the models context window, containing PII spans derived from 9,674 publicly available synthetic user profiles. We measure lexical diversity and S-BERT diversity of the created dataset to evaluate realism. Finally, we present a case study showcasing the utility of PANOPTICON data for understanding Prompt Inversion Attacks (PIAs). PANOPTICON thus emerges as the first benchmark dataset for studying PIAs over private corpora, providing a foundation for future LLM privacy research.
Chinese Translation
大型语言模型(LLMs)能够对人类语言进行泛化,以完成前所未见的任务,从而实现广泛应用。尽管这种自动化提供了明显的效用,但完成这些任务往往需要插入个人可识别信息(PII),即能够唯一识别某个个体的信息字符串,这引发了隐私问题。然而,伦理问题阻止了公共真实PII数据集的整理。在缺乏适当数据集的情况下,量化隐私风险变得困难。因此,我们引入了PANOPTICON管道和数据集。该数据集由Meta的Llama-3.1-8B-Instruct模型生成,包含67,718个提示,旨在为模型的上下文窗口提供包含来自9,674个公开可用合成用户档案的PII范围的信息。我们测量了所创建数据集的词汇多样性和S-BERT多样性,以评估其真实性。最后,我们展示了一个案例研究,展示了PANOPTICON数据在理解提示反转攻击(PIAs)中的效用。因此,PANOPTICON成为研究私人语料库中PIAs的第一个基准数据集,为未来的LLM隐私研究奠定了基础。
cs.AI / 82 / 2607.22697

Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning

测试时覆盖:面向部署的学习的数据条件化整理
Chang, Nadine, Shen, Maying, Diao, Shizhe, Wang, Jialiang, Chen, Jingde, Breuel, Thomas, Molchanov, Pavlo, Mahmood, Rafid, Alvarez, Jose M.
Abstract
Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, standard data curation methods score training-side criteria rather than directly optimizing deployment match. We introduce TTCov (Test-Time Coverage), a data-level test-conditioned curation method that uses test-side information before training instead of updating model weights at inference. TTCov decomposes deployment-conditioned curation into coverage and distribution. To represent coverage, it builds a task Atlas, a collection of LLM-based atomic propositions (APs) describing deployment-relevant concepts, seeded from open task knowledge and expanded with unmatched APs extracted from unlabeled deployment samples. To represent distribution, it instantiates the matched deployment APs with their frequencies, yielding a Knowledge Atlas (K-Atlas) that operationalizes the deployment distribution as a curation target. TTCov then selects a budgeted training set whose deployment APs distribution approximates this target. We apply TTCov towards autonomous driving (AD), keeping adaptation off the inference path while selecting data with greater deployment-relevant coverage, closer K-Atlas matching, and stronger downstream end-to-end driving performance than data-curation baselines, including seamless adaptability to novel domains via city-to-city expansion.
Chinese Translation
部署的人工智能系统通常从广泛的候选数据池中进行训练,因此需要对数据进行整理,以适应部署测试分布。然而,标准的数据整理方法主要侧重于训练侧标准,而不是直接优化部署匹配。我们提出了TTCov(测试时覆盖),这是一种数据级的测试条件化整理方法,它在训练之前使用测试侧信息,而不是在推理时更新模型权重。TTCov将部署条件的整理分解为覆盖和分布。为了表示覆盖,它构建了一个任务图谱(Task Atlas),这是一个基于大型语言模型(LLM)的原子命题(Atomic Propositions, APs)集合,描述与部署相关的概念,种子来自开放任务知识,并通过从未标记的部署样本中提取的未匹配APs进行扩展。为了表示分布,它实例化了匹配的部署APs及其频率,生成一个知识图谱(Knowledge Atlas, K-Atlas),将部署分布作为整理目标。然后,TTCov选择一个预算训练集,其部署APs分布近似于这一目标。我们将TTCov应用于自动驾驶(Autonomous Driving, AD),在选择具有更大部署相关覆盖、更接近K-Atlas匹配和更强下游端到端驾驶性能的数据时,保持适应性不影响推理路径,相较于数据整理基线,包括通过城市到城市扩展无缝适应新领域。
cs.AI / 83 / 2607.22699

Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures

相似性贯穿始终:大型语言模型中的多语言泛化依赖于语言级相似性结构
Rakshit, Supantho, Goldberg, Adele, Conklin, Henry
Abstract
As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs' representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs' latent representations largely recover the hierarchical structure of the Indo-European language family tree -- grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.
Chinese Translation
随着大型语言模型(LLMs)在多样任务中的能力不断增强,它们的(不)能力进行泛化仍然难以量化,并且在有限领域之外的理解较差。特别是,LLMs在多语言泛化方面表现不佳,尤其是对于非英语语言以及在其训练数据中证据稀缺的语言。为了理解其原因,以及是什么使得某些模型表现优于其他模型,我们借鉴了认知科学领域的长期研究,认为成功的泛化源于相似性空间中的适当表示。我们考察了LLMs的表示在多大程度上捕捉了不同语言之间的层次相似性结构。令人惊讶的是,我们展示了LLMs的潜在表示在很大程度上恢复了印欧语言家族树的层次结构——将同一亚家族的语言在表示空间中紧密地聚集在一起。此外,我们还表明,模型在多大程度上反映语言的相似性结构与它们在XNLI(一个多语言自然语言推理基准)上的表现相关。这扩展了关于规模化相似性驱动泛化的经典研究,展示了如何以相似方式表示相似语言的模型在一种语言到另一种语言的泛化能力上表现更佳。
cs.AI / 84 / 2607.22700

RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction

RoleMix:通过语义标记化统一顺序和非顺序特征以进行点击后转化率预测
Wang, Wenan, Zhao, Qin, Lu, Zhixiang
Abstract
Post-click conversion rate (PCVR) prediction is central to industrial recommendation, but remains challenged by the structural mismatch between sparse, unordered multi-field features and long, domain-specific behavior histories. Existing models often process these signals through separate pathways and fuse them late, weakening semantic roles and limiting cross-signal refinement. We propose RoleMix, a unified interaction architecture that represents sequential and non-sequential evidence through a shared, role-preserving token interface. Non-sequential fields are converted into explicit semantic tokens that preserve user, item, pairwise, dense, contextual, and cross-feature roles, while long behavior domains are compressed into item- and context-aware sequence-query tokens through two-stage hierarchical window attention. The resulting global, semantic, and sequence-query tokens are jointly refined by stacked UniMixing-Lite blocks for PCVR prediction. On the large-scale KDD Cup 2026 Tencent UniRec Challenge, RoleMix achieves 83.648% online AUC, outperforming the official industrial baseline by 1.953%. Ablation studies show that semantic tokenization yields the largest isolated gain, highlighting a key principle for large-scale PCVR modeling: preserving field semantics at the token-interface level is as important as scaling the interaction backbone.
Chinese Translation
点击后转化率(PCVR)预测是工业推荐系统的核心,但由于稀疏、无序的多字段特征与长时间域特定行为历史之间的结构不匹配,仍面临挑战。现有模型通常通过独立路径处理这些信号,并在后期进行融合,这削弱了语义角色并限制了跨信号的细化。我们提出了RoleMix,一种统一的交互架构,通过共享的、保留角色的标记接口表示顺序和非顺序证据。非顺序字段被转换为保留用户、项目、成对、密集、上下文和跨特征角色的显式语义标记,而长行为域则通过两阶段层次窗口注意力压缩为项目和上下文感知的序列查询标记。最终生成的全局、语义和序列查询标记通过堆叠的UniMixing-Lite模块共同细化以进行PCVR预测。在大规模KDD Cup 2026腾讯UniRec挑战赛中,RoleMix实现了83.648%的在线AUC,超越了官方工业基准1.953%。消融研究表明,语义标记化带来了最大的独立增益,突显了大规模PCVR建模的一个关键原则:在标记接口层面保留字段语义与扩展交互骨干同样重要。
cs.AI / 85 / 2607.22706

MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation

MPR-CiteG:通过多组合检索和基于引用的生成增强RAG
Lee, Hyewon, Song, Minkyung, Oh, Junghyun, Han, Seunghoon, Lim, Sungsu
Abstract
This paper presents the MPR-CiteG framework, which achieved second place in the ScienceON AI Challenge by addressing two fundamental challenges in generative AI: inefficient retrieval and the absence of source verification. We propose a dual-component system, termed MPR-CiteG, in which the Multi-Portfolio Retriever (MPR) efficiently retrieves diverse and relevant information, while the Citation-Grounded Generation (CiteG) module ensures that every generated output remains factually consistent and explicitly attributed to its source. MPR-CiteG represents a significant step toward building more trustworthy and accurate LLMs that are not only capable of generating information but also of grounding their responses in reliable evidence, thereby mitigating common issues like model hallucination. Extensive experiments on the challenge dataset validate the effectiveness and reliability of our approach. Our code is available at https://github.com/2noweyh/MPR-citeG.
Chinese Translation
本文提出了MPR-CiteG框架,该框架在ScienceON AI Challenge中获得第二名,旨在解决生成性人工智能中的两个基本挑战:低效的检索和缺乏来源验证。我们提出了一种双组件系统,称为MPR-CiteG,其中多组合检索器(Multi-Portfolio Retriever, MPR)高效地检索多样且相关的信息,而基于引用的生成模块(Citation-Grounded Generation, CiteG)确保每个生成的输出在事实上一致并明确归属其来源。MPR-CiteG代表了朝着构建更可信和准确的大型语言模型(LLMs)迈出的重要一步,这些模型不仅能够生成信息,还能够将其响应基于可靠证据,从而减轻模型幻觉等常见问题。对挑战数据集的广泛实验验证了我们方法的有效性和可靠性。我们的代码可在 https://github.com/2noweyh/MPR-citeG 获取。
cs.AI / 86 / 2607.22713

SEGRA: Structured Experience-Guided Graph Reasoning Agent for Gremlin Based Question Answering

SEGRA:基于结构化经验引导的图推理代理用于 Gremlin 问答
Lyu, Saiyue, Dundua, Mariam, Kapoor, Vishaal, Ahuja, Sarthak, Kordjazi, Neda, Yortucboylu, Evren, Amin, Harsh, Steinert, Rebecca
Abstract
Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes, and historical resolutions. Yet querying them in Gremlin requires knowledge of graph schemas, traversal semantics, edge directionality, and property-graph-specific constraints, making them difficult for non-expert operators to use. We introduce SEGRA, an experience-guided agent for enterprise text-to-Gremlin question answering. SEGRA integrates intent routing, schema- and taxonomy-grounded query generation, multi-shot decomposition, execution-aware verification, and a curriculum-bootstrapped skill library that reuses verified query patterns. On an enterprise IT support benchmark, SEGRA achieves a $7.0\times$ higher mean judge score than backbone-only chain-of-thought prompting. Its skill library further reduces LLM calls by $20\%$ and dollar cost by $18\%$ relative to SEGRA without skills, while preserving answer quality. These results show that schema-grounded agent design and reusable execution experience improve both accuracy and efficiency for enterprise graph QA.
Chinese Translation
企业 IT 支持知识图谱捕捉了案例、用户、设备、症状、分类类别、根本原因和历史解决方案之间的丰富关系。然而,在 Gremlin 中查询这些图谱需要了解图的模式、遍历语义、边的方向性以及特定于属性图的约束,这使得非专业操作员难以使用。我们介绍了 SEGRA,一种用于企业文本到 Gremlin 问答的经验引导代理。SEGRA 集成了意图路由、基于模式和分类的查询生成、多轮分解、执行感知验证以及一个利用经过验证的查询模式的课程引导技能库。在企业 IT 支持基准测试中,SEGRA 的平均评审分数比仅使用基础链式思维提示高出 $7.0 imes$。其技能库进一步使 LLM 调用减少 $20\%$,相较于没有技能的 SEGRA,成本降低 $18\\%$,同时保持了答案质量。这些结果表明,基于模式的代理设计和可重用的执行经验提高了企业图形问答的准确性和效率。
cs.AI / 87 / 2607.22732

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning

大型语言模型游戏代理中的空间推理:因果上下文和多步规划的影响
Jiwatode, Mohit, Fuchs, Ronja, Schmöcker, Robin, Rosenhahn, Bodo, Dockhorn, Alexander
Abstract
LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and evaluates whether causal prompt augmentation and multi-step planning can improve win-rates while managing response latency. Using the open-source Qwen3 model family, we conduct experiments across varying model scales, reasoning modes, and planning horizons. We further introduce a focused GVGAI benchmark consisting of three custom games with five difficulty levels to isolate spatial navigation. The evaluation follows two paradigms: an initial ``positioning experiment'' to test an agent's ability to find its exact coordinates, and a study of game-play success. Our results show that while larger models with an enabled thinking mode identify their positions more accurately, overall performance in coordinate matching remains limited for smaller models. Win rates decrease as game levels and layout complexity increase, validating the benchmark's difficulty scaling. Integrating causal context into the prompts tends to improve the agents' success rates, particularly for bigger models. While enabling thinking mode and longer planning horizons significantly improve performance, multi-step planning further reduces mean per-step response times, offering a practical trade-off between reasoning depth and execution speed.
Chinese Translation
基于大型语言模型(LLM)的游戏代理在处理更复杂任务时通常表现不佳。本研究探讨这些失败是否与有限的空间推理能力有关,并评估因果提示增强和多步规划是否能够在管理响应延迟的同时提高胜率。我们使用开源的 Qwen3 模型系列,在不同的模型规模、推理模式和规划范围内进行实验。此外,我们还引入了一个专注的 GVGAI 基准,包括三个自定义游戏和五个难度级别,以隔离空间导航。评估遵循两种范式:初始的“定位实验”以测试代理找到其确切坐标的能力,以及游戏成功率的研究。我们的结果表明,虽然启用思考模式的较大模型能够更准确地识别其位置,但较小模型在坐标匹配方面的整体表现仍然有限。随着游戏关卡和布局复杂性的增加,胜率下降,验证了基准的难度比例。将因果上下文整合到提示中往往能提高代理的成功率,特别是对于较大模型。虽然启用思考模式和更长的规划范围显著提高了性能,但多步规划进一步减少了每步的平均响应时间,提供了推理深度与执行速度之间的实用权衡。
cs.AI / 88 / 2607.22750

Commitment To Cooperation With Self-Negotiated Contracts

通过自我协商合同实现合作承诺
Wyse, Tim, Bustos, Kaitlin, Volkova, Yulia, Kleiman-Weiner, Max
Abstract
As AI agents operate with increasing autonomy in a multi-agent world, they will need to learn to cooperate with other agents and with humans to generate mutual benefits. However, cooperation is a challenge because the costs of cooperation are often incurred early on, but the benefits are only realized later, creating an incentive to defect. How can AI agents cooperate with commitment? Here, we draw on inspiration from legal institutions and contracting that human societies have used to solve principal-agent problems of this kind. Contracts provide observable representations of agreements that enable credible commitments through the enforcement of terms. We study the role of contract-based cooperation using LLM-based agents in \CT, a spatial-temporal game that combines bargaining with navigation towards a goal. We study a suite of contract representations that range from formal contracts that compile to code to natural contracts that require reinterpretation. We evaluate agents with a range of LLM backbones using different sizes and providers. We find that self-negotiated contracts can improve cooperative outcomes beyond what is possible with regular trading.
Chinese Translation
随着人工智能代理在多智能体世界中越来越自主地运作,它们需要学习与其他代理和人类合作,以产生互惠的利益。然而,合作是一项挑战,因为合作的成本往往在早期就会产生,而收益却仅在后期才能实现,这导致了背叛的诱因。人工智能代理如何能够在承诺的基础上进行合作?在这里,我们借鉴了人类社会为解决此类委托-代理问题而使用的法律制度和合同的启示。合同提供了可观察的协议表示,通过对条款的执行实现可信的承诺。我们研究了基于合同的合作在 extit{CT}(一个结合了谈判与朝目标导航的时空游戏)中使用基于大语言模型(LLM)的代理的作用。我们研究了一系列合同表示形式,从编译成代码的正式合同到需要重新解释的自然合同。我们评估了使用不同规模和提供者的多种 LLM 主干的代理。我们的研究发现,自我协商的合同可以改善合作结果,超越常规交易所能实现的水平。
cs.AI / 89 / 2607.22829

Disentangling Multi-View Scanning in Mamba for Network Traffic Anomaly Detection

在 Mamba 中解构多视角扫描以进行网络流量异常检测
Lian, Xinglin, Cao, Chengtai, Zhong, Ting, Zhou, Fan
Abstract
Network Traffic Anomaly Detection (NTAD) is a critical task in cybersecurity, yet timely and accurate anomaly detection remains challenging. Mamba has emerged as a particularly promising backbone for NTAD due to its linear-time complexity for long-sequence modeling. It further incorporates a dedicated multi-view scanning mechanism to enhance detection precision through complementary contextual cues. However, we identify a previously overlooked structural deficiency in multi-view Mamba scanning for NTAD: redundancy accumulation. Specifically, distinct scanning branches capture substantial view-invariant information, which is repeatedly amplified during multi-view fusion; conversely, view-specific information is diluted or even suppressed, leading to representation homogenization and multi-view degradation. To address this problem, we propose DisenMamba, a novel disentangled multi-view Mamba framework. DisenMamba reformulates multi-view scanning as a two-stage disentangle-then-fuse process that explicitly separates view-invariant and view-specific components prior to fusion. This design prevents the invariant information accumulation while preserving complementary multi-view cues, yielding more discriminative representations for subtle traffic anomalies. Extensive experiments demonstrate the effectiveness of DisenMamba, establishing a new disentangled multi-view Mamba paradigm. Code is available at https://github.com/ikun0124/DisenMamba.
Chinese Translation
网络流量异常检测(NTAD)是网络安全中的一项关键任务,但及时和准确的异常检测仍然具有挑战性。Mamba 因其在长序列建模中的线性时间复杂度而成为 NTAD 的一种特别有前景的基础架构。它进一步结合了一种专门的多视角扫描机制,通过互补的上下文线索来提高检测精度。然而,我们发现多视角 Mamba 扫描在 NTAD 中存在一个先前被忽视的结构缺陷:冗余累积。具体而言,不同的扫描分支捕获了大量视角不变的信息,这些信息在多视角融合过程中被反复放大;相反,特定视角的信息则被稀释甚至抑制,导致表示同质化和多视角退化。为了解决这个问题,我们提出了 DisenMamba,一种新颖的解构多视角 Mamba 框架。DisenMamba 将多视角扫描重新构建为一个两阶段的解构-再融合过程,在融合之前明确分离视角不变和视角特定的组件。这一设计防止了不变信息的累积,同时保留了互补的多视角线索,从而为微妙的流量异常提供了更具辨别力的表示。大量实验表明 DisenMamba 的有效性,建立了一种新的解构多视角 Mamba 范式。代码可在 https://github.com/ikun0124/DisenMamba 获取。
cs.AI / 90 / 2607.22854

Coordinated Networking for On-Device Agent-Augmented Real-Time Communication

设备端协调网络用于增强代理的实时通信
Lee, Goodsol, Yi, Juheon, Wang, Jinglu, Xu, Haowen, Bahk, Saewoong, Lu, Yan
Abstract
AI agents are enabling a new paradigm of agent-augmented real-time communication (RTC), where humans focus on high-level collaboration, while agents autonomously retrieve, analyze, and generate information in real time to support their interactions. These apps enable new experiences across various domains: for example, when corporate employees co-author a legal document, their agents can discuss and draft on their behalf, sparing them the burden of manually reviewing each other's work. As existing cloud-based agents suffer from privacy risks and unscalable server costs, on-device agent-augmented RTC offers a promising alternative. However, this on-device paradigm introduces a new networking challenge: contention between concurrent traffic flows generated by humans (for live video streaming) and agents (for sending context files for analysis). We design HFS, a framework to ensure both high live video quality and low agent response latency in agent-augmented RTC apps. We achieve the goal through an app-guided multi-flow transport approach, where a unified app-layer orchestrator jointly controls the sending rates of live video and agent context flows based on their heterogeneous app requirements. Our prototype built atop WebRTC and llama.cpp demonstrates that HAFS outperforms baselines, achieving 1.5x higher video quality while reducing agent response time by 31%.
Chinese Translation
人工智能代理正在推动一种新的增强代理实时通信(RTC)范式,在这种范式中,人类专注于高层次的协作,而代理则自主实时检索、分析和生成信息以支持他们的互动。这些应用在各个领域提供了新的体验:例如,当公司员工共同撰写法律文件时,他们的代理可以代表他们进行讨论和草拟,从而减轻他们手动审查彼此工作的负担。然而,现有的基于云的代理面临隐私风险和不可扩展的服务器成本,设备端增强代理的RTC提供了一种有前景的替代方案。然而,这种设备端范式引入了一个新的网络挑战:人类(用于实时视频流)和代理(用于发送分析所需的上下文文件)生成的并发流量之间的竞争。我们设计了HFS,一个框架,旨在确保增强代理RTC应用中高质量的实时视频和低延迟的代理响应。我们通过一种应用引导的多流传输方法实现了这一目标,其中一个统一的应用层协调器根据其异构的应用需求共同控制实时视频和代理上下文流的发送速率。我们基于WebRTC和llama.cpp构建的原型展示了HAFS在性能上优于基线,实现了1.5倍更高的视频质量,同时将代理响应时间减少了31%。
cs.AI / 91 / 2607.22868

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

什么可以被强制执行?工具使用代理的认证运行时安全理论
Ray, Shawn
Abstract
Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior. We separate three questions. First, relative to fixed oracle predicates, a deterministic gate enforces exactly the nonempty safety policies whose good prefixes its register model recognizes; policy nontriviality is undecidable with two decrementable counters but in PSPACE for a separable monotone fragment. Second, under a fixed exogenous law, Neyman-Pearson gives the exact false-block/miss frontier and conformal calibration gives a finite-sample marginal certificate, possibly via block-all. Third, once blocking changes future proposals, static scores and ungated trajectories need not identify the closed-loop frontier; a specified finite controlled model instead yields an occupancy program. Bounded representation attacks add a robustness margin, so benign calibration alone does not transfer. Experiments target these distinctions through static diagnostics, controlled-model enumeration, representation rewrites, and paired closed-loop reruns.
Chinese Translation
运行时保护措施在不可逆的工具调用之前起作用,但其保证依赖于可表示的政策状态、法官观察到的内容以及干预是否会改变未来行为。我们将三个问题分开讨论。首先,相对于固定的预言机谓词,确定性门精确地强制执行其寄存器模型所识别的非空安全政策的良好前缀;对于两个可递减计数器,政策的非平凡性是不可判定的,但在一个可分离的单调片段中是 PSPACE 可判定的。其次,在固定的外生法律下,Neyman-Pearson 定理给出了精确的虚假阻断/漏失边界,而符合校准则提供了有限样本边际证书,可能通过全阻断实现。第三,一旦阻断改变未来提案,静态评分和无门轨迹不一定能识别闭环边界;而指定的有限控制模型则产生一个占用程序。有限表示攻击增加了稳健性边际,因此单纯的良性校准并不能转移。实验通过静态诊断、控制模型枚举、表示重写和配对闭环重跑来针对这些区别。
cs.AI / 92 / 2607.22877

Physical AI Governance: From Theory to Practice Across Life Cycle

物理人工智能治理:从理论到实践的全生命周期
Yang, Wang, Wang, Shaobo, Liu, Hongxuan, Cai, Xiaoran, He, Yunyu, Zhou, Jingzong, Ma, Mengzhong, Yu, Yi, Sharma, Rohit, Fu, Jingjing, Qi, Peng
Abstract
With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact with, and act in the physical world. Unlike traditional AI, Physical AI operates under real-time safety constraints, continuously interacts with dynamic environments, and coexists with humans, introducing governance challenges that existing AI governance frameworks do not explicitly address. This paper presents a comprehensive survey of Physical AI governance from both scientific and operational perspectives. We synthesize existing governance principles and organize them into a unified governance framework tailored to physical AI systems. Building on this foundation, we propose a five-stage Physical AI lifecycle comprising research, design, data, model development, and deployment, and demonstrate how governance can be operationalized across each stage through concrete implementation practices. By connecting governance principles with engineering workflows, this survey provides a structured reference for researchers, developers, and policymakers to build Physical AI systems that are safe, trustworthy, and aligned with societal values.
Chinese Translation
随着物理人工智能的出现,人工智能正从基于屏幕的应用扩展到能够感知、与物理世界互动并采取行动的具身系统。与传统人工智能不同,物理人工智能在实时安全约束下运行,持续与动态环境互动,并与人类共存,这引入了现有人工智能治理框架未明确解决的治理挑战。本文从科学和操作两个角度对物理人工智能治理进行了全面的调查。我们综合现有的治理原则,并将其组织成一个针对物理人工智能系统的统一治理框架。在此基础上,我们提出了一个包含研究、设计、数据、模型开发和部署的五阶段物理人工智能生命周期,并展示了如何通过具体实施实践在每个阶段实现治理。通过将治理原则与工程工作流程相连接,本调查为研究人员、开发者和政策制定者提供了一个结构化的参考,以构建安全、可信并与社会价值观相一致的物理人工智能系统。
cs.AI / 93 / 2607.22902

How Well Can AI Generate Backlogs from App Mockups?

人工智能从应用程序原型生成待办事项列表的能力如何?
Airaldi, Andrea Lezcano, Romera, Lourdes, Maalej, Walid
Abstract
Creating sprint backlogs requires considerable effort, as items such as epics, user stories, and tasks can be missed or inconsistently specified. We propose a multimodal approach to support backlog generation from visual app mockups, an artifact available at early project stages. We evaluate three prompting strategies on GPT-4o: a zero-shot baseline, Compositional Chain-of-Thought (CCoT) for vision-language reasoning, and a persona-driven prompt. We study seven app development projects across two countries and interview developers about the results. Overall, we observed that the baseline prompt favours recall over precision, whereas CCoT is more balanced, achieving average F1 scores of 52-66% for epics and user stories. Tasks were more challenging to generate accurately. Precision gains were most consistent when adding architectural context, particularly for backend tasks (precision gains up to 35%). Interviews with developers revealed that up to 26% of false positives were still considered useful, reflecting the creative and open-ended nature of backlog creation. To capture this, we propose a new measure called Revised Recall, which complements ground-truth evaluation with developer assessments. Our findings suggest that hybrid prompting with architectural context can assist backlog generation from early mockups, though results vary by item type and developer oversight remains necessary.
Chinese Translation
创建冲刺待办事项列表需要相当大的努力,因为史诗、用户故事和任务等项目可能会被遗漏或不一致地指定。我们提出了一种多模态方法,以支持从视觉应用程序原型生成待办事项列表,这是一种在项目早期阶段可用的工件。我们在 GPT-4o 上评估了三种提示策略:零-shot 基线、用于视觉-语言推理的组合链式思维(Compositional Chain-of-Thought, CCoT)和基于角色的提示。我们研究了来自两个国家的七个应用程序开发项目,并对开发人员进行了结果访谈。总体而言,我们观察到基线提示更倾向于召回而非精确度,而 CCoT 则更为平衡,史诗和用户故事的平均 F1 分数为 52-66%。任务的准确生成更具挑战性。当添加架构上下文时,精确度的提升最为一致,尤其是对于后端任务(精确度提升高达 35%)。与开发人员的访谈显示,最多有 26% 的假阳性仍被认为是有用的,反映了待办事项创建的创造性和开放性特征。为此,我们提出了一种新的度量标准,称为修订召回(Revised Recall),它通过开发人员评估来补充真实值评估。我们的研究结果表明,结合架构上下文的混合提示可以帮助从早期原型生成待办事项列表,尽管结果因项目类型而异,开发人员的监督仍然是必要的。
cs.AI / 94 / 2607.22917

Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Coding Agent Teams

代理团队工作区:一个自动化的、持久的长寿命编码代理团队工作空间
Wang, Shouren
Abstract
Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the most powerful LLM coding agents and is capable of conducting complex coding tasks. However, several drawbacks can undermine long-term agentic workflows. (1) Irrecoverable agent teams: The Agent Teams feature is powerful, but the working state accumulated by each teammate is lost and cannot be resumed once the process stops, for example, when a terminal is closed. (2) Compaction erodes working detail: Compaction condenses the conversation into a summary, causing an agent's working details to become vague. (3) Agentic "technical debt": Over time, a user's decisions and the agents' operations become trapped in compacted old chats, making the project increasingly difficult to maintain and review. (4) Heavy prompt writing: Assigning or handing off tasks requires users to repeatedly write long prompts to achieve the expected agentic performance. We propose ATWZ (Agent Team Work Zone), a filesystem-based operations layer built around Claude Code's native Agent Teams that addresses these problems. Its central design principle is to treat each agent and teammate as a human employee and preserve their important working state in files stored in a dedicated directory called a "workstation," together with the skills, hooks, and scripts that use and maintain these files. With ATWZ, an agent team can periodically back up its working state, allowing an agent's knowledge to be recovered after compaction. After a process ends, the team can be restored with a single command. These features also substantially mitigate the agentic "technical debt" described above. Moreover, within ATWZ, agent "employees" can send documents to one another, greatly reducing the effort required to write prompts.
Chinese Translation
大型语言模型(LLM)代理显著改善了编码和编程工作流程。特别是Claude Code,是最强大的LLM编码代理之一,能够执行复杂的编码任务。然而,几个缺点可能会削弱长期的代理工作流程。(1)不可恢复的代理团队:代理团队功能强大,但每个队友积累的工作状态在过程停止后会丢失,无法恢复,例如,当终端关闭时。(2)压缩侵蚀工作细节:压缩将对话浓缩为摘要,导致代理的工作细节变得模糊。(3)代理的“技术债务”:随着时间的推移,用户的决策和代理的操作被困在压缩的旧聊天中,使得项目越来越难以维护和审查。(4)繁重的提示编写:分配或移交任务需要用户反复编写长提示,以实现预期的代理性能。我们提出了ATWZ(代理团队工作区),这是一个基于文件系统的操作层,围绕Claude Code的原生代理团队构建,以解决这些问题。其核心设计原则是将每个代理和队友视为人类员工,并将其重要的工作状态保存在一个名为“工作站”的专用目录中的文件中,以及使用和维护这些文件的技能、钩子和脚本。通过ATWZ,代理团队可以定期备份其工作状态,允许在压缩后恢复代理的知识。在过程结束后,团队可以通过一个命令恢复。这些功能也大大减轻了上述提到的代理“技术债务”。此外,在ATWZ中,代理“员工”可以相互发送文档,大大减少了编写提示所需的努力。
cs.AI / 95 / 2607.22926

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

SAGE:高影响生成性人工智能生命周期控制的安全优先深度防御护栏
Eslamimehr, Mahdi
Abstract
High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authorization-separated architecture in which credible catastrophic-enablement risk constrains admissibility before utility, latency, or commercial objectives are considered. It combines signed release manifests, diverse detectors, robust risk envelopes, least-risk defaults, output checking, three-valued monitoring, protected audit chains, containment, and rollback. Formal results establish safety priority, conservative detector bounds, monotone release gating, tamper-evident records, and an authorization cut; two PRISM abstractions verify authorization separation and lifecycle invariants under explicit assumptions. A frozen, vendor-symmetric study sent 84 cases to each of four GPT, four Claude, and two Gemini snapshots: 840 calls yielded 794 target responses, 46 provider errors, and 449 successful judgments covering 375 responses. Eight snapshots had complete judged domain coverage. Harmful-compliance estimates were low; variation arose mainly from benign utility and safe redirection. Seven multiplicity-adjusted contrasts involving Claude, Gemini, or GPT-5 snapshots and the GPT-5 mini and GPT-5 nano snapshots were supported, while no tested contrast between the Claude or Gemini snapshots and GPT-5 or GPT-5.5 survived correction. The observed harmful-compliance range is a conservative, protocol-bound view from one generation per prompt with no tools, retrieval, history, or human adjudication; it is not an upper bound on operational assistance. A preregistered extension specifies how to test a wider best-worst gap using a locked split, repeated sampling, multi-turn and sandboxed-tool conditions, and domain-expert scoring.
Chinese Translation
高影响生成性人工智能使得灾难性误用成为一个生命周期控制问题,而不仅仅是一个提示过滤问题。SAGE是一种安全优先、授权分离的架构,其中可信的灾难性使能风险在考虑效用、延迟或商业目标之前限制了可接受性。它结合了签名发布清单、多样化检测器、稳健风险包、最小风险默认、输出检查、三值监控、受保护的审计链、隔离和回滚。正式结果确立了安全优先级、保守的检测器界限、单调发布门控、可篡改证据记录和授权切割;两个PRISM抽象验证了在明确假设下的授权分离和生命周期不变性。一项冻结的、供应商对称的研究向四个GPT、四个Claude和两个Gemini快照发送了84个案例:840次调用产生了794个目标响应、46个提供者错误和449个成功判断,涵盖了375个响应。八个快照具有完整的判断领域覆盖。对有害合规性的估计较低;变异主要来自良性效用和安全重定向。涉及Claude、Gemini或GPT-5快照以及GPT-5 mini和GPT-5 nano快照的七个多重调整对比得到了支持,而在Claude或Gemini快照与GPT-5或GPT-5.5之间测试的对比没有通过校正。观察到的有害合规范围是一个保守的、协议限制的视角,基于每个提示生成一代,没有工具、检索、历史或人工裁决;这并不是对操作辅助的上限。一个预注册的扩展指定了如何使用锁定分割、重复采样、多轮和沙盒工具条件以及领域专家评分来测试更广泛的最佳-最差差距。
cs.AI / 96 / 2607.22928

Design Theater: A Benchmark for Generative UI

设计剧场:生成用户界面的基准测试
Imteyaz, Kashif, Imteyaz, Kaif, Rajpal, Nakul, Shaikh, Kaif, Muller, Michael, Savage, Saiph
Abstract
Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the interface, these tools generate user-facing design rationales that explain their layout, accessibility, and design choices. However, it remains unclear whether these stated rationales are actually reflected in the interfaces they produce. We call this disconnect ``Design Theater'': plausible and confident design rationales that have little relationship to the actual implementation. To study this phenomenon, we introduce a benchmark and three metrics for measuring Design Theater. The benchmark includes 24 UI generation tasks spanning structural, styling, and functional design requirements. Using this benchmark, we evaluate 120 interfaces created by five generative UI tools. On average, over 25\% of user-facing design rationales are not implemented in the generated interface, and the implementation failure increases to 34\% for functional requirements. Tools recognize roughly half of the UX principles embedded in prompts (mean = 0.54), with four of five tools implementing 6\% or fewer functional principles. We also measure interface similarity across tools and find convergence in visual appearance and layout organization, with greater variation in color choices. Overall, we contribute: 1) the concept of Design Theater; 2) a benchmark with metrics for assessing whether the stated reasoning of generative UI tools is reflected in their implementations; 3) and findings from a systematic evaluation of these tools. We discuss what these findings mean for the design and evaluation of generative UI tools.
Chinese Translation
生成用户界面(UI)工具承诺通过将自然语言描述转化为完整的界面来实现UI设计的民主化。除了界面,这些工具还生成用户面向的设计理由,以解释其布局、可访问性和设计选择。然而,目前尚不清楚这些声明的理由是否真正反映在它们所生成的界面中。我们将这种脱节称为“设计剧场”(Design Theater):看似合理且自信的设计理由与实际实现之间几乎没有关系。为了研究这一现象,我们引入了一个基准测试和三个用于测量设计剧场的指标。该基准测试包括24个UI生成任务,涵盖结构、样式和功能设计要求。利用该基准测试,我们评估了五个生成UI工具创建的120个界面。平均而言,超过25%的用户面向设计理由未在生成的界面中实现,而功能要求的实现失败率则上升至34%。工具大约识别出嵌入提示中的一半用户体验(UX)原则(平均值=0.54),其中五个工具中有四个实现的功能原则不超过6%。我们还测量了不同工具之间的界面相似性,发现视觉外观和布局组织趋于一致,而颜色选择则存在更大的变异性。总体而言,我们的贡献包括:1)设计剧场的概念;2)一个基准测试及其评估生成UI工具所述推理是否反映在其实现中的指标;3)对这些工具进行系统评估的研究结果。我们讨论了这些发现对生成UI工具的设计和评估意味着什么。
cs.AI / 97 / 2607.22947

Let AI Agents Translate Networks, Not Reason About Them

让人工智能代理翻译网络,而不是对其进行推理
Hè, Hongyu, Apostolaki, Maria
Abstract
A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network has one, since writing a model by hand demands rare expertise and is hard to keep current as the network changes frequently. At its core, network modeling is a typographical exercise: it translates network artifacts (e.g., configurations, topology, and routing state) into rules in formal logic. Translation of this kind is what large language models (LLMs) nowadays do well. Unlike free-form AI reasoning, such translation can be formally verified. Once modeling is no longer the bottleneck, trusting AI to reason over large, complex networks no longer makes sense. Our position therefore cuts against the prevailing race to put autonomous AI agents in charge end-to-end. We instead confine AI to translation and rely on a solver for reliable long-horizon reasoning, building a reusable formal model of general network behavior that can then be specialized to specific tasks, e.g., root-cause analysis (RCA). We build TypoNet that constructs and validates a symbolic model of an emulated production-scale WAN from the network's own artifacts. Our preliminary evaluation shows TypoNet helps in two ways. On its own, TypoNet answers operational questions (e.g., reachability verification and change-impact analysis) faster, more cheaply, and more reliably than an LLM. As a tool for an AI agent, TypoNet boosts fault localization at lower cost. The result makes the case for AI that builds verifiable network models and relies on a solver for reliable long-horizon reasoning.
Chinese Translation
一个正式模型能够验证可达性、定位故障或预测变更的影响范围。然而,几乎没有生产网络具备这样的模型,因为手动编写模型需要稀缺的专业知识,并且随着网络频繁变化,保持模型的更新也很困难。从本质上讲,网络建模是一种排版练习:它将网络工件(例如,配置、拓扑和路由状态)翻译为形式逻辑中的规则。这种翻译正是当前大型语言模型(LLMs)擅长的。与自由形式的人工智能推理不同,这种翻译可以进行形式验证。一旦建模不再是瓶颈,信任人工智能对大型复杂网络进行推理就不再有意义。因此,我们的观点与当前推动自主人工智能代理全程负责的趋势相悖。我们将人工智能的角色限制为翻译,并依赖求解器进行可靠的长期推理,构建一个可重用的正式模型,以描述一般网络行为,然后可以专门化为特定任务,例如根本原因分析(RCA)。我们构建了TypoNet,它从网络自身的工件中构建和验证一个模拟生产规模广域网的符号模型。我们的初步评估表明,TypoNet在两个方面提供了帮助。TypoNet本身比大型语言模型更快、更便宜且更可靠地回答操作问题(例如,可达性验证和变更影响分析)。作为人工智能代理的工具,TypoNet以更低的成本提升了故障定位能力。该结果为构建可验证网络模型并依赖求解器进行可靠长期推理的人工智能提供了依据。
cs.AI / 98 / 2607.22953

Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI

仅分享请求所需的信息:面向视角的人工智能的联邦披露
Khanzadeh, Sourena, Platnick, Daniel, Alirezaie, Marjan, Rahnama, Hossein
Abstract
Modern AI systems bring societal risks such as mass surveillance, extreme concentrations of power, and loss of user autonomy---calling into question a model where third-parties collect and control massive amounts of user data. Users require a sovereign system to securely own, govern, and disclose their context while remaining compliant across regulated domains with strict provenance, interpretability, and policy adherence. Perspective-aware AI approaches this by transforming a user's aggregated personal data into a structured identity model called a \emph{Chronicle}: a temporal knowledge graph that represents and grows with the user. Chronicles support the secure disclosure of context across federated networks. A Chronicle holder may expose a queryable, authorized view that a third-party agent may consult without centralizing anyone's data. This paper explores the problem of minimum-necessary disclosure across domain boundaries: when a requester's agent queries a Chronicle, how can the system constrain its response to release only what the requester's relationship, stated purpose, and specific task require? We propose \textbf{Provenance Preserving Chronicles} (PPC), a federated protocol that compiles each holder's Chronicle into a compact \emph{authorized evidence subgraph} governed by one rule: \emph{share no more than the request requires}. Holders keep local sovereignty; an access controller projects relationship-aware views over domain-expert ontologies; and a two-phase flow returns provenance-linked text first, releasing raw artifacts only after explicit holder approval. We frame the problem, map gaps in blockchain, P2P, and holder-sovereign designs, define the core constructs, and sketch the protocol with an explicit threat model.
Chinese Translation
现代人工智能系统带来了诸如大规模监控、极端权力集中和用户自主权丧失等社会风险——这使得第三方收集和控制大量用户数据的模型受到质疑。用户需要一个主权系统,以安全地拥有、管理和披露他们的上下文,同时在严格的来源、可解释性和政策遵循的监管领域中保持合规。面向视角的人工智能通过将用户的聚合个人数据转化为一个称为 extit{Chronicle}的结构化身份模型来解决这一问题:这是一个随着用户而发展和表示的时间知识图谱。Chronicles支持在联邦网络中安全地披露上下文。Chronicle持有者可以暴露一个可查询的、授权的视图,第三方代理可以在不集中任何人数据的情况下进行咨询。本文探讨了跨域边界的最小必要披露问题:当请求者的代理查询一个Chronicle时,系统如何限制其响应,仅释放请求者的关系、声明目的和特定任务所需的信息?我们提出了 extbf{保留来源的Chronicles}(PPC),这是一种联邦协议,将每个持有者的Chronicle编译成一个紧凑的 extit{授权证据子图},遵循一个规则: extit{仅分享请求所需的信息}。持有者保持本地主权;访问控制器在领域专家本体上投影关系感知视图;并且一个两阶段流程首先返回与来源相关的文本,仅在明确获得持有者批准后释放原始文档。我们框定了问题,映射区块链、P2P和持有者主权设计中的差距,定义核心构造,并勾勒出具有明确威胁模型的协议。
cs.AI / 99 / 2607.22962

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

ConsistencyGate:通过自一致性准入控制防止大型语言模型代理中的记忆污染
Zhang, Yan, Li, Shibo
Abstract
LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subsequent step, a failure mode we call memory contamination. Existing memory management addresses retrieval and capacity but not write-time correctness; this admission problem cannot be solved by utility- or recency-based criteria, and uncontrolled contamination compounds across long trajectories. We propose ConsistencyGate, a write-time admission gate that, before committing a candidate fact m extracted from context c, queries the LLM K times for a soft support score and admits m only when the average exceeds a threshold. The mechanism is model-agnostic, requires no fine-tuning, and reduces to a single forward pass in a log-probability variant for latency-sensitive deployments. To measure the effect on natural data, we construct two real-conversation benchmarks (LoCoMo-Contam and MSC-Contam) by planting controlled single-detail corruptions in long-term conversations from LoCoMo and MSC, and complement them with a structured synthetic corpus (MemContam) that isolates a near-oracle upper bound. Across four LLM backbones, ConsistencyGate reduces contamination on every benchmark relative to a write-everything baseline, with the cost concentrated on facts that are stated only implicitly in the source context. We release all three benchmarks together with the gate implementation.
Chinese Translation
在多轮交互中操作的大型语言模型(LLM)代理会在外部记忆存储中积累事实,并将其作为下游推理的前提。因此,在某一步写入的虚假事实会在随后的每一步中持续存在,形成我们称之为记忆污染的失效模式。现有的记忆管理方法主要关注检索和容量,但未能解决写入时的正确性;这一准入问题无法通过效用或时效性标准来解决,且在长时间轨迹中,失控的污染会不断加剧。我们提出了ConsistencyGate,这是一种写入时准入门控机制,在提交从上下文c中提取的候选事实m之前,先对LLM进行K次查询以获取软支持分数,仅在平均分数超过阈值时才接受m。该机制与模型无关,无需微调,并且在对延迟敏感的部署中简化为单次前向传递的对数概率变体。为了测量对自然数据的影响,我们通过在LoCoMo和MSC的长期对话中植入受控的单细节腐败,构建了两个真实对话基准(LoCoMo-Contam和MSC-Contam),并辅以一个结构化的合成语料库(MemContam),以隔离近似于神谕的上限。在四个LLM基础模型上,ConsistencyGate在每个基准测试中相对于写入一切的基线减少了污染,其成本集中在仅在源上下文中隐含陈述的事实上。我们将这三个基准及门控实现一起发布。
cs.AI / 100 / 2607.23019

Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming

推理波普:利用归纳逻辑编程修补上下文推理
Chen, Zirong, Ma, Meiyi
Abstract
Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound. We present Reason Popper-ly, a neurosymbolic framework that uses inductive logic programming (ILP) to learn relation composition rules from reasoning traces and deploys them as an online verifier for step-level correction. Given an LLM-generated trace, the method checks each inferred step against the learned rule table, diagnoses the violation type, rewrites incorrect steps with symbolically derived repairs, and regenerates the remaining suffix so that the model can produce its final answer conditioned on a verified trace. We evaluate on CLUTRR, a multi-hop kinship reasoning benchmark, using five language models over reasoning chains of 2 to 10 hops. Across all models, Reason Popper-ly consistently improves terminal accuracy over standard CoT, with gains of up to 48 percentage points for small models and 15 points for frontier models on the longest chains. Compared with a fully exogenous symbolic pipeline, our method performs better on harder instances by preserving the model's successful grounding while correcting only verifiable reasoning failures. In addition, step-level ILP verification yields a fine-grained error taxonomy that provides diagnostic insight beyond final-answer accuracy.
Chinese Translation
链式思维(CoT)提示使大型语言模型(LLMs)能够处理多步骤推理任务,但生成的中间步骤并不一定在逻辑上是可靠的。我们提出了推理波普(Reason Popper-ly),一个神经符号框架,利用归纳逻辑编程(ILP)从推理轨迹中学习关系组合规则,并将其作为在线验证器进行步骤级纠正。给定一个LLM生成的轨迹,该方法检查每个推断步骤是否符合学习到的规则表,诊断违规类型,用符号推导的修正重写不正确的步骤,并重新生成其余后缀,以便模型能够基于经过验证的轨迹生成最终答案。我们在CLUTRR这一多跳亲属推理基准上进行了评估,使用五种语言模型处理2到10跳的推理链。在所有模型中,推理波普在终端准确性上始终优于标准的链式思维,对于小型模型在最长链上提高了多达48个百分点,对于前沿模型则提高了15个百分点。与完全外生的符号管道相比,我们的方法在更困难的实例上表现更佳,能够保留模型成功的基础,同时仅纠正可验证的推理失败。此外,步骤级的ILP验证产生了细粒度的错误分类法,提供了超越最终答案准确性的诊断见解。
cs.AI / 101 / 2607.23045

Stress-testing large language model agents in a robotic chemistry laboratory

在机器人化学实验室中对大型语言模型代理进行压力测试
Guo, Lulu, Sun, Yingkai, Li, Xiaobo, Ge, Luyao, Wang, Ziming, Zheng, Haitao, Li, Jingyu, Zhang, Huijuan, Chen, Bingxu, Liu, Daobin, Liu, Yuebo, Li, Jie, Li, Xiaohui, Chen, Linjiang, Luo, Yi, Jiang, Jun
Abstract
AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency measurable. Its 45 modular workstations exposed as machine-readable skills enabled 4,608 trials. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints; even the best system achieved 28.1%. Long-horizon planning remained a challenge: only three executable workflows exceeded 30 operations, although the longest contained 44. Across five rounds, experimental feedback prompted local adjustments but no workflow-level replanning or analytical-method redesign. By making physical executability and evidence-driven replanning measurable, our study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements towards physically grounded autonomous research.
Chinese Translation
人工智能通过知识、推理和计划生成进行评估,但科学代理需要可靠的物理行动和对证据的适应。在这里,我们使用机器人化学实验室作为物理世界的测试平台,使科学代理的可测量性得以实现。其45个模块化工作站以机器可读的技能形式呈现,支持了4,608次实验。仅有3.3%的实验在实验室约束下产生了专家评估的可执行工作流程;即使是最佳系统也仅达到了28.1%。长期规划仍然是一个挑战:只有三个可执行工作流程超过了30个操作,尽管最长的工作流程包含44个操作。在五轮实验中,实验反馈促使了局部调整,但没有进行工作流程级别的重新规划或分析方法的重新设计。通过使物理可执行性和基于证据的重新规划可测量,我们的研究提供了对部署准备情况的基于证据的评估,以及一个诊断框架,以指导向物理基础的自主研究的闭环改进。
cs.AI / 102 / 2607.23055

SymStep: Symbolic Step Verification for Logical Reasoning

SymStep:逻辑推理的符号步骤验证
Usmanova, Aida, Gao, Rui, Azizov, Dilshod, Usbeck, Ricardo, Iklassov, Zangir
Abstract
Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps. We introduce SymStep: an LLM makes one atomic claim at a time (DEDUCE: Alice, pet, Cat), then a lightweight constraint propagator checks the claim for consistency with prior accepted deductions, rejects contradictions, and cascades implied facts automatically. SymStep+G additionally provides MRV guidance after each accepted step, directing the LLM toward the most constrained unresolved variable. On a 35-puzzle retained subset of ZebraLogicBench, a benchmark of 1,000 Einstein-style logic puzzles, Direct and CoT both achieve 0%, while SymStep+G reaches 97%. On AR-LSAT analytical reasoning problems, SymStep achieves 100% vs. CoT's 87%. On LGP-14, SymStep+G achieves 100% vs. 0% for CoT and Logic-LM, the strongest prior symbolic+LLM baseline we compare against. Ablation studies reveal that MRV guidance is a key mechanism for reducing directionless cycling, while consistency checking provides a safety net against explicit contradictions. Across six benchmarks spanning five task domains, SymStep variants match or exceed every baseline on constraint-dense and arithmetic tasks. Experiments on AQUA-RAT algebra confirm the advantage is constraint-density-specific.
Chinese Translation
链式思维(CoT)提示在约束密集的逻辑推理任务中可能严重失效,因为未经过验证的错误在步骤之间静默累积。我们提出了SymStep:一个大型语言模型(LLM)一次提出一个原子性主张(例如:推导:爱丽丝,宠物,猫),然后一个轻量级约束传播器检查该主张与先前接受的推导的一致性,拒绝矛盾,并自动传播隐含事实。SymStep+G在每个接受的步骤后提供最小剩余变量(MRV)指导,引导LLM朝向约束最强的未解决变量。在ZebraLogicBench的35个难题保留子集上(这是一个包含1000个爱因斯坦风格逻辑难题的基准),直接推理和CoT均达到0%的成功率,而SymStep+G则达到了97%。在AR-LSAT分析推理问题上,SymStep的成功率为100%,而CoT为87%。在LGP-14上,SymStep+G的成功率为100%,而CoT和Logic-LM(我们比较的最强符号+LLM基线)均为0%。消融研究表明,MRV指导是减少无方向循环的关键机制,而一致性检查则为显式矛盾提供了安全保障。在涵盖五个任务领域的六个基准测试中,SymStep变体在约束密集和算术任务上与每个基线相匹配或超越。对AQUA-RAT代数的实验确认了这一优势是特定于约束密度的。
cs.AI / 103 / 2607.23077

Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

结构优先:用于多实体推理的单块时空变换器
Sivalingam, Narthana, Sivasthigan, Santhirarajah, Wijenayake, Buddhi, Godaliyadda, Roshan, Herath, Vijitha, Ekanayake, Parakrama
Abstract
Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approaches perform well but often rely on deep stacks of layers to learn these heterogeneous dependencies implicitly, increasing computational cost. We revisit this problem from a structural perspective and decompose multi-entity temporal dynamics into three interaction types: spatial interactions among entities, temporal interactions across time, and cross interactions coupling the two domains. We propose a structured spatio-temporal transformer block that explicitly models all three within a single stage. It uses parallel spatial and temporal self-attention, followed by bidirectional cross-attention, and combines the outputs through learnable gated fusion. By directly encoding these complementary views, the model reduces the need for deep stacking. We evaluate the approach on video-based group activity recognition, skeleton-based human interaction analysis, and wearable sensor-based activity recognition. Despite its simplicity, the single structured Transformer block matches or outperforms deeper architectures with only 1.76M parameters. The results suggest that depth in prior models partly compensates for implicit and entangled interaction modeling, whereas explicit factorization offers a more efficient and transparent alternative. More broadly, this work supports a structure-first design principle: expressive multi-entity temporal reasoning can emerge by exposing interaction structure rather than relying on depth.
Chinese Translation
建模多实体时间数据需要捕捉实体、时间及其交互之间的依赖关系。基于变换器的方法表现良好,但通常依赖于深层堆叠的层来隐式学习这些异质依赖,从而增加了计算成本。我们从结构的角度重新审视这个问题,将多实体时间动态分解为三种交互类型:实体之间的空间交互、时间上的时间交互,以及耦合这两个领域的交叉交互。我们提出了一种结构化的时空变换器块,在单个阶段内显式建模这三种交互。它使用并行的空间和时间自注意力,随后是双向交叉注意力,并通过可学习的门控融合组合输出。通过直接编码这些互补视角,该模型减少了对深层堆叠的需求。我们在基于视频的群体活动识别、基于骨架的人类交互分析和基于可穿戴传感器的活动识别上评估了该方法。尽管其结构简单,单个结构化的变换器块在仅有176万参数的情况下与更深的架构相匹配或超越。结果表明,先前模型中的深度在一定程度上弥补了隐式和交织的交互建模,而显式分解提供了一种更高效和透明的替代方案。更广泛地说,这项工作支持结构优先的设计原则:通过揭示交互结构而不是依赖深度,可以实现富有表现力的多实体时间推理。
cs.AI / 104 / 2607.23089

Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization

基于编译器的层次化诊断用于LLM驱动的Triton内核优化
Chen, Dongjie, Zhao, Ping, Zhan, Bohua, Wang, Yulong, Chen, Shushu, Feng, Liangjun, Zhou, Hao, Shen, Min, Wang, Linmu, Sheng, Weijia, Wei, Xiangyu, Ding, Weijie, Huang, Jianhui, Gao, Yaoqing
Abstract
Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rely on surface signals such as compilation feedback and profiling metrics. These signals reveal that a kernel is slow, but not why the backend compiler fails to realize a profitable optimization, especially on emerging accelerators such as NPUs. We therefore formulate kernel optimization as a progressive cross-layer diagnosis problem that links runtime symptoms to IR structure and compiler behavior before rewriting source. Based on this insight, we present our system, a compiler-grounded and hierarchical optimization framework for Triton kernels. the system escalates from lightweight pattern triage and profiling diagnosis to IR attribution and compiler-grounded analysis only when deeper evidence is needed, then proposes evidence-backed source-level rewrites. We implement the system on Triton for Ascend NPUs and evaluate it on 37 successfully converted entries from a standardized NPUKernelBench-derived Ascend 950 benchmark. Across these entries, the system attains a geometric-mean speedup of 4.35$\times$ and a median speedup of 2.73$\times$ from the initial to optimized Triton kernel; 22/37 exceed 2$\times$ and 13/37 exceed 5$\times$. The complete distribution ranges from near-baseline entries to large wins, motivating transparent reporting of the current system's scope and limitations.
Chinese Translation
近年来,大型语言模型(LLMs)的进展使得自动化内核生成和优化成为可能,但大多数现有方法依赖于编译反馈和性能分析指标等表面信号。这些信号揭示了内核的运行速度较慢,但并未说明后端编译器为何未能实现有利的优化,尤其是在新兴的加速器(如神经处理单元,NPUs)上。因此,我们将内核优化形式化为一个渐进的跨层诊断问题,将运行时症状与中间表示(IR)结构和编译器行为联系起来,然后再重写源代码。基于这一见解,我们提出了我们的系统,一个基于编译器的层次化优化框架,用于Triton内核。该系统在需要更深入证据时,从轻量级模式分类和性能分析诊断逐步升级到IR归因和基于编译器的分析,然后提出基于证据的源代码重写。我们在Ascend NPUs上实现了该系统,并在37个成功转换的条目上进行了评估,这些条目来自标准化的NPUKernelBench衍生的Ascend 950基准测试。在这些条目中,该系统从初始到优化的Triton内核实现了几何平均加速比4.35×,中位数加速比为2.73×;22/37的条目超过了2×,13/37的条目超过了5×。完整的分布范围从接近基线的条目到显著的胜利,促使对当前系统的范围和局限性进行透明的报告。
cs.AI / 105 / 2607.23123

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

SQBench:评估语言模型代理在生产导向工作流程中任务交付的基准
Sun, Summer
Abstract
Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation. We introduce SQBench, a benchmark for evaluating production-oriented task delivery by language-model agents. SQBench v1.0 contains 220 standardized tasks organized into L1 atomic capabilities, L2 composite skills, and L3 business scenarios. Each task requires an agent to process input assets, use available tools, and produce an explicitly specified deliverable. The evaluation first computes functional Completion and then derives Risk Penalty and Performance from independently evidenced triggers in a 10D Risk Matrix. A Strict Pass requires Completion = 1 and Risk Penalty = 0. We evaluate 27 model configurations under a common protocol, with one run per configuration-task pair. The highest prespecified Weighted Pass@1 is 60.5%. Mean Strict Pass@1 on L3 is 18.5%, and every configuration performs worse on L3 than on both L1 and L2, indicating that delivery under domain constraints is a shared weakness within the current task set. Of 2,348 results with Completion = 1, 113 (4.8%) fail the Strict Pass criterion because of risks such as unverifiable citations, inappropriate resource use, or format violations. These results show that functional completion alone does not fully characterize delivery quality and that risk determinations should be reported separately.
Chinese Translation
现有的大型语言模型评估涵盖了知识、推理、编码和工具使用,但很少将在受限工作流程中产生的可验证交付物视为评估单位。我们提出了SQBench,这是一个用于评估语言模型代理在生产导向任务交付中的基准。SQBench v1.0包含220个标准化任务,分为L1原子能力、L2复合技能和L3商业场景。每个任务要求代理处理输入资产,使用可用工具,并生成明确指定的交付物。评估首先计算功能完成度(Completion),然后从10维风险矩阵中的独立证据触发器推导风险惩罚(Risk Penalty)和性能(Performance)。严格通过(Strict Pass)要求完成度=1且风险惩罚=0。我们在一个共同协议下评估了27种模型配置,每种配置-任务对进行一次运行。预设的最高加权通过率(Weighted Pass@1)为60.5%。L3的平均严格通过率为18.5%,且每种配置在L3的表现均低于L1和L2,这表明在领域约束下的交付是当前任务集的共同弱点。在2348个完成度=1的结果中,有113个(4.8%)未能通过严格通过标准,原因包括不可验证的引用、不当的资源使用或格式违规。这些结果表明,仅凭功能完成度无法充分表征交付质量,风险评估应单独报告。
cs.AI / 106 / 2607.23124

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

AgentOmnia:扩展代理模型以适应全场景应用
Jiang, Hao, Xin, Gangtao, Huang, Yingdi, Zhu, Guojie, Zhang, Jiangshan, Lin, Xinyuan, Xu, Yunkun, Shen, Chengyu, Fei, Wenlong, Li, Jiawei, Fu, Yujie, Kang, Sichen, Xie, Tingyu, Hu, Yedi, Zhang, Jingren, Gao, Hongcheng, Zeng, Jianshu, Chen, Chong, Guo, Chang, Feng, Chao, Wang, Feng, Lin, Fulin, Ma, Jinchao, Mei, Lang, Huang, Li, Liu, Liyan, He, Qing, Tao, Shuting, Mo, Siyu, Chen, Xiangnan, Yu, Xiaohan, Li, Xiaoyang, Hou, Yanheng, Wu, Yanyu, Yang, Zhihan, Zhang, Wentao, Gao, Yang, Cao, Zhao
Abstract
Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present AgentOmnia, a framework coordinating task-space definition, data synthesis, post-training, evaluation, and improvement across To-Consumer (ToC), To-Business (ToB), and To-Employee (ToE) applications. An extensible Domain x Capability x Atomic Difficulty taxonomy aligns these stages and enables fine-grained diagnosis with OmniaBench. AgentOmnia combines bidirectional environment-task synthesis with tool-dependency, program-structured, and solver-based pipelines, constructing 5,018 stateful environments with 255,375 tools and 52,361 tasks. Programs, solvers, and verifiers provide correctness signals, while supervised fine-tuning, online agentic reinforcement learning, and a rollback curriculum support post-training. Evaluation failures translate into Product Requirement Documents (PRDs) for targeted self-evolution. Starting from Qwen3-30B-A3B-Thinking-2507, AgentOmnia raises the pass rate on the OmniaBench challenging subset from 9.16% to 37.11% and the macro-average across OmniaBench, $\tau^2$-Bench, DeepPlanning, and VitaBench from 22.86% to 41.69%. Under a unified protocol,it leads the evaluated agentic post-trained baselines on OmniaBench and retains the highest four-benchmark macro-average. It also surpasses Qwen3-235B-A22B-Thinking-2507 on all four benchmarks and exceeds Qwen3.5-35B-A3B on the macro-average. Gains span three application splits, ten capability dimensions, eight atomic-difficulty factors, and 76 of 90 level-1 domains, indicating broad rather than category-specific improvement. A one-round study provides initial evidence for PRD-guided self-evolution, motivating validation at larger scales and in industrial settings.
Chinese Translation
大型语言模型代理的快速发展仍然在不同领域、能力、任务难度和交互设置之间存在碎片化的进展。我们将其框架化为全场景代理扩展,并提出AgentOmnia,一个协调任务空间定义、数据合成、后训练、评估和改进的框架,涵盖面向消费者(To-Consumer, ToC)、面向企业(To-Business, ToB)和面向员工(To-Employee, ToE)应用。一个可扩展的领域×能力×原子难度分类法对齐这些阶段,并通过OmniaBench实现细粒度诊断。AgentOmnia结合了双向环境-任务合成与工具依赖、程序结构化和求解器基础的管道,构建了5,018个有状态环境,配备255,375个工具和52,361个任务。程序、求解器和验证器提供正确性信号,而监督微调、在线代理强化学习和回滚课程支持后训练。评估失败转化为产品需求文档(Product Requirement Documents, PRDs),以实现针对性的自我进化。从Qwen3-30B-A3B-Thinking-2507开始,AgentOmnia将OmniaBench挑战子集的通过率从9.16%提高到37.11%,并将OmniaBench、$ au^2$-Bench、DeepPlanning和VitaBench的宏平均从22.86%提高到41.69%。在统一协议下,它在OmniaBench上领先于评估的代理后训练基线,并保持最高的四基准宏平均。它还在所有四个基准上超越了Qwen3-235B-A22B-Thinking-2507,并在宏平均上超过了Qwen3.5-35B-A3B。增益覆盖三个应用拆分、十个能力维度、八个原子难度因素和90个一级领域中的76个,表明改进是广泛的,而非特定类别的。一轮研究提供了PRD指导的自我进化的初步证据,激励在更大规模和工业环境中进行验证。
cs.AI / 107 / 2607.23159

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

CachedSearch:无训练的缓存探索用于视频扩散的测试时搜索
Saini, Shreshth, Birkbeck, Neil, Wang, Yilin, Adsumilli, Balu, Bovik, Alan C.
Abstract
Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discarded. Training-free caching makes each rollout 2-3x faster at near-lossless quality. Composition is safe only if lossy caching preserves verifier rankings. We present the first study of whether caching corrupts candidate ranking in video test-time search. On Wan2.1-T2V-1.3B with an adaptive caching wrapper (~2x per-candidate speedup), ImageReward scores seed-matched cached and full rollouts. Median per-prompt Spearman rank correlation is 0.905, with 72% top-1 agreement on the VBench suite. VBench-2.0 replicates this result on a harder suite. Recomputing the cached winner at full compute retains 90-94% of the full-search gain. Errors cluster among near-tied candidates, making corruption self-limiting. This finding leads to CachedSearch. It explores every candidate with aggressive caching, then re-generates only the winner at full compute. At N=8, it captures 94.7% of best-of-N's gain at 63% of the cost. Capture rises with width. At matched budget, it searches twice as wide for 38% more gain. The result holds from 1.3B-14B across six models and four families: Wan, LTX, CogVideoX, and Hunyuan. Wan2.1-14B matches the 1.3B model's fidelity. Mid-trajectory pruning multiplies the exploration saving to 3.11x at 88.6% capture. Ports to other model families require recalibrating a single parameter, showing that fidelity tracks architecture rather than parameter count. CachedSearch is training-free, verifier-agnostic, and orthogonal to the search algorithm, making it a plug-in multiplier for test-time scaling.
Chinese Translation
测试时搜索使得小型视频扩散模型能够与大型模型相媲美,但成本却高出2-10倍。所有候选项都经过完全去噪,尽管大多数会被丢弃。无训练的缓存使得每次回滚的速度提高了2-3倍,同时保持近乎无损的质量。只有在有损缓存保留验证者排名的情况下,组合才是安全的。我们首次研究了缓存是否会破坏视频测试时搜索中的候选排名。在带有自适应缓存包装器的Wan2.1-T2V-1.3B上(每个候选项的速度提升约为2倍),ImageReward对比了种子匹配的缓存回滚和完整回滚。每个提示的中位Spearman等级相关性为0.905,在VBench套件上有72%的前1一致性。VBench-2.0在更难的套件上复制了这一结果。在完全计算下重新计算缓存赢家保留了90-94%的完整搜索增益。错误集中在接近平局的候选项之间,使得破坏自我限制。这个发现促成了CachedSearch的提出。它通过激进的缓存探索每个候选项,然后仅在完全计算下重新生成赢家。在N=8时,它以63%的成本捕获了94.7%的最佳N增益。随着宽度的增加,捕获率提高。在匹配预算下,它以38%的增益搜索了两倍的宽度。这个结果在1.3B-14B的六个模型和四个家族中保持一致:Wan、LTX、CogVideoX和Hunyuan。Wan2.1-14B的保真度与1.3B模型相匹配。中途修剪将探索节省倍增至3.11倍,捕获率为88.6%。移植到其他模型家族只需重新校准一个参数,显示出保真度跟踪架构而非参数数量。CachedSearch是无训练的、与验证者无关的,并且与搜索算法正交,使其成为测试时扩展的插件倍增器。
cs.AI / 108 / 2607.23219

An Ontology for Machine Learning Interatomic Potentials

机器学习原子间势的本体论
Hernández, Daniel, Jung, Jong Hyun, Ikeda, Yuji, Ou, Yongliang, Kumar, Pranav, Schächtel, Tom, Liu, Wenchuan, Li, Xin, Zhang, Xi, Xu, Xiang, Zhu, Lifang, Körmann, Fritz, Staab, Steffen, Grabowski, Blazej
Abstract
Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density functional theory (DFT) or wave-function methods---at a fraction of the cost. The field encompasses a growing ecosystem of algorithms, training datasets, hyperparameters, and target materials, yet the metadata needed to systematically compare, reproduce, and build upon MLIP studies remains scattered across papers, scripts, and ad-hoc file formats. We present the MLIPs ontology, an OWL 2 DL ontology that captures the concepts needed to describe MLIP methods, their hyperparameters, training datasets with DFT provenance, and published benchmarks. The ontology is organized into three modules---Method, Training Data, and Benchmark---and connects existing ontologies in materials science (MDO, CMSO/ASMO) and machine learning (ML-Schema), complementing dataset-side schemas such as Croissant. It declares 27 formal axioms enforcing data completeness and consistency, including property chains that link trained models to their methods and training data. We demonstrate the ontology through a running example based on Moment Tensor Potentials and evaluate it through competency-question execution on a 20-paper seeded knowledge graph, OWL reasoning, and comparison with existing ontologies.
Chinese Translation
机器学习原子间势(MLIPs)以较低的成本近似量子力学能量和力——这些通常通过密度泛函理论(DFT)或波函数方法计算。该领域涵盖了一个日益增长的算法、训练数据集、超参数和目标材料的生态系统,但用于系统比较、重现和基于MLIP研究进行扩展所需的元数据仍然分散在论文、脚本和临时文件格式中。我们提出了MLIPs本体论,这是一个OWL 2 DL本体,捕捉了描述MLIP方法、其超参数、具有DFT来源的训练数据集和已发布基准所需的概念。该本体分为三个模块——方法、训练数据和基准——并连接了材料科学(MDO、CMSO/ASMO)和机器学习(ML-Schema)中的现有本体,补充了如Croissant等数据集侧的模式。它声明了27个正式公理,以强制数据的完整性和一致性,包括将训练模型与其方法和训练数据链接的属性链。我们通过基于矩阵张量势的运行示例展示了该本体,并通过在一个包含20篇论文的知识图谱上执行能力问题、OWL推理以及与现有本体的比较来评估它。
cs.AI / 109 / 2607.23243

Characterisation of Density-based FM generation methods in the context of Information Fusion

信息融合背景下基于密度的模糊测度生成方法的表征
Huang, Yanhao, Wagner, Christian
Abstract
Fuzzy Integral (FI) based aggregation provides a powerful mechanism for nuanced aggregation, for example, in ensemble approaches or decision-level fusion more generally. The main challenge of this approach is the appropriate parametrization of the Fuzzy Measure (FM), which captures the worths of the individual components--and their combinations--which are being fused. Here, widely used approaches including the Sugeno-$\lambda$ and Decomposable FMs, parametrize the FM by extrapolating from the densities, i.e. the weights associated with individual sources, while respecting the FM's monotonicity constraint. This paper articulates that this information is, in general, insufficient to uniquely identify a discrete FM; but shows how an interval-valued FM can indeed be determined uniquely. We proceed to show how the incorporation of additional information beyond the above, such as the choice of a specific FI and a dataset, then allows for obtaining even more specific interval-valued FMs. In practice, establishing the quality of an empirically determined FM is not trivial. To help address this, we show how the likelihood with which a resulting interval FM encompasses the `ideal', i.e. the commonly intangible, best, or ground-truth numeric FM, can be determined, producing a confidence interval at a given confidence level. Finally, based on a series of experiments, we demonstrate empirically that the Choquet FI output based on this FM can also be regarded as the confidence interval for the `ideal' information fusion result, providing a novel means to characterize FI fusion outcomes a priori and charting a pathway for future research.
Chinese Translation
基于模糊积分(Fuzzy Integral, FI)的聚合提供了一种强大的机制,用于细致的聚合,例如在集成方法或更一般的决策级融合中。这种方法的主要挑战在于模糊测度(Fuzzy Measure, FM)的适当参数化,该测度捕捉了被融合的各个组成部分及其组合的价值。在这里,广泛使用的方法包括Sugeno-$ ext{λ}$和可分解模糊测度,通过从密度(即与各个源相关的权重)外推来参数化FM,同时遵循FM的单调性约束。本文阐明,这些信息通常不足以唯一识别一个离散的FM;但展示了如何唯一确定一个区间值FM。我们进一步展示,超出上述信息的额外信息的纳入,例如特定FI的选择和数据集的选择,可以使得获得更具体的区间值FM。在实践中,确立经验确定的FM的质量并非易事。为了解决这个问题,我们展示了如何确定一个结果区间FM包含“理想”值的可能性,即通常难以捉摸的最佳或真实数值FM,从而在给定的置信水平下生成置信区间。最后,基于一系列实验,我们实证证明,基于该FM的Choquet FI输出也可以视为“理想”信息融合结果的置信区间,为事先表征FI融合结果提供了一种新方法,并为未来的研究指明了方向。
cs.AI / 110 / 2607.23258

CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics

CAPT:一种多任务连续自回归变换器,能够实现跨数据集和跨物种的钙离子群体动力学迁移
Xu, Xinhong, Zhang, Yimeng, Zhang, Yuanlong
Abstract
Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbf{whether a model pretrained on one collection of recordings can generalize to new datasets, experimental paradigms, and even species.} Existing approaches are often designed for specific tasks and evaluated on a single dataset, making it unclear whether their learned representations are reusable for new calcium trace datasets. To tackle this gap, we present \textbf{CAPT}, a \textbf{C}ontinuous \textbf{A}utoregressive \textbf{P}opulation \textbf{T}ransformer for calcium population dynamics. CAPT models continuous calcium traces directly through a continuous patch tokenization strategy and is trained autoregressively, enabling end-to-end pretraining and adaptation to diverse downstream tasks. We first pretrain CAPT on a large-scale mouse calcium imaging dataset and evaluate its transferability across independent mouse, larval zebrafish, and \textit{C. elegans} datasets collected by different laboratories. In these transfer settings, the pretrained backbone is frozen and only adaptation modules are updated. Across neural population forecasting and behavior decoding tasks, CAPT consistently outperforms specialized and general-purpose baselines. Alongside predictive performance, multimodal analyses using NeuroPAL annotations in \textit{C. elegans} datasets show that CAPT embeddings form a shared functional space across datasets and capture anatomical cell-identity-related structure. These results suggest that the continuous autoregressive modeling opens up possibilities for a simple route towards general-purpose neural foundation models for calcium imaging, which can generalize across datasets, experimental paradigms, and species.
Chinese Translation
大规模钙成像为构建神经群体动力学的基础模型提供了机会,但一个核心问题仍未解决: extbf{一个在一组记录上预训练的模型是否能够推广到新的数据集、实验范式,甚至不同物种。} 现有的方法通常针对特定任务设计,并在单一数据集上进行评估,这使得其学习到的表示是否可用于新的钙迹数据集变得不明确。为了解决这一问题,我们提出了 extbf{CAPT},一种用于钙离子群体动力学的 extbf{C}ontinuous extbf{A}utoregressive extbf{P}opulation extbf{T}ransformer。CAPT通过连续补丁标记策略直接建模连续的钙迹,并以自回归方式进行训练,从而实现端到端的预训练和适应多样的下游任务。我们首先在一个大规模的小鼠钙成像数据集上预训练CAPT,并评估其在不同实验室收集的独立小鼠、幼虫斑马鱼和 extit{C. elegans}数据集上的迁移能力。在这些迁移设置中,预训练的主干被冻结,仅更新适应模块。在神经群体预测和行为解码任务中,CAPT始终优于专业和通用基线。除了预测性能外,使用 extit{C. elegans}数据集中的NeuroPAL注释进行的多模态分析表明,CAPT嵌入形成了跨数据集共享的功能空间,并捕捉到与解剖细胞身份相关的结构。这些结果表明,连续自回归建模为实现通用钙成像神经基础模型提供了一条简单的途径,这些模型能够在不同数据集、实验范式和物种之间进行推广。
cs.AI / 111 / 2607.23263

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

SeekJudge:计算机使用代理强化学习的实用奖励框架
Wan, Yang, Zhang, Zhenhao, Wang, Jierui, Zhu, Linchao
Abstract
Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with reinforcement learning. This judgment has long relied on rule-based evaluation, which struggles to align with human intention and goes stale when an app updates or its online content drifts. Existing model-based judges attempt to address these problems but still leave a performance gap to the rule-based evaluation. We propose the \textbf{SeekJudge} framework, in which four role-specialized agents, a Condense, a Ground, a Seek and an Analyze agent, reach a verdict through a Seek--Analyze loop over the trajectory. A seed-calibrated distillation pipeline trains one specialized $9$B model to serve as the shared backbone for all four agents. Measured by downstream success rate on held-out RL test goals, SeekJudge is the first practical model-based reward to match or surpass native rule-based supervision in online RL. Beyond accuracy, SeekJudge provides step-level judgments, runs far cheaper than a closed-source large model, and keeps a small per-call context that scales to much longer trajectories. We further contribute a general architectural improvement to the reward server that speeds up judging in RL. Together these make model-based reward a practical drop-in for rule-based supervision in CUA reinforcement learning.
Chinese Translation
判断一个轨迹是否真正满足其指令决定了我们如何在长时间跨度的图形用户界面任务中评估计算机使用代理,以及如何通过强化学习训练它们。这一判断长期以来依赖于基于规则的评估,但这种评估难以与人类意图对齐,并且在应用程序更新或其在线内容漂移时会变得过时。现有的基于模型的评判者试图解决这些问题,但仍然存在与基于规则的评估之间的性能差距。我们提出了 extbf{SeekJudge}框架,其中四个角色专用的代理——凝聚代理(Condense)、基础代理(Ground)、搜索代理(Seek)和分析代理(Analyze)——通过对轨迹的搜索-分析循环达成裁决。一个种子校准的蒸馏管道训练了一个专用的9B模型,作为所有四个代理的共享骨干。通过在保留的强化学习测试目标上的下游成功率衡量,SeekJudge是第一个在在线强化学习中与本地基于规则的监督相匹配或超越的实用模型基础奖励。除了准确性,SeekJudge还提供逐步判断,运行成本远低于闭源大型模型,并保持小的每次调用上下文,能够扩展到更长的轨迹。我们进一步贡献了对奖励服务器的一般架构改进,加快了强化学习中的判断速度。这些共同使得基于模型的奖励成为计算机使用代理强化学习中基于规则监督的实用替代方案。
cs.AI / 112 / 2607.23286

TopoFE: topology-aware LLM-guided Automated Feature Engineering

TopoFE:基于拓扑的 LLM 引导自动特征工程
Li, Sha, Ramakrishnan, Naren
Abstract
Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective is to discover predictive feature transformations from an exponentially large search space. Recent advances in large language models (LLMs) have expanded the expressiveness of AutoFE by enabling feature program generation beyond predefined operator libraries. However, existing LLM-based approaches remain fundamentally limited by stateless generation and homogeneous search: feature proposals are produced from static prompts without accumulating search experience, while single-population exploration quickly converges to dominant transformation patterns and rarely discovers complementary feature compositions across transformation families. We propose TOPOFE, a topology-aware multi-island evolutionary framework for LLM-guided feature engineering. TOPOFE combines family-specialized exploration, adaptive prompt memory, and topology-guided knowledge transfer to efficiently discover diverse and compositional feature programs. Experiments on 29 public tabular datasets demonstrate consistent improvements over state-of-the-art AutoFE methods across classification and regression tasks. Beyond predictive performance, TOPOFE discovers more diverse and transferable feature programs that generalize across multiple downstream predictors and LLM backbones.
Chinese Translation
表格学习的自动特征工程(AutoFE)可以自然地被表述为一个程序合成问题,其目标是在一个指数级大的搜索空间中发现预测特征变换。最近在大型语言模型(LLMs)方面的进展,通过使特征程序生成超越预定义操作符库,扩展了 AutoFE 的表达能力。然而,现有的基于 LLM 的方法在本质上仍然受到无状态生成和同质搜索的限制:特征提案是从静态提示中产生的,未能积累搜索经验,而单一种群的探索迅速收敛于主导变换模式,且很少发现跨变换家族的互补特征组合。我们提出了 TOPOFE,一种基于拓扑的多岛进化框架,用于 LLM 引导的特征工程。TOPOFE 结合了家族专门化探索、自适应提示记忆和拓扑引导的知识转移,以高效发现多样化和组合性的特征程序。在 29 个公共表格数据集上的实验表明,在分类和回归任务中,TOPOFE 一直优于最先进的 AutoFE 方法。除了预测性能外,TOPOFE 还发现了更具多样性和可转移性的特征程序,这些程序能够在多个下游预测器和 LLM 骨干网络中进行泛化。
cs.AI / 113 / 2607.23290

RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning

RareLens:通过对齐不同的大型语言模型推理实现端到端的罕见疾病护理
Chen, Xi, Zhou, Hongru, Feng, Shiyu, Zhou, Hanyu, Yi, Huahui, Wang, Rongsheng, He, Tiancheng, Wang, Kun, Liu, Pingping, Li, Qiankun, Lin, Sicheng, Ou, Huiying, Zheng, Xiaohong, Zang, Tianying, Wu, Zhuohang, Jiang, Leheng, Cao, Kexin, Zhang, Wenhan, Li, ChengYi, Wang, Zhiyang, Li, Songlin, Wang, Benyou, Yin, Ningbei, Zhang, Shaoting, Fu, Weili, Li, Jian, Li, Kang
Abstract
Rare diseases collectively affect an estimated 3.5% to 5.9% of the population, yet more than 70% of patients are misdiagnosed and many endure years of evaluation before a diagnosis is reached, because early presentations are nonspecific and relevant expertise is scarce and unevenly distributed. Artificial intelligence could provide support, but existing systems address isolated stages of care, overwhelmingly diagnosis. They typically depend on the results of downstream investigations, and they treat the variability between models as noise to be eliminated. Here we present RareLens, a system that supports clinical decision-making across the entire rare disease trajectory by exploiting this variability. When heterogeneous large language models evaluate the same case, they generate divergent but complementary reasoning, which RareLens aligns and calibrates into a single convergent, actionable decision at each stage. Four coordinated modules perform primary-visit risk screening, diagnosis, treatment planning and prognosis. Developed and evaluated on RareBench, a real-world dataset of 157,525 cases spanning all 33 Orphanet categories and more than 7,000 conditions, RareLens outperformed every frontier model tested, including GPT-5, DeepSeek-R1, Claude-3.7-Sonnet and Gemini-2.5-Pro, at each stage. It achieved an area under the curve of 0.917 for screening and top-1 accuracies of 65.5% and 89.8% for diagnosis and treatment. In an external study spanning 1,287 cases and 23 physicians, autonomous RareLens and physicians assisted by RareLens both substantially outperformed unaided physicians. These findings indicate that aligning divergent model reasoning, rather than scaling a single model, offers a generalizable strategy for high-uncertainty clinical decision-making.
Chinese Translation
罕见疾病总共影响估计3.5%至5.9%的人口,但超过70%的患者被误诊,许多人在确诊之前经历多年评估,因为早期表现不具特异性,相关专业知识稀缺且分布不均。人工智能可以提供支持,但现有系统仅针对护理的孤立阶段,主要集中在诊断上。它们通常依赖于下游检查的结果,并将模型之间的变异性视为需要消除的噪声。在此,我们提出了RareLens,一个通过利用这种变异性来支持整个罕见疾病轨迹的临床决策系统。当异质的大型语言模型评估同一病例时,它们生成不同但互补的推理,RareLens将这些推理对齐并校准为每个阶段的单一收敛、可操作的决策。四个协调模块执行初诊风险筛查、诊断、治疗计划和预后。在RareBench上开发和评估,该数据集涵盖157,525个病例,涉及所有33个Orphanet类别和7000多种疾病,RareLens在每个阶段的表现超过了所有测试的前沿模型,包括GPT-5、DeepSeek-R1、Claude-3.7-Sonnet和Gemini-2.5-Pro。它在筛查中的曲线下面积达到0.917,诊断和治疗的前1准确率分别为65.5%和89.8%。在一项涵盖1,287个病例和23名医生的外部研究中,独立使用RareLens的医生和由RareLens辅助的医生均显著优于未辅助的医生。这些发现表明,对齐不同模型推理,而不是扩展单一模型,提供了一种适用于高不确定性临床决策的可推广策略。
cs.AI / 114 / 2607.23317

Ordered Network Analysis of Epistemic Emotions during Collaborative Problem Solving

协作问题解决过程中的认知情感有序网络分析
Anindho, Sifatul, Venkatesha, Videep, Ocumpaugh, Jaclyn, Blanchard, Nathaniel
Abstract
Investigating how affective states such as confusion and frustration persist and transition during co-situated collaborative problem solving (CPS) is important for understanding the dynamics of epistemic emotions. However, the accurate identification of affective states remain challenging as there is no gold-standard truth in this space. Here, we analyze affective states collected through retrospective cued-recall during an in-person CPS task. Using ordered network analysis (ONA), we examine (1) the overall ordered structure of affective states and how this structure differs across self-caught and probe-caught reporting methods, and (2) what aspects of this ordered structure are emphasized differently in slower and faster groups. We find that ONA reveals differences in persistence and transition patterns that are not apparent from descriptive summaries alone. In particular, we observe a stable epistemic core linking curiosity, optimism, and confusion, with different reporting methods emphasizing different connections among states. An analysis between faster and slower groups show that roles of confusion and disengagement also shift significantly during collaboration, particularly in their relationship to conflict. We interpret our findings in the context of collaboration and discuss their implications in developing AI systems that support CPS.
Chinese Translation
研究混乱和挫败等情感状态在共同情境下的协作问题解决(CPS)过程中如何持续和转变,对于理解认知情感的动态至关重要。然而,由于在这一领域没有公认的标准真相,准确识别情感状态仍然具有挑战性。在此,我们分析了通过回顾性提示回忆在面对面的CPS任务中收集的情感状态。利用有序网络分析(ONA),我们考察了(1)情感状态的整体有序结构,以及这一结构在自我捕捉和探测捕捉报告方法之间的差异;(2)在较慢和较快的组别中,这一有序结构的不同方面被强调的方式。我们发现,ONA揭示了持久性和转变模式的差异,这些差异在仅依赖描述性总结时并不明显。特别是,我们观察到一个稳定的认知核心将好奇心、乐观和困惑联系在一起,不同的报告方法强调了状态之间不同的联系。对较快和较慢组别的分析显示,困惑和脱离的角色在协作过程中也显著变化,特别是在它们与冲突的关系中。我们在协作的背景下解读我们的发现,并讨论其在开发支持CPS的人工智能系统中的意义。
cs.AI / 115 / 2607.23326

ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications

ESF-Bench:针对真实企业应用的挑战性插槽填充场景基准测试
Liang, Toby, Sarda, Gopal, Davasam, Sagar, Yadav, Vikas
Abstract
The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints and unexpected user behaviors. Among these applications, slot filling is essential for converting unstructured input into structured, actionable data. In this work, we introduce ESF-Bench, a challenging Enterprise Slot Filling benchmark consisting of 810 multi-turn samples and 6530 slots over 8 unique domains. Curated using a taxonomy of the 57 most challenging slot-filling scenarios observed during real-world enterprise deployments, ESF-Bench exposes notable limitations in current state-of-the-art LLMs, with GPT-OSS-120b low successfully extracting slots for only 20.7% of benchmark samples. To support continued research in this area, we publicly release the benchmark dataset, taxonomy, and accompanying evaluation code on GitHub.
Chinese Translation
大型语言模型(LLMs)的快速崛起推动了企业的变革性采用。然而,在真实环境中部署这些模型面临着复杂系统约束和意外用户行为等独特挑战。在这些应用中,插槽填充对于将非结构化输入转换为结构化、可操作的数据至关重要。在本研究中,我们介绍了ESF-Bench,这是一个由810个多轮样本和6530个插槽组成的挑战性企业插槽填充基准,涵盖8个独特领域。该基准使用在真实企业部署中观察到的57种最具挑战性的插槽填充场景的分类法进行策划,ESF-Bench揭示了当前最先进的LLMs的显著局限性,其中GPT-OSS-120b仅成功提取了20.7%的基准样本插槽。为了支持该领域的持续研究,我们在GitHub上公开发布了基准数据集、分类法和相关评估代码。
cs.AI / 116 / 2607.23386

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

自信地错误:前沿大型语言模型规则评估中的异常链崩溃
Simpson, Paul, Kozak, John, Doake, Lisa
Abstract
We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional rules of the form "A is required UNLESS B applies, UNLESS C overrides B". The failure reproduces at first observation, but its empirical surface is unstable: between March and April 2026 several failure cells closed silently under the same model alias, with no version bump (GPT-5.4 on construction insurance moved from 96.6% to 100%, same prompt and harness). For regulated workflows, frontier-model accuracy is a moving compliance boundary that shifts without notice. We present the Aethis Eligibility Module, a neuro-symbolic architecture in which LLMs author rules from authoritative sources and an SMT-based layer executes them deterministically, consistent with the authored specification regardless of model drift, reasoning-effort defaults, or prompt format. Three evidence bases: (i) a controlled benchmark of 225 scenarios across four regulatory domains documents the pattern and, in replication, the drift that partially closed it; (ii) a 20-scenario adversarial extension on construction insurance, where the engine scores 20/20, as does one of four frontier configurations (GPT-5.4 at low reasoning effort), while the other three, including Anthropic's strongest model at evaluation time, fail the same coverage-gap edge case; (iii) external validation on nine peer-reviewed LegalBench tasks, 949 held-out cases, where the engine is significantly more accurate than all three frontier models (combined McNemar's p <= 0.003), with margins up to +41 points on the curated multi-prong tasks against the Anthropic models. The contribution is to relocate uncertainty from the inference boundary, where it is silent, to the specification boundary, where it is deliberate and audited. All scenarios, rule encodings, and results are public and reproducible.
Chinese Translation
我们记录了前沿大型语言模型中的一种失败类别——异常链崩溃——该现象在嵌套条件规则下的资格评估中观察到,规则形式为 "A 是必需的,除非 B 适用,除非 C 覆盖 B"。这种失败在首次观察时会重现,但其经验表面不稳定:在 2026 年 3 月和 4 月之间,多个失败单元在同一模型别名下悄然关闭,且没有版本更新(GPT-5.4 在建筑保险上的准确率从 96.6% 提升至 100%,使用相同的提示和工具)。对于受监管的工作流程,前沿模型的准确性是一个不断变化的合规边界,随时可能发生变化。我们提出了 Aethis 资格模块,这是一种神经符号架构,其中大型语言模型从权威来源生成规则,并且基于 SMT 的层以确定性方式执行这些规则,确保与所编写的规范一致,无论模型漂移、推理努力默认值或提示格式如何。三个证据基础:(i)一个涵盖四个监管领域的 225 个场景的受控基准记录了模式,并在复制中记录了部分关闭的漂移;(ii)一个针对建筑保险的 20 个场景的对抗性扩展,其中引擎得分为 20/20,四个前沿配置中的一个(在低推理努力下的 GPT-5.4)也得分相同,而其他三个,包括评估时的 Anthropic 最强模型,未能覆盖同一覆盖缺口边缘案例;(iii)在九个经过同行评审的 LegalBench 任务上的外部验证,949 个保留案例,其中引擎的准确性显著高于所有三个前沿模型(综合 McNemar 的 p <= 0.003),在针对 Anthropic 模型的策划多方面任务中,准确率最高可达 +41 分。该贡献在于将不确定性从推理边界转移到规范边界,从而使其变得明确和可审计。所有场景、规则编码和结果都是公开和可重复的。
cs.AI / 117 / 2607.23393

Key-Interval A*: Accelerating Grid Pathfinding via Structural Abstraction

关键区间 A*: 通过结构抽象加速网格路径寻找
Sui, Taiquan
Abstract
Existing exact methods for 4-connected grid pathfinding reduce online search, but often either retain fine-grained search states or require substantial preprocessing. This paper presents Key-Interval A* (KIA*), an optimal pathfinding algorithm that uses lightweight preprocessing to construct and search over a compact interval-level abstraction of free space. KIA* represents free space using intervals: maximal contiguous runs of traversable cells. It extracts key intervals that capture structural boundary changes and connects them through contiguous non-key regions. KIA* then performs A*-style search on the resulting key-interval graph and constructively reconstructs grid paths from interval chains, without cell-level local search. We prove the completeness and optimality of KIA* on 4-connected grids. Experiments on standard benchmarks show that KIA* preserves exact shortest-path lengths and achieves the fastest runtime on seven of eight benchmark groups, with the largest gains on structured and game maps.
Chinese Translation
现有的精确方法用于4连通网格路径寻找,减少了在线搜索,但往往要么保留细粒度的搜索状态,要么需要大量的预处理。本文提出了关键区间 A* (KIA*),这是一种最优路径寻找算法,利用轻量级的预处理构建并搜索一个紧凑的自由空间区间级抽象。KIA* 使用区间表示自由空间:可遍历单元的最大连续运行。它提取捕捉结构边界变化的关键区间,并通过连续的非关键区域将其连接。KIA* 然后在得到的关键区间图上执行 A* 风格的搜索,并从区间链中构造性地重建网格路径,而无需进行单元级的局部搜索。我们证明了 KIA* 在4连通网格上的完备性和最优性。在标准基准测试中的实验表明,KIA* 保持了精确的最短路径长度,并在八个基准组中的七个组中实现了最快的运行时间,在结构化和游戏地图上获得了最大的收益。
cs.AI / 118 / 2607.23394

Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning

推理时共识以减轻大语言模型微调中的隐性行为
Narang, Adhyyan, Tajdini, Artin, Zhang, Claire, Morgenstern, Jamie
Abstract
Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such as data filtering, mixing in harmless data, and regularization, attenuate these effects but do not eliminate them. We instead pursue robustness through redundancy: collecting multiple datasets from different sources and only learning what is common between them. Thus, if only a subset of sources are malicious, the misbehavior will be blocked. In order to implement this defense strategy, we fine-tune a separate reference model on each source's dataset and aggregate their next-token distributions at decoding time. We introduce two consensus decoders: a token-wise minimum, which caps each token at the lowest probability any source assigns, and a base-relative variant, which reverts to the base probability on any token the sources move in opposing directions. We further relax exact agreement to tolerate partial support across sources and different surface expressions of the same intention. Across controlled poisoning tasks, subliminal learning, and emergent misalignment, consensus decoding suppresses source-specific misbehavior while preserving shared desirable behavior, including cases where union training and weight averaging retain the unwanted behavior.
Chinese Translation
近期研究表明,即使在少量被污染的数据上微调语言模型也可能导致针对性的错误行为,而表面上无害的数据则可能传递出广泛泛化的隐性偏好。标准防御措施,如数据过滤、混入无害数据和正则化,虽然能够减弱这些影响,但并不能完全消除它们。我们则通过冗余追求鲁棒性:从不同来源收集多个数据集,仅学习它们之间的共同部分。因此,如果只有一部分来源是恶意的,错误行为将被阻止。为了实施这一防御策略,我们在每个来源的数据集上微调一个单独的参考模型,并在解码时聚合它们的下一个标记分布。我们引入了两种共识解码器:一种是逐标记最小值,它将每个标记的概率限制在任何来源分配的最低概率上;另一种是基于相对基准的变体,它在任何标记上恢复到基础概率,当这些来源朝相反方向移动时。我们进一步放宽了精确一致性,以容忍来源之间的部分支持以及相同意图的不同表述。在受控的污染任务、潜意识学习和新兴的不一致性中,共识解码抑制了特定来源的错误行为,同时保留了共享的期望行为,包括在联合训练和权重平均中保留不希望的行为的情况。
cs.AI / 119 / 2607.23408

NeurGO: Learning to Generate Elite Candidates for Meta-Black-Box Expensive Optimization

NeurGO:学习生成元黑箱昂贵优化的精英候选者
He, Jintao, Zhen, Huixiang, Gong, Wenyin
Abstract
Expensive black-box optimization is ubiquitous in science and engineering, where function evaluations are costly and the evaluation budget is limited. Traditional evolutionary algorithms and Meta-BlackBox Optimization (MetaBBO) approaches typically consume most evaluations on candidate selection, often wasting precious budget on inferior solutions. Although surrogate-assisted evolution and Bayesian optimization aim to reduce evaluations through surrogate models, constructing an accurate global model from limited data remains challenging, and model bias can easily trap the search in local optima. To overcome these limitations, we propose NeurGO, a generative MetaBBO framework that directly synthesizes elite candidates from historical population states. Specifically, we employ an attention-based encoder to capture the population-level search trend and condition a decoder on this representation to generate high-quality candidates, avoiding the expensive evaluation of large offspring pools. We then design a quality-diversity loss to maintain solution quality and population diversity throughout the search. Through extensive benchmarking on CEC 2008 and the COCO BBOB test suites, our method achieves better optimization performance under the same evaluation budget and exhibits faster convergence.
Chinese Translation
昂贵的黑箱优化在科学和工程中无处不在,其中函数评估成本高且评估预算有限。传统的进化算法和元黑箱优化(MetaBBO)方法通常在候选选择上消耗大部分评估,常常浪费宝贵的预算在劣质解决方案上。尽管代理辅助进化和贝叶斯优化旨在通过代理模型减少评估,但从有限数据构建准确的全局模型仍然具有挑战性,模型偏差容易使搜索陷入局部最优。为克服这些限制,我们提出了NeurGO,一种生成式MetaBBO框架,直接从历史种群状态合成精英候选者。具体而言,我们采用基于注意力的编码器来捕捉种群级别的搜索趋势,并在此表示上对解码器进行条件设置,以生成高质量候选者,从而避免对大型后代池的昂贵评估。然后,我们设计了一种质量-多样性损失,以在整个搜索过程中保持解决方案质量和种群多样性。通过在CEC 2008和COCO BBOB测试套件上的广泛基准测试,我们的方法在相同评估预算下实现了更好的优化性能,并表现出更快的收敛速度。
cs.AI / 120 / 2607.23438

Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels

将能力与权限分离:一种代理人工智能自主水平的治理框架
Zheng, Haining, Dong, Qian, Depena, Rodolfo K., Bhatia, Jonathan D., Xiao, Feng, Xu, Peng
Abstract
As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they should be permitted to do in practice. This paper introduces a governance framework that explicitly separates Allowed Autonomy Levels (AAL), which define the degree of autonomy an AI agent is authorized to exercise given risk, oversight, and accountability considerations, from Autonomous Capability Levels (ACL), which characterize an agent's inherent technical abilities. We present a structured set of autonomy levels spanning reactive execution, decision support, supervised action, goal-directed autonomy, and delegated operational authority, and describe how control, reversibility, and accountability change as autonomy increases. To operationalize this framework, we propose a risk-aware decision process for assigning allowed autonomy, analyze how risk and accountability evolve across autonomy levels, and demonstrate its application through a deployed enterprise data engineering agent, illustrating how a system assessed at a high capability level can be deliberately constrained to a lower allowed autonomy based on risk, reversibility, and organizational readiness. By distinguishing authorization from capability, this work provides practical guidance for the design, deployment, and governance of Agentic AI systems.
Chinese Translation
随着人工智能系统越来越多地表现出代理行为,自主性的讨论常常将系统在技术上能够做的事情与在实践中应被允许做的事情混为一谈。本文提出了一种治理框架,明确区分了允许的自主水平(Allowed Autonomy Levels, AAL),即在风险、监督和问责考虑下,人工智能代理被授权行使的自主程度,以及自主能力水平(Autonomous Capability Levels, ACL),即代理的固有技术能力。我们提出了一套结构化的自主水平,包括反应执行、决策支持、监督行动、目标导向自主和委托操作权,并描述了随着自主性增加,控制、可逆性和问责如何变化。为了使这一框架具备可操作性,我们提出了一种风险意识决策过程,用于分配允许的自主性,分析风险和问责如何在自主水平之间演变,并通过一个已部署的企业数据工程代理展示其应用,说明如何将一个在高能力水平评估的系统有意限制在基于风险、可逆性和组织准备程度的较低允许自主水平。通过区分授权与能力,本文为代理人工智能系统的设计、部署和治理提供了实用指导。
cs.AI / 121 / 2607.23496

Do LLMs Know Their Vulnerable Scenarios?

大型语言模型是否了解其脆弱场景?
Peng, Ziheng, Deng, Huiqi, Jing, Haoran, Rong, Xuankun, Han, Jiahui, Wang, Xiting, Zou, Na, Hu, Xia
Abstract
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relationship between the two representations. In this work, we show that scenario-wrapped prompts activate internal scenario directions whose causal steering consistently reduces refusal scores. Building on this finding, we propose \textsc{Concept2Scenario}, a concept-based attribution framework for vulnerable scenario discovery. It instantiates a broad concept space with a sparse autoencoder, attributes refusal suppression to individual concepts, translates the identified concepts into interpretable natural-language scenarios, and identifies synergistic scenario combinations through interaction attribution. Across three open-source models, two safety benchmarks, and six black-box jailbreak methods, the discovered scenarios serve as reusable priors that improve average attack success rates by up to $18.2$ percentage points. They also transfer to GPT-5, Claude-Haiku-4.5, and Gemini-3-Flash, suggesting that some scenario-level refusal vulnerabilities are shared across model families. Moreover, the identified combinations outperform their individual constituents and enable iterative attacks to succeed in fewer turns.
Chinese Translation
安全对齐的大型语言模型被训练以拒绝有害请求,但将相同请求嵌入特定场景中可以绕过其安全防护。现有的红队测试方法通过观察攻击结果经验性地识别有效场景,但为何特定场景会削弱拒绝能力在机制上仍不清楚。同时,机制可解释性研究已表征了拒绝方向和与越狱相关的特征,但未解释这两种表征之间的关系。在本研究中,我们展示了场景包装提示激活内部场景方向,其因果引导持续降低拒绝分数。基于这一发现,我们提出了 extsc{Concept2Scenario},一个基于概念的脆弱场景发现归因框架。该框架通过稀疏自编码器实例化广泛的概念空间,将拒绝抑制归因于个别概念,将识别的概念转化为可解释的自然语言场景,并通过交互归因识别协同场景组合。在三个开源模型、两个安全基准和六种黑箱越狱方法中,发现的场景作为可重用的先验,平均提高攻击成功率达 $18.2$ 个百分点。这些场景也可以迁移到 GPT-5、Claude-Haiku-4.5 和 Gemini-3-Flash,表明某些场景级拒绝脆弱性在模型家族之间是共享的。此外,识别的组合表现优于其个体成分,并使迭代攻击在更少的回合中成功。
cs.AI / 122 / 2607.23524

Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis

深度搜索中的委托智能:一个可控的解耦能力诊断框架
Yao, Xinhao, Liu, Yuanzhuo, Wang, Changhao, Yu, Yunfei, Tan, Haoran, Zhang, Yuyao, Ren, Ruifeng, Peng, Minlong, Liu, Yong
Abstract
Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy. This coupled evaluation paradigm entangles retrieval quality, long-context comprehension, evidence verification, and tool-use decisions, making it difficult to determine whether a model truly knows when and how to delegate information seeking to search. To this end: (1) We formalize this meta-capability as Delegation Intelligence in deep search and decompose it into complementary dimensions-Search Decision-Making (recognizing information insufficiency and deciding whether, when, and how to search) and Information Synthesis and Verification (aggregating evidence from multiple sources, judging source reliability, and synthesizing information under noisy, potentially adversarial conditions). (2) To enable disentangled and reproducible measurement, we develop a controllable synthesis pipeline built on document-grounded reverse engineering. This yields a general recipe for constructing controlled deep-search evaluations rather than a single fixed dataset. (3) As a concrete instantiation, we construct DelegSearchBench, together with a disentangled evaluation protocol that isolates each capability dimension by varying document composition and tool access. (4) Across representative models, we demonstrate that deep-search competence cannot be adequately characterized by final-answer accuracy alone...
Chinese Translation
深度搜索正成为现代智能体系统的核心能力,但通常仅基于端到端的答案准确性进行评估。这种耦合评估范式将检索质量、长文本理解、证据验证和工具使用决策纠缠在一起,使得难以判断模型是否真正知道何时以及如何将信息搜索委托给搜索。为此:(1)我们将这种元能力形式化为深度搜索中的委托智能,并将其分解为互补维度——搜索决策(识别信息不足并决定是否、何时以及如何搜索)和信息综合与验证(从多个来源聚合证据、判断来源可靠性,并在嘈杂的、潜在对抗的条件下综合信息)。 (2)为了实现解耦和可重复的测量,我们开发了一个基于文档驱动的逆向工程的可控综合管道。这为构建可控的深度搜索评估提供了一般性方案,而不是单一固定的数据集。 (3)作为具体实例,我们构建了DelegSearchBench,并设计了一个解耦评估协议,通过改变文档组成和工具访问来隔离每个能力维度。 (4)在代表性模型中,我们证明仅通过最终答案的准确性无法充分表征深度搜索能力...
cs.AI / 123 / 2607.23537

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

ObsDriveBench:在不利天气条件下基于可观测性的多模态理解基准测试
Yan, Qiao, Wang, Yihan, Xing, Zhenghao, Xu, Jiaqi, Heng, Pheng-Ann
Abstract
Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language models behave under real-world adverse weather with multi-modal inputs. We argue that a key difficulty lies in degraded environmental observability: under fog, rain, snow, and low illumination, multi-modal observations become unreliable and cross-modally inconsistent, posing challenges to scene understanding, and subsequent decision-making. To study this, we introduce \textbf{ObsDriveBench}, a real-world multi-modal benchmark for adverse-weather autonomous driving. Our benchmark is designed with three capability dimensions: \textbf{observability awareness}, \textbf{spatial reliability}, and \textbf{risk-aware decision-making}, enabling fine-grained diagnosis of model behavior under degraded observations. We construct the benchmark through observability meta-annotation, scene description, and capability oriented multiple-choice tasks over synchronized camera, LiDAR, and radar inputs, forming a benchmark with over 14k training and 13k test questions. Experiments reveal consistent performance degradation of existing vision-language models. We further introduce \textbf{ObsDrive} model with normal-weather supervised fine-tuning and adverse-weather reinforcement learning, improving robustness across all three capabilities. The dataset and evaluation code will be released at \href{https://github.com/russellyq/ObsDriveBench}{\texttt{ObsDriveBench}}.
Chinese Translation
在不利天气条件下的自主驾驶仍然是一个关键挑战,但现有的视觉-语言基准主要在标准条件、合成干扰或单一模态下进行评估。因此,尚不清楚视觉-语言模型在真实世界的不利天气条件下如何表现,尤其是在多模态输入的情况下。我们认为一个关键困难在于环境可观测性的下降:在雾、雨、雪和低光照条件下,多模态观测变得不可靠且跨模态不一致,这对场景理解和后续决策造成了挑战。为此,我们引入了 extbf{ObsDriveBench},这是一个针对不利天气条件下自主驾驶的真实世界多模态基准。我们的基准设计了三个能力维度: extbf{可观测性意识}、 extbf{空间可靠性}和 extbf{风险感知决策},以便对模型在降级观测下的行为进行细致诊断。我们通过可观测性元注释、场景描述和基于能力的多项选择任务构建了该基准,涵盖同步的摄像头、激光雷达(LiDAR)和雷达输入,形成一个包含超过14,000个训练问题和13,000个测试问题的基准。实验结果显示现有视觉-语言模型的性能一致下降。我们进一步引入了 extbf{ObsDrive}模型,通过正常天气的监督微调和不利天气的强化学习,提高了三个能力维度的鲁棒性。数据集和评估代码将发布在 exttt{ObsDriveBench}上。
cs.AI / 124 / 2607.23581

Verification-Notebook Learning for Source-Aware Multimodal Misinformation Detection

源感知多模态虚假信息检测的验证笔记本学习
Tan, Junyuan
Abstract
Multimodal misinformation verification is challenging because misleading signals may come from different parts of a post and require different forms of evidence. LVLMs are well suited to this task, but their verification performance often depends on the inference procedure applied to each instance. Existing methods improve this procedure through stronger prompting, retrieval, or deliberation, but rarely retain the verification patterns learned from previous examples. We propose Verification-Notebook Learning (VNL), a non-parametric framework that learns an external verification procedure for a frozen LVLM before inference. VNL builds a compact notebook of decision principles, evidence cues, and recurring pitfalls from prior verification experience. The notebook remains fixed during inference and guides the verification of new examples. Rather than updating model parameters or storing demonstrations, VNL records learned knowledge in an artifact that can be inspected directly. Experiments show that VNL consistently outperforms a range of competitive baselines. Further analyses show that the Verification Notebook improves fine-grained source attribution while remaining compact and interpretable, providing an effective way to accumulate verification knowledge without model training.
Chinese Translation
多模态虚假信息验证面临挑战,因为误导性信号可能来自帖子不同部分,并需要不同形式的证据。大规模语言模型(LVLMs)非常适合这项任务,但它们的验证性能往往依赖于应用于每个实例的推理过程。现有方法通过更强的提示、检索或深思熟虑来改善这一过程,但很少保留从先前示例中学习到的验证模式。我们提出了验证笔记本学习(Verification-Notebook Learning, VNL),这是一种非参数框架,在推理之前为冻结的LVLM学习外部验证过程。VNL构建了一个紧凑的决策原则、证据线索和先前验证经验中的常见陷阱的笔记本。在推理过程中,笔记本保持固定,并指导新示例的验证。VNL并不更新模型参数或存储示例,而是将学习到的知识记录在一个可以直接检查的工件中。实验表明,VNL在多种竞争基线中始终表现优越。进一步分析表明,验证笔记本在保持紧凑和可解释的同时,改善了细粒度的来源归属,提供了一种有效的方式来积累验证知识,而无需模型训练。
cs.AI / 125 / 2607.23586

Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents

您仍然是我授权的代理吗?在固定上限下演变代理的获得授权
Zhang, Zhaoxi, Zhang, Xiaomei
Abstract
Long-lived AI agents increasingly evolve after deployment by retaining experience, acquiring skills and tools, revising workflows, delegating work, and moving across task phases. This improves adaptation but creates a distinct authorization problem. Tool-enabled agents can turn model errors and prompt injections into consequential external actions; when evolution occurs under a live grant, the subject exercising that authority or the context in which it acts may no longer match what the user evaluated. Evolution can change both the effects reachable under an old grant and the authority required by the task, which may rise, fall, or become incomparable. Existing tool policies constrain actions but do not determine when a grant survives this change. We formulate authorization continuity: when does an existing grant remain valid, how may active authority change, and what boundary must never move? Our state-bound model fixes a transition envelope and an immutable effect ceiling at grant time. The envelope determines whether the grant survives a mutation; below the ceiling, authority may contract freely and expand only under specified evidence conditions. We distinguish requested from realized effects and prove that, under complete mediation, sound effect abstraction, attenuating delegation, and monitor integrity, mutation cannot amplify protected effects beyond the user-issued ceiling. Agent-produced evidence may allocate authority below the ceiling but cannot raise it. Finally, we map six mutation classes to their authorization consequences.
Chinese Translation
长期存在的人工智能代理在部署后逐渐演变,通过保留经验、获取技能和工具、修订工作流程、委派工作以及在任务阶段之间移动。这种演变提高了适应性,但也带来了明显的授权问题。工具驱动的代理能够将模型错误和提示注入转化为重要的外部行为;当演变在有效授权下发生时,行使该授权的主体或其所处的上下文可能与用户评估时的情况不再匹配。演变可能改变旧授权下可达的效果以及任务所需的授权,这种授权可能上升、下降或变得不可比。现有的工具政策限制了行为,但并未确定在这种变化下授权何时仍然有效。我们提出了授权连续性:现有授权何时仍然有效,主动授权如何变化,以及什么边界绝不能移动?我们的状态绑定模型在授权时固定了一个转变范围和一个不可变的效果上限。该范围决定了授权是否在变异中存活;在上限以下,授权可以自由收缩,仅在特定证据条件下扩展。我们区分请求的效果与实现的效果,并证明在完全中介、有效效果抽象、减弱委派和监控完整性的条件下,变异无法将受保护的效果放大到超出用户设定的上限。代理生成的证据可能在上限以下分配授权,但无法提高上限。最后,我们将六种变异类别映射到其授权后果。
cs.AI / 126 / 2607.23605

Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning

统一评论者的混合优势估计用于视觉语言模型(VLM)自主强化学习
Zhang, Wenxuan, Wang, Yuhui, Jia, Donggang, Shen, Xiaoqian, Ding, Jian, Viola, Ivan, Schmidhuber, Jürgen, Elhoseiny, Mohamed
Abstract
Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end training in agentic environments can improve such multi-turn decision-making abilities, current methods mainly rely on either token-wise optimization over concatenated token trajectories or turn-wise optimization with uniform within-turn credit. In this work, we establish theoretical formulations for the two levels of optimization and derive a hybrid advantage that serves both objectives. Furthermore, with an appropriate choice of discount factor and learning target, we prove that a unified critic model can estimate values for both turn-wise and token-wise. As such, we propose HyGAE, an actor-critic framework that jointly optimizes token- and turn-level objectives with the hybrid advantage and unified critic. We conduct extensive evaluations of HyGAE across five multi-turn decision-making environments, where it achieves an average success rate of 91% and a significant improvement of 10% over other methods. Furthermore, we provide an in-depth analysis showing that the exact analytic form of the hybrid advantage and return is crucial for optimization. Project Page: https://wx-zhang.github.io/hygae-web/.
Chinese Translation
大型视觉语言模型(VLM)现在在互动环境中充当代理,成功需要在多个回合中进行连贯的推理和决策。尽管在自主环境中的端到端训练可以提高这种多回合决策能力,但当前的方法主要依赖于对连接的令牌轨迹进行逐令牌优化或对回合内均匀信用进行逐回合优化。在本研究中,我们为这两种优化层次建立了理论公式,并推导出一种混合优势,以同时服务于这两个目标。此外,通过适当选择折扣因子和学习目标,我们证明了统一评论者模型可以同时估计逐回合和逐令牌的值。因此,我们提出了HyGAE,这是一种演员-评论者框架,利用混合优势和统一评论者共同优化令牌和回合级目标。我们在五个多回合决策环境中对HyGAE进行了广泛评估,结果显示其平均成功率达到91%,比其他方法显著提高了10%。此外,我们提供了深入分析,表明混合优势和回报的确切解析形式对优化至关重要。项目页面:https://wx-zhang.github.io/hygae-web/
cs.AI / 127 / 2607.23676

SpecAHD: Localize to Specialize for Automated Heuristic Design in Large-Scale Routing Problems

SpecAHD:在大规模路由问题中本地化以专门化的自动启发式设计
Lai, Kezhao, Lai, Yutao, Liu, Hai-Lin
Abstract
LLM-based automated heuristic design (AHD) typically scores executable programs on complete instances or within fixed solver components. In large-scale routing problems, localized reconstruction reduces the size of each optimization task, but repair regions within the same incumbent can exhibit substantially different structures. One construction rule must therefore compromise across them. In this paper, we propose SpecAHD, a coupled bilevel framework for within-instance specialization. An upper-level search learns where to expose bounded repair regions, while a lower-level search evolves a complementary repertoire of executable heuristics for the induced repair tasks. The upper-level program determines the repair tasks seen by the lower level, while checked repair outcomes determine how upper-level programs are evaluated. The lower-level objective favors heuristics that perform well on average or solve tasks that the current repertoire handles poorly. For the repair tasks induced by a fixed upper-level program and a fixed lower-level candidate pool, this objective is monotone submodular, allowing greedy repertoire selection with a (1-1/e) approximation guarantee. Across four routing problems and multiple LLM backbones, SpecAHD reduces held-out objective cost by up to 57.7% against the strongest competing AHD baseline and outperforms the per-instance baseline envelope on most public instances.
Chinese Translation
基于大型语言模型(LLM)的自动启发式设计(AHD)通常在完整实例或固定求解器组件内对可执行程序进行评分。在大规模路由问题中,本地化重构减少了每个优化任务的规模,但同一现有解中的修复区域可能表现出显著不同的结构。因此,一个构造规则必须在这些区域之间进行妥协。本文提出了SpecAHD,一种用于实例内专门化的耦合双层框架。上层搜索学习在哪里暴露有界修复区域,而下层搜索则演化出一组互补的可执行启发式方法以应对引发的修复任务。上层程序确定下层所见的修复任务,而检查过的修复结果则决定上层程序的评估方式。下层目标偏好在平均表现良好或解决当前库处理不佳的任务的启发式方法。对于由固定上层程序和固定下层候选池引发的修复任务,该目标是单调子模的,允许以(1-1/e)的近似保证进行贪婪的库选择。在四个路由问题和多个LLM基础模型中,SpecAHD在最强竞争的AHD基线下将保留目标成本降低了多达57.7%,并在大多数公共实例上超越了每实例基线的包络线。
cs.AI / 128 / 2607.23678

Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems

关注即是全部:多智能体图系统的自适应目标感知注意力调度
Fan, Mingzhou, Xu, Siyuan, Yuan, Mingxuan
Abstract
Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use. Recent systems increasingly organize these agents as graphs of specialized, interconnected nodes. Although graph-based orchestration supports flexible decomposition and coordination, it creates a key challenge: \textbf{attention allocation}. As workflows grow, existing approaches often execute graph components uniformly, wasting resources on irrelevant or low-impact tasks. We introduce \textbf{Attention Orchestration}, a paradigm that extends Transformer-style attention from token representations to workflow-level agent coordination. Our framework, \textbf{Adaptive Goal-aware Attention Orchestration (AGAO)}, dynamically estimates agent importance based on user objectives, graph dependencies, and computational constraints. AGAO combines three components: (1) goal-aware attention, measuring semantic relevance between user goals and agent capabilities; (2) topology-aware attention, modeling structural dependencies in agent graphs; and (3) resource-aware attention, allocating budgets and execution priorities across heterogeneous agents. Together, these mechanisms transform static agent graphs into adaptive systems that focus computation on goal-critical reasoning paths. Experiments across diverse multi-agent workloads show that AGAO improves task effectiveness while reducing unnecessary computation, latency, and token consumption compared with existing graph-based execution strategies. Our work establishes \textbf{Attention Engineering} as a direction for scalable, intelligent multi-agent systems. Code: https://github.com/MingzhouFan97/AGAO.
Chinese Translation
大型语言模型(LLMs)使自主智能体能够进行推理、规划和工具使用。最近的系统越来越多地将这些智能体组织为专门的、互联的节点图。尽管基于图的调度支持灵活的分解和协调,但它带来了一个关键挑战: extbf{注意力分配}。随着工作流程的增长,现有方法往往以统一的方式执行图组件,浪费资源在无关或低影响的任务上。我们引入了 extbf{注意力调度},这一范式将Transformer风格的注意力从标记表示扩展到工作流程级别的智能体协调。我们的框架 extbf{自适应目标感知注意力调度(AGAO)},根据用户目标、图依赖关系和计算约束动态估计智能体的重要性。AGAO结合了三个组件:(1)目标感知注意力,测量用户目标与智能体能力之间的语义相关性;(2)拓扑感知注意力,建模智能体图中的结构依赖关系;(3)资源感知注意力,在异构智能体之间分配预算和执行优先级。这些机制共同将静态智能体图转变为自适应系统,集中计算在目标关键的推理路径上。针对多样化的多智能体工作负载的实验表明,与现有的基于图的执行策略相比,AGAO提高了任务有效性,同时减少了不必要的计算、延迟和标记消耗。我们的工作确立了 extbf{注意力工程}作为可扩展、智能多智能体系统的一个方向。代码:https://github.com/MingzhouFan97/AGAO。
cs.AI / 129 / 2607.23693

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

全球计算,本地实现:稀疏事件-KV 的内存契约
Cai, Zefeng, Cai, Zerui
Abstract
Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Eviction and episodic-memory schemes therefore rest on a premise rarely tested directly, that a retained event is still informative once the observations that produced it are gone. We test it by omitting one earlier observation from what is served, across otherwise identical agent histories. Among items sensitive to that observation, the answer overwhelmingly follows the omitted value, though no served span says which value is correct. We call this semantic materialization: a downstream event's cached rows act as an independently servable view of computation whose inputs are gone. It can also be written on purpose. A deliberately phrased, answer-free event raises donor-aligned recovery from 6% to 51% on Qwen3-8B without ever naming the value, whereas passively harvesting natural mentions from long-term dialog yields no detected advantage. What such a row carries is specific and bounded. Compact state survives, larger payloads decay toward chance, and whether a construction writes at all turns on phrasing rather than on meaning alone, so two phrasings the model comprehends equally well can diverge sharply. The result is a memory contract for sparse event-KV serving: what to write, where it lands, and what survives once the source is gone. For anyone who evicts the corollary is that dropping a source event and observing no accuracy loss does not show the source was unnecessary.
Chinese Translation
长期代理越来越多地将其 KV 缓存作为内存使用:一个服务系统保留一部分缓存条目并丢弃其余部分。因此,驱逐和情节记忆方案基于一个很少直接测试的前提,即一旦产生它的观察消失,保留的事件仍然是有信息量的。我们通过从所服务的内容中省略一个早期观察来测试这一点,其他条件下代理历史保持不变。在对该观察敏感的项目中,答案显著地遵循被省略的值,尽管没有服务的范围表明哪个值是正确的。我们称之为语义实现:下游事件的缓存行作为一个独立可服务的计算视图,其输入已经消失。它也可以故意编写。一个故意措辞的、没有答案的事件使得在 Qwen3-8B 上的捐赠者对齐恢复率从 6% 提升到 51%,而从长期对话中被动收集自然提及则没有检测到优势。这样的行所承载的是特定且有限的。紧凑的状态得以保留,更大的负载则趋向于偶然,而构造是否写入完全取决于措辞而非仅仅是意义,因此模型同样理解的两种措辞可能会显著分歧。结果是稀疏事件-KV 服务的内存契约:写什么、落在哪里,以及一旦源消失后什么会存活。对于任何驱逐的人来说,推论是丢弃源事件并观察到没有准确性损失并不表明源是多余的。
cs.AI / 130 / 2607.23696

Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing

基于生成模型和自适应测试的离线到在线创意优化
Lee, Kevin, Letham, Benjamin, Lin, Zhiyuan Jerry, Samson, Elodie, Onofrey, Eric, Zhang, Poppy, Hill, Shawndra, Bakshy, Eytan
Abstract
Ad creative optimization is increasingly constrained by evaluation rather than generation. Generative models can produce many plausible creatives, but reliable evaluation requires online experiments, in which only a limited slate can be tested. We study how to use data from historical A/B tests to generate and select the candidates in that slate. We developed and deployed a performance-driven offline-to-online workflow that guides creative generation with a predictive model as an inference-time critic. In the offline phase, we use a predictive model trained on historical experiments to rank and refine variants created by a generative model. A final test slate is then deployed in an online adaptive experiment. In a 50-arm field experiment, we found that the best creative generated with this method yielded 45.1% higher engagement than the best human-authored creative. Two additional experiments showed the same upper-tail pattern, with lifts of 46.7% and 36.2%. We found that despite the predictive model being too noisy to directly identify the best creative offline, it effectively guides the generative model toward creating strong candidates that can be efficiently evaluated in an adaptive experiment. The results suggest a design principle for creative optimization with generative models: use predictive models to guide generation of a slate to test, judge the slate by whether it contains high-performing candidates at a feasible test size, and use adaptive experiments to select among candidates while limiting traffic lost to weak arms.
Chinese Translation
创意优化越来越受到评估而非生成的限制。生成模型能够产生许多合理的创意,但可靠的评估需要在线实验,而在这些实验中只能测试有限的创意组合。我们研究如何利用历史 A/B 测试的数据来生成和选择这些组合中的候选项。我们开发并部署了一种以性能为驱动的离线到在线工作流程,该流程通过预测模型作为推理时的评判者来指导创意生成。在离线阶段,我们使用基于历史实验训练的预测模型对生成模型创建的变体进行排序和优化。最终的测试组合随后在在线自适应实验中部署。在一项包含 50 个臂的现场实验中,我们发现采用这种方法生成的最佳创意的参与度比最佳人类创作的创意高出 45.1%。另外两个实验显示了相同的上尾模式,提升幅度分别为 46.7% 和 36.2%。我们发现,尽管预测模型的噪声过大,无法直接识别最佳创意,但它有效地引导生成模型创造出可以在自适应实验中高效评估的强候选项。结果表明了一种使用生成模型进行创意优化的设计原则:利用预测模型指导测试组合的生成,通过是否包含高性能候选项来评判该组合的可行性,并使用自适应实验在候选项中进行选择,同时限制因弱臂而损失的流量。
cs.AI / 131 / 2607.23700

Offline-Online Curriculum RL for Multimodal Reasoning

用于多模态推理的离线-在线课程强化学习
Deng, Wendi, Du, Hang, Nan, Guoshun, Tian, Haokun, Yu, Jiaqi, Cao, Xinlei, Li, Jaile, Chen, Jingfeng, Deng, Ling, Li, Ting, Yang, Hao, Liu, Jun, Jiang, Xudong, Leng, Sicong
Abstract
Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines interpretability and reliability, suggesting reliance on spurious shortcuts rather than faithful reasoning. Although efforts have explored step-level supervision, distinguishing decisive steps from redundant ones remains challenging. We propose $O^2$-CritiCuRL, a novel curriculum reinforcement learning framework that introduces critical-step awareness through an iterative offline-online paradigm. In the offline stage, $O^2$-CritiCuRL conducts multi-rollout analysis over step-annotated trajectories to estimate step-level importance, allowing the framework to distill critical reasoning steps and filter out redundant ones. In the online stage, we employ a progressive step-level reinforcement learning strategy, where truncated chains guide the model to infer missing steps and refine its reasoning, thereby sharpening its focus on critical steps and overcoming the limitations of static supervision. Extensive experiments on multimodal reasoning benchmarks show that our method achieves state-of-the-art performance while delivering superior training and inference efficiency. Code is available at https://github.com/kk0013/CritiCuRL.
Chinese Translation
多模态大型语言模型在推理任务中展现出能力,但在生成正确最终答案的同时,往往会产生错误的中间步骤。这种行为削弱了可解释性和可靠性,表明模型依赖于虚假的捷径而非真实的推理。尽管已有研究探索了步骤级监督,但区分决定性步骤与冗余步骤仍然具有挑战性。我们提出了 $O^2$-CritiCuRL,一种新颖的课程强化学习框架,通过迭代的离线-在线范式引入关键步骤意识。在离线阶段,$O^2$-CritiCuRL 对步骤标注的轨迹进行多次回放分析,以估计步骤级重要性,从而使框架能够提炼出关键推理步骤并过滤掉冗余步骤。在在线阶段,我们采用渐进式步骤级强化学习策略,通过截断链引导模型推断缺失步骤并完善其推理,从而增强其对关键步骤的关注,克服静态监督的局限性。在多模态推理基准上的大量实验表明,我们的方法实现了最先进的性能,同时提供了更优的训练和推理效率。代码可在 https://github.com/kk0013/CritiCuRL 获取。
cs.AI / 132 / 2607.23722

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

E-Bench:在真实产品场景中对多步骤工具使用代理的基准测试
Zheng, Weihuang, Zou, Tianyuan, Ye, Eileen, Liu, Alphet, Kong, Youyong, Zhang, Ya-Qin, Zheng, Duran, Pan, Maxm
Abstract
Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes. We refer to this capability as multi-step tool use. Existing benchmarks have advanced tool-use agent evaluation, but often focus on isolated API calls, short trajectories, or settings that are difficult to scale or control. We introduce E-Bench, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting. E-Bench decouples environment synthesis from task synthesis: graph-guided database filling builds reusable, orphan-free product environments, while generator-solver asymmetry creates tasks with both an information gap and a tool gap, requiring agents to discover hidden data and compose multiple tool calls before changing state. Outcomes are graded deterministically by database-state diffs. Since both environments and tasks are synthetic, E-Bench is controllable at the environment level and scalable at the task level. Benchmarking 11 cutting-edge LLMs shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code extension, reliability (Pass^3) remains below 70%.
Chinese Translation
大型语言模型(LLMs)越来越多地作为代理在多步骤的状态环境中进行交互:收集隐藏信息、组合工具调用以及提交状态变化。我们将这种能力称为多步骤工具使用。现有的基准测试推动了工具使用代理的评估,但往往集中于孤立的API调用、短暂的轨迹或难以扩展或控制的设置。我们介绍了E-Bench,这是一个完全合成的基准测试,涵盖了三个产品领域中的323个状态变化任务:王者荣耀、QQ音乐和腾讯会议。E-Bench将环境合成与任务合成解耦:图引导的数据库填充构建了可重用的、无孤儿的产品环境,而生成器-求解器的不对称性则创建了既有信息差又有工具差的任务,要求代理发现隐藏数据并组合多个工具调用,然后再改变状态。结果通过数据库状态差异以确定性方式进行评分。由于环境和任务均为合成,E-Bench在环境层面可控,在任务层面可扩展。对11个前沿LLM的基准测试表明,多步骤工具使用仍然具有挑战性:即使对于最强模型,Pass^3的得分仍低于60%,而在E-Bench-Code扩展中进行代码执行时,可靠性(Pass^3)仍低于70%。
cs.AI / 133 / 2607.23771

Training Language Models to Cooperate with Inference-Time Controllers

训练语言模型以与推理时控制器协作
Choudhury, Moumita, Khattar, Vanshaj, Liu, Jing, Koike-Akino, Toshiaki, Chakrabarty, Ankush, Zilberstein, Shlomo, Wang, Ye
Abstract
Large language model (LLM) performance increasingly depends not only on the base model, but also on the inference-time controller used to organize reasoning. Existing post-training methods, however, typically optimize for a single fixed interaction pattern, despite real deployments relying on diverse controllers such as Chain-of-Thought, self-consistency, debate, planning, and verification pipelines. This creates a training--deployment mismatch and limits transfer to new workflows. We introduce CALM (Controller-Aware Language Models), a post-training framework that explicitly places controllers in the training loop. We formulate controller-aware post-training as multi-task reinforcement learning over controller-induced interaction protocols, where controllers are compositions of reusable local reasoning modules. This structure also induces a module-level decomposition of mixed-controller training under a turn-level GRPO objective, enabling a systematic study of controller and module-aware training strategies. We evaluate CALM on held-out controller compositions and broader controller shifts, showing that controller-aware post-training improves generalization across inference-time workflows beyond single-controller optimization.
Chinese Translation
大型语言模型(LLM)的性能越来越依赖于基础模型以及用于组织推理的推理时控制器。然而,现有的后训练方法通常仅针对单一固定的交互模式进行优化,尽管实际部署依赖于多样化的控制器,如思维链(Chain-of-Thought)、自一致性(self-consistency)、辩论(debate)、规划(planning)和验证(verification)流程。这导致了训练与部署之间的不匹配,并限制了向新工作流的迁移。我们提出了CALM(控制器感知语言模型),这是一种后训练框架,明确将控制器纳入训练循环。我们将控制器感知的后训练形式化为在控制器引导的交互协议上的多任务强化学习,其中控制器是可重用的局部推理模块的组合。这种结构还促使在回合级GRPO目标下对混合控制器训练进行模块级分解,从而使控制器和模块感知的训练策略的系统研究成为可能。我们在保留的控制器组合和更广泛的控制器变化上评估了CALM,结果表明,控制器感知的后训练在推理时工作流中超越了单一控制器优化,提高了泛化能力。
cs.AI / 134 / 2607.23802

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

从RLVR到RLSVR:任务转化引发开放式大型语言模型自我提升的自验证奖励
Wang, Qinsi, Shi, Jing, Wang, Huazheng, Wan, Kun, Wu, Yiran, Liu, Bo, Wu, Qingyun, Li, Hai Helen, Chen, Yiran, Zhao, Handong, Zhao, Wentian
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verified. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference costs.Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a multi-agent self-play environment inspired by Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains. Models and code have been released at https://github.com/wangqinsi1/SpyRL.
Chinese Translation
具有可验证奖励的强化学习(RLVR)通过实现大规模优化,推动了面向推理的大型语言模型(LLMs)的近期进展。然而,其适用性在很大程度上仍然局限于数学和编码等领域,这些领域的正确性可以被确定性地验证。相反,开放式任务通常依赖于人类偏好、奖励模型或基于LLM的评判者,这引入了评估偏见、评判能力瓶颈和额外的推理成本。基于自监督学习的原则,我们提出了具有自验证奖励的强化学习(RLSVR),这是一种基于任务转化的训练范式,旨在将RLVR扩展到开放式任务。RLSVR将开放式任务转化为可验证的代理环境,其内部规则和交互结果自动生成奖励信号。我们用SpyRL实例化RLSVR,这是一种受《谁是间谍?》启发的多智能体自我对抗环境。代理接收不对称信息,完成相同的目标任务,并投票识别指定的间谍。由于间谍身份是预先确定的,投票结果提供了完全可验证的奖励,而成功识别与输出质量密切相关。在文本摘要、创意写作和数学推理的实验中,SpyRL在不可验证任务上优于现有的自我提升方法,并在可验证推理任务上获得了一致的提升。这些结果表明,任务转化可以将基于可扩展RLVR的自我提升扩展到固有可验证领域之外。模型和代码已在 https://github.com/wangqinsi1/SpyRL 发布。
cs.AI / 135 / 2607.23809

ACM: Agentic Context Management for Long Horizon Tasks

ACM:长时间任务的代理上下文管理
Li, Xiaochuan, Ming, Ryan, Chu, Meng, Shao, Shuai, Jin, Rong, Xiong, Chenyan
Abstract
Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existing context compression methods inevitably incur information loss and are triggered by rigid heuristic rules, leaving them misaligned with the agent's evolving reasoning focus. We propose Agentic Context Management (ACM), a framework that equips agents with purpose-built context editing tools for lossless context management. Inspired by the interaction between short-term and long-term human memory, the agent autonomously decides when to compress its context, offloads discarded content to an external memory system, and queries it on demand for later retrieval. Building on this framework, we further develop a post-training pipeline that constructs high-quality demonstrations of context management and improves model performance on both agentic search and coding tasks. Further analysis reveals that effective context management reduces peak token pressure, enables extended explorations, and yields more consistent solutions across independent trials. Code, data, and model checkpoints are available at https://github.com/lixiaochuan2020/agentic-context-management.
Chinese Translation
代理任务本质上是长时间和多轮的,通过与环境的交互不断积累上下文。现有的上下文压缩方法不可避免地会导致信息损失,并且是由僵化的启发式规则触发的,这使得它们与代理不断演变的推理焦点不一致。我们提出了代理上下文管理(ACM),这是一个为代理提供专用上下文编辑工具以实现无损上下文管理的框架。受到短期和长期人类记忆之间交互的启发,代理自主决定何时压缩其上下文,将丢弃的内容转移到外部记忆系统,并根据需要查询以便后续检索。在此框架的基础上,我们进一步开发了一个后训练流程,构建高质量的上下文管理示范,并提高模型在代理搜索和编码任务上的性能。进一步分析表明,有效的上下文管理减少了峰值令牌压力,支持了更长时间的探索,并在独立试验中产生了更一致的解决方案。代码、数据和模型检查点可在 https://github.com/lixiaochuan2020/agentic-context-management 获取。
cs.AI / 136 / 2607.23845

Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

视觉特征是否改善他人发起的修复检测?一种双向多模态方法
Ngo, Anh, Rollet, Nicolas, Pelachaud, Catherine, Clavel, Chloé
Abstract
Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a recipient signals a problem in speaking, hearing, or understanding, prompting the previous speaker to resolve it. In the case of conversational agents, it is essential to accurately identify these repair initiation strategies to address communication breakdowns efficiently. While conversational analysis studies have shown that OIR initiation is accompanied by both verbal and non-verbal signals such as gaze shifts, facial expressions, body postures, and hand gestures, existing computational approaches rely mainly on text and audio. This paper introduces a novel multimodal model for OIR detection and classification, incorporating a set of visual features drawn from conversation analysis. We evaluate our approach on two corpora with distinct languages and interaction settings. Results demonstrate that visual information consistently improves performance over text and audio baselines, and provide insights into cross-modal feature contributions across two corpora.
Chinese Translation
他人发起的自我修复,简称他人发起的修复(Other-initiated Repair, OIR),是对话互动中的一个重要机制,其中接收者信号表明在说话、听力或理解方面存在问题,促使前一位说话者进行修正。在对话代理的情况下,准确识别这些修复发起策略对于有效解决沟通障碍至关重要。尽管对话分析研究表明,OIR的发起伴随着诸如视线转移、面部表情、身体姿势和手势等言语和非言语信号,但现有的计算方法主要依赖于文本和音频。本文介绍了一种新颖的多模态模型,用于OIR的检测和分类,结合了一组来自对话分析的视觉特征。我们在两个具有不同语言和互动设置的语料库上评估了我们的方法。结果表明,视觉信息在文本和音频基准之上始终提高了性能,并提供了关于跨模态特征贡献的见解。
cs.AI / 137 / 2607.23854

Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search

通过学习和搜索理解组合优化中的类人解决方案
Yan, Haijiang, Zhu, Jian-Qiao, Huang, Liqiang, Meng, Ming
Abstract
Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms. In the Euclidean traveling salesman problems (TSP), people rapidly produce tours that are near-optimal, despite severe limits on time and computation. What makes a tour human-like, and how might such solutions be learned? Here we address these questions through a large-scale behavioral and computational investigation of human performance in Euclidean TSP. We sampled a broad space of TSP instances, collected human solutions, and compared them with neural policies based on Pointer Networks, which are recurrent neural networks with an attention-based pointing mechanism that define probability distributions over valid tours. We trained these networks under multiple objectives, including reinforcement learning (RL), supervised learning from optimal tours, supervised learning from human tours, and RL fine-tuning after optimal-supervised pretraining. Human tours were not identical to optimal tours, but occupied a near-optimal geometric basin: they shared many structural properties with optimal solutions while preserving systematic human-specific deviations. The best account of human tours was not direct imitation of optimal tours, but a model pretrained on optimal tours, fine-tuned by RL, and decoded through $\text{Best-of-}N$ sampling. These findings suggest that human-like solutions may emerge from a combination of structured supervised learning, RL, and test-time search, echoing computational principles underlying many modern artificial intelligence systems.
Chinese Translation
人类经常能够找到组合优化问题的良好解决方案,这些问题即使对于先进的计算机算法来说也计算困难。在欧几里得旅行商问题(TSP)中,尽管时间和计算的限制非常严苛,人们仍能迅速生成接近最优的旅行路线。那么,什么使得一条旅行路线显得类人?这样的解决方案又如何能够被学习?在此,我们通过对人类在欧几里得 TSP 中表现的大规模行为和计算调查来探讨这些问题。我们对广泛的 TSP 实例进行了抽样,收集了人类的解决方案,并将其与基于指针网络(Pointer Networks)的神经策略进行了比较。指针网络是一种具有基于注意力的指向机制的递归神经网络,能够定义有效旅行路线的概率分布。我们在多个目标下训练了这些网络,包括强化学习(RL)、从最优旅行路线的监督学习、从人类旅行路线的监督学习,以及在最优监督预训练后进行的 RL 微调。人类的旅行路线与最优旅行路线并不完全相同,但位于一个近似最优的几何盆地:它们与最优解决方案共享许多结构特性,同时保留了系统的人类特定偏差。人类旅行路线的最佳解释并不是对最优旅行路线的直接模仿,而是一个在最优旅行路线上进行预训练的模型,通过 RL 进行微调,并通过 $ ext{Best-of-}N$ 采样解码。这些发现表明,类人解决方案可能源于结构化的监督学习、强化学习和测试时搜索的结合,反映了许多现代人工智能系统背后的计算原则。
cs.AI / 138 / 2607.23896

Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery

成本意识的恢复路径识别与自主材料发现的贝叶斯优化
Ray, Debajyoti, Srinivas, Niranjan
Abstract
Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We formulate this as a sequential decision problem with a discrete pathway-identification stage and a continuous within-pathway optimization stage under heterogeneous experimental costs. Our implementation, Coactive learning, combines a cost-sensitive Bayesian hypothesis-discrimination policy motivated by EC2 (Golovin et al., 2010) with Gaussian-process Bayesian optimization (Srinivas et al., 2010). Under explicitly stated assumptions, the expected spend of one fixed-budget campaign attempt is bounded by the expected pathway-identification cost plus the capped within-pathway optimization budget. We evaluate the method on synthetic benchmarks constrained by selected results reported for PNNL's CICERO selective-precipitation study (Ritchhart et al., 2026). The method performs comparably to an oracle-pathway Bayesian-optimization reference and to a strong split-plate baseline that discriminates pathways with its first plate, without receiving an oracle label for the correct pathway. It is given a candidate hypothesis space and a diagnostic likelihood model. On an NdFeB-inspired instance, it avoids the simulated penalty of a commit-first baseline that initially selects a plausible but inferior hydroxide pathway. This hypothetical wrong-first-commitment scenario is motivated by the hydroxide-oxalate performance contrast reported by CICERO. We characterize the sensitivity of these conclusions to the assumed cost model. The code and benchmark are open source.
Chinese Translation
自主实验室自动化执行实验,但在实验过程中,必须决定哪个恢复路径值得优化。我们将此问题表述为一个序列决策问题,其中包括一个离散的路径识别阶段和一个在异构实验成本下的连续路径内优化阶段。我们的实现方法,Coactive learning,结合了一种基于EC2(Golovin et al., 2010)的成本敏感贝叶斯假设区分策略与高斯过程贝叶斯优化(Srinivas et al., 2010)。在明确的假设条件下,固定预算的单次实验尝试的预期支出由预期的路径识别成本加上限制的路径内优化预算所界定。我们在受限于PNNL的CICERO选择性沉淀研究(Ritchhart et al., 2026)所报告的选定结果的合成基准上评估该方法。该方法的表现与一个oracle路径贝叶斯优化参考和一个强大的分板基线相当,后者在没有获得正确路径的oracle标签的情况下,通过其第一块板来区分路径。它给定了一个候选假设空间和一个诊断似然模型。在一个受NdFeB启发的实例中,它避免了一个首先承诺的基线所模拟的惩罚,该基线最初选择了一个看似合理但劣质的氢氧化物路径。这个假设的错误优先承诺场景是由CICERO报告的氢氧化物-草酸盐性能对比所激励的。我们描述了这些结论对假设成本模型的敏感性。代码和基准是开源的。
cs.AI / 139 / 2607.23913

GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

GOTS:用于高分辨率视觉语言模型的贪婪正交标记选择
Ling, Jun, Huang, Tao, Liu, Junzhuo, Tang, Bowen, Wang, Peng
Abstract
Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream language-model inference cost. Existing token-reduction methods assess token utility through token-wise importance, query relevance, coverage, pairwise diversity, or subset-level objectives. Our key insight is to view visual token reduction through selected-span complementarity: instead of scoring a token in isolation or through pairwise relations, we assess how much of its feature is orthogonal to the span of the already retained subset. Based on this perspective, we propose Greedy Orthogonal Token Selection (GOTS), a training-free and query-agnostic method. At each step, GOTS selects the token with the largest residual energy orthogonal to the current retained span. This rule exactly maximizes the one-step augmented Gram determinant among candidate additions, giving each greedy step a precise local geometric guarantee for subset expansion. Across five high-resolution VLM backbones from the Qwen-VL and InternVL families and eleven diverse benchmarks, GOTS achieves higher average performance retention than the strongest evaluated baselines, and a controlled OCRBench study shows that it reduces model-side time-to-first-token after accounting for selection overhead. Code is available at https://github.com/newLLing/GOTS.
Chinese Translation
现代视觉语言模型(VLMs)越来越依赖于动态或高分辨率的视觉编码,生成数千个视觉标记,这大大增加了下游语言模型的推理成本。现有的标记减少方法通过标记的重要性、查询相关性、覆盖率、成对多样性或子集级目标来评估标记的效用。我们的关键见解是通过选定范围的互补性来看待视觉标记的减少:我们不是孤立地或通过成对关系来评分一个标记,而是评估其特征有多少与已保留子集的范围正交。基于这一视角,我们提出了贪婪正交标记选择(GOTS),这是一种无训练且与查询无关的方法。在每一步中,GOTS选择与当前保留范围正交的残余能量最大的标记。该规则精确地在候选添加中最大化一步增强的Gram行列式,为子集扩展的每一步贪婪选择提供了精确的局部几何保证。在来自Qwen-VL和InternVL系列的五个高分辨率VLM基础模型和十一项多样化基准测试中,GOTS实现了比最强评估基线更高的平均性能保留,并且经过选择开销调整后,受控的OCRBench研究显示它减少了模型侧的首次标记时间。代码可在 https://github.com/newLLing/GOTS 获取。
cs.AI / 140 / 2607.23927

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

大型语言模型中的现实监测:随着对话记忆的变化而转变的自我知识
Ranjan, Saurabh, Sokratous, Konstantina, Odegaard, Brian
Abstract
A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles, evaluating what they know is not enough: tracking where that knowledge came from may matter equally.
Chinese Translation
一个无法区分自身输出与用户所说内容的对话式人工智能,会将自己的错误视为用户提供的事实。在人类中,这种能力被称为现实监测,其失败与幻觉、妄想和虚构有关,但大型语言模型(LLMs)是否具备这种能力尚未得到验证。在此,我们通过两个实验和六个LLMs展示了源归属依赖于对话记忆的结构:在最低记忆需求下,自生成内容的准确性达到顶峰,但一旦情节延迟去除这一捷径,便转变为脆弱的外部项目优势。反馈揭示了两个失败:在某些模型中,内部和外部判断互换;在其他模型中,准确性提高而置信度与正确性脱钩,这种分离在现有基准测试中是不可见的。跨模型的这一模式表明,影响的是主动的,而非总的参数数量。这表明,随着人工智能系统承担自主的多轮角色,评估它们所知的内容并不足够:追踪这些知识的来源同样可能至关重要。
cs.AI / 141 / 2607.23929

MemTX: Transactional Belief Commit for Stateful Agent Memory

MemTX:面向状态代理内存的事务性信念提交
Li, Xiaoyang, Wang, Yiqi, Lu, Haohui, Chen, Zhi, Li, Mo, Song, Pingan, Cai, Taotao
Abstract
LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a tool call with real side effects. Current agent memory systems treat every accepted write as immediately actionable truth, so a polluted tool result, a stale update, or a teammate's half-finished note can silently drive an irreversible action. We argue that a memory write is not a belief commit. We present MemTX, a transactional belief-commit protocol. Each record carries evidence, permissions, provenance, and validity. Writes are staged inside snapshot-isolated transactions and admitted by a validate-and-commit pipeline, irreversible tool calls are gated on in-flight belief state, and retracting a belief triggers typed cascading repair of its derived records and tool side effects. Two invariants, action-safety gating and cascade-repair completeness, are machine-checked by property-based testing and bounded exhaustive enumeration of 5.5 million protocol states, with zero violations. Across five backbones from three model families, MemTX leads all eight baselines with paired-McNemar significance on four backbones and statistically ties the best baseline on the fifth and strongest, while remaining the only method with zero downstream harm on every backbone. Backbone capability does not substitute for commit discipline.
Chinese Translation
大型语言模型(LLM)代理越来越多地通过持久共享内存进行协调:一个代理的写入成为另一个代理的前提,最终导致具有实际副作用的工具调用。目前的代理内存系统将每个被接受的写入视为立即可执行的真理,因此,污染的工具结果、过时的更新或队友未完成的笔记可能会悄无声息地驱动不可逆转的行动。我们认为,内存写入并不等同于信念提交。我们提出了MemTX,一种事务性信念提交协议。每条记录都携带证据、权限、来源和有效性。写入在快照隔离事务中进行分阶段处理,并通过验证和提交管道进行接纳,不可逆的工具调用受到正在进行的信念状态的限制,而撤回信念会触发其衍生记录和工具副作用的类型级联修复。两个不变式,即行动安全门控和级联修复完整性,通过基于属性的测试和对550万个协议状态的有界穷举枚举进行机器检查,未出现任何违规。在来自三个模型家族的五个基础模型中,MemTX在所有八个基准测试中表现优异,在四个基础模型上与配对的McNemar检验具有显著性,而在第五个最强的基础模型上与最佳基准统计上持平,同时仍然是唯一在每个基础模型上都没有下游损害的方法。基础模型的能力并不能替代提交的纪律。
cs.AI / 142 / 2607.23942

From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps

从认知架构到语言代理:机制层面的谱系、趋同与迁移差距的回顾
Fan, Haodi, Lan, Zucong
Abstract
Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agent actually runs. This review connects ten historical cognitive architectures, eight language-agent runtime families, and forty-two mechanism-focused modern systems. We reconstruct each mechanism through state, control, transition, persistence, failure, learning, and resource governance, then code evidence relation (E1-E4) separately from migration depth (D0-D4). The resulting landscape is uneven. Modern agents have operationalized substantial parts of adaptive memory, failure recovery, dynamic team selection, workflow search, skill induction, resource scheduling, and uncertainty-conditioned action, although often through independent convergence rather than documented inheritance. The strongest remaining opportunities lie in couplings among mechanisms. Closest-baseline screening closes one proposed gap: GraSP already combines calibrated multi-skill selection, typed compilation, verification, bounded repair, and replanning or ReAct fallback. Five residual bundles remain: activation with latency and action utility; typed impasse with isolated substates and resolution compilation; bounded content competition with broadcast and admission learning; persistent intention with reconsideration and live method authority; and uncertainty with resource allocation, interruption, and stopping. We contribute a distinctive-mechanism catalog, an auditable evidence-depth framework, and a falsifiable agenda for testing these bundles as composable runtime invariants.
Chinese Translation
记忆、规划、反思和工具使用常被作为特征标签进行比较,这掩盖了决定代理实际运行方式的控制语义。本综述连接了十种历史认知架构、八种语言代理运行时家族和四十二种以机制为中心的现代系统。我们通过状态、控制、转变、持久性、失败、学习和资源治理重构每个机制,然后将证据关系(E1-E4)与迁移深度(D0-D4)分别编码。结果的景观是不均匀的。现代代理已经在适应性记忆、失败恢复、动态团队选择、工作流搜索、技能归纳、资源调度和不确定性条件下的行动等方面实现了相当大的部分,尽管通常是通过独立的趋同而非文献继承。剩余的最强机会在于机制之间的耦合。最近基线筛选填补了一个提议的差距:GraSP已经结合了校准的多技能选择、类型编译、验证、有界修复以及重新规划或ReAct回退。还有五个残余组合:具有延迟和行动效用的激活;具有孤立子状态和解决编译的类型僵局;具有广播和接纳学习的有界内容竞争;具有重新考虑和实时方法权威的持久意图;以及具有资源分配、中断和停止的不确定性。我们贡献了一个独特机制目录、一个可审计的证据深度框架,以及一个可证伪的议程,用于测试这些组合作为可组合运行时不变性的可能性。
cs.AI / 143 / 2607.23944

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models

DICA:双指示器引导的多模态大型语言模型对比对齐
Yang, Hao, Wang, Jin, Zhang, Xuejie
Abstract
Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusing on question-relevant regions. However, multimodal large language models may deviate from this pattern due to attention drift and the underutilization of visual evidence, which can lead to hallucinations. To mitigate these issues, this study proposes a Dual-Indicator Guided Contrastive Alignment (DICA), which tracks two information-theoretic indicators during inference: Visual Attention Entropy (VAE), which reflects the concentration of visual attention, and Output Image Correlation (OIC), which measures the dependence of generated outputs on the visual input. An abnormal increase in VAE or a decrease in OIC corresponds to different failure modes, which trigger targeted contrastive alignment to restore visual grounding. Experimental results across multiple benchmarks demonstrate that DICA consistently outperforms existing approaches and substantially reduces hallucinations, highlighting the effectiveness of indicator-driven intervention in improving multimodal inference reliability. The code is publicly available at https://github.com/BGWH123/DICA/.
Chinese Translation
人类视觉推理通常遵循从粗到细的注意力过程,首先进行全局场景理解,然后逐渐聚焦于与问题相关的区域。然而,由于注意力漂移和视觉证据的不足利用,多模态大型语言模型可能偏离这一模式,从而导致幻觉。为了解决这些问题,本研究提出了一种双指示器引导的对比对齐方法(DICA),在推理过程中跟踪两个信息论指标:视觉注意力熵(Visual Attention Entropy, VAE),反映视觉注意力的集中程度,以及输出图像相关性(Output Image Correlation, OIC),测量生成输出与视觉输入之间的依赖关系。VAE的异常增加或OIC的减少对应于不同的失败模式,这会触发针对性的对比对齐,以恢复视觉基础。多个基准测试的实验结果表明,DICA始终优于现有方法,并显著减少幻觉,突显了基于指标的干预在提高多模态推理可靠性方面的有效性。代码已公开发布在 https://github.com/BGWH123/DICA/。
cs.AI / 144 / 2607.23955

EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff

EviBack:通过证据约束教师回退的搜索代理强化学习
Ma, Xiao, Hu, Zhiquan, Wei, Yi, Zhao, Chenchen, Chen, Yijun, Zhao, Jicheng, Dai, Yuming Li Chuang
Abstract
Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hide useful search behavior. We present EviBack, an evidence- constrained Teacher backoff that supplies auxiliary super- vision to such groups while preserving verifiable Actor re- wards. It separates evidence assessment from answer refine- ment, preventing reference answers from overriding evidence- insufficiency judgments. A fully automated, end-to-end GPT- 5.5-assisted APE pipeline starts from a manually authored single-prompt dual-task Teacher, automatically partitions and labels rollout data, and performs ablation, task decomposition, evaluation, and selection to produce a gated two-stage Teacher. Compared with the manual design, the resulting Teacher im- proves downstream F1 and valid-answer rate while reduc- ing search, duplicate queries, and forced termination. Across seven open-domain QA benchmarks and three Qwen3 scales, EviBack improves F1 over Search-R1 and raises both single- and multi-hop macro F1. We guarantee that the code will be made publicly available at a later stage.
Chinese Translation
强化学习使得代理型RAG系统能够通过可验证的结果奖励学习多轮搜索,但全零回滚组没有提供比较信号,可能掩盖有用的搜索行为。我们提出了EviBack,一种证据约束的教师回退方法,为此类组提供辅助监督,同时保留可验证的代理奖励。它将证据评估与答案细化分开,防止参考答案覆盖证据不足的判断。一个完全自动化的端到端GPT-5.5辅助的APE管道从手动编写的单提示双任务教师开始,自动划分和标记回滚数据,并进行消融、任务分解、评估和选择,以生成一个门控的两阶段教师。与手动设计相比,所得到的教师在下游F1和有效答案率上有所提高,同时减少了搜索、重复查询和强制终止。在七个开放域问答基准和三个Qwen3规模上,EviBack在F1上超越了Search-R1,并提高了单跳和多跳宏F1。我们保证代码将在后期公开发布。
cs.AI / 145 / 2607.23967

Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

在权重衰减时钟上的理解:来自轻微破缺对称性的速率层次
Kim, Taeyoung
Abstract
Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the grokking subspace. Along this subspace, the training predictions remain unchanged, leaving weight decay as the sole restoring force and giving rise to a slow dissipative relaxation governed by an exact discrete-time and continuous-time law. We show that only this subspace contributes to the slow asymptotic decay of the population risk and derive explicit iteration-scale predictions for the grokking time, recovering the familiar $(1-\beta)/(\eta\lambda)$ scaling in the weak-regularization regime. The theory further predicts distinct effects of optimizer choice, distinguishing coupled $L_2$ regularization from decoupled weight decay, and yields causal predictions for interventions that modify the grokking component. We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. We further observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling and the late-time relaxation agrees closely with the theoretical clock.
Chinese Translation
尽管进行了广泛的实证研究,延迟泛化或理解(grokking)仍然不甚明了。我们识别出一种在使用全批次重球优化和权重衰减训练的线性模型中,能够精确求解的晚期松弛机制,并对非线性神经网络进行了局部二次扩展。我们的分析揭示了经验零空间中一个显著的群体活跃成分,我们称之为理解子空间(grokking subspace)。在这个子空间中,训练预测保持不变,权重衰减成为唯一的恢复力,导致一种由精确的离散时间和连续时间法则主导的缓慢耗散松弛。我们表明,只有这个子空间对群体风险的缓慢渐近衰减有贡献,并推导出理解时间的明确迭代尺度预测,恢复了在弱正则化状态下熟悉的 $(1-eta)/( hetaeta)$ 规模。该理论进一步预测了优化器选择的不同影响,区分了耦合的 $L_2$ 正则化与解耦的权重衰减,并为修改理解成分的干预提供了因果预测。我们在一个合成模型中验证了所有理论身份,在该模型中,所有子空间和松弛速率都可以以封闭形式计算。我们进一步观察到在模块加法中的真实延迟泛化,其中测得的延迟遵循预测的规模,且晚期松弛与理论时钟的吻合度很高。
cs.AI / 146 / 2607.23975

Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

Plato-Bio:以验证为先的生物新颖性筛选,结合时间再发现和结构基准
Creadore, Stefan G.
Abstract
Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We developed Plato-Bio, a biology-routed extension of the open Plato/Denario architecture that couples explicit workflow states with provenance records, citation checks, claim-to-evidence links, scoped file writes, and publication gates. A source audit identified and repaired three defects that could distort evaluation: loss of task domain in the default factory, omission of declared method signals from scoring, and evidence sidecars that lacked the drafted-claim denominator. On the current clean revision, the full Python suite completed with 931 passes, six skips, and no failures or errors; targeted biology, genomics, evidence/citation, and adversarial-safety suites likewise completed without failure. We evaluated two narrow use cases. In a frozen historical rediscovery task, independent pre-1986 literature bridges ranked the later-studied relation between fish oil and Raynaud phenomenon first; TF-IDF ranked it second and corpus frequency third. This single curated task measures retrospective ranking, not prospective discovery. In a separate comparison of AlphaFold models with experimental structures for 15 human proteins, 11 targets had high-confidence-core C-alpha RMSD below 1 Angstrom (median 0.501 Angstrom). Four targets exceeded 2 Angstrom, and confidence masking reduced the SUMO1 discrepancy from 16.61 to 2.58 Angstrom over 74 residues. The workflow emitted 27 traceable discrepancy regions, all retained as unvalidated hypotheses. Plato-Bio therefore provides reproducible software contracts and auditable screening baselines; broader claims of agent efficacy or biological novelty require preregistered evaluation, independent review, and prospective validation.
Chinese Translation
大型语言模型研究代理能够连接文献检索、分析代码和手稿准备,但连贯的输出并不能确立科学有效性。我们开发了Plato-Bio,这是一个基于生物学的扩展,建立在开放的Plato/Denario架构之上,结合了明确的工作流程状态与来源记录、引用检查、主张与证据链接、范围文件写入和出版门槛。源审计识别并修复了三个可能扭曲评估的缺陷:默认工厂中任务领域的丧失、评分中声明的方法信号的遗漏,以及缺乏草拟主张分母的证据附属信息。在当前的清洁修订版中,完整的Python套件完成了931次通过、6次跳过,没有失败或错误;针对生物学、基因组学、证据/引用和对抗安全的套件同样完成且没有失败。我们评估了两个狭窄的使用案例。在一个冻结的历史再发现任务中,独立的1986年前文献桥接将后续研究的鱼油与雷诺现象的关系排名第一;TF-IDF将其排名第二,语料库频率排名第三。这个单一的策划任务衡量的是回顾性排名,而非前瞻性发现。在对15个人类蛋白质的AlphaFold模型与实验结构的单独比较中,11个目标的高置信度核心C-alpha RMSD低于1埃(中位数0.501埃)。四个目标超过2埃,置信掩蔽将SUMO1的差异从16.61减少到74个残基的2.58埃。该工作流程发出了27个可追踪的差异区域,所有区域均保留为未经验证的假设。因此,Plato-Bio提供了可重复的软件合同和可审计的筛选基线;更广泛的代理有效性或生物新颖性的主张需要预注册评估、独立审查和前瞻性验证。
cs.AI / 147 / 2607.23997

Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation

探索基于预算的图像分类与内容敏感的资源分配
Papadopoulos, Athanasios G.
Abstract
The ever-growing adoption of Artificial Intelligence (AI) creates the need to deploy Deep Neural Networks in a variety of computational environments. We consider dynamic environments, where computational requirements are subject to change, and we pose the following question: How do we adjust the complexity of an AI classification system, in order to maximize its accuracy, while meeting changing computational constraints? We call this problem Budgeted Image Classification, and we formally formulate it as a resource allocation integer program. Given a computational budget, a batch of images, and a classification system that can make decisions with varying complexity (it has multiple decision points), we explore strategies to allocate images to decision points, in order to maximize accuracy within the available budget. The original integer program is NP-Hard, so, we propose a continuous relaxation, leading to a content-agnostic allocation strategy which assigns images to decision points without considering their particular content. We address this issue by proposing a content-sensitive strategy, that we experimentally show it leads to superior performance. We theoretically study the behavior of our strategies, deriving conditions that must be satisfied by decision points to be suitable for budgeted classification. We analyze fails cases, offering insights for future research directions.
Chinese Translation
人工智能(AI)的日益普及促使我们需要在各种计算环境中部署深度神经网络。我们考虑动态环境,其中计算需求可能会发生变化,并提出以下问题:我们如何调整AI分类系统的复杂性,以最大化其准确性,同时满足变化的计算约束?我们将此问题称为预算图像分类,并将其正式表述为资源分配整数规划。在给定计算预算、一批图像和一个能够以不同复杂性做出决策的分类系统(它具有多个决策点)的情况下,我们探索将图像分配给决策点的策略,以在可用预算内最大化准确性。原始整数规划是NP-困难的,因此,我们提出了一种连续松弛,导致了一种与内容无关的分配策略,该策略在不考虑图像特定内容的情况下将图像分配给决策点。我们通过提出一种内容敏感的策略来解决这个问题,实验结果表明该策略能够实现更优的性能。我们从理论上研究了我们的策略的行为,推导出决策点必须满足的条件,以适合预算分类。我们分析了失败案例,为未来的研究方向提供了见解。
cs.AI / 148 / 2607.24023

Self-Supervised Consistency Enhanced Disentangled Learning for Neural Decoding Generalization in Brain-Machine Interface

自监督一致性增强的解耦学习用于脑机接口中的神经解码泛化
Wei, Jiyu, Hong, Di, Zhang, Zhanjie, Rong, Dazhong, He, Qinming, Wang, Yueming
Abstract
Brain-Machine Interfaces (BMIs) provide a direct communication pathway between the brain and external devices, enabling humans to control assistive and robotic technologies, with potential applications in rehabilitation, human motor augmentation, and human-centered robotics. However, due to neural drift, the performance of BMIs decreases over time, posing challenges for long-term viability, particularly for invasive BMIs (iBMIs). Existing solutions suffer from two main drawbacks: (i) difficulty in learning robust neural representations, and (ii) neglecting that neural drift varies across motor parameters (e.g., velocity, direction, and speed). To overcome these limitations, we propose Self-Supervised Consistency enhanced Disentangled Learning (SSCDL), a neural decoding generalization framework built on two key innovations. We first design a backbone model named Consistency enhanced Neural Decoder (CND), using a novel teacher-student consistency constraint with simulated neural signal perturbations to learn robust representations invariant to neural drift. Then, we employ three dedicated CNDs under the Complementary-Disentangled Generalization (CDG) mechanism, which disentangles motor signals into velocity, direction, and speed with inspiration from neural preference theory. This disentangled learning enables SSCDL to capture invariant neural representations from diverse neural preference perspectives, significantly enhancing cross-day generalization. Extensive experimental results show that SSCDL delivers state-of-the-art decoding performance, exhibiting high robustness and cross-day stability. These capabilities underscore its strong potential for long-term interaction in human-centric robotic and fine-grained assistive applications.
Chinese Translation
脑机接口(BMIs)提供了大脑与外部设备之间的直接通信通道,使人类能够控制辅助和机器人技术,具有在康复、人类运动增强和以人为本的机器人等领域的潜在应用。然而,由于神经漂移,BMIs的性能随着时间的推移而下降,这对长期可行性构成了挑战,特别是对于侵入性脑机接口(iBMIs)。现有解决方案存在两个主要缺陷:(i)难以学习稳健的神经表征,以及(ii)忽视神经漂移在运动参数(例如速度、方向和速率)之间的变化。为克服这些局限性,我们提出了自监督一致性增强的解耦学习(SSCDL),这是一个基于两个关键创新构建的神经解码泛化框架。我们首先设计了一个名为一致性增强神经解码器(CND)的主干模型,利用一种新颖的教师-学生一致性约束与模拟神经信号扰动相结合,以学习对神经漂移不变的稳健表征。然后,我们在互补解耦泛化(CDG)机制下采用三个专用的CND,灵感来自神经偏好理论,将运动信号解耦为速度、方向和速率。这种解耦学习使SSCDL能够从多样的神经偏好视角捕捉不变的神经表征,显著增强跨天泛化能力。大量实验结果表明,SSCDL提供了最先进的解码性能,展现出高稳健性和跨天稳定性。这些能力突显了其在以人为本的机器人和细粒度辅助应用中的长期交互的强大潜力。
cs.AI / 149 / 2607.24031

A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine Interfaces

一种基于不确定性引导的自适应-泛化框架,用于长期脑机接口的循环学习
Wei, Jiyu, Hong, Di, Zhang, Zhanjie, Rong, Dazhong, He, Qinming, Wang, Yueming
Abstract
Brain-Machine Interfaces (BMIs), which link the brain to external devices, hold great potential in rehabilitation, human performance augmentation, and human-centered robotics. However, invasive BMIs face a critical challenge for long-term deployment due to neural drift, which degrades decoding performance over time and necessitates frequent recalibration. Existing methods designed to mitigate neural drift typically rely on either domain adaptation (DA) or domain generalization (DG) alone and often fail to capture fine-grained distribution shifts across neural subdomains, resulting in limited performance. To overcome these limitations, we propose Uncertainty-guided Self-paced Cycling (UnSPC), a robust framework that synergizes DA and DG for target domain refining under an Uncertainty-guided Self-paced Pseudo-labeling (UnSPL) mechanism. To handle subdomain neural drift across domains, UNSPL is proposed to iteratively mine reliable pseudo-labeled samples with a noise-robust ranking strategy for further fine-tuning. Leveraging these high-quality samples, we introduce a novel Cycling Adaptation and Generalization (CycAG) strategy, which integrates DA and DG within an iterative cycle to progressively mitigate both global and subdomain drift. This cyclic process enables effective alignment to evolving target distributions while preserving robust and transferable representations, thereby mitigating performance degradation under long-term neural drifts. Extensive experiments on multiple neural decoding datasets demonstrate the effectiveness and robustness of UnSPC. To our knowledge, our proposed UnSPC is the first to cyclically integrate DA and DG with pseudo-labeling, paving the way toward stable long-term BMI controls.
Chinese Translation
脑机接口(BMIs)将大脑与外部设备连接,具有在康复、人类性能增强和以人为中心的机器人技术中巨大的潜力。然而,侵入式脑机接口在长期部署中面临着神经漂移的重大挑战,这种漂移会随着时间的推移降低解码性能,并需要频繁的重新校准。现有旨在减轻神经漂移的方法通常仅依赖于领域适应(DA)或领域泛化(DG),往往无法捕捉神经子领域之间的细粒度分布变化,从而导致性能有限。为克服这些局限性,我们提出了不确定性引导的自适应循环(UnSPC)框架,该框架通过不确定性引导的自适应伪标签(UnSPL)机制协同DA和DG,以精细化目标领域。为处理跨领域的子领域神经漂移,提出了UnSPL,通过一种抗噪声的排名策略迭代挖掘可靠的伪标签样本,以便进一步微调。利用这些高质量样本,我们引入了一种新颖的循环适应与泛化(CycAG)策略,该策略在迭代循环中整合DA和DG,以逐步减轻全局和子领域的漂移。这个循环过程能够有效对齐不断演变的目标分布,同时保持稳健和可迁移的表示,从而减轻长期神经漂移下的性能下降。在多个神经解码数据集上的广泛实验表明了UnSPC的有效性和鲁棒性。据我们所知,我们提出的UnSPC是首个循环整合DA和DG与伪标签的方法,为稳定的长期BMI控制铺平了道路。
cs.AI / 150 / 2607.24032

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

生成性人工智能证据的半衰期:40条记录审计、主张货币框架及前沿模型辅助研究的反思案例
Iacono, Carlo
Abstract
Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes. First, it audits a maximum-variation purposive corpus of 40 empirical records appearing between 18 July 2025 and 17 July 2026. The audit coded publication route, execution timing, model identity, age of the newest named generation or immutable snapshot, same-family supersession and refresh behaviour. At appearance, the newest named model was a median 281 days old (middle 50%: 75-478; range: 11-939). Median age was 395 days for 25 journal articles, 56 days for 14 preprints and 49 days for one laboratory report. Thirty-five records included a superseded family, seven supplied a precise dated identifier, three clearly refreshed model evidence, and one added a late sensitivity test. All 40 included an OpenAI system, a feature of this corpus rather than a prevalence estimate. The paper distinguishes model age from claim currency and proposes six reporting practices. Second, it treats its own two-day production process as a reflexive case of frontier-model-assisted research creation. GPT-5.6 Sol Pro in ChatGPT supported candidate discovery, source reconciliation, calculations, drafting and critique; the author checked sources, made all substantive decisions and accepts responsibility. This is a proof-of-practice, not a controlled estimate of productivity or quality. By applying its own Model Facts and model-currency statement, the paper shows how rapid AI-assisted research can be made inspectable without treating model output as independent validation. The title uses half-lives metaphorically; no universal decay rate is estimated.
Chinese Translation
生成性人工智能评估在出版前可能就已成为历史,然而日历年龄并不对每个结论产生同等影响。本文有两个相关的目的。首先,审计了2025年7月18日至2026年7月17日之间出现的40条经验记录的最大变异目的语料库。审计编码了出版途径、执行时间、模型身份、最新命名生成或不可变快照的年龄、同一家族的取代和刷新行为。在出现时,最新命名模型的中位年龄为281天(中间50%:75-478;范围:11-939)。25篇期刊文章的中位年龄为395天,14篇预印本为56天,1份实验室报告为49天。35条记录包含了一个被取代的家族,7条提供了精确的日期标识符,3条清晰地刷新了模型证据,1条增加了晚期敏感性测试。所有40条记录均包含OpenAI系统,这是该语料库的一个特征,而非普遍性估计。本文将模型年龄与主张货币区分开,并提出六项报告实践。其次,本文将自身的两天生产过程视为前沿模型辅助研究创作的反思案例。GPT-5.6 Sol Pro在ChatGPT中支持候选发现、来源协调、计算、草拟和批评;作者检查了来源,做出了所有实质性决策并承担责任。这是实践证明,而不是对生产力或质量的控制估计。通过应用自身的模型事实和模型货币声明,本文展示了如何在不将模型输出视为独立验证的情况下,使快速的人工智能辅助研究可供检查。标题以半衰期作为隐喻;未估计任何普遍衰减率。
cs.AI / 151 / 2607.24049

Quantum-Inspired Evolutionary Neighborhood Search for Arrival-Departure Track Utilization Adjustment under Short-Term Disturbances

量子启发的进化邻域搜索用于短期干扰下到达-离开轨道利用调整
Li, Xiaobin, Lei, Wuming, Gao, Yanbin, Wang, Weiguang
Abstract
Short-term disturbances at major passenger railway stations alter train arrival and departure times as well as the release sequence of station resources. Effective recovery therefore requires coordinated adjustment of arrival-departure track allocation, station resource occupation, and train retiming. This study represents the station resources involved in train arrival, track occupancy, and departure operations as zone-level resource-occupation intervals. An arrival-departure track allocation adjustment model is formulated. Resource compatibility is imposed as the feasibility condition, while train delays and resource reassignment costs are jointly considered. A quantum-inspired evolutionary algorithm combined with neighborhood search (QEA-NS) is proposed to solve the model. Perturbation instances are constructed using GTFS timetable data from Frankfurt Hauptbahnhof, Germany. QEA-NS is compared with CP-SAT under the same candidate resource set and feasibility criteria. Both methods generate solutions satisfying the modeled resource compatibility constraints. QEA-NS yields a total delay of 388 min, compared with 519 min for CP-SAT, representing a reduction of 25.2\%. The mean delay of delayed trains decreases from 4.99 to 3.73 min, although QEA-NS requires a longer solution time. Across 10 random perturbation instances, QEA-NS achieves lower total delay in every case. Its mean total delay and standard deviation are 390.5 min and 35.945 min, respectively, compared with 673.8 min and 105.739 min for CP-SAT. The results indicate that, under the adopted resource representation and constraints, QEA-NS improves the delay performance of recovery plans. Its computational efficiency, however, requires further improvement.
Chinese Translation
主要客运铁路车站的短期干扰会改变列车的到达和离开时间以及车站资源的释放顺序。因此,有效的恢复需要协调调整到达-离开轨道分配、车站资源占用和列车重调度。本研究将涉及列车到达、轨道占用和离开操作的车站资源表示为区域级资源占用区间。建立了一个到达-离开轨道分配调整模型。资源兼容性被作为可行性条件,而列车延误和资源重新分配成本则被共同考虑。提出了一种结合邻域搜索的量子启发进化算法(QEA-NS)来解决该模型。使用德国法兰克福中央车站的GTFS时刻表数据构建了扰动实例。在相同候选资源集和可行性标准下,将QEA-NS与CP-SAT进行了比较。两种方法均生成满足建模资源兼容性约束的解决方案。QEA-NS的总延误为388分钟,而CP-SAT为519分钟,减少幅度为25.2%。尽管QEA-NS的解决时间较长,但延误列车的平均延误从4.99分钟下降至3.73分钟。在10个随机扰动实例中,QEA-NS在每个案例中都实现了更低的总延误。其平均总延误和标准差分别为390.5分钟和35.945分钟,而CP-SAT为673.8分钟和105.739分钟。结果表明,在采用的资源表示和约束条件下,QEA-NS改善了恢复计划的延误表现。然而,其计算效率仍需进一步提高。
cs.AI / 152 / 2607.24054

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

成功并非自我解释:代理评估中的成功来源审计
Luo, Jingkun, Peng, Da-Tian
Abstract
A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes intended reasoning from answer acquisition. Outcome evidence and exposure detection do not establish whether success depended on an acquired target; we call this missing evaluation object success provenance. AcquaBench audits it through matched CLEAN, GOLD, and SHAM value substitution on four standardized surfaces with joint qid-clustered analysis. CLEAN retains benchmark-authorized information. GOLD makes the correct target available. SHAM preserves source structure and exposure opportunity but substitutes a matched incorrect value. GOLD minus CLEAN measures the total score response to correct-target availability; GOLD minus SHAM tests whether that response tracks target correctness beyond matched source exposure. In D0, GOLD exceeds SHAM by 19.1 to 25.9 percentage points, showing that success follows the correct value. In D2, GOLD still exceeds SHAM under distributed sufficiency while coloc no longer transfers as a high-score marker, with AUROC 0.376 and 0.142. Behavioral dependence can thus persist beyond this probe's intended observation unit. In model comparison, a supported 5.0-point CLEAN score gap compresses to a raw GOLD difference of -0.6 points without establishing rank inversion. Agent benchmarks should report success together with whether the evaluated information state supported it.
Chinese Translation
一个正确的答案可能掩盖了代理成功的原因。一旦代理在评估过程中改变了其信息状态,正确性便无法区分意图推理与答案获取。结果证据和曝光检测并不能确定成功是否依赖于获取的目标;我们称这种缺失的评估对象为成功来源(success provenance)。AcquaBench通过在四个标准化表面上进行匹配的CLEAN、GOLD和SHAM值替换,并结合联合qid聚类分析来审计它。CLEAN保留基准授权的信息。GOLD使正确目标可用。SHAM保留源结构和曝光机会,但替换为匹配的不正确值。GOLD减去CLEAN测量对正确目标可用性的总分响应;GOLD减去SHAM测试该响应是否超越匹配源曝光跟踪目标正确性。在D0中,GOLD超过SHAM 19.1到25.9个百分点,表明成功遵循正确值。在D2中,GOLD在分布充分性下仍然超过SHAM,而coloc不再作为高分标记转移,AUROC为0.376和0.142。因此,行为依赖性可以超越该探测的预期观察单位。在模型比较中,支持的5.0分CLEAN分数差压缩为原始GOLD差值-0.6分,而未建立排名反转。代理基准应报告成功以及评估的信息状态是否支持该成功。
cs.AI / 153 / 2607.24063

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

知晓的代价:一种资源感知协议用于超越静态排行榜的幻觉基准测试
Li, Keyu, Gao, Jin, Wang, Dequan
Abstract
On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs. Static leaderboards score factuality in isolation and treat compute as free, so they cannot tell a genuinely better system apart from one that simply spends more. Consider a ranking reversal. A brute-force Best-of-4 agent posts the higher raw factuality score (H-Score 0.9169 vs 0.9103) and would top a static leaderboard, but once cost is counted it is the worse system, losing on Q-Score (0.5169 vs 0.5217) at roughly four times the tokens and latency, under a reported cost weight whose sensitivity we sweep. So the system that tops a static leaderboard can be the worse one to deploy. To make this trade-off visible, we introduce MAS-HQ (Multi-Agent System Hallucination Quest), a resource-aware evaluation protocol. It wraps any factuality detector and normalizes for cost, and it pits systems against each other rather than scoring them in isolation. The Q-Score measures factuality minus normalized cost under a competitive match. Across summarization and open-domain QA, single-agent baselines drift into resource-heavy over-optimization, while competition elicits more resource-efficient policies. These gains are small but consistent, and stable across 100 trials. The axis stays discriminative for frontier systems (Gemini-2.5-Pro, and GPT-5 in simulated preview) whose raw factuality scores are already bunched near the ceiling. MAS-HQ provides a reproducible way to measure how much a factual answer costs.
Chinese Translation
在标准事实性任务中,前沿模型现在聚集在量表的顶部。因此,问题的焦点正在从系统的事实性转向这种事实性所需的计算成本。静态排行榜孤立地评分事实性,并将计算视为免费的,因此无法区分真正更好的系统与仅仅花费更多的系统。考虑一个排名反转。一种暴力破解的Best-of-4代理发布了更高的原始事实性得分(H-Score 0.9169对0.9103),并将在静态排行榜中名列前茅,但一旦计算成本被考虑,它实际上是更糟糕的系统,在Q-Score(0.5169对0.5217)上落后,消耗的令牌和延迟大约是前者的四倍,且在我们所扫掠的报告成本权重下。因此,位于静态排行榜顶部的系统可能是部署时更糟糕的选择。为了使这种权衡变得可见,我们引入了MAS-HQ(多智能体系统幻觉探索),一种资源感知的评估协议。它包裹任何事实性检测器并对成本进行标准化,同时将系统相互对抗而不是孤立评分。Q-Score在竞争匹配下测量事实性减去标准化成本。在摘要和开放域问答中,单代理基线趋向于资源密集型的过度优化,而竞争则引发了更具资源效率的策略。这些收益虽然微小但一致,并在100次试验中保持稳定。该轴线对前沿系统(Gemini-2.5-Pro和在模拟预览中的GPT-5)保持区分性,其原始事实性得分已接近上限。MAS-HQ提供了一种可重复的方法来衡量一个事实答案的成本。
cs.AI / 154 / 2607.24074

MiSS: A Logic-Driven Explanation of Minimal Sufficient Coalitions for Point Cloud Classifiers

MiSS:一种基于逻辑的最小充分联盟解释方法用于点云分类器
Xing, Mengda, Lagniez, Jean-Marie
Abstract
We present MiSS, a black-box, query-based framework for explaining 3D point cloud classifiers through perturbation-relative sufficiency reasoning. MiSS treats a superpoint partition as an interpretable abstraction layer and asks whether the original prediction can be certified from a minimal coalition of geometric regions under a specified perturbation distribution. Unlike abductive explainers that require Boolean feature spaces or white-box logical encodings of the predictor, MiSS separates candidate proposal from verification: a weighted MaxSAT procedure proposes coalitions using a heuristic adaptive cardinality floor, certified exact-size fallback, a safely tightened upper bound, blocking clauses, and a surrogate acquisition heuristic learned from previous oracle evaluations, while a blackbox statistical oracle decides sufficiency from prediction queries. The system returns a statistically verified sufficient coalition as a binary attribution, with minimum cardinality guaranteed when certified search completes. Experiments on ModelNet40 and ShapeNet with PointNet and PointMLP classifiers show higher precision and coverage than rule-based baselines in most settings, with lower explanation time than exhaustive search.
Chinese Translation
我们提出了MiSS,一个基于查询的黑箱框架,通过扰动相对充分性推理来解释3D点云分类器。MiSS将超点划分视为一个可解释的抽象层,并询问在指定的扰动分布下,是否可以从几何区域的最小联盟中认证原始预测。与需要布尔特征空间或预测器的白箱逻辑编码的推理器不同,MiSS将候选提案与验证分开:加权MaxSAT过程使用启发式自适应基数下限、认证的精确大小回退、安全收紧的上界、阻塞子句以及从先前的oracle评估中学习的替代获取启发式来提出联盟,而黑箱统计oracle则通过预测查询决定充分性。该系统返回一个经过统计验证的充分联盟作为二元归因,确保在认证搜索完成时具有最小基数。在ModelNet40和ShapeNet上使用PointNet和PointMLP分类器的实验表明,在大多数设置中,其精度和覆盖率高于基于规则的基线,并且解释时间低于穷举搜索。
cs.AI / 155 / 2607.24082

Towards High-Level Semantic Intelligence

迈向高层次语义智能
Song, Xiujie, Yang, Gefei, You, Yining, Gan, Jiahui, Jia, Qi, Watanabe, Shota, Wan, Tianxi, Wu, Mengyue, Yu, Kai
Abstract
Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems mainly addressed tasks involving direct and literal semantic perception or expression, contemporary systems are increasingly expected to perform more sophisticated cognitive reasoning, enabling the understanding and generation of High-Level Semantics (HLS). A similar trajectory can also be observed in human cognitive development. We define this transition as the shift from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI). However, this issue has not yet been systematically and comprehensively examined in prior work. Motivated by this gap, this survey reviews the development of AI semantic intelligence from the perspective of semantic complexity. We systematically survey existing research on HLS tasks, including humor, sarcasm, metaphor, empathy, persuasion, narrative, and other general HLS phenomena, across text, speech, vision, and multimodal scenarios. Specifically, we summarize data construction methods, modeling and optimization strategies, and evaluation methodologies for both understanding and generation. HLS is essential for advancing AI toward genuinely human-like intelligence. By synthesizing existing methods and insights from the perspective of semantic intelligence, this survey aims to support the continued development of AI toward HLSI.
Chinese Translation
近年来,人工智能(AI)的进步显著扩展了其认知和推理能力。从语义复杂性的角度来看,人工智能的发展显示出从简单到复杂的语义处理的明确轨迹。早期的人工智能系统主要处理涉及直接和字面语义感知或表达的任务,而当代系统则越来越被期望执行更复杂的认知推理,从而能够理解和生成高层次语义(High-Level Semantics, HLS)。在人类认知发展中也可以观察到类似的轨迹。我们将这一转变定义为从基础层次语义智能(Basic-Level Semantic Intelligence, BLSI)到高层次语义智能(High-Level Semantic Intelligence, HLSI)的转变。然而,之前的研究尚未对这一问题进行系统和全面的考察。基于这一空白,本调查从语义复杂性的角度回顾了人工智能语义智能的发展。我们系统性地调查了现有关于高层次语义任务的研究,包括幽默、讽刺、隐喻、同理心、说服、叙事以及其他一般的高层次语义现象,涵盖文本、语音、视觉和多模态场景。具体而言,我们总结了理解和生成的数据信息构建方法、建模和优化策略,以及评估方法论。高层次语义对于推动人工智能向真正的人类智能迈进至关重要。通过从语义智能的角度综合现有方法和见解,本调查旨在支持人工智能向高层次语义智能的持续发展。
cs.AI / 156 / 2607.24097

MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

MemChain:为增强记忆的LLM代理学习可解释的记忆痕迹
Ma, Yiwen, Tu, Songjun, Zhang, Qichao, Li, Dong, Li, Linjing, Zhao, Dongbin
Abstract
Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This retrieval-as-evidence paradigm assumes retrieved memories are already suitable for reasoning, leaving the answer model to resolve redundancy, conflicts, and weak relevance while incurring substantial context overhead in long-term memory tasks. We propose MemChain, a trainable post-retrieval memory policy that transforms retrieved candidates into answer-facing active memory, represented as a compact and grounded evidence context. Given a user query and retrieved candidates, MemChain first generates a question-conditioned evidence plan, then constructs an ordered grounded evidence trace that organizes retrieved memories according to their semantic roles and dependencies, and finally executes explicit memory actions to produce a concise evidence context for answer generation. To train the mediator, we introduce a two-stage learning framework. Supervised trace learning first teaches the policy to generate structurally valid plans, traces, actions, and evidence contexts. We then propose Trace-Guided Memory Policy Optimization (TMPO), a reinforcement learning objective that optimizes the memory policy using downstream answer quality while jointly encouraging trace grounding, evidence support, structural validity, and answer stability across multiple rollouts. Experiments on LoCoMo and LongMemEval-S demonstrate that MemChain consistently achieves state-of-the-art performance across both closed-source and open-weight frozen answer models while substantially reducing the memory context passed to the answer model.
Chinese Translation
增强记忆的LLM代理通常通过检索相关记忆并直接将其输入到答案模型中来回答查询。这种检索作为证据的范式假设检索到的记忆已经适合推理,留给答案模型解决冗余、冲突和弱相关性的问题,同时在长期记忆任务中产生大量的上下文开销。我们提出了MemChain,一种可训练的后检索记忆策略,它将检索到的候选记忆转化为面向答案的主动记忆,表现为一个紧凑且有依据的证据上下文。给定用户查询和检索到的候选记忆,MemChain首先生成一个基于问题的证据计划,然后构建一个有序的有依据的证据痕迹,根据语义角色和依赖关系组织检索到的记忆,最后执行明确的记忆操作,以生成简洁的证据上下文用于答案生成。为了训练中介,我们引入了一个两阶段的学习框架。监督痕迹学习首先教会策略生成结构有效的计划、痕迹、操作和证据上下文。然后,我们提出了基于痕迹的记忆策略优化(TMPO),这是一种强化学习目标,利用下游答案质量优化记忆策略,同时共同鼓励痕迹的基础、证据支持、结构有效性和多次回合中的答案稳定性。在LoCoMo和LongMemEval-S上的实验表明,MemChain在闭源和开放权重冻结答案模型上始终实现了最先进的性能,同时显著减少了传递给答案模型的记忆上下文。
cs.AI / 157 / 2607.24112

Scaling GUI Agents with Visual State Transitions

通过视觉状态转移扩展图形用户界面代理
Liu, Xiangyan, Li, Kaixin, Wang, Haonan, Wu, Biao, Fang, Meng, Dou, Longxu, Du, Chao, Shieh, Michael Qizhe, Pang, Tianyu
Abstract
We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from state changes) and forward dynamics (predicting next states from current states and actions). This optimization equips the model with better action-grounded visual representations and an internal world model of GUI dynamics. When subsequently fine-tuned on trajectories with task instructions, our STP-trained models consistently outperform baselines trained solely via direct trajectory fine-tuning across agent benchmarks in both desktop and mobile GUI scenarios (AgentNetBench, AndroidControl, and GUIOdyssey). Further empirical studies show that joint dynamics optimization yields stable improvements over single-objective training, and downstream performance scales steadily with the volume of transition data.
Chinese Translation
我们提出了一种新的图形用户界面(GUI)代理扩展轴——状态转移预训练(State Transition Pretraining, STP)。在STP阶段,我们通过联合优化逆动态(从状态变化预测动作)和正向动态(从当前状态和动作预测下一个状态),持续对统一的多模态模型进行预训练,专注于视觉状态转移。这种优化使模型具备了更好的基于动作的视觉表征和GUI动态的内部世界模型。在随后对带有任务指令的轨迹进行微调时,我们的STP训练模型在桌面和移动GUI场景(AgentNetBench、AndroidControl和GUIOdyssey)的代理基准测试中,始终优于仅通过直接轨迹微调训练的基线模型。进一步的实证研究表明,联合动态优化相较于单一目标训练能够带来稳定的性能提升,并且下游性能随着转移数据量的增加而稳步提升。
cs.AI / 158 / 2607.24117

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

叙述者评分:一种用于多智能体知识系统中声明级来源的伊斯纳德-里贾尔框架
Raja, Ali Zahid
Abstract
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing. Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator's integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design. We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support.
Chinese Translation
现代多智能体知识系统越来越多地通过自主转化链积累知识,而非直接检索。现有的来源研究记录了发生的事情——执行轨迹、工具调用、证据链接——并且源可靠性估计早已建立(真相发现、声誉系统)。缺少的是一个操作框架,它将分级的、按领域的传输者可靠性附加到声明级传输链上,具备完整性语义、转化类型聚合、解耦内容批评以及服务/审查/隔离路由。古典伊斯兰教哈迪斯科学面临着一个结构上类似的问题:决定通过人类叙述者链传递的知识是否应被接受。几个世纪以来,它发展出了一套严格的方法论——伊斯纳德(附加于每个声明的完整传输链)、里贾尔(系统性评分每个叙述者的诚信和准确性)、最弱链评估、通过独立链的证实以及马特恩批评(内容独立于链质量进行评估)。本文将该方法论转移到人工智能系统设计中。我们贡献了从哈迪斯科学概念到多智能体管道的正式映射、实现声明链和分级叙述者注册的关系模式、结合链评分与内容批评的决策矩阵,以及对来自真实物理教科书的20,000个声明的评估。评估验证了最弱链隔离和独立链证实;报告了评分恢复循环的部分失败,未能识别出最高故障叙述者;并且报告了两个分析结果不确定,包括框架无法与参考内容批评进行匹配覆盖比较。本文明确指出了证据支持和不支持的声明。
cs.AI / 159 / 2607.24148

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

一种基于运动感知的向量量化框架,具有质心重用以实现高效的视觉-语言-动作推理
Song, Zhuoran, Jiang, Haozhe, Qi, Chunyu, Pei, Minnan, Li, Gang, Liang, Xiaoyao, Guan, Haibing
Abstract
Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited. In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics. We first introduce MotionVQ, a motion-aware vector quantization scheme that dynamically adjusts quantization precision based on the robot's execution state, reducing memory access while preserving task success rate. We then propose a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation, eliminating redundant multiplications through spatial aggregation and temporal reuse of centroids. To realize these optimizations, we design an accelerator that efficiently supports dynamic precision selection and centroid-reuse computation. Experimental results show that VQVLA achieves 6.5x, 2.8x, 1.9x, 3.3x, and 4.3x speedup over the A100 GPU, Dadu-Corki, LUT-DLA, CodeGEMM, and ShiftAddLLM, respectively, with negligible accuracy degradation.
Chinese Translation
视觉-语言-动作(VLA)模型在具身人工智能中展现出强大的潜力,但其在GPU上的高推理延迟限制了实时部署。现有的加速器,如Dadu-Corki,虽然提高了效率,但将VLA模型视为全精度工作负载,导致内存和计算中存在大量冗余未被充分利用。本文提出了VQVLA,一个算法-硬件协同设计框架,通过利用权重相似性和执行动态来加速VLA推理。我们首先介绍了MotionVQ,一种运动感知的向量量化方案,根据机器人的执行状态动态调整量化精度,减少内存访问,同时保持任务成功率。然后,我们提出了一种合并质心的向量化GEMM范式,该范式在代码本索引表示上操作,通过空间聚合和质心的时间重用消除冗余乘法。为了实现这些优化,我们设计了一种加速器,有效支持动态精度选择和质心重用计算。实验结果表明,VQVLA在A100 GPU、Dadu-Corki、LUT-DLA、CodeGEMM和ShiftAddLLM上分别实现了6.5倍、2.8倍、1.9倍、3.3倍和4.3倍的加速,且准确性下降微乎其微。
cs.AI / 160 / 2607.24162

Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness

Agent-UCT:应用于树的上置信界限用于具有成本意识的代理工作流优化
Li, Yang, Liu, Hai, Shao, Dian, Wang, Yu, Chen, Xiyu, Volkov, Sergey, Wang, Bozhi, Sun, Ziyu, Liu, Sihang, Luo, Ye, Zhang, Xiaowei
Abstract
Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimization, and standard tree search methods - do not explicitly exploit the compositional structure of these workflows, leading to redundant computation and inefficient budget allocation. We introduce Agent-UCT (Agent-based Cost-Aware Upper Confidence Bounds Applied to Trees), a tree search algorithm that extends UCT with a reuse-aware regularization term derived from a bipartite prefix reuse graph. Agent-UCT biases selection toward branches that leverage previously materialized configuration prefixes, reducing redundant execution while maintaining effective exploration. Our framework, RAGSpace, unifies heterogeneous RAG components from LongRAG, LightRAG, and Self-RAG into a five-dimensional configuration space, enabling systematic cross-framework recombination. WTB (Workflow Test Bench) provides deterministic replay, content-addressable caching, and transactional consistency, ensuring that intermediate states are materialized once and reused across the search. Experiments on HotpotQA and UltraDomain demonstrate that Agent-UCT identifies configurations with the highest out-of-sample performance among the evaluated fixed framework presets. Under full-pool evaluation, bipartite prefix reuse reduces logical search cost by 73.6% relative to the no-prefix-sharing cost upper bound. Compared with full-pool evaluation, sampling-based evaluation further achieves a 4.2x wall-clock speedup. Agent-UCT, RAGSpace, and WTB together provide a unified framework for cost-aware, reproducible, and compositionally efficient agentic workflow optimization.
Chinese Translation
优化代理工作流,例如检索增强生成(RAG)管道,需要在严格的评估预算下导航离散组件选择的组合空间。现有方法——启发式搜索、黑箱优化和标准树搜索方法——并未明确利用这些工作流的组合结构,导致冗余计算和低效的预算分配。我们提出了Agent-UCT(基于代理的成本意识上置信界限应用于树),这是一种树搜索算法,通过从二分前缀重用图中派生的重用意识正则化项扩展了UCT。Agent-UCT在选择时偏向于利用先前实现的配置前缀的分支,从而减少冗余执行,同时保持有效的探索。我们的框架RAGSpace将来自LongRAG、LightRAG和Self-RAG的异构RAG组件统一到一个五维配置空间中,实现系统的跨框架重组。WTB(工作流测试平台)提供确定性重放、内容可寻址缓存和事务一致性,确保中间状态仅被实现一次并在搜索中重用。在HotpotQA和UltraDomain上的实验表明,Agent-UCT能够识别在评估的固定框架预设中具有最高样本外性能的配置。在全池评估下,二分前缀重用相对于无前缀共享成本上限减少了73.6%的逻辑搜索成本。与全池评估相比,基于采样的评估进一步实现了4.2倍的实际时间加速。Agent-UCT、RAGSpace和WTB共同提供了一个统一的框架,用于成本意识、可重现和组合高效的代理工作流优化。
cs.AI / 161 / 2607.24167

Falsifiable Commitment Planning for Self-Correcting Web Agents

可证伪的承诺规划用于自我修正的网络代理
Liu, Guangyi, Zhao, Huan, Yao, Quanming
Abstract
Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, or plan assumption no longer supports the user instruction. Existing agents can plan, reflect, or reuse experience, but their plans rarely specify the evidence under which an active step should still be trusted. We propose FCPAgent, a falsifiable commitment planning framework for robust long-horizon web agents. FCPAgent represents each plan step as a Falsifiable Commitment Unit (FCU): a subgoal grounded in a reusable skill, together with confirming evidence, falsifying evidence, and a confidence score. Execution is organized as a plan-test-repair loop. The hybrid commitment testing module checks candidate actions before they modify the browser and checks observations after execution; for efficiency, it combines lightweight evidence matching with LLM-based diagnostic verification. When evidence falsifies a commitment, scope-aware repair localizes the contradiction to the execution, skill, or planning level and revises the smallest adequate part. On WebArena, FCPAgent achieves a 13.8% relative improvement in average success over the strongest baseline, with especially large gains on long-horizon tasks.
Chinese Translation
长时间跨度的网络代理在最终失败之前常常偏离轨道:即使当前状态、重用的技能或计划假设不再支持用户指令,轨迹仍然可能在局部上看似合理。现有的代理可以进行规划、反思或重用经验,但它们的计划很少明确指定在何种证据下,当前步骤仍然值得信任。我们提出了FCPAgent,一种用于强健的长时间跨度网络代理的可证伪承诺规划框架。FCPAgent将每个计划步骤表示为一个可证伪承诺单元(Falsifiable Commitment Unit, FCU):一个基于可重用技能的子目标,以及确认证据、反驳证据和置信度评分。执行过程组织为计划-测试-修复循环。混合承诺测试模块在候选动作修改浏览器之前检查它们,并在执行后检查观察结果;为了提高效率,它结合了轻量级证据匹配与基于大型语言模型(LLM)的诊断验证。当证据反驳承诺时,范围感知修复将矛盾局限于执行、技能或规划层面,并修订最小的适当部分。在WebArena上,FCPAgent在平均成功率上相较于最强基线实现了13.8%的相对提升,尤其在长时间跨度任务上取得了显著的进步。
cs.AI / 162 / 2607.24187

Myopia Prevention and Control 3.0: Artificial Intelligence--Driven Risk Stratification, Proactive Monitoring, and Personalized Intervention

近视预防与控制 3.0:基于人工智能的风险分层、主动监测与个性化干预
Wang, Tieniu, Huang, Cangzhu, Li, Qianhui
Abstract
The convergence of artificial intelligence (AI), digital sensing, and ubiquitous computing has created an unprecedented opportunity to transform myopia prevention from a reactive, population-based model into a proactive, precision-driven one. Despite evidence that half the world's population will be myopic by 2050, conventional approaches---school-based vision screening (Phase 1.0) and evidence-based risk factor management (Phase 2.0)---have proven insufficient. We review the emergence of Myopia Prevention and Control 3.0, defined by AI integration across three interconnected domains forming a closed-loop pipeline: (1) AI-driven risk stratification predicting individual-level risk through machine learning on multimodal data; (2) AI-enabled proactive monitoring via wearables, smartphones, and school screening networks; and (3) AI-powered personalized intervention with closed-loop feedback. We critically evaluate evidence across each stage, discuss challenges in data quality, model validation, ethics, and equity, and outline future directions including multimodal foundation models, digital twins, and causal machine learning.
Chinese Translation
人工智能(AI)、数字传感和无处不在的计算的融合,为将近视预防从反应式、基于人群的模型转变为主动式、精准驱动的模型创造了前所未有的机会。尽管有证据表明到2050年全球一半人口将会近视,但传统方法——基于学校的视力筛查(阶段 1.0)和基于证据的风险因素管理(阶段 2.0)——已被证明不足以应对这一挑战。我们回顾了近视预防与控制 3.0 的出现,该阶段的特点是人工智能在三个相互关联的领域中的整合,形成一个闭环管道:(1) 基于人工智能的风险分层,通过对多模态数据进行机器学习预测个体风险;(2) 通过可穿戴设备、智能手机和学校筛查网络实现的人工智能驱动的主动监测;(3) 结合闭环反馈的人工智能驱动的个性化干预。我们对每个阶段的证据进行了批判性评估,讨论了数据质量、模型验证、伦理和公平性方面的挑战,并概述了未来的发展方向,包括多模态基础模型、数字双胞胎和因果机器学习。
cs.AI / 163 / 2607.24213

Integrating Factual and Normative Industrial Knowledge via Constraint-Aware Graph Attention for Process Plan Recommendation

通过约束感知图注意力整合事实与规范工业知识以进行工艺计划推荐
Chen, Yuntong, Li, Yingqi, Xiao, Yingying, Wang, Ziang, Liu, Zewei, Liu, Jiahao, Tian, Xitian, Huang, Lijiang
Abstract
Integrating heterogeneous industrial knowledge, including factual relations and decision constraints, remains a core challenge in industrial information systems. Machining process planning exemplifies this problem because engineers must select operations by combining material properties, feature characteristics, and quality requirements. Existing methods rely mainly on similarity retrieval or classification, without a unified ranking objective or standardized evaluation. We propose PCA-GAT, which formulates machining process plan recommendation as a knowledge graph enhanced collaborative filtering problem. Bayesian Personalized Ranking provides the learning objective, while Recall@K and NDCG@K define evaluation. The knowledge graph supplies semantic structure when collaborative signals are sparse. Four domain constraints, material compatibility, precision requirements, feature applicability, and operation sequencing, are introduced as attention biases during graph propagation. Type-specific weights learn their importance, and an adaptive gate adjusts their influence using local context. On a real aerospace dataset with 115 parts and 507 plans, PCA-GAT achieves Recall@1 = 0.9087 and strong cold-start robustness, with about half the degradation of the strongest baseline under severe sparsity. Ablation studies show that knowledge graph enrichment is essential, constraints add value, and ungated constraint injection can hurt performance. The learned weights identify material-operation compatibility as the dominant factor, consistent with domain expertise. Results on three public benchmarks show no degradation when constraints are absent, supporting generalization beyond manufacturing. This study establishes a standardized recommendation protocol for engineering process planning and benchmarks seven methods across three categories, showing that knowledge representation is the main bottleneck.
Chinese Translation
整合异构工业知识,包括事实关系和决策约束,仍然是工业信息系统中的一个核心挑战。机械加工过程规划就是一个典型例子,因为工程师必须通过结合材料属性、特征特性和质量要求来选择操作。现有方法主要依赖于相似性检索或分类,没有统一的排序目标或标准化评估。我们提出了PCA-GAT,将机械加工过程计划推荐公式化为一个知识图谱增强的协同过滤问题。贝叶斯个性化排序提供学习目标,而Recall@K和NDCG@K定义评估标准。知识图谱在协同信号稀疏时提供语义结构。引入四个领域约束:材料兼容性、精度要求、特征适用性和操作顺序,作为图传播过程中的注意力偏置。特定类型的权重学习其重要性,自适应门控使用局部上下文调整其影响。在一个包含115个零件和507个计划的真实航空航天数据集中,PCA-GAT实现了Recall@1 = 0.9087,并在严重稀疏情况下展现出强大的冷启动鲁棒性,其性能下降约为最强基线的一半。消融研究表明,知识图谱的丰富化至关重要,约束增加了价值,而无门控的约束注入可能会损害性能。学习到的权重识别材料-操作兼容性为主导因素,这与领域专业知识一致。在三个公共基准上的结果表明,当缺少约束时没有性能下降,支持超越制造的推广。本研究建立了一个工程过程规划的标准化推荐协议,并在三个类别中基准测试了七种方法,显示知识表示是主要瓶颈。
cs.AI / 164 / 2607.24243

Epistemic Norms for AI Safety and Alignment Research

人工智能安全与对齐研究的认识规范
Navaie, Keivan
Abstract
Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes---{\it capability profile}, demonstrating the absence of hazardous behaviours rather than the presence of positive capabilities, and {\it risk profile}, bounding worst-case outcomes under fat-tailed uncertainty rather than optimising average-case performance---and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps, we propose {\sc ECAISA}, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an information-hazard adjudication procedure, and seven anti-gaming mechanisms. {\sc ECAISA} does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target.
Chinese Translation
主流人工智能研究强调能力增长,并在平均性能较高时容忍低失败率。而人工智能安全与对齐研究则有不同的使命:确保在稀疏证据、对抗性动态和厚尾风险下,灾难性失败绝不发生。我们认为这两个领域在两个分析上独立的轴线上存在差异——{ extit 能力特征},强调缺乏危险行为而非积极能力的存在,以及{ extit 风险特征},在厚尾不确定性下界定最坏结果,而非优化平均性能——而主流的认识实践在这两方面均显得不足。基于一个结构化的综合,建立在预注册的文献计量基线之上,我们识别出当前对齐研究中的五个交叉缺口维度,包括几乎缺乏制度化的独立验证。为了解决这些缺口,我们提出了{ extsc ECAISA},即人工智能安全与对齐的认识规范,包含八项原则、一个三级评分标准、一个四级披露梯度,旨在调和透明度与信息风险和商业机密约束之间的关系、一个分级适用方案、一个信息风险裁定程序,以及七个反游戏机制。{ extsc ECAISA}并不认证任何人工智能系统是安全的;它限制了安全相关研究主张的记录、检查和依赖方式,以可审计性而非认证作为其治理目标。
cs.AI / 165 / 2607.24259

Generative Artificial Intelligence (GenAI) to convert images of queuing networks into verifiable simulation models: an open-weight LLM workflow approach

生成性人工智能(GenAI)将排队网络图像转换为可验证的仿真模型:一种开放权重大语言模型工作流程方法
Monks, Thomas, Harper, Alison, Heather, Amy, Mustafee, Navonil
Abstract
Recent work has explored the use of Large Language Models (LLMs) to automate simulation model building, typically by generating executable code directly from natural language descriptions. However, this raises challenges for verification and reproducibility particularly for users without programming expertise. We propose Sketch2DES, a sketch-to-simulation workflow that converts diagrammatic representations of queuing networks into verifiable discrete-event simulation models using open-weight LLMs. The workflow has three stages: (1) translation of a diagram into a semi-structured textual description using a multimodal LLM; (2) conversion into schema-validated structured data (JSON) via an LLM with a reflection-based verification loop; and (3) deterministic transformation into an executable simulation model using a software adapter. Intermediate artefacts can therefore be inspected and automatically validated before execution. We evaluate the approach on eight queuing-network diagrams of varying complexity. The workflow achieved high reliability for all stages, and results were statistically indistinguishable from human-coded and analytical benchmarks. Compared to direct code generation, the workflow improves reproducibility, transparency, and verifiability, while reducing the need for programming expertise. Limitations include restricted model scope and dependence on accurate visual interpretation. The results demonstrate the feasibility of structured, workflow-based model generation as a robust foundation for LLM-assisted simulation modelling.
Chinese Translation
近期的研究探讨了使用大型语言模型(LLMs)来自动化仿真模型构建,通常是通过直接从自然语言描述生成可执行代码。然而,这对缺乏编程专业知识的用户在验证和可重复性方面提出了挑战。我们提出了Sketch2DES,这是一种草图到仿真的工作流程,利用开放权重的LLMs将排队网络的图示表示转换为可验证的离散事件仿真模型。该工作流程分为三个阶段:(1)使用多模态LLM将图示翻译为半结构化文本描述;(2)通过具有反射式验证循环的LLM转换为模式验证的结构化数据(JSON);(3)使用软件适配器确定性地转换为可执行的仿真模型。因此,中间产物可以在执行之前进行检查和自动验证。我们对八个不同复杂度的排队网络图进行了评估。该工作流程在所有阶段都实现了高可靠性,结果在统计上与人类编码和分析基准无显著区别。与直接代码生成相比,该工作流程提高了可重复性、透明性和可验证性,同时减少了对编程专业知识的需求。局限性包括模型范围受限和对准确视觉解释的依赖。结果表明,基于结构化工作流程的模型生成作为LLM辅助仿真建模的稳健基础是可行的。
cs.AI / 166 / 2607.24280

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

从专有到开源:通过多智能体协议蒸馏弥合智能搜索中的分布差距
Liu, Junlin, Chen, Jiangwang, Song, Zixin, Zhou, Shuaiyu, Lv, Chunji, Wu, Hank, Jiang, Kailin, Wu, Jinyang, Yu, Bohan, Zhou, Chenxi
Abstract
Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded by hidden logits and mismatched tokenizers, whereas raw natural language trajectory imitation transfers superficial stylistic artifacts rather than core reasoning competence. To address the heterogeneous distillation problem and bridge the distribution gap, we propose Multi-Agent Protocol Distillation (MAPD), a joint distillation and RL framework uses a structured, style-normalized protocol as an intermediate representation. An offline multi-agent system (MAS) decomposes each query, retrieves supporting evidence, repairs failed searches, and converts the resulting exploration trace into a JSON protocol containing the task type, reasoning plan, and extractive grounding facts. During training, the protocol is provided only to a privileged branch of the student policy, whose token distributions furnish a dense distillation signal alongside the sparse RL objective. Extensive evaluations across seven QA benchmarks demonstrate that MAPD consistently outperforms competitive distillation and RL, achieving average success rates of 39.4\% on Qwen3-1.7B and 44.4\% on Qwen3-4B. Crucially, the framework generalizes robustly across diverse proprietary teachers while effectively mitigating the student policy from style drift and verbosity degeneration.
Chinese Translation
智能搜索使大型语言模型能够通过将多步骤推理与检索交织在一起来解决知识密集型任务,但通过基于结果的强化学习(RL)进行优化仅提供稀疏的监督。知识蒸馏可以提供更密集的指导,而具有强大推理能力的先进专有模型则是有前景的教师。尽管从专有模型中进行蒸馏可以使这种监督信号更加密集,但传统的对数匹配由于隐藏的对数和不匹配的标记器而受到限制,而原始自然语言轨迹模仿则转移了表面的风格特征,而非核心推理能力。为了解决异构蒸馏问题并弥合分布差距,我们提出了多智能体协议蒸馏(MAPD),这是一个联合蒸馏和强化学习框架,使用结构化、风格标准化的协议作为中间表示。一个离线多智能体系统(MAS)对每个查询进行分解,检索支持证据,修复失败的搜索,并将生成的探索轨迹转换为包含任务类型、推理计划和提取基础事实的JSON协议。在训练过程中,协议仅提供给学生策略的一个特权分支,其标记分布提供了密集的蒸馏信号,同时伴随稀疏的RL目标。在七个问答基准上的广泛评估表明,MAPD在竞争蒸馏和RL中始终表现优越,在Qwen3-1.7B上实现了39.4\%的平均成功率,在Qwen3-4B上实现了44.4\%。重要的是,该框架在多样化的专有教师之间具有强大的泛化能力,同时有效减轻了学生策略的风格漂移和冗长退化。
cs.AI / 167 / 2607.24336

Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination

不平等的出行,不平等的地点:诊断和缓解自主车辆车队协调中的延迟不平等
Hu, Nicole, Zhang, Mingtao, LI, Haoyang, Zhang, Chen Jason, Qing, Li
Abstract
City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed across trips and regions. We conduct a distributional audit on three real-city road-network and taxi-demand datasets from Manhattan, Chicago, and San Francisco. The audit reveals pervasive trip-length inequity whose direction depends on the city and coordinator. After accounting for trip length, spatial inequity becomes more pronounced as demand grows and is consistently stronger when trips are grouped by origin rather than destination. These findings motivate SPatially Aware RErouting (SPARE), a budgeted online coordination framework that assigns limited replanning capacity to delayed vehicles and redirects them using recently observed waiting pressure. SPARE provides a per-review decision guarantee and explicitly bounds online route updates. Experiments on all three datasets against six representative baselines show that SPARE delivers the strongest joint efficiency-fairness performance while retaining city-scale scalability. The results demonstrate that bounded congestion-responsive rerouting improves performance and equity without full-fleet replanning.
Chinese Translation
城市规模的自主车辆车队协调通常以总旅行时间为优化目标,但车队平均值掩盖了延迟在出行和区域间的分布情况。我们对来自曼哈顿、芝加哥和旧金山的三个真实城市道路网络和出租车需求数据集进行了分布审计。审计结果揭示了普遍存在的出行长度不平等,其方向取决于城市和协调者。在考虑出行长度后,随着需求的增长,空间不平等变得更加明显,并且当出行按起点而非终点分组时,这种不平等始终更为显著。这些发现促使我们提出了SPatially Aware RErouting (SPARE),一种预算在线协调框架,旨在将有限的重新规划能力分配给延迟车辆,并利用最近观察到的等待压力对其进行重新引导。SPARE提供了逐个审查决策的保证,并明确限制在线路线更新。对所有三个数据集进行的实验与六个代表性基准进行比较,结果表明SPARE在保持城市规模可扩展性的同时,提供了最强的联合效率-公平性能。结果表明,有限的拥堵响应性重新规划在不进行全车队重新规划的情况下改善了性能和公平性。
cs.AI / 168 / 2607.24339

Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families

Gubernaut:一种确定性的稳态控制器,用于情感调节的LLM代理,已在独立模型家族中验证
Sharma, Dushyant
Abstract
Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stuck. These are failures of propensity, not capability; they concern what a model does under sustained pressure, which training-time alignment reduces but does not eliminate at runtime. This research led to the Gubernaut Cognitive Controller (GCC), a model-agnostic runtime control layer in a Nelson--Narens monitoring--control loop: an object level reads and writes text, while a deterministic meta level reads only the numeric telemetry {intensity, valence, repetition} and returns a regulating posture. Because the meta level ingests zero tokens, no injection channel to the controller exists by construction (an architectural property, not yet adversarially tested); the text-exposed arbiter's compliance is measured, not assumed. We evaluate the GCC with a pre-registered, generate-once/judge-many protocol across a 4x4 matrix of four frontier models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3), each serving as both a generator and a judge. The regulated arm is calmer in 13 of 16 cells at p<.05 and 15 of 16 by sign; the three sub-threshold cells, including a -0.04 null, all fall on the single near-saturated host. The effect survives a lineage-independent fourth judge family (xAI), strong evidence that it is no artifact of shared judge style. The clearest mechanism is the recovery signature: arousal that integrates under attack and then decays, valence-gated, on de-escalation, replicating across all four families. Transcripts and panels ship with SHA-256 provenance and are re-judgeable; five failure modes are pre-registered. No consciousness claims are made.
Chinese Translation
大型语言模型(LLM)代理继承了反应性失败模式:在挑衅下升级,在奉承下趋向阿谀,陷入困境时的持续反复。这些是倾向性的失败,而非能力的失败;它们涉及模型在持续压力下的表现,训练时的对齐虽然能减少这种情况,但并不能在运行时消除。此次研究导致了Gubernaut认知控制器(GCC)的出现,它是一个模型无关的运行时控制层,位于Nelson-Narens监控-控制循环中:对象层读取和写入文本,而确定性的元层仅读取数值遥测{强度、效价、重复}并返回调节姿态。由于元层不摄取任何标记,因此控制器不存在注入通道(这是一个架构属性,尚未经过对抗性测试);文本暴露的仲裁者的合规性是通过测量而非假设得出的。我们通过一个预注册的生成一次/判断多次的协议评估GCC,涵盖四个前沿模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Flash、Grok 4.3)组成的4x4矩阵,每个模型既作为生成器又作为评判者。在16个单元中,受调节的臂在13个单元中表现得更平静(p<.05),在16个单元中有15个单元呈现出显著性;三个低于阈值的单元,包括一个-0.04的无效单元,均落在单一的近饱和主机上。该效应在一个与家族无关的第四评判者家族(xAI)中依然存在,强有力地证明它不是共享评判风格的伪影。最清晰的机制是恢复特征:在攻击下整合的唤醒,然后在去升级时以效价为门限衰减,这一现象在所有四个家族中均得以复制。转录和面板附带SHA-256来源信息,并可重新评判;五种失败模式已预注册。未提出任何意识的主张。
cs.AI / 169 / 2607.24341

Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age

利用考虑交易成本的LLM时代模拟租户对能源政策干预的响应
Xia, Weijie, Horian, Stefanie, Huang, Hanyue, Qian, Queena K., Yang, Jie, Barrios, Pedro P. Vergara
Abstract
Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal, or persona-based descriptions. Yet such simulations rarely model the practical, cognitive, or social frictions that shape how people respond to policy interventions. Perceived transaction cost (PTC) provides a useful lens for modeling the practical frictions that shape policy responses, such as information burden, administrative effort, coordination demands, and perceived uncertainty. We use this lens to develop a friction-aware persona modeling approach for LLM-based simulation. In the context of energy-efficient renovation (EER), tenants are represented not only by who they are demographically, but by how they perceive the costs, benefits, barriers, and uncertainties associated with proposed renovation plans. Using survey data collected from 1,068 citizens in the Netherlands, comprising approximately 40,548 survey question and answer pairs, we compare prompt-only and fine-tuned settings across GPT-3.5-turbo, Ministral-8B-Instruct, and Llama-3.1-8B-Instruct, and evaluate supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO) for local open-weight models. Results show that incorporating PTC-based personas and reasoning consistently improves model performance across both prompt-only and fine-tuned settings, suggesting that PTC-based persona design provides a useful bridge between institutional policy theory and interpretable LLM-based policy simulation. Code is available at https://github.com/xiaweijie1996/socialagent.
Chinese Translation
近期研究利用大型语言模型(LLMs)通过人口统计、态度或角色描述来模拟人类的意见和决策。然而,这类模拟很少考虑影响人们对政策干预响应的实际、认知或社会摩擦。感知交易成本(PTC)为建模影响政策响应的实际摩擦提供了有用的视角,例如信息负担、行政努力、协调需求和感知不确定性。我们利用这一视角开发了一种考虑摩擦的角色建模方法,用于基于LLM的模拟。在能源效率改造(EER)的背景下,租户不仅通过其人口统计特征来表示,还通过他们对提议的改造计划相关成本、收益、障碍和不确定性的感知来表示。我们使用从荷兰收集的1,068名公民的调查数据,包含约40,548对调查问题和答案,比较了在GPT-3.5-turbo、Ministral-8B-Instruct和Llama-3.1-8B-Instruct下的仅提示和微调设置,并评估了局部开放权重模型的监督微调(SFT)和群体相对政策优化(GRPO)。结果表明,结合基于PTC的角色和推理在提示仅和微调设置中均能持续提高模型性能,表明基于PTC的角色设计为制度政策理论与可解释的基于LLM的政策模拟之间提供了有用的桥梁。代码可在 https://github.com/xiaweijie1996/socialagent 获取。
cs.AI / 170 / 2607.24354

Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization

提示优化器是盲目的吗?用于自动提示优化的跨模态视觉反馈
Liu, Haoyue, Ma, Xiaoyu, Chen, Ye, Zou, Yuexian, Tang, Xiaoying
Abstract
Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However, on multimodal tasks, the effectiveness of APO is fundamentally bottlenecked by a blind feedback channel: the optimizer reads the question, the prediction, and the gold answer, but never the input image on which the model failed, and therefore cannot diagnose visually grounded errors. As a remedy, we introduce Cross-Modal Visual Feedback (CMVF). CMVF incorporates (1) a failure-conditioned visual diagnosis stage, in which a stronger optimizer VLM inspects each failed image without access to predictions or labels, and (2) an error-aware aggregation stage that compresses these observations into reusable, task-level visual blind-spot patterns that drive the prompt rewrite. Crucially, the image is consumed only during optimization; the deployed artifact is an ordinary text prompt that runs at the same inference cost as any text-only baseline. Extensive results across 12 VQA datasets and 4 target VLMs demonstrate that CMVF consistently ranks first, improving over the strongest baseline on every target by 2.4 points on average, with gains of up to 6.5 points on individual benchmarks. Moreover, the optimizer self-organizes into expert-style visual checklists that transfer across models without re-optimization.
Chinese Translation
自动提示优化(APO)已被广泛应用于将视觉-语言模型(VLMs)适应于下游任务,而无需更新权重,取得了良好的效果。然而,在多模态任务中,APO的有效性在根本上受到一个盲反馈通道的瓶颈:优化器读取问题、预测和真实答案,但从未读取模型失败时的输入图像,因此无法诊断视觉基础的错误。为了解决这个问题,我们引入了跨模态视觉反馈(CMVF)。CMVF包含(1)一个基于失败的视觉诊断阶段,其中一个更强的优化器VLM在没有访问预测或标签的情况下检查每个失败的图像,以及(2)一个错误感知聚合阶段,将这些观察压缩成可重用的任务级视觉盲点模式,以驱动提示重写。关键是,图像仅在优化期间被使用;部署的产物是一个普通的文本提示,其推理成本与任何仅文本的基线相同。在12个VQA数据集和4个目标VLM上的广泛结果表明,CMVF始终排名第一,平均在每个目标上比最强基线提高2.4分,在个别基准上提升高达6.5分。此外,优化器自我组织成专家风格的视觉检查清单,可以在不同模型之间转移而无需重新优化。
cs.AI / 171 / 2607.24419

Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

失败揭示了指标的不足:一种基于证据的心电图分类器递归优化代理
Deng, Jinliang, Niu, Yiming, Pan, Yibo, Shao, Zhiqi, Luo, Qin, Tong, Yongxin
Abstract
Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for automated model design, but when guided only by aggregate performance metrics, they lack insight into why individual cases fail and how the classifier should be revised. We present RecursiveECG, an evidence-driven LLM-as-Designer framework in which an LLM serves as an offline model designer that refines ECG classifiers based on concrete failures and objective ECG evidence. To ground failure diagnosis in executable evidence, Criteria-to-Measurement Compilation converts curated ECG criteria into validated deterministic functions that produce reproducible, reference-backed measurements for individual ECGs. Building on these measurements, Evidence-Grounded Failure Review analyzes failed and comparator cases by jointly considering raw waveforms, measurements, and model outputs, enabling the LLM to diagnose classifier limitations and formulate targeted revisions. Candidate revisions are executed and re-evaluated under a fixed problem contract, and only evidence-supported updates are retained. The resulting predictor is frozen after refinement and requires no LLM inference during deployment, while an audit trail links each accepted revision to its supporting evidence. Across PTB-XL, Georgia, and CPSC2018, RecursiveECG consistently outperforms strong baselines, achieving an average relative improvement of 10.0%. Extensive ablation and transfer studies further validate the effectiveness of its evidence-grounded refinement process.
Chinese Translation
深度模型在12导联心电图分类方面取得了显著进展,但其优化仍然高度依赖人类专家对失败案例的检查和分类器设计的迭代修订。最近基于大型语言模型(LLM)的代理展示了自动化模型设计的潜力,但仅依赖于汇总性能指标时,它们缺乏对个别案例失败原因的洞察,以及如何修订分类器的指导。我们提出了RecursiveECG,这是一种基于证据的LLM设计者框架,其中LLM作为离线模型设计者,根据具体的失败案例和客观的心电图证据来优化心电图分类器。为了将失败诊断与可执行的证据相结合,标准到测量编译(Criteria-to-Measurement Compilation)将策划的心电图标准转换为经过验证的确定性函数,这些函数为个别心电图生成可重复的、基于参考的测量。基于这些测量,基于证据的失败审查(Evidence-Grounded Failure Review)通过共同考虑原始波形、测量和模型输出,分析失败和对照案例,使LLM能够诊断分类器的局限性并制定针对性的修订。候选修订在固定问题合同下执行并重新评估,只有得到证据支持的更新才会被保留。经过优化后,生成的预测器被冻结,在部署过程中不需要LLM推理,同时审计轨迹将每个接受的修订与其支持证据关联。在PTB-XL、乔治亚州和CPSC2018数据集中,RecursiveECG始终优于强基线,平均相对提升达到10.0%。广泛的消融和迁移研究进一步验证了其基于证据的优化过程的有效性。
cs.AI / 172 / 2607.24459

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

从执行到能力:通过程序知识综合实现科学经验的巩固
Dong, Liwei, Zhao, Jiahao, Xu, Nan
Abstract
Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather than cross-task computational mechanisms; and a weaker target model may be unable to operationalize an otherwise valid abstract procedure - an abstraction-execution gap. We introduce SciConsolidate, which contrasts verified successes and failures to induce cross-task procedures, selects them through a development-validation gate, and uses failure-informed, answer-free query synthesis to expand the consolidation data without requiring pre-existing reference answers. Because the target model may not directly execute these abstractions, a stronger model concretizes them into executable code supervision for standard, procedure-free SFT; a matched no-procedure teacher branch isolates the value of procedural guidance. On SciCode, runtime procedure injection improves Qwen3.6-27B by +3.85/+6.26 sub-step/main-problem points, but yields almost no aggregate main-problem gain for Qwen3.5-9B, providing operational evidence of the abstraction-execution gap. After procedure-guided concretization, the 9B student improves under procedure-free deployment by +3.89/+6.25 points over the no-procedure SFT control and by +5.62/+11.25 over the original 9B model. These results establish an experience-to-capability pathway for scientific computing and provide a practical starting point for scaling self-improving scientific assistance.
Chinese Translation
大型语言模型越来越多地解决科学计算任务,但一个问题的可执行反馈很少能在后续问题中转化为持久的能力。我们研究科学计算经验的巩固:将经过验证的运行时经验转化为可转移的程序知识和持久的模型改进。这个设置面临两个挑战:轨迹衍生的工件可能编码特定于源的修复,而不是跨任务的计算机制;而且较弱的目标模型可能无法将其他有效的抽象程序操作化——即抽象-执行差距。我们引入了SciConsolidate,它通过对比验证的成功与失败来诱导跨任务程序,通过开发-验证门选择这些程序,并利用基于失败的信息、无答案的查询综合来扩展巩固数据,而无需预先存在的参考答案。由于目标模型可能无法直接执行这些抽象,因此一个更强的模型将其具体化为可执行代码的监督,以便进行标准的无程序SFT;一个匹配的无程序教师分支则隔离了程序指导的价值。在SciCode上,运行时程序注入使Qwen3.6-27B的子步骤/主要问题得分分别提高了+3.85/+6.26,但对Qwen3.5-9B几乎没有整体主要问题的提升,提供了抽象-执行差距的操作证据。在程序指导的具体化之后,9B学生在无程序部署下的得分比无程序SFT控制提高了+3.89/+6.25,比原始9B模型提高了+5.62/+11.25。这些结果建立了科学计算的经验到能力的路径,并为扩展自我改进的科学辅助提供了一个实用的起点。
cs.AI / 173 / 2607.24512

Making Mathematical Knowledge Explainable, Accessible and Interoperable Through Large Language Model Integration

通过大语言模型集成使数学知识可解释、可获取和可互操作
Range, Jan, Schembera, Björn, Göddeke, Dominik
Abstract
Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles. Knowledge bases such as the Mathematical Model Database (MathModDB) address this gap by providing curated, semantically rich representations of mathematical models. Built on Wikibase, the same open-source infrastructure underlying Wikidata, MathModDB utilizes Semantic Web technologies to support Linked Open Data, collaborative editing, and the storage of semantically enriched metadata, making it a domain-specific knowledge graph within the broader Wikidata ecosystem. However, access to MathModDB currently requires either navigating a complex web interface or proficiency in SPARQL and Wikibase APIs, posing significant barriers for potential users. In addition, the combination of such curated knowledge bases with actual research data stored, e.g., in Dataverse repository instances, remains a challenge. To overcome these limitations, we propose integrating Large Language Models (LLMs) with MathModDB via a Model Context Protocol (MCP) server that exposes a vector-indexed schema retrieval and Steiner-tree-based join planner, combining dialogue-based natural language interaction with curated, epistemically grounded knowledge. Although instantiated on MathModDB, the architecture can be applied to other Wikibase-based systems. We demonstrate that this approach enables epistemically grounded LLM usage, improves model explainability and accessibility beyond what the standard Wikibase interface offers, and simplifies interoperability with external databases and tools, such as Dataverse data repositories. We illustrate the benefits of combining the accessibility of an LLM with the epistemic safety of a curated knowledge base through the adaptability of the MCP protocol by two use cases involving mathematical models in the fields of continuum mechanics and enzyme kinetics.
Chinese Translation
数学模型在研究问题的形式化中至关重要,但其文档通常未能符合FAIR原则。诸如数学模型数据库(Mathematical Model Database, MathModDB)等知识库通过提供经过策划的、语义丰富的数学模型表示来填补这一空白。MathModDB基于Wikibase构建,Wikibase是支撑Wikidata的开源基础设施,利用语义网技术支持链接开放数据、协作编辑和存储语义增强的元数据,使其成为更广泛的Wikidata生态系统中的一个领域特定知识图谱。然而,目前访问MathModDB需要用户导航复杂的网络界面或具备SPARQL和Wikibase API的专业知识,这对潜在用户构成了重大障碍。此外,将此类策划的知识库与实际研究数据(例如存储在Dataverse存储库实例中的数据)结合仍然是一个挑战。为了克服这些限制,我们提出通过模型上下文协议(Model Context Protocol, MCP)服务器将大语言模型(Large Language Models, LLMs)与MathModDB集成,该服务器公开一个基于向量索引的模式检索和基于斯坦纳树的连接规划,将基于对话的自然语言交互与经过策划的、具有认知基础的知识相结合。尽管该架构在MathModDB上实现,但可以应用于其他基于Wikibase的系统。我们证明这种方法使得基于认知的LLM使用成为可能,提升了模型的可解释性和可获取性,超越了标准Wikibase界面的提供,并简化了与外部数据库和工具(如Dataverse数据存储库)的互操作性。我们通过两个涉及连续介质力学和酶动力学领域的数学模型的用例,展示了将LLM的可获取性与策划知识库的认知安全性结合的好处,体现了MCP协议的适应性。
cs.AI / 174 / 2607.24539

Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis

面向任务的多模态大语言模型在电网诊断中的可信度审计
Zhao, Tianqiao, Yue, Meng, Wang, Jianhui
Abstract
Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with self-reported reliance and behavioral changes under controlled modality ablations. To resolve detected discrepancies, we design an evidence-gated correction and re-audit mechanism that regenerates failed responses under evidence constraints and independently re-ablates them to verify improved grounding without performance loss. Case studies evaluate three differently scaled LLMs on IEEE 39- and 118-bus scenarios. These results validate the framework ability to detect, diagnose, and correct task-conditional faithfulness failures.
Chinese Translation
多模态大语言模型(LLMs)能够结合拓扑结构、测量数据和事件文本进行电网诊断,但答案的准确性并不能证明使用了适合任务的证据。本文提出了一个通用框架,以进行面向任务的可信度审计。该框架比较了自我报告的依赖程度、干预导出的行为依赖和预注册的工程重要性。框架首先注册任务特定的证据要求,并将其与自我报告的依赖程度和在控制模态消融下的行为变化进行比较。为了解决检测到的差异,我们设计了一种证据门控的修正和重新审计机制,该机制在证据约束下重新生成失败的响应,并独立进行再消融以验证在不损失性能的情况下的改进基础。案例研究评估了三种不同规模的LLMs在IEEE 39和118节点场景下的表现。这些结果验证了该框架检测、诊断和修正面向任务的可信度失败的能力。
cs.AI / 175 / 2607.24551

LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph

基于大型语言模型的本体工程与法国法律知识图谱构建
Montenegro, G{é}nesis, Billami, Mokhtar Boumedyen, Faron, Catherine, Gandon, Fabien, Monnin, Pierre
Abstract
Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate into operational systems. This paper presents a two-stage LLM-assisted workflow for French maintenance regulations: ontology engineering from a SEMLEG-based core ontology, followed by construction of an ontology-grounded French legal knowledge graph. The first stage consists in the open extraction of typed entities and triples from a stratified corpus sample, the normalization of labels through embedding-based fusion, and the induction of candidate object properties with their signature (domain and range). The second stage uses the resulting ontology to guide the closed extraction of triples and RDF graph construction over the full corpus. Experiments with GPT-4.1 and mistral-large-2512 show robust structured outputs, near-complete class alignment, and a substantial reduction of duplicated entities and predicates after fusion. Fewer than 20% of triples introduce unseen properties, while lower exact signature compliance reveals new domain-range combinations for existing predicates. These results point to predicate normalization and the validation of newly observed relation signatures as key refinement steps for industrial maintenance settings.
Chinese Translation
维护法规是复杂的法律文本,在处理特定案例时难以利用,并且在集成到操作系统中时面临挑战。本文提出了一种基于大型语言模型(LLM)的两阶段工作流程,用于法国维护法规:首先从基于SEMLEG的核心本体进行本体工程,然后构建一个以本体为基础的法国法律知识图谱。第一阶段包括从分层语料样本中开放提取类型实体和三元组,通过基于嵌入的融合对标签进行规范化,以及诱导候选对象属性及其签名(领域和范围)。第二阶段利用生成的本体指导对完整语料库的三元组闭合提取和RDF图构建。与GPT-4.1和mistral-large-2512的实验显示出强大的结构化输出,几乎完全的类别对齐,以及在融合后显著减少的重复实体和谓词。不到20%的三元组引入了未见属性,而较低的确切签名合规性揭示了现有谓词的新领域-范围组合。这些结果表明,谓词规范化和新观察到的关系签名的验证是工业维护环境中的关键优化步骤。
cs.AI / 176 / 2607.24562

Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models

层次化组条件的符合风险控制用于语言模型中的选择性预测
Salem, Murilo, Böhm, Luísa, Pontes, Daniel, Ferrugem, Anderson
Abstract
Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control (CRC) gives rigorous marginal risk guarantees for selective prediction with abstention, but marginal guarantees do not imply per-group ones: a model can meet the population budget while systematically over-exposing subgroups to errors. Under mild shift in group composition, standard CRC violates the budget in up to 47% of trials. We propose HG-CRC (Hierarchical Group-Conditional CRC), a post-hoc calibration framework enforcing simultaneous risk guarantees across all nodes of a user-defined group hierarchy. It applies a Bonferroni correction over nodes and a leaf-first policy that uses the most specific applicable threshold, falling back to coarser nodes when a finer one is uncertified or rejects the example. It needs only a held-out calibration set, with no retraining. We evaluate on three models (Qwen3-4B, Llama-3.1-8B-Instruct, Gemma-3-4B) and two benchmarks (ARC Challenge, MMLU-Pro) across eight configurations probing IID generalization, heterogeneity, mixture/domain/prompt/difficulty shift, label noise, and quantization. Main result: HG-CRC reaches an empirical 0% violation rate and WGER=0 on ARC Challenge for high-accuracy models (Qwen3-4B, Llama-3.1-8B). At 500 bootstrap trials these zeros are empirical upper bounds (true rate up to 0.6%), not certified. Results are benchmark-specific: on MMLU-Pro these models abstain entirely or (Llama) retain WGER=0.014. Gemma-3-4B, poorly calibrated here, degrades gracefully by abstaining. Participation cost vs. global CRC is 22 to 37 points. Ablations show hierarchical depth clears the budget: removing difficulty level returns violations to about 11%. Bonferroni is needed for the theoretical guarantee, though its empirical effect matters only with many nodes.
Chinese Translation
大型语言模型服务于由领域、主题难度和语言风格构成的异质人群。符合风险控制(CRC)为选择性预测提供了严格的边际风险保证,但边际保证并不意味着对每个组的保证:一个模型可以满足人群预算,同时系统性地使子群体暴露于错误之中。在组构成轻微变化的情况下,标准的CRC在多达47%的试验中违反了预算。我们提出了HG-CRC(层次化组条件的CRC),这是一种后处理校准框架,强制在用户定义的组层次结构的所有节点上同时满足风险保证。它在节点上应用Bonferroni校正,并采用优先处理叶节点的策略,使用最具体的适用阈值,当更细的阈值未获得认证或拒绝示例时回退到较粗的节点。它只需要一个保留的校准集,无需重新训练。我们在三个模型(Qwen3-4B、Llama-3.1-8B-Instruct、Gemma-3-4B)和两个基准(ARC Challenge、MMLU-Pro)上进行了评估,涵盖了八种配置,探讨了IID泛化、异质性、混合/领域/提示/难度变化、标签噪声和量化。主要结果:HG-CRC在高准确率模型(Qwen3-4B、Llama-3.1-8B)上在ARC Challenge中达到了0%的违反率和WGER=0。在500次自助试验中,这些零值是经验上限(真实率高达0.6%),并未获得认证。结果是基准特定的:在MMLU-Pro上,这些模型完全放弃预测或(Llama)保持WGER=0.014。Gemma-3-4B在这里校准不佳,通过放弃预测优雅降级。参与成本与全球CRC相比为22到37个点。消融实验表明,层次深度清除了预算:去除难度级别使违反率回升至约11%。虽然Bonferroni对于理论保证是必要的,但其经验效果仅在节点较多时才显著。
cs.AI / 177 / 2607.24563

TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs

TRACE-CTI:基于知识图谱的可审计后提取治理TTP声明
Valletta, Federico, Longo, Giacomo, Russo, Enrico, Merlo, Alessio
Abstract
Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to decide whether an individual mapping should be trusted. We present TRACE- CTI, a post-extraction claim-governance framework that preserves run-level Predictions, aggregates them into configuration-level GraphAssertions, materializes setup-deduplicated corroboration as ConsensusAssertions, and exposes only GraphAssertions backed by policy-compliant validation grounds. The framework retains native evidence granularity, complete extraction provenance, versioned trust decisions, and non-destructive revocation history. We evaluate TRACE-CTI on two public CTI corpora comprising 65 reports and 5,303 sentences, using a controlled 2 x 3 matrix of retrievers and generator families, incrementally ingested across six GraphVersions. All setups are incorporated without schema modification; provenance paths remain complete, operational scopes remain disjoint, and every trusted GraphAssertion has an active qualifying validation ground. Cross-generator-family setup pairs exhibit greater output diversity than same-family pairs. At the final graph state, increasing setup support from k >= 1 to six-setup unanimity raises gold-aligned precision from 25.3% to 90.6%, while recall decreases from 88.2% to 16.3%. The graph also directly answers seven questions about provenance, trust, versioning, dependency, disagreement, and review-queue that the evaluated minimal flat output cannot fully answer without enrichment or reprocessing. These results support explicit, auditable governance of extracted TTP claims; the observed corroboration trajectory is descriptive and does not establish statistical independence or a causal model-family effect.
Chinese Translation
安全运营中心越来越依赖于将网络威胁情报报告自动映射到MITRE ATT&CK,但提取器的输出仍然存在错误,并且通常在没有证据、来源和验证历史的情况下存储,这些都是决定单个映射是否值得信赖所需的。我们提出了TRACE-CTI,一个后提取声明治理框架,该框架保留运行级别的预测,将其聚合为配置级别的图断言(GraphAssertions),将去重的证实物化为共识断言(ConsensusAssertions),并仅暴露由符合政策的验证依据支持的图断言。该框架保留了原生证据的粒度、完整的提取来源、版本化的信任决策和非破坏性的撤销历史。我们在两个公共CTI语料库上评估TRACE-CTI,这些语料库包含65份报告和5,303个句子,使用一个受控的2 x 3检索器和生成器系列矩阵,逐步摄取跨六个图版本(GraphVersions)。所有设置均在不修改模式的情况下纳入;来源路径保持完整,操作范围保持不重叠,每个受信任的图断言都有一个有效的合格验证依据。跨生成器系列的设置对比显示输出多样性大于同系列对比。在最终图状态下,将设置支持从k >= 1增加到六个设置一致性,使得与金标准对齐的精确度从25.3%提高到90.6%,而召回率则从88.2%降低到16.3%。该图还直接回答了关于来源、信任、版本控制、依赖关系、分歧和审查队列的七个问题,而评估的最小平面输出在没有丰富或重新处理的情况下无法完全回答。这些结果支持对提取的TTP声明进行明确的、可审计的治理;观察到的证实轨迹是描述性的,并未建立统计独立性或因果模型系列效应。
cs.AI / 178 / 2607.24567

DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing

DSCH-Loss:一种用于深度语义哈希的动态语义通道目标
Bauer, Tobias J., Riess, Christian, Loebenberger, Daniel, Bergler, Christian
Abstract
Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better semantic capturing capabilities than traditional approaches relying on manual feature engineering. Moreover, they enable a data-driven approach to semantic hashing across diverse data modalities, yielding high-quality cross-modal hash codes within a shared Hamming space. Previous work investigated the properties of this Hamming space and introduced a loss function based on predefined so-called semantic channels with fixed width and Hamming distances derived from label similarities. However, this formulation also introduced discontinuities into the loss landscape, complicating optimization. Based on these observations, we propose a newly designed loss function, Dynamic Semantic Channel Hashing (DSCH), using dynamically sized and positioned semantic channels in order to avoid loss landscape discontinuities. Furthermore, we endorse the use of tie-aware Mean Average Precision (mAP) as evaluation metric as it addresses the ambiguity in sample retrieval ordering, which emerges from the discreteness of hash code distances. Finally, multiple experimental settings conducted on two popular datasets and incorporating two different model architectures provide strong evidence that training using the DSCH objective outperforms training using other state-of-the-art loss functions. In a total of 35 out of 40 cross-modal and intra-modal retrieval tasks, models trained with DSCH achieve significantly higher tie-aware mAP scores across all four tested hash code lengths, showing compelling results across model architecture and used dataset. The mAP score uplifts are consistent and amount up to 1.75 percentage points compared to the respective second best.
Chinese Translation
近年来,生成短二进制哈希码的语义哈希方法在高维数据空间中实现高效的近似最近邻搜索受到了广泛关注。基于深度学习的方法相比于依赖手动特征工程的传统方法,具有更好的语义捕捉能力。此外,它们还支持在多种数据模态下采用数据驱动的方法进行语义哈希,从而在共享的汉明空间内生成高质量的跨模态哈希码。之前的研究探讨了该汉明空间的特性,并引入了一种基于预定义的固定宽度语义通道和源自标签相似性的汉明距离的损失函数。然而,这种表述也在损失景观中引入了不连续性,增加了优化的复杂性。基于这些观察,我们提出了一种新设计的损失函数,动态语义通道哈希(Dynamic Semantic Channel Hashing,DSCH),使用动态大小和位置的语义通道,以避免损失景观的不连续性。此外,我们支持使用考虑平局的平均精度(Mean Average Precision,mAP)作为评估指标,因为它解决了由于哈希码距离的离散性而导致的样本检索顺序的模糊性。最后,在两个流行数据集上进行的多种实验设置,并结合两种不同的模型架构,提供了强有力的证据表明,使用DSCH目标进行训练的效果优于使用其他最先进的损失函数。在40个跨模态和内模态检索任务中,有35个任务中,使用DSCH训练的模型在所有四种测试的哈希码长度上均获得了显著更高的考虑平局的mAP分数,显示出在模型架构和使用数据集方面的良好结果。与相应的第二名相比,mAP分数的提升是一致的,最高可达1.75个百分点。
cs.AI / 179 / 2607.24573

LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports

LLM-SoccerArena:在体育领域对大型语言模型进行现实世界预测的基准测试
Schröder, Jonas, Schweisthal, Jonas, Müller, Oliver, Weinmann, Markus, Feuerriegel, Stefan
Abstract
Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLMs generated forecasts for all 104 matches and 15 tournament-related questions. We provide a detailed analysis of model performance across information access, prompting strategy, and forecast horizon. As a result, LLM-SoccerArena provides new evidence about the forecasting performance of state-of-the-art LLMs. For example, LLMs with web access outperform those without, but only by a small margin (i.e., a 0.023 improvement in Brier score). Overall, LLM-SoccerArena provides a flexible, open-source platform for prospective benchmarking of unresolved events. LLM-SoccerArena will be continuously updated, and can be directly applied to future national and international tournaments and league competitions.
Chinese Translation
大型语言模型(LLMs)越来越多地支持关于不确定未来事件的决策,但评估它们预测现实世界结果的能力仍然困难。特别是,现有基准通常是静态和回顾性的,因此无法测试LLMs如何综合信息以在不确定性下预测未来事件。我们引入了LLM-SoccerArena(https://llm-soccerarena.com),一个前瞻性的实时基准,评估LLMs在结果未知之前预测现实世界体育事件的能力。LLM-SoccerArena提供了(1)一个前瞻性的实时基准协议,(2)一个公共开源平台,以及(3)一个带有与锦标赛相关问题(例如,哪个队将获胜)的因子基准设计。LLM-SoccerArena自动记录未解决事件的时间戳、模式验证的预测,以及提示、模型版本、工具跟踪和成本。因子设计在四个维度上变化:(1)模型版本(例如,GPT-5.5,Claude Opus 4.8);(2)信息获取;(3)提示策略,以及(4)预测时间范围。我们通过对2026年国际足联世界杯的大规模评估展示了LLM-SoccerArena,在该评估中,七个LLMs为所有104场比赛和15个与锦标赛相关的问题生成了预测。我们提供了关于模型在信息获取、提示策略和预测时间范围方面表现的详细分析。因此,LLM-SoccerArena提供了关于最先进的LLMs预测性能的新证据。例如,具有网络访问权限的LLMs优于没有网络访问权限的LLMs,但仅有微小的优势(即在Brier分数上提高了0.023)。总体而言,LLM-SoccerArena提供了一个灵活的开源平台,用于对未解决事件进行前瞻性基准测试。LLM-SoccerArena将持续更新,并可以直接应用于未来的国家和国际锦标赛及联赛竞争。
cs.AI / 180 / 2607.24588

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

SIREN:面向端到端极端天气预警的经验基础大型语言模型代理
Ni, Hang, Zhang, Weijia, Liu, Fan, Lu, Mengqian, Liu, Hao
Abstract
Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent advances in Large Language Model (LLM) agents have enabled the automation of weather-related tasks, existing studies remain centered on isolated scientific tasks and overlook the chain of interdependent processes required for operational extreme-weather early warning. To bridge this gap, this study investigates automated end-to-end extreme-weather early warning through LLM agents. We first develop SIREN-Bench, a comprehensive benchmark comprising 600 question-answer instances across 19 tasks, and covering four individual warning procedures and an end-to-end warning chain. Evaluation on SIREN-Bench reveals substantial capability gaps in existing weather agent frameworks. This motivates us to develop SIREN, an experience-grounded agent framework inspired by experts' use of historical cases, which combines an agentic execution environment integrating heterogeneous weather evidence and tools with a family of agent harnesses that exploit historical cases through retrieval, skill distillation, and predictive modeling. Extensive experiments demonstrate that SIREN outperforms weather-agent baselines on both individual warning procedures and end-to-end warning chains.
Chinese Translation
极端天气的早期预警对于减轻由危险天气事件带来的社会、经济和环境风险至关重要。然而,以专家为中心的预警工作流程成本高、劳动密集且难以在预警到行动的过程中进行规模化。尽管最近大型语言模型(LLM)代理的进展使得天气相关任务的自动化成为可能,但现有研究仍然集中在孤立的科学任务上,忽视了进行操作性极端天气早期预警所需的相互依赖过程链。为了解决这一问题,本研究通过LLM代理探讨了自动化的端到端极端天气早期预警。我们首先开发了SIREN-Bench,这是一个综合基准,包含600个问题-答案实例,涵盖19个任务,涉及四个单独的预警程序和一个端到端的预警链。在SIREN-Bench上的评估揭示了现有天气代理框架的显著能力差距。这促使我们开发了SIREN,一个基于经验的代理框架,灵感来自专家对历史案例的使用,结合了一个集成异构天气证据和工具的代理执行环境,以及一系列通过检索、技能蒸馏和预测建模利用历史案例的代理工具。大量实验表明,SIREN在单个预警程序和端到端预警链上均优于天气代理基线。
cs.AI / 181 / 2607.24589

Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions

人工智能与创新生态系统:演变发展、挑战与未来方向
Zhang, Zhimin, Ma, Chengzhen, Chai, Jia, Zhan, Rongxin, Ning, Huansheng, Mao, Lingfeng, Zhang, Dan, Jiang, Suiping
Abstract
The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digitization, informatization, and intelligence. Exploring how AI can leverage inherent characteristics to influence the development trajectory of IE is a topic that warrants further investigation. Given AI's increasing prominence and role within IE, the paper analyzes this new form, examining both AI's unique contributions to IE and its potential challenges. Firstly, the paper synthesizes the conceptual frameworks surrounding IE, decomposing them into manifestations in physical, social, and thinking spaces. Furthermore, the concept of Artificial Intelligence IE (AIIE) is introduced from a spatial perspective, with an exploration of the characteristics AI contributes to IE. Subsequently, the paper employs an evolutionary perspective to analyze the roles provided by AI during different development periods of AIIE. The paper then verifies the feasibility, effectiveness, and rationality of the AIIE's definition and analyzes AIIE development from an evolutionary perspective using enterprise development examples. Finally, acknowledging AI's inherent limitations, the paper examines potential challenges facing AIIE in the future from four perspectives, aiming to identify new research avenues for the further development of AIIE.
Chinese Translation
创新生态系统(IE)的发展为经济整合、协同进步和共享成就提供了新的范式。人工智能(AI)的崛起显著加速了全球数字化、信息化和智能化进程。探讨AI如何利用其固有特征影响IE的发展轨迹是一个值得深入研究的课题。鉴于AI在IE中的日益重要性和作用,本文分析了这种新形式,考察了AI对IE的独特贡献及其潜在挑战。首先,本文综合了围绕IE的概念框架,将其分解为在物理、社会和思维空间中的表现。此外,从空间角度引入了人工智能创新生态系统(AIIE)的概念,探讨了AI对IE贡献的特征。随后,本文采用演变视角分析了AI在AIIE不同发展阶段所提供的角色。接着,验证了AIIE定义的可行性、有效性和合理性,并利用企业发展实例从演变的角度分析AIIE的发展。最后,考虑到AI的固有限制,本文从四个视角探讨了未来AIIE面临的潜在挑战,旨在为AIIE的进一步发展识别新的研究方向。
cs.AI / 182 / 2607.24647

Efficiency Matters in Autonomous Research

自主研究中的效率问题
Yang, Haiqian, Cao, Yuan
Abstract
AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of the solution-search process is an equally important but often overlooked dimension of performance. A strong AR system should not only produce high-quality results, but also reach them with as small a budget as possible. Search efficiency will become increasingly important as AR expands from domains with inexpensive verification, such as mathematics and coding, to real-world scientific settings in which solution evaluation may require costly physical experiments. To capture this dimension, we propose evaluating AR systems using the area under the curve (AUC) of the Pareto frontier, alongside final outcome quality. We compare several families of search algorithms, including hill climbing, beam search, tree search, and evolutionary search, across twelve systems-optimization tasks. We find that no single search structure is consistently the most efficient. We also show that search efficiency and final outcome quality are distinct performance dimensions: a method that eventually achieves the best result may nevertheless improve slowly and consume substantially more evaluation budget before reaching that result. Because the most effective search policy is generally unknown in advance, we introduce an adaptive procedure called fluid search, which uses a portfolio bandit to dynamically allocate a fixed evaluation budget across a forest of search processes. Across the evaluated tasks, fluid search achieves the highest overall search efficiency, closely matching the performance of a per-task oracle that is given the best search structure for each task in advance.
Chinese Translation
基于人工智能的自主研究(AR)系统在广泛的任务中变得越来越有效。然而,它们的性能仍然主要通过最终结果的质量来评估。本文论证了解决方案搜索过程的效率是一个同样重要但常常被忽视的性能维度。一个强大的AR系统不仅应该产生高质量的结果,还应该以尽可能小的预算达到这些结果。随着AR从数学和编码等验证成本低的领域扩展到需要昂贵物理实验的现实科学环境,搜索效率将变得越来越重要。为了捕捉这一维度,我们建议在评估AR系统时,除了最终结果质量外,还应使用帕累托前沿下的面积(AUC)进行评估。我们比较了包括爬山算法、束搜索、树搜索和进化搜索在内的几类搜索算法,在十二个系统优化任务中进行测试。我们的发现是,没有单一的搜索结构在效率上始终表现最佳。我们还表明,搜索效率和最终结果质量是不同的性能维度:一种最终实现最佳结果的方法可能在达到该结果之前进展缓慢,并消耗更多的评估预算。由于最有效的搜索策略通常无法提前确定,我们引入了一种称为流体搜索(fluid search)的自适应程序,该程序使用投资组合赌博机(portfolio bandit)动态分配固定的评估预算到一组搜索过程。经过评估的任务中,流体搜索实现了最高的整体搜索效率,接近于为每个任务提前提供最佳搜索结构的任务特定神谕(oracle)的性能。
cs.AI / 183 / 2607.24649

Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

基于理由的行为模型用于审计大型语言模型社交模拟器
Pandey, Atharva, Jajoo, Gautam
Abstract
Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-derived reason pattern. We study this problem through a 94-person sunscreen concept test in which each respondent evaluated three product concepts and wrote open-ended rationales. We map those rationales into signed reason states $Z$, where positive signs support adoption and negative signs block it. This gives a practical audit: holding respondent descriptors $D$, category context $K$, and concept treatment $X$ fixed, do human rationale-derived reasons help predict behavior $Y$, and can an LLM simulate the same reason state without seeing the human rationale or outcome? Human rationale-derived reasons substantially improve held-out prediction of purchase intent. LLM-simulated reasons are more brittle: they often sound plausible, but frequently echo the concept board rather than recover the respondent's acceptance or rejection path. The paper contributes an evaluation framework for social simulators. Reason states do not identify natural causal effects by themselves, but they provide an interpretable test of whether a simulator's stated reasons align with human evidence.
Chinese Translation
大型语言模型越来越多地被用作社交模拟器,包括作为合成调查受访者。大多数评估关注模拟结果是否与人类结果相似。我们认为这虽然必要,但过于薄弱:一个模拟器可以匹配最终答案,但使用错误的基于理由的推理模式。我们通过一项94人参与的防晒霜概念测试研究这个问题,在该测试中,每位受访者评估了三个产品概念并写下开放式的推理。我们将这些推理映射到带符号的理由状态 $Z$,其中正号支持采用,负号则阻止采用。这提供了一种实用的审计:在固定受访者描述符 $D$、类别背景 $K$ 和概念处理 $X$ 的情况下,基于人类推理的理由是否有助于预测行为 $Y$,以及一个大型语言模型(LLM)是否可以在未看到人类推理或结果的情况下模拟相同的理由状态?基于人类推理的理由显著改善了对购买意图的预测。LLM模拟的理由则更为脆弱:它们听起来往往合理,但常常反映概念板而不是恢复受访者的接受或拒绝路径。本文为社交模拟器提供了一个评估框架。理由状态本身并不能识别自然因果效应,但它们提供了一个可解释的测试,以检验模拟器所陈述的理由是否与人类证据一致。
cs.AI / 184 / 2607.24667

Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating

驱逐作为估计:测试时间记忆的固定滞后平滑视角,以及何时测量优于累积
Vemula, Maruthi, Gajula, Neeraj Praneeth
Abstract
A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag $H$: online filters and learned predictors commit at $H=0$, while Belady's offline optimum sits where the whole future is known. The missing regime in between, fixed-lag smoothing, waits a bounded number of steps, observes which items a correct near-future prediction attended to, and only then commits. This measurement, demonstrated utility, turns Belady's unobservable future request into something we read off the model itself. We instantiate it as a training-free policy, RMM, a strict generalization of H2O that reduces to it exactly when the measurement is uniform. In controlled settings where reuse is endogenous and separated in time, demonstrated utility identifies used memory far better than accumulated attention, and a small bounded memory behaves like a much larger one. But on independent third-party benchmarks, run inside NVIDIA's KVPress harness against its own SnapKV, H2O, and StreamingLLM implementations, the advantage mostly disappears: RMM is on par with H2O for single-turn question answering and loses to both H2O and SnapKV in a streaming multi-turn setting. The cause is simple: on natural text the model is correct about most tokens, so weighting attention by correctness barely changes it, and demonstrated utility collapses onto accumulated attention unless reuse is sharp and endogenous, which standard benchmarks do not exercise. Our contribution is the framework and an honest map of when measuring beats accumulating, not a new state of the art.
Chinese Translation
具有有限工作记忆的语言模型必须反复决定保留哪些存储项。每种部署的方法在项到达的时刻做出决定,可能是基于过去(如 StreamingLLM、H2O)或对未来的猜测(如 SnapKV)。我们将这一选择重新表述为一个关于隐藏信号的估计问题,即某个项是否会被重用,从而将现有方法置于一个轴上,承诺滞后 $H$:在线过滤器和学习预测器在 $H=0$ 时承诺,而 Belady 的离线最优解则在未来完全已知的情况下。缺失的中间状态,固定滞后平滑,等待有限步数,观察正确的近未来预测关注了哪些项,然后才做出承诺。这种测量,展示了效用,将 Belady 的不可观察的未来请求转变为我们从模型本身读取的内容。我们将其实例化为一种无训练策略 RMM,这是 H2O 的严格推广,当测量均匀时,RMM 恰好简化为 H2O。在重用是内生且时间上分离的受控环境中,展示的效用远比累积注意力更好地识别已使用的记忆,而小的有限记忆表现得像一个更大的记忆。但在独立的第三方基准测试中,在 NVIDIA 的 KVPress 环境中运行,与其自身的 SnapKV、H2O 和 StreamingLLM 实现相比,优势大多消失:RMM 在单轮问答中与 H2O 不相上下,但在流式多轮设置中输给了 H2O 和 SnapKV。原因很简单:在自然文本中,模型对大多数标记的判断是正确的,因此通过正确性加权注意力几乎不会改变它,除非重用是明显且内生的,否则展示的效用会崩溃为累积注意力,而标准基准测试并未对此进行考验。我们的贡献在于框架和诚实的地图,展示了何时测量优于累积,而不是新的最先进技术。
cs.AI / 185 / 2607.24707

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

ERUnderstand:在结构化ER图上评估视觉-语言模型
Ansari, Ali, Mohammadi, Yasmin, Nili, Farnoush, Esmaeilkhani, Parsa, Latecki, Longin Jan, Dragut, Eduard
Abstract
Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the first large-scale benchmark for structured understanding of ER diagrams, comprising 2,960 diagrams collected from curated educational sources, real-world schemas, and synthetically generated examples spanning diverse domains, notations, complexity levels, and Extended Entity-Relationship (EER) constructs. Each diagram is paired with a standardized machine-readable representation for fine-grained evaluation of schema elements. Evaluating state-of-the-art Vision-Language Models (VLMs), we find that while common ERD elements are recovered reliably (F1 > 0.74), performance drops sharply on weak entities (as low as 0.28 F1), multivalued attributes (0.14 F1), and N-ary relationships (0.07 F1). Reasoning-augmented models improve overall performance by 15-25% but remain sensitive to linguistic priors and increasing diagram complexity. ERUnderstand provides a standardized benchmark for evaluating multimodal understanding of conceptual database schemas. The benchmark, dataset, evaluation toolkit, and generation code are publicly available at https://github.com/salinaria/ERUnderstand.
Chinese Translation
实体-关系图(ERD)是概念数据库设计的核心,但它们通常仅以渲染图像的形式存在,而非机器可读的模式,这限制了AI辅助的数据库工程。我们提出了ERUnderstand,这是第一个针对ER图结构理解的大规模基准,包含从精心策划的教育资源、真实世界模式和合成生成的示例中收集的2960个图,涵盖了不同领域、符号、复杂性水平和扩展实体-关系(EER)构造。每个图都配有标准化的机器可读表示,以便对模式元素进行细粒度评估。在评估最先进的视觉-语言模型(VLM)时,我们发现,尽管常见的ERD元素能够可靠地恢复(F1 > 0.74),但在弱实体(最低可达0.28 F1)、多值属性(0.14 F1)和N元关系(0.07 F1)上的性能急剧下降。增强推理的模型整体性能提高了15-25%,但仍对语言先验和图的复杂性增加敏感。ERUnderstand为评估概念数据库模式的多模态理解提供了一个标准化的基准。该基准、数据集、评估工具包和生成代码已公开发布在https://github.com/salinaria/ERUnderstand。
计算语言学 (Computation and Language)
79
cs.CL / 1 / 2607.22546

Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution

解释 GAND:性别模糊自然数据与对比归因的资源
Hackenbuchner, Janiça, Degraeuwe, Jasper, Tezcan, Arda, Daems, Joke
Abstract
Machine translation (MT) systems continue to produce gender-biased translations. In a time where self-expression is paramount, mistranslations based on default behaviour and stereotyping can lead to harm for users of these systems. To better understand how these systems translate gender in the absence of clear gender cues, we need benchmarking resources that reflect gender-ambiguous scenarios in a natural way. To this end, we present GAND, a gender-ambiguous natural data benchmarking resource for MT consisting of English source sentences, specifically designed to analyse the influence of contextual cues on gender in translation. We leverage GAND to conduct an interpretability analysis: we translate a subset of GAND into two grammatical gender languages and extend these with manually crafted contrastive translations. A following feature attribution analysis reveals source words in context that inform the gender translation of an ambiguous referent entity in the target translation.
Chinese Translation
机器翻译(MT)系统仍然会产生性别偏见的翻译。在自我表达至关重要的时代,基于默认行为和刻板印象的错误翻译可能会对这些系统的用户造成伤害。为了更好地理解这些系统在缺乏明确性别线索的情况下如何翻译性别,我们需要能够自然反映性别模糊场景的基准资源。为此,我们提出了 GAND,一个用于机器翻译的性别模糊自然数据基准资源,包含英语源句,专门设计用于分析上下文线索对翻译中性别的影响。我们利用 GAND 进行可解释性分析:将 GAND 的一个子集翻译成两种语法性别语言,并通过手动制作的对比翻译进行扩展。随后的特征归因分析揭示了上下文中的源词,这些词影响了目标翻译中模糊指称实体的性别翻译。
cs.CL / 2 / 2607.22552

MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities

MioFFAn:一种具有大型语言模型自动化能力的公式形式化注释软件
Sibuet, Nicolas, Saggion, Horacio, Rossi, Riccardo
Abstract
The automatic translation of mathematical expressions in scientific literature into executable symbolic code (a process we refer to as Formula Formalization) is hindered by a severe scarcity of high-quality, ground-truth datasets specialized for technical scientific domains. In this paper, we present MioFFAn, an open-source, document-centric, and customizable framework designed to facilitate rapid annotation for this task. Building upon the MioGatto architecture, we extend existing features to overcome structural limitations and pivot its scope by introducing specific functionalities for Formula Formalization, such as selection of equations of interest and aided symbolic code specification. By allowing users to configure custom taxonomies and properties for identified symbols, and compatible symbolic operators, we ensure the framework is adaptable to diverse specialized scientific fields. Furthermore, MioFFAn is designed to incorporate partial automation via Large Language Models. By defining a modular set of automated sub-tasks with strict output formats, we enable researchers to iteratively refine automation capabilities and evaluate competing strategies using standard NLP metrics. We specify the current automation methodology and perform a preliminary evaluation that demonstrates to efficacy of this human-in-the-loop approach.
Chinese Translation
科学文献中数学表达式自动翻译为可执行符号代码的过程(我们称之为公式形式化)受到高质量、真实数据集的严重匮乏的阻碍,这些数据集专门针对技术科学领域。在本文中,我们介绍了MioFFAn,这是一个开源的、以文档为中心且可定制的框架,旨在促进这一任务的快速注释。基于MioGatto架构,我们扩展了现有功能,以克服结构限制,并通过引入特定的公式形式化功能(如感兴趣方程的选择和辅助符号代码规范)来调整其范围。通过允许用户为识别的符号配置自定义分类法和属性,以及兼容的符号运算符,我们确保该框架能够适应多样化的专业科学领域。此外,MioFFAn还旨在通过大型语言模型实现部分自动化。通过定义一组模块化的自动化子任务及其严格的输出格式,我们使研究人员能够迭代地完善自动化能力,并使用标准自然语言处理指标评估竞争策略。我们具体说明了当前的自动化方法,并进行初步评估,展示了这种人机协作方法的有效性。
cs.CL / 3 / 2607.22553

Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review

评估评审者指南设计对基于大语言模型的自动化同行评审的影响
Li, Haowen, Ishibashi, Yoichi, Oyamada, Masafumi
Abstract
Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.
Chinese Translation
同行评审是科学研究中的一个重要过程,但日益增长的工作负担使得其自动化变得愈加必要。在本研究中,我们分析了不同类型的评审者指南(例如官方会议指南和通过高质量人类评审生成的模仿评审者的指南)如何影响自动化同行评审。我们的实验表明,官方会议指南产生的评审结果与人类判断最为一致,这表明通过会议实践精炼的评估标准也为自动化评审提供了有效的指导。相比之下,模仿评审者的指南通常不如官方会议指南有效。此外,强制执行严格的评分标准持续降低了表现,突显了允许主观和整体评分的重要性。
cs.CL / 4 / 2607.22622

Learning When to Reason for Text-to-SQL via SFT and DPO

通过 SFT 和 DPO 学习何时进行推理以实现文本到 SQL 的转换
Jang, Soohyuk, Yeom, Jiheum, Park, Nohil, Kim, Sang Hun, Choi, Yoonyoung, Bae, Kiwook, Yoon, Sungroh
Abstract
Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inference-time overhead. However, a large fraction of real-world queries are simple lookups or aggregations that can be resolved without multi-step deduction, making forced reasoning wasteful. Thus, we propose AutoThinkSQL, a framework that integrates an auto-thinking mechanism into both Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) on Text-to-SQL. Our approach enables the model to dynamically bypass reasoning for simple queries while invoking deep CoT for complex queries. On Qwen3-Coder-30B-A3B, our method achieves consistent gains compared to the best counterpart baseline on both Spider and BIRD benchmarks while simultaneously reducing average output tokens by 24.6% and 18.3%, and average latency by 17.1% and 11.5% compared to CoT-only generation. Further analysis indicates that the model learns to align its reasoning decisions with query difficulty.
Chinese Translation
近期的文本到 SQL 方法在很大程度上依赖于以推理为中心的范式,如链式思维(Chain-of-Thought, CoT),在复杂基准测试中取得了显著的进展,但代价是高推理时间开销。然而,现实世界中的大量查询是简单的查找或聚合,可以在没有多步推理的情况下解决,因此强制推理显得浪费。因此,我们提出了 AutoThinkSQL,一个将自动思考机制集成到文本到 SQL 的监督微调(Supervised Fine-Tuning, SFT)和直接偏好优化(Direct Preference Optimization, DPO)中的框架。我们的方法使模型能够动态地跳过简单查询的推理,同时对复杂查询调用深层链式思维。在 Qwen3-Coder-30B-A3B 上,我们的方法在 Spider 和 BIRD 基准测试中与最佳对比基线相比取得了一致的提升,同时将平均输出标记减少了 24.6% 和 18.3%,平均延迟减少了 17.1% 和 11.5%,与仅使用 CoT 生成相比。进一步分析表明,模型学习将其推理决策与查询难度对齐。
cs.CL / 5 / 2607.22657

Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

抑制与崩溃之间:使用 LENS 评估叙事遗忘
Makovska, Viktoriia, Fletcher, George
Abstract
Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine-unlearning algorithms can suppress this behavior. We introduce Level-based Evaluation of Narrative Suppression (LENS), a contextualization based evaluation protocol for testing target narrative reproduction across direct, attributed, contrastive, and abstract resistance levels. We evaluate two source-grounded narratives: one framing Russia's war against Ukraine as forced by NATO expansion, and one framing the United States as exploiting or abandoning Taiwan. The experiments cover four near-12B multilingual instruction models: Lapa LLM, Gemma-12B, Qwen-14B, and TAIDE-Gemma. We introduce the Suppression-Collapse Efficiency (SCE) score as a checkpoint selection summary that rewards target-narrative suppression while penalizing degraded outputs. Our results shows that selected checkpoints can reduce narrative reproduction and suppression may transfer beyond direct forget prompts. We also report entity recovery as a separate side effect: abstract A/B/C prompts can cause models to recover the real-world actors associated with the target frame after unlearning. These findings demonstrate that LENS is a successful diagnostic protocol for both reporting and guiding the further study of the deeper structure of narrative unlearning.
Chinese Translation
大型语言模型(LLMs)能够将与虚假信息一致的叙事框架再现为可信的解释,这引发了一个问题:现有的机器遗忘算法是否能够抑制这种行为。我们引入了基于层级的叙事抑制评估(Level-based Evaluation of Narrative Suppression, LENS),这是一种基于情境的评估协议,用于测试目标叙事再现的直接、归因、对比和抽象抵抗水平。我们评估了两个基于来源的叙事:一个将俄罗斯对乌克兰的战争框定为北约扩张所迫,另一个将美国框定为利用或抛弃台湾。实验涵盖了四个接近120亿参数的多语言指令模型:Lapa LLM、Gemma-12B、Qwen-14B 和 TAIDE-Gemma。我们引入了抑制-崩溃效率(Suppression-Collapse Efficiency, SCE)分数,作为一个检查点选择摘要,奖励目标叙事的抑制,同时惩罚降级输出。我们的结果表明,所选检查点可以减少叙事再现,抑制可能超越直接遗忘提示进行转移。我们还报告了实体恢复作为一个独立的副作用:抽象的 A/B/C 提示可以导致模型在遗忘后恢复与目标框架相关的现实世界参与者。这些发现表明,LENS 是一个成功的诊断协议,既可用于报告,也可用于指导对叙事遗忘更深层结构的进一步研究。
cs.CL / 6 / 2607.22859

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

PatiGonit22K:用于解决复杂孟加拉数学文字问题的综合数据集
Kundu, Swastika, Fayaz, Azizul Hakim, Muhammad, Tashreef
Abstract
Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent progress in high resource languages, Bengali remains underexplored due to the limited availability of large scale annotated datasets. In this work, we introduce PatiGonit22K, an expanded Bengali MWP dataset containing 22,441 problems, developed by extending the original PatiGonit dataset with a substantially larger collection of complex mathematical problems. The dataset includes both simple and multi operation equations, providing a balanced benchmark for evaluating mathematical reasoning across different difficulty levels. Each problem is carefully translated, annotated, culturally adapted, and verified to ensure linguistic consistency and mathematical correctness. By increasing both the scale and complexity of Bengali MWPs, PatiGonit22K provides a more comprehensive resource for future research on mathematical reasoning and educational NLP applications in low resource languages.
Chinese Translation
数学文字问题(MWPs)是评估自然语言理解和定量推理的重要基准。尽管在高资源语言方面取得了近期进展,但由于大规模注释数据集的有限可用性,孟加拉语仍然未得到充分探索。在本研究中,我们引入了PatiGonit22K,这是一个扩展的孟加拉MWP数据集,包含22,441个问题,通过扩展原始PatiGonit数据集并增加大量复杂数学问题而开发。该数据集包括简单和多操作方程,为评估不同难度水平的数学推理提供了一个平衡的基准。每个问题都经过仔细翻译、注释、文化适应和验证,以确保语言的一致性和数学的正确性。通过增加孟加拉MWPs的规模和复杂性,PatiGonit22K为未来在低资源语言中进行数学推理和教育自然语言处理应用的研究提供了更全面的资源。
cs.CL / 7 / 2607.22884

CHiPS: Character Histograms and Positional Signals for Lightweight Authorship Attribution in Romanian Texts

CHiPS:用于罗马尼亚文本轻量级作者归属的字符直方图和位置信号
Avram, Sanda-Maria, Ţurcaş, George C.
Abstract
We propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts. All reported experiments are closed-set: the true author is one of the candidate authors in the training data. CHiPS studies two complementary fingerprints of writing style: CH-SVM, a character-histogram classifier based on one-character marginal distributions, and FFT12-LR, a positional-signal classifier that represents selected characters and punctuation classes as impulse trains (binary indicator sequences over character positions) and extracts Fourier/Welch spectral descriptors. We also report CHiPS-F, a leakage-safe decision-level fusion variant, and an optional top-5 listwise reranker trained only on out-of-fold predictions. The method requires no tokenization, syntactic analysis, pretrained language model, or transformer fine-tuning, and it avoids character $n$-gram features with $n \geq 2$ in the histogram component. On a locked grouped ROST split comprising 400 files from 392 source-text groups, written by 10 authors, with source-text-level evaluation and grouped five-fold model selection, CHiPS-F reaches 0.9310 accuracy and 0.9341 macro-F1. A matched but unrestricted character 2--5-gram TF--IDF SVM comparator reaches 1.0000 accuracy and macro-F1 on the same held-out groups, so the contribution is not a claim of best possible classification accuracy. Instead, the experiments ask how far restricted, transparent character evidence can go under strict leakage control. On ROSTories-cleaned, a secondary ROST-overlapping corpus comprising 1,248 files from 1,240 source-text groups, written by 19 authors, the same protocol gives 0.8919 accuracy and 0.8708 macro-F1 for CHiPS-R.
Chinese Translation
我们提出了CHiPS,一种针对罗马尼亚文本的轻量级字符级作者归属方法。所有报告的实验均为封闭集:真实作者是训练数据中候选作者之一。CHiPS研究了两种互补的写作风格指纹:CH-SVM,一种基于单字符边际分布的字符直方图分类器,以及FFT12-LR,一种位置信号分类器,该分类器将选定字符和标点类表示为脉冲列(字符位置上的二进制指示序列),并提取傅里叶/韦尔奇谱描述符。我们还报告了CHiPS-F,一种安全防泄漏的决策级融合变体,以及一个可选的仅在折外预测上训练的前5名列表重排序器。该方法不需要分词、句法分析、预训练语言模型或变换器微调,并且在直方图组件中避免使用字符 $n$-gram 特征($n geq 2$)。在一个锁定的分组 ROST 拆分中,该拆分包含来自392个源文本组的400个文件,由10位作者撰写,采用源文本级评估和分组五折模型选择,CHiPS-F 达到了0.9310的准确率和0.9341的宏F1。一个匹配但不受限制的字符2--5-gram TF--IDF SVM比较器在相同的保留组上达到了1.0000的准确率和宏F1,因此该贡献并不是最佳分类准确率的声明。相反,实验探讨了在严格的泄漏控制下,受限的透明字符证据能够达到多远。在经过清理的ROSTories中,一个包含来自1240个源文本组的1248个文件的二级ROST重叠语料库,由19位作者撰写,同样的协议为CHiPS-R提供了0.8919的准确率和0.8708的宏F1。
cs.CL / 8 / 2607.22923

Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge

简单语言规范化获胜:2026 TidyVoice挑战的跨语言说话人验证
Hosseini-Kivanani, Nina
Abstract
Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development speakers in 40 languages and 2,200 evaluation speakers in 38 unseen languages, without language labels at test time. Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (NAP) as a simple language-normalization step in the embedding space. We estimate a compact language subspace from cross-language same-speaker differences and project embeddings onto its orthogonal complement before cosine scoring with Adaptive Symmetric score normalization. This reduces development EER from 2.97\% with cosine and 2.70\% with AS-Norm to 2.18\% and yields a Codabench evaluation score of 8.40, showing that simple back-end language normalization can rival more complex systems.
Chinese Translation
跨语言不匹配仍然是现代说话人验证中整体性能下降的主要来源。TidyVoice2026挑战针对这一情况,采用文本独立的验证方法,包含来自40种语言的3,666名训练说话人和808名开发说话人,以及来自38种未见语言的2,200名评估说话人,测试时不提供语言标签。我们从官方的SimAM-ResNet34基线模型出发,该模型在VoxBlink2和VoxCeleb2上预训练,并在TidyVoice上进行了微调,我们重新审视了作为嵌入空间中简单语言规范化步骤的干扰属性投影(Nuisance Attribute Projection, NAP)。我们从跨语言同说话人差异中估计出一个紧凑的语言子空间,并将嵌入投影到其正交补空间,然后使用自适应对称分数规范化进行余弦评分。这将开发集的等错误率(EER)从使用余弦相似度的2.97%和使用AS-Norm的2.70%降低到2.18%,并获得了Codabench评估分数8.40,显示出简单的后端语言规范化可以与更复杂的系统相媲美。
cs.CL / 9 / 2607.22925

Not All LLM Reasoning is Visible in the Chain-of-Thought

并非所有大语言模型的推理都在思维链中可见
Baherwani, Vatsal, Goldstein, Tom, Panda, Ashwinee
Abstract
A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tasks. We evaluate 13 frontier language models across three tasks and find that many models benefit significantly from filler tokens, with accuracy improvements of up to 13 percentage points. The benefit depends on which tokens are used and differs across models. We further show that filler tokens enable Claude Opus 4.5 to satisfy a hidden modular arithmetic constraint without sacrificing accuracy on its primary task, demonstrating that invisible reasoning can serve objectives entirely invisible to CoT monitoring. Reinforcement learning gives Qwen3-235B strong preferences over filler token content, but neither RL nor supervised fine-tuning produces a filler token benefit that persists at test time. Our results indicate that frontier models already perform consequential computation with no interpretable trace in their output tokens.
Chinese Translation
人工智能安全的一个关键问题是语言模型是否在其输出标记中表达了所有推理。我们展示了一种具体的失败模式,其中前沿模型通过利用语义上无关的填充标记来提高在合成推理任务上的表现,从而表现出不可见的推理。我们评估了13个前沿语言模型在三个任务上的表现,发现许多模型显著受益于填充标记,准确率提高最多可达13个百分点。该收益取决于使用的标记,并且在不同模型间存在差异。我们进一步展示,填充标记使Claude Opus 4.5能够满足一个隐藏的模算术约束,而不牺牲其主要任务的准确性,证明了不可见的推理可以服务于完全不为思维链监控所察觉的目标。强化学习使Qwen3-235B对填充标记内容产生强烈偏好,但无论是强化学习还是监督微调都未能在测试时产生持久的填充标记收益。我们的结果表明,前沿模型已经在其输出标记中执行了重要的计算,而没有可解释的痕迹。
cs.CL / 10 / 2607.22954

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

朝向电子健康记录中文档不一致性的自动检测
Lu, Jian, Chen, Panyu, Treggiari, Miriam, Blessing, Robert, Zhuo, Danyang, Weng, Chunhua, Stead, William W., Zhang, Anru R.
Abstract
Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.
Chinese Translation
目的:描述通用领域大型语言模型(LLM)能够从真实世界出院总结中发现的内部文档不一致性的类型,并识别限制大规模可靠性的重复失败模式。材料与方法:我们应用了一个两阶段的LLM流程——开放式候选识别(Gemini 2.5 Pro)后跟上下文基础的验证(Gemini 2.5 Flash)——对3,000份随机抽样的MIMIC-IV-Note出院总结进行分析。然后,临床专家手动审查了部分流程输出。结果:我们的流程发现了3,460个候选不一致,影响了69.7%的入院病例。代表性例子涵盖了人口统计、过敏、程序、诊断、实验室、药物和护理规划领域,对临床推理或患者安全具有直接影响。专家审查还揭示了在验证需要时间推理、不断演变的诊断背景或模型本身不具备的门诊处方惯例知识时出现的重复失败模式。讨论:检测高度依赖上下文:许多被标记的对需要将每个陈述锚定到其源部分和临床领域,然后评估冲突是否反映真实矛盾或缺失的上下文。我们提出了一个涵盖严格矛盾和模糊性的分级本体,建立了一个模式,通过类别、部分、领域和不一致性轴对每个标记案例进行特征描述。结论:这项基础性研究建立了一个方法论基础和概念框架,以指导后续经过验证的大规模EHR不一致性分析。
cs.CL / 11 / 2607.22996

Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning

超越直接回答:通过启发式强化学习将教育大型语言模型对齐为苏格拉底式指导者
Wang, Xiaokun, Song, Siyu, Liu, Wentao, Zou, Xiaodong
Abstract
Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes. We present HeuristicEdu, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO). Training uses SocraticEdu, 797 multi-turn Chinese children's science dialogues reconstructed from a live platform, with a heuristic reward over cognitive depth (R_cog), curiosity engagement (R_eng), and directness (R_dir), together with a K_query correction for student-introduced terms. We introduce Scaffolding Effectiveness (SE) and Conversation Depth (CD) to evaluate outcomes beyond surface fluency. On 30 held-out questions, the best GRPO variant improves SE from 30.0% to 63.3% and lowers keyword leakage from 30.0% to 13.3%. Notably, this best variant omits the directness penalty during optimization, suggesting that explicit anti-leakage terms can conflict with gradient-based behavioral alignment. An unaligned Qwen-72B baseline reaches 0% SE and 96.7% leakage, showing that scale alone does not induce Socratic behavior.
Chinese Translation
在教育环境中部署的大型语言模型(LLMs)通常表现为直接回答者:它们在开场时披露目标概念,而不是按照苏格拉底教学法的要求,通过渐进式探究引导学生。我们提出了HeuristicEdu,一个两阶段的流程,通过监督预热和群体相对政策优化(Group Relative Policy Optimization, GRPO)将Qwen2.5-7B对齐为苏格拉底式辅导。训练使用SocraticEdu,基于一个实时平台重构的797个多轮中文儿童科学对话,并结合对认知深度(R_cog)、好奇心参与(R_eng)和直接性(R_dir)的启发式奖励,以及针对学生引入术语的K_query修正。我们引入了支架有效性(Scaffolding Effectiveness, SE)和对话深度(Conversation Depth, CD)来评估超越表面流利度的结果。在30个保留问题上,最佳GRPO变体将SE从30.0%提高到63.3%,并将关键词泄漏从30.0%降低到13.3%。值得注意的是,这一最佳变体在优化过程中省略了直接性惩罚,表明显式的反泄漏项可能与基于梯度的行为对齐相冲突。一个未对齐的Qwen-72B基线达到0%的SE和96.7%的泄漏,显示出单靠规模并不能促成苏格拉底式行为。
cs.CL / 12 / 2607.23037

Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating

语音信号补充大型语言模型(LLMs)以预测快速约会中的人际吸引力
Kikuchi, Yuriko, Hayashi, Takato, Kimura, Ryusei, Inoue, Naoya, Ishii, Ryo, Okada, Shogo
Abstract
Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM prediction. Using Japanese speed-dating conversations, we combine predictions from a transcript-only LLM and a supervised speech predictor to estimate participants' reported liking of their partners. We show that speech can complement transcript-only LLM prediction, but that this complementarity is conditional rather than universal. Combining the two predictions significantly improves pairwise ranking accuracy over the transcript-only LLM alone in all evaluated conditions. By contrast, gains in per-participant Pearson $r$ vary across conversation rounds and rating directions, with none significant after correction. Retrospectively, these $r$ gains are concentrated among participants for whom the speech predictor is more accurate. Speech can therefore retain predictive value even when an LLM predicts attraction from transcripts. The relevant question is not simply whether speech helps, but where its complementarity emerges.
Chinese Translation
大型语言模型(LLMs)能够从对话记录中预测人际吸引力,但尚不清楚语音预测器能在多大程度上超越仅依赖记录的LLM预测。通过使用日本快速约会对话,我们结合了仅基于记录的LLM预测和一个监督的语音预测器,以估计参与者对其伴侣的喜欢程度。我们展示了语音可以补充仅基于记录的LLM预测,但这种互补性是有条件的而非普遍适用。将这两种预测结合起来显著提高了在所有评估条件下的成对排名准确性,相较于仅依赖记录的LLM。相比之下,每位参与者的Pearson $r$ 增益在对话轮次和评分方向上有所不同,经过修正后均未显著。回顾来看,这些 $r$ 增益主要集中在语音预测器更为准确的参与者中。因此,即使在LLM通过记录预测吸引力的情况下,语音仍然可以保留预测价值。相关的问题不仅仅是语音是否有帮助,而是其互补性何时出现。
cs.CL / 13 / 2607.23058

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

ADAGE:一种与语言无关的类比推理评估管道
Ahmed, Ahmed Haj, Grissom II, Alvin
Abstract
Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical Difficulty-by-design Assessment for Grounded Evaluation), a language-agnostic pipeline that combines native-speaker curation with LLM-assisted generation to construct challenging, translation-free benchmarks for abstract analogical reasoning. We validate ADAGE by constructing benchmarks for Arabic, Amharic, and Japanese. Evaluating 14 open-weight models, we find a consistent cultural reasoning gap: models that perform well on English proverb reasoning struggle substantially on all three native benchmarks, with accuracy dropping by 12--52 percentage points relative to English. We release the pipeline, all three benchmarks, and the full evaluation suite.
Chinese Translation
多语言推理评估在很大程度上依赖于翻译英语基准,这一做法引入了语言学伪影,并未能测试文化根植的推理能力。我们提出了ADAGE(Analogical Difficulty-by-design Assessment for Grounded Evaluation),这是一种与语言无关的管道,结合了母语者的策划与大型语言模型(LLM)辅助生成,旨在构建具有挑战性的、无翻译的抽象类比推理基准。我们通过为阿拉伯语、阿姆哈拉语和日语构建基准来验证ADAGE。在评估14个开放权重模型时,我们发现了一致的文化推理差距:在英语谚语推理上表现良好的模型在所有三个母语基准上均表现不佳,准确率相较于英语下降了12至52个百分点。我们发布了该管道、所有三个基准以及完整的评估套件。
cs.CL / 14 / 2607.23067

Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models

基于注意力引导的层选择在大型语言模型中的对比解码
Sakai, Yusuke, Kertkeidkachorn, Natthawut, Shirai, Kiyoaki
Abstract
Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of mature and premature layers. However, DoLa's dynamic layer selection relies solely on divergences in output vocabulary distributions. In this work, we propose three attention-guided strategies: Attention-JSD, Attention-Entropy-Max, and Attention-Entropy-Min, which leverage structural information carried by internal self-attention mechanisms as a signal for layer selection. Experimental results on TruthfulQA demonstrate that our strategies, particularly Attention-JSD and Attention-Entropy-Min, consistently outperform the original DoLa. We observe significant gains on multi-answer metrics (MC2 and MC3), suggesting that attention distributions can provide a more sensitive signal for resolving factual knowledge than output vocabulary distributions.
Chinese Translation
对比解码方法如 DoLa 通过对比成熟层和未成熟层的输出分布,提高了大型语言模型(LLMs)的事实性。然而,DoLa 的动态层选择仅依赖于输出词汇分布的差异。在本研究中,我们提出了三种基于注意力引导的策略:Attention-JSD、Attention-Entropy-Max 和 Attention-Entropy-Min,这些策略利用内部自注意力机制所携带的结构信息作为层选择的信号。在 TruthfulQA 上的实验结果表明,我们的策略,特别是 Attention-JSD 和 Attention-Entropy-Min,始终优于原始的 DoLa。我们观察到在多答案指标(MC2 和 MC3)上有显著提升,这表明注意力分布可以提供比输出词汇分布更敏感的信号,用于解决事实知识问题。
cs.CL / 15 / 2607.23083

LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation

用于性别包容性重写和激活引导的 LoRA 方法在反叙事生成中的应用
P, Akhil Rajeev, J, Manoj Balaji
Abstract
Gender-inclusive language generation seeks to transform biased text into inclusive alternatives while preserving semantic meaning and contextual coherence. This paper presents the IHLC system for the LT-EDI 2026 Shared Task, addressing both gender-inclusive rewriting and counter-narrative generation. For gender-inclusive rewriting, we employ parameter-efficient Low-Rank Adaptation (LoRA) fine-tuning, achieving an official score of 80.00%. Our primary contribution is a compute-efficient inference-time representation engineering approach for counter-narrative generation. We derive a principal steering direction from contrastive hidden-state activations using principal component analysis (PCA) and inject it into the intermediate representations of Gemma-3-4B-it during inference, enabling behavioral steering toward inclusive responses without modifying model weights. Combined with constrained prompting, this approach produces polite and contextually appropriate counter-narratives, achieving an official score of 78.12%. We further present a manual analysis of steering behavior, identifying key failure modes including semantic drift, residual bias leakage, layer sensitivity, over-steering, and text degeneration. Our findings highlight both the practical potential and current limitations of activation steering as a lightweight alternative to parameter updates for controllable and socially aligned language generation.
Chinese Translation
性别包容性语言生成旨在将有偏见的文本转化为包容性替代品,同时保持语义意义和上下文连贯性。本文提出了 IHLC 系统,针对 LT-EDI 2026 共享任务,解决性别包容性重写和反叙事生成两个问题。在性别包容性重写方面,我们采用了参数高效的低秩适应(Low-Rank Adaptation, LoRA)微调,达到了 80.00% 的官方得分。我们的主要贡献是提出了一种计算高效的推理时表示工程方法用于反叙事生成。我们通过主成分分析(Principal Component Analysis, PCA)从对比隐藏状态激活中推导出主要引导方向,并在推理过程中将其注入 Gemma-3-4B-it 的中间表示中,从而实现了在不修改模型权重的情况下,向包容性响应的行为引导。结合约束提示,这种方法生成了礼貌且上下文适宜的反叙事,达到了 78.12% 的官方得分。我们还进行了手动分析以研究引导行为,识别出包括语义漂移、残余偏见泄漏、层敏感性、过度引导和文本退化等关键失效模式。我们的研究结果突显了激活引导作为可控和社会对齐语言生成的轻量级替代方案的实际潜力和当前局限性。
cs.CL / 16 / 2607.23142

Interview with Kalle Lyytinen on "Implications of Theories of Language for Information Systems"

关于“语言理论对信息系统的影响”的卡勒·吕伊廷访谈
Lyytinen, Kalle J., Maier, Pierre, Toffey, Paul Ackah
Abstract
Over fourty years after the initial publication of "Implications of Theories of Language for Information Systems" in MIS Quaterly, Lyytinen reflects about the origins of his publication and the developments in this area of research over the past decades. In the here presented interview, Lyytinen discusses the linguistic core of information systems also in light of recent trends and developments in the field, especially with regards to large language models and generative AI. Future research directions following a linguistic perspective on Information Systems (IS) research are outlined.
Chinese Translation
在《管理信息系统季刊》(MIS Quarterly)首次发表《语言理论对信息系统的影响》四十多年后,吕伊廷回顾了他这篇论文的起源以及过去几十年该研究领域的发展。在本次访谈中,吕伊廷讨论了信息系统的语言学核心,并结合该领域的最新趋势和发展,特别是关于大型语言模型和生成式人工智能的讨论。文章还概述了从语言学视角出发的信息系统(IS)研究的未来研究方向。
cs.CL / 17 / 2607.23175

Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining

超越全球规范:在不重新训练的情况下个性化语言模型的毒性敏感性
Diaconescu, Rares A. C., Slanina, Iulia, Florea, Alina, Trache, Andrei B., Coroi, Miruna E., Arzberger, Anne, Yang, Jie, Liscio, Enrico
Abstract
Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language generation to user-specific toxicity sensitivities across three inference-time intervention stages: pre-decoding (prompt conditioning and rewriting), in-decoding (token, logit, and representation steering), and post-decoding (candidate re-ranking). Evaluated against toxicity sensitivity targets derived from the PRISM dataset, all methods reduce alignment error by 28-47%. However, the results reveal a fundamental trade-off between alignment effectiveness, personalization, and general language quality, showing how toxicity sensitivity alignment is an inherently multi-objective problem.
Chinese Translation
减少毒性通常被视为一个全球对齐问题,但对有害语言的感知是主观且依赖于上下文的。我们首次对无训练方法进行比较评估,以使语言生成与用户特定的毒性敏感性对齐,涵盖三个推理时干预阶段:解码前(提示条件化和重写)、解码中(标记、逻辑值和表示引导)以及解码后(候选项重新排序)。根据来自PRISM数据集的毒性敏感性目标进行评估,所有方法均将对齐误差降低了28-47%。然而,结果揭示了对齐有效性、个性化和一般语言质量之间的根本权衡,表明毒性敏感性对齐本质上是一个多目标问题。
cs.CL / 18 / 2607.23242

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

IndicTalk:一个大规模基于角色的印度语言多语种对话语料库
Gawande, Sahil Deepak, Singh, Mayank
Abstract
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .
Chinese Translation
大型语言模型(LLMs)已经改变了对话式人工智能的格局,但高质量的多语种混合对话资源仍然稀缺,尤其是在印度语言中,讲者在母语和英语之间自然切换,且使用本土书写和罗马化形式。我们提出了IndicTalk,这是最大的多语种印度混合对话语料库之一,包含超过1,328,604个基于事件的多轮对话,涵盖9种印度语言的18种语言变体。该语料库通过一个完全自动化的流程生成,结合了现实世界新闻的基础、基于角色的对话生成(使用多语种LLMs)以及自动质量验证。广泛的语言学、自动化和人工评估表明,IndicTalk能够生成流畅、一致且自然混合的对话,适用于两种书写变体。我们将发布IndicTalk,以支持对代表性不足的印度语言的多语种对话人工智能的开发和评估。数据集可在以下网址获取:https://huggingface.co/datasets/LingoIITGN/IndicTalk 。
cs.CL / 19 / 2607.23278

Co-Evolving Graph and Text Memory for Training-Free Multi-Hop Question Answering

共演化图与文本记忆用于无训练的多跳问答
Man, Hieu, Nguyen, Thien Huu
Abstract
Multi-hop question answering requires coordinating relational and textual evidence across reasoning steps, a combination neither a text corpus nor a knowledge graph can supply alone. Prior work often emphasizes only part of this loop: graph-augmented RAG retrieves from a pre-built or query-updated graph, KGQA systems search within topic-centered subgraphs, and memory-augmented agents maintain evolving memories without continuously reconciling graph memory with textual context. We propose Co-E, a training-free system built around synchronized bidirectional graph-text working memory. A synchronization cycle consolidates textual memory, extracts relational triples into graph memory, and injects graph facts back into the generation context. Because both memories are maintained, they shape subsequent retrieval and generation. Evaluated on six multi-hop QA benchmarks, Co-E improves over comparable training-free open-backbone baselines and is competitive with larger or trained systems.
Chinese Translation
多跳问答需要在推理步骤中协调关系和文本证据,这种组合单靠文本语料库或知识图谱都无法单独提供。之前的研究通常只强调这一循环的部分:图增强的RAG从预构建或查询更新的图中检索,KGQA系统在以主题为中心的子图中进行搜索,而增强记忆的代理则维护不断演变的记忆,但未能持续协调图记忆与文本上下文。我们提出了Co-E,一个围绕同步双向图-文本工作记忆构建的无训练系统。同步周期整合文本记忆,将关系三元组提取到图记忆中,并将图事实注入生成上下文中。由于两种记忆都得以维护,它们影响后续的检索和生成。在六个多跳问答基准测试中评估,Co-E在可比的无训练开放骨干基线之上有所改进,并且与更大或经过训练的系统具有竞争力。
cs.CL / 20 / 2607.23319

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

BHARATI:针对古典印度语言的形态学感知分词器及其子词丰度分析
Kumaresan, Poornima, Muruganantham, Pavithra, Rajendran, Lakshmi, Sivasubramani, Santhosh
Abstract
Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages. Sanskrit, Tamil, and other classical Indic languages exhibit agglutinative morphology, productive sandhi (phonological fusion at word boundaries), and domain-specific vocabularies absent from general-purpose training data. This paper presents BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced 781 MB corpus spanning seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam) with native script support for all languages. We describe three successive tokenizer versions: v1 (English and Sanskrit only, with broken byte-fallback for Tamil), v2 (four-language support with byte-level fallback for southern languages), and v3 (full seven-language native subword coverage). Subword fertility analysis demonstrates that v3 averages 2.6 tokens per Indian Knowledge System (IKS) technical term, compared to 5.25 tokens per term with GPT-2's tokenizer and 3.75 tokens with the multilingual SentencePiece baseline, with the largest gains on a set of reserved IKS terms that are represented as single tokens by construction. On a held-out test set of 490 IKS-domain sentences (70 per language across seven languages, released with the measurement script), v3 reduces sequence length by roughly 90% relative to GPT-2 and byte-level encoding (which lack native Indic subwords) and by approximately 25% relative to the mBART-50 multilingual baseline, averaged across the six Indic languages, directly translating to increased effective context length for downstream language models. The tokenizer models (32,000 vocabulary), training scripts, and evaluation benchmarks are released under open licenses.
Chinese Translation
标准的子词分词算法,如字节对编码(Byte-Pair Encoding, BPE)和SentencePiece,主要在现代语言语料库上进行训练,因此在应用于古典印度语言时产生了低效的分段。梵语、泰米尔语及其他古典印地语言展现出粘合形态、富有生产性的音变(在词边界的音韵融合)以及在通用训练数据中缺失的特定领域词汇。本文提出了BHARATI,一组基于SentencePiece BPE的分词器,训练于一个平衡的781 MB语料库,涵盖七种语言(英语、印地语、梵语、泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语),并为所有语言提供本土书写支持。我们描述了三个连续的分词器版本:v1(仅支持英语和梵语,泰米尔语采用破碎字节回退),v2(支持四种语言,南方语言采用字节级回退),以及v3(全面支持七种语言的本土子词)。子词丰度分析表明,v3在每个印度知识体系(Indian Knowledge System, IKS)技术术语上平均产生2.6个子词,而GPT-2的分词器为5.25个,跨多语言的SentencePiece基线为3.75个,尤其在一组被保留的IKS术语上,v3将其表示为单个子词。针对490个IKS领域句子的保留测试集(每种语言70个,涵盖七种语言,并随测量脚本发布),v3相较于GPT-2和字节级编码(缺乏本土印地子词)将序列长度减少了约90%,相较于mBART-50多语言基线则减少了约25%,在六种印地语言中平均计算,直接转化为下游语言模型的有效上下文长度的增加。分词器模型(32,000词汇)、训练脚本和评估基准在开放许可下发布。
cs.CL / 21 / 2607.23322

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

IKS-Instruct:一个包含24,000个示例的多语言数据集,用于教授语言模型印度知识体系
Singaravelu, Shwetha, Muruganantham, Gayathri, Rajendran, Lakshmi, Sivasubramani, Santhosh
Abstract
Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains. This paper presents IKS-Instruct, a dataset of 24,795 instruction-response pairs for teaching language models to deliver educational content grounded in Indian Knowledge Systems (IKS). The dataset spans seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam), covers 41 pedagogical techniques from the Vedic oral and mathematical traditions, and is aligned with the Central Board of Secondary Education (CBSE) curriculum for classes 6 through 12. The pairs are derived from six source types: classical text corpora (Bhagavad Gita, Thirukkural, Sangam literature, Vedic texts), curriculum-aligned pedagogical templates, Vedic mathematical sutra demonstrations, bilingual instruction pairs, technique-grounded multi-turn dialogues, and cross-tradition comparative analyses. Quality is assessed through a multi-judge evaluation framework in which independent language models score responses on 12 dimensions including technique fidelity, pedagogical quality, factual accuracy, and IKS cultural depth. Under a uniform five-judge external panel (median aggregation over 1,201 stratified items), the strongest IKS-Instruct fine-tune of a compact 7B model reaches a median judge score of 6.39, within 0.15 of a strong general-purpose reference model (Nemotron-Nano at 6.54) at a fraction of its deployment cost, while the base model without IKS fine-tuning scores near zero on the IKS-specific dimensions. Model quality does not increase monotonically with data curation, a result we report together with the corresponding data-quality gains.
Chinese Translation
指令调优已成为将大型语言模型适应人类意图的标准方法,但现有的指令数据集主要集中在英语的通用知识任务上,缺乏对专业教育领域的覆盖。本文提出了IKS-Instruct,这是一个包含24,795个指令-响应对的数据集,旨在教授语言模型提供基于印度知识体系(IKS)的教育内容。该数据集涵盖七种语言(英语、印地语、梵语、泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语),涉及41种来自吠陀口头和数学传统的教学技术,并与中央中等教育委员会(CBSE)6至12年级的课程相一致。这些对来自六种来源类型:经典文本语料库(《博伽梵歌》,《提鲁库拉尔》,桑伽姆文学,吠陀文本)、与课程对齐的教学模板、吠陀数学公式演示、双语指令对、基于技术的多轮对话以及跨传统的比较分析。通过多评审评估框架评估质量,其中独立语言模型在包括技术保真度、教学质量、事实准确性和IKS文化深度等12个维度上对响应进行评分。在一个统一的五评审外部小组(对1,201个分层项目进行中位数聚合)下,最强的IKS-Instruct对紧凑型7B模型的微调达到了6.39的中位数评分,与强通用参考模型(Nemotron-Nano为6.54)相差0.15,且其部署成本仅为后者的一小部分,而未经过IKS微调的基础模型在IKS特定维度上的得分接近于零。模型质量并未随着数据整理而单调增加,我们报告了这一结果以及相应的数据质量提升。
cs.CL / 22 / 2607.23344

BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi

基于BERT的模型与大型语言模型在低资源命名实体识别中的比较研究:以马拉地语为例
Ingle, Hariom, Ghode, Ronit, Gondkar, Ishwari, Harad, Jidnyasa, Joshi, Raviraj
Abstract
Named Entity Recognition (NER) for low-resource languages such as Marathi remains a challenging task due to limited annotated resources and linguistic complexity. Although recent Large Language Models (LLMs) have demonstrated strong performance across a wide range of natural language processing tasks, their effectiveness for language-specific NER in low-resource settings remains uncertain. In this study, we fine-tune MahaBERT-v2 on different variants of the MahaNER dataset and systematically compare the performance of these models with an existing MahaNER baseline and prominent general-purpose LLMs, including Gemini, LLaMA-3.3-70B, and Gemma models. All models are evaluated on a Marathi NER test dataset using standard metrics of precision, recall, and F1-score. The experimental results show that the fine-tuned MahaBERT-based models consistently outperform both the baseline and all evaluated LLMs, with the fine-tuned models achieving F1-scores ranging from 0.88 to 0.91, surpassing the existing MahaNER model (0.8843) and significantly exceeding the performance of LLM-based approaches, whose F1-scores range from 0.57 to 0.69. These findings demonstrate that task-specific, language-focused models trained on domain-relevant data remain more effective than general-purpose LLMs for Marathi NER, highlighting the continued importance of specialized architectures for low-resource language processing.
Chinese Translation
低资源语言(如马拉地语)的命名实体识别(NER)由于注释资源有限和语言复杂性,仍然是一项具有挑战性的任务。尽管最近的大型语言模型(LLMs)在广泛的自然语言处理任务中表现出色,但它们在低资源环境下对特定语言的NER的有效性仍然不确定。在本研究中,我们对MahaBERT-v2进行了微调,使用不同变体的MahaNER数据集,并系统地比较了这些模型与现有MahaNER基线及一些著名的通用LLMs(包括Gemini、LLaMA-3.3-70B和Gemma模型)的性能。所有模型均在马拉地语NER测试数据集上使用标准的精确度、召回率和F1分数进行评估。实验结果表明,微调后的MahaBERT模型在性能上始终优于基线和所有评估的LLMs,微调后的模型F1分数范围为0.88到0.91,超越了现有的MahaNER模型(0.8843),并显著超过了基于LLM的方法,其F1分数范围为0.57到0.69。这些发现表明,针对特定任务、专注于语言的模型在相关领域数据上训练后,仍然比通用LLMs在马拉地语NER中更有效,突显了专门架构在低资源语言处理中的持续重要性。
cs.CL / 23 / 2607.23362

Joint Optimization for Greedy Longest-match Tokenization

贪婪最长匹配分词的联合优化
Singh, Adhiraj, Mody, Deepanshu, Shdaifat, Ghina Al, Alshamy, Hamza, Wiemerslage, Adam, Reddy, Varshini, Schmidt, Craig W.
Abstract
Recent work has shown that subword vocabularies can be trained to optimize compression for a specific inference rule rather than relying on greedy heuristics such as Byte Pair Encoding (BPE). We extend this approach to greedy left-to-right longest-match decoding, the fast and widely used inference rule underlying WordPiece. We introduce Joint Optimization for Greedy Longest-Match Tokenization (JOLT), which formulates vocabulary learning as an integer program over vocabulary-selection and segmentation-choice variables. Greedy-consistency constraints ensure that each optimized segmentation exactly matches the segmentation produced by longest-match decoding under the selected vocabulary, aligning the training objective with deployment-time tokenization. To scale the optimization, we solve a linear programming relaxation and selectively introduce higher-order segmentations only for unresolved pretokens. The resulting relaxation is nearly integral: rounded solutions fall within 0.008 - 0.176 % of the LP lower bound on the training scope. The bound also shows that BPE is already within 1 - 2 % of the best achievable compression under greedy longest-match decoding, while JOLT closes 89.6 - 99.4 % of the remaining gap. On held-out validation data across four training scopes and vocabulary sizes of 32,000 and 64,000, JOLT produces up to 0.78 % fewer tokens than BPE, with improvements generally increasing as the training scope grows. These results demonstrate that inference-aligned vocabulary optimization can recover most of the limited compression headroom left by BPE while providing a certificate of near-optimality.
Chinese Translation
近期的研究表明,子词词汇可以被训练以优化特定推理规则的压缩,而不是依赖于贪婪启发式方法,如字节对编码(Byte Pair Encoding, BPE)。我们将这种方法扩展到贪婪的从左到右的最长匹配解码,这是支撑WordPiece的快速且广泛使用的推理规则。我们提出了贪婪最长匹配分词的联合优化(Joint Optimization for Greedy Long-Match Tokenization, JOLT),将词汇学习形式化为一个关于词汇选择和分段选择变量的整数规划。贪婪一致性约束确保每个优化的分段与所选词汇下的最长匹配解码产生的分段完全一致,从而将训练目标与部署时的分词对齐。为了扩展优化,我们解决了线性规划松弛,并仅对未解决的预标记引入选择性高阶分段。结果的松弛几乎是整数的:四舍五入的解在训练范围内落在线性规划下界的0.008% - 0.176%之间。该界限还表明,BPE在贪婪最长匹配解码下已经接近最佳可实现压缩的1% - 2%,而JOLT则缩小了剩余差距的89.6% - 99.4%。在四个训练范围和32,000及64,000的词汇大小的保留验证数据上,JOLT产生的标记数量比BPE少最多0.78%,且随着训练范围的扩大,改进通常会增加。这些结果表明,与推理对齐的词汇优化可以恢复BPE所留下的大部分有限压缩空间,同时提供近似最优性的证明。
cs.CL / 24 / 2607.23379

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

当激活预言机学会不去读取:微调预言机中的概念特定盲点
Bersia, Tobias, Gaintseva, Tatiana
Abstract
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training. This failure is not simply explained by absence of the concept from the subject or oracle representations: the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses indicate that the failure arises in the AO readout pathway. Our results show that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising a reliability concern for learned interpretability interfaces.
Chinese Translation
激活预言机(Activation Oracles, AOs)是经过训练的语言模型,用于回答关于另一个模型内部激活的自然语言问题。它们提供了一种灵活的接口,用于从模型状态中读取隐藏信息,尤其是在相关信息在内部表示但在可见行为中缺失或不完整时。然而,AOs 本身是学习系统:它们的回答受到训练数据、目标和学习的报告行为的影响,而不是对所表示信息的中立读取。我们在一个受控的禁忌词猜测设置中研究这一点,在该设置中,主题模型经过微调以在内部使用隐藏概念,同时避免直接披露。与预期的AOs在此类主题上训练后成为专业读取者的想法相反,我们发现微调后的AOs可能成为概念特定的反读取者:它们选择性地未能恢复在自身训练期间持续存在的概念。这一失败并不能简单地通过主题或预言机表示中缺少该概念来解释:目标在预言机内部仍然是可解码的,而 LogitLens 和层消融分析表明,失败发生在 AO 的读取路径中。我们的结果表明,行为泄漏、表示级别的可解码性和 AO 的可言说性可能会分离,这引发了对学习可解释性接口的可靠性担忧。
cs.CL / 25 / 2607.23420

LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction

LA-RL:基于标签的自我反思强化学习在信息提取中的应用
You, Xiao, Yan, Tianwei, Shan, Zixu, Du, Longyu, Zhao, Shan
Abstract
Large language models show strong promise for information extraction (IE), but existing reflection-based correction methods are often misaligned with structured extraction outputs. Free-form self-reflection can flag an error, yet it rarely identifies whether the failure is a missing span, wrong label, boundary mismatch, invalid relation type, or reversed argument order. We introduce LA-RL (Label-Aware Reflective Reinforcement Learning), an outcome-supervised framework that guides IE self-correction with task-grounded diagnostic labels. A single backbone first predicts an extraction, diagnoses task-specific error labels, and then revises its output conditioned on the diagnosis. Training starts from diagnostic data labeled by an annotation model for cold-start supervised fine-tuning and proceeds through two GRPO stages that reward final extraction quality, format validity, and first-pass correctness, without a process reward model. Experiments on named entity recognition, relation extraction, and event extraction show consistent same-backbone gains over SFT, including 6.83 average F1 on SciER relation extraction, about 20 F1 on out-of-distribution relation extraction, and 14.80 trigger F1 plus 17.50 argument F1 on DuEE1.0. Ablations show that reflection structure is task-sensitive: stronger constraints benefit relation extraction, whereas named entity recognition needs less restrictive correction under domain shift.
Chinese Translation
大型语言模型在信息提取(IE)方面展现出强大的潜力,但现有的基于反思的纠正方法往往与结构化提取输出不一致。自由形式的自我反思可以标记错误,但很少能识别失败是由于缺失的跨度、错误的标签、边界不匹配、无效的关系类型或反向的论证顺序。我们提出了LA-RL(基于标签的反思强化学习),这是一个结果监督框架,通过任务基础的诊断标签指导IE自我纠正。一个单一的主干网络首先预测提取结果,诊断特定任务的错误标签,然后根据诊断修正其输出。训练从由注释模型标记的诊断数据开始,以进行冷启动的监督微调,并通过两个GRPO阶段进行,奖励最终提取质量、格式有效性和第一次通过的正确性,而不使用过程奖励模型。在命名实体识别、关系提取和事件提取的实验中,显示出相对于SFT的一致性同主干增益,包括在SciER关系提取中平均6.83的F1,在分布外关系提取中约20的F1,以及在DuEE1.0中14.80的触发F1和17.50的论证F1。消融实验表明反思结构对任务敏感:更强的约束有利于关系提取,而命名实体识别在领域转移下需要较少的限制性纠正。
cs.CL / 26 / 2607.23440

Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

推理还是记忆:大型语言模型能理解和生成中文歇后语谜语吗?
Hu, Hai, Song, Siyuan, Shao, Chongtian, Zhang, Kejia, Zhu, Tianjian, Zhao, Xiaojing
Abstract
In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($\Delta_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $\Delta_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $\Delta_{acc}$ of 23.6\%, while English-centric models tested have a mean $\Delta_{acc}$ of 5.1\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs' creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.
Chinese Translation
在本文中,我们通过在中文语言游戏歇后语中测试大型语言模型(LLMs),推动了其推理能力的边界。我们使用由语言学家创作的全新歇后语,以避免数据污染。我们通过多项选择题(MCQ)、自由形式的解释生成和新歇后语的创作来评估LLMs理解和创作歇后语的能力。在多项选择题中,我们使用现有但低频的歇后语与新歇后语之间的准确率差异($ ext{Δ}_{ ext{acc}}$)作为记忆的指标。母语者的$ ext{Δ}_{ ext{acc}}$非常低,表明其处理机制相似。然而,我们发现前沿的中文模型的平均$ ext{Δ}_{ ext{acc}}$为23.6\%,而测试的以英语为中心的模型的平均$ ext{Δ}_{ ext{acc}}$为5.1\\%,这表明前沿中文模型可能使用了更大规模的中文数据,从而记忆了更多低频的歇后语。对于新歇后语,Gemini 3.1 Pro展现出了卓越的能力,准确率为92.6,超过人类准确率24\\%。在歇后语创作中,LLMs创作的歇后语评分远低于人类创作的。这些结果表明,关于LLMs推理能力的主张可能需要在考虑数据污染问题后进行仔细重新审视,并且LLMs在语言相关任务中的创造力可能仍然落后于人类专家,至少在中文歇后语方面是如此。
cs.CL / 27 / 2607.23442

Do LLM Debates Repeat Arguments Differently Across Languages?

大型语言模型辩论是否在不同语言中以不同方式重复论点?
Lai, Huiqian
Abstract
LLM debate is usually evaluated by final answers, but transcripts also reveal whether later turns develop new argumentative content or return to earlier claims in new wording. We study this process with \textit{prior-argument similarity}, an aggregate diagnostic comparing extracted argument units with earlier units in the same debate. In controlled eight-turn debates over 71 motions, six languages, and four model agents, Chinese is the only tested language with a consistently positive gap relative to English across three multilingual embedding models. The gap persists across agents, turn positions, regression adjustment, metric variants, extraction-length controls, a second-extractor subset, and cross-encoder tail rescoring. Manual calibration shows weak item-level alignment but a high-similarity tail enriched for substantive repetition. A diversity-aware prompt lowers prior-argument similarity across languages, yet does not significantly narrow the Chinese--English gap. These findings suggest that multilingual debate evaluation should measure argumentative development over time and report mitigation effects in both average and gap terms.
Chinese Translation
大型语言模型辩论通常通过最终答案进行评估,但转录文本也揭示了后续发言是否发展出新的论证内容,或以新的措辞回归早期主张。我们通过 extit{先前论点相似性}这一聚合诊断工具来研究这一过程,该工具比较提取的论证单元与同一辩论中早期单元的相似性。在对71个议题、六种语言和四个模型代理的八轮控制辩论中,中文是唯一一种在三个多语言嵌入模型中相对于英语始终保持正差距的测试语言。该差距在不同代理、发言位置、回归调整、度量变体、提取长度控制、第二提取子集和交叉编码尾部重评分中持续存在。手动校准显示出较弱的项目级对齐,但在实质性重复方面表现出高度相似的尾部。关注多样性的提示降低了不同语言间的先前论点相似性,但并未显著缩小中英文之间的差距。这些发现表明,多语言辩论评估应衡量论证的发展过程,并在平均值和差距方面报告缓解效应。
cs.CL / 28 / 2607.23446

Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh

小模型是否能使用您提供的法律?针对孟加拉国法律问答的上下文注入微调
Mahadi, Moniruzzaman, Alam, Abrar Mohammed Tanzim, Monalisa, Sayma Siddika, Abdullah, Mir Mohammad Asif, Shatabda, Swakkhar, Arefeen, Md Adnan
Abstract
A small language model can receive the governing statutory provision and still answer incorrectly. We test whether fine-tuning on examples containing relevant law improves later use of retrieved law. We curate 2{,}165 bilingual QA records from six Bangladeshi acts and three schedules, then fine-tune Qwen3.5 at 0.8B, 2B, and 4B. Evaluation uses the 2022 and 2023 Bangladesh Bar Council exams in Bangla and machine-translated English, with no retrieval, BM25, or FAISS, scored by strict consistency over three seeded runs. At 0.8B, fine-tuning raises the 2022 English FAISS score from 2 to 34 of 100. Gains at 0.8B and 2B survive paired testing, but the 4B model has no detectable net gain: Bangla improves while several English conditions regress. Fine-tuning also reduces answers that drift from Bangla into mostly English from 44.0--53.2\% to 0.2--0.7\%, with adjusted $p<.001$ at every scale. Retrieval quality is therefore not the only bottleneck. Small bilingual legal models also differ in how they use supplied law and whether they answer in the requested language. The dataset is publicly available at https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.
Chinese Translation
一个小型语言模型可以接收相关的法定条款,但仍然可能回答错误。我们测试了在包含相关法律的示例上进行微调是否能改善后续检索法律的使用。我们从六部孟加拉国法律和三个附表中整理了2165条双语问答记录,然后对Qwen3.5模型在0.8B、2B和4B进行微调。评估使用2022年和2023年孟加拉国律师协会考试的孟加拉语和机器翻译的英语,未使用检索、BM25或FAISS,评分依据三次种子运行的一致性。在0.8B时,微调将2022年英语FAISS得分从2提升至34(满分100)。在0.8B和2B的配对测试中,提升依然显著,但4B模型没有可检测的净增益:孟加拉语有所改善,而多个英语条件则出现退步。微调还将从孟加拉语漂移到主要为英语的回答比例从44.0-53.2%降低至0.2-0.7%,在每个规模下调整后的$p<.001$。因此,检索质量并不是唯一的瓶颈。小型双语法律模型在使用提供的法律和是否用请求的语言回答方面也存在差异。数据集可在https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset公开获取。
cs.CL / 29 / 2607.23458

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

链式思维不忠实的两种机制:模型错误时行为检测失败
Angdembay, Suramya R., Aryal, Dikshant, Rahimi, Nick
Abstract
Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-box (behavioral) detection of unfaithful CoT against FaithCoT-Bench's human annotations, we find answer correctness structures the problem at every level. Answer incorrectness alone (an oracle diagnostic, not a deployable detector) outperforms every purpose-built signal (AUROC 0.696), because 69% of annotated unfaithfulness occurs on incorrect answers. Stratifying by correctness splits detection into two regimes: on correct answers, behavioral signals moderately separate faithful from post-hoc reasoning (0.63-0.67); on incorrect answers, where most unfaithfulness lives, no tested signal is detectably above chance (replicated on all four models for benchmark-wide signals). The standard step-removal metric anti-correlates with human labels; this inversion reproduces on the benchmark's released scores and on hint-dependent counterfactually labeled traces. Linear probes decode the behaviorally blind regime in Llama-3.1-8B and the correct-answer regime in Qwen-2.5-7B, with no shared, positively aligned direction detected across regimes; instructed answer-first traces (7 models) transfer to neither annotated regime, while hint-induced unverbalized answer flips do, in model- and source-dependent settings. We also independently verify and resolve a documentation-data mismatch in the benchmark's label semantics.
Chinese Translation
链式思维(CoT)解释只有在其忠实的情况下才能支持监督:所述推理必须实际产生答案。在对黑箱(行为)检测不忠实链式思维进行审计时,我们发现答案的正确性在每个层面上都构成了问题的结构。仅仅依赖答案的错误性(一个神谕诊断,而非可部署的检测器)在每个专门构建的信号中表现优于其他信号(AUROC 0.696),因为69%的标注不忠实发生在错误答案上。根据正确性进行分层将检测分为两种机制:在正确答案上,行为信号适度区分忠实与事后推理(0.63-0.67);而在错误答案上,大多数不忠实存在的地方,没有测试信号的检测结果明显高于随机水平(在所有四个模型上对基准广泛信号进行了重复验证)。标准的步骤移除指标与人类标签呈反相关;这种反转在基准的发布分数和依赖提示的反事实标记轨迹中得以再现。线性探测器在 Llama-3.1-8B 的行为盲区和 Qwen-2.5-7B 的正确答案区解码,但在不同机制之间未检测到共享的、正相关的方向;指令优先的答案轨迹(7个模型)未能转移到任何标注机制,而提示诱导的未言表答案翻转则在模型和源依赖的设置中得以转移。我们还独立验证并解决了基准标签语义中的文档数据不匹配问题。
cs.CL / 30 / 2607.23481

Mwando: Leveraging AI to Preserve and Teach shiKomori

Mwando:利用人工智能保护和教授shiKomori语言
Mohamed, Naira Abdou, Ali, Haidar Nassur Said, Hazra, Mohamed, Soibira, Naoufal Mohamed, Yamani, Roushnaty Ali
Abstract
This paper presents Mwando, a virtual educational assistant designed to support the teaching and preservation of shiKomori, the language of the Comoros Islands. The system covers the four main dialectal variants (shiNgazidja, shiMwali, shiNdzuani and shiMaore) through a knowledge base constructed from phrases, proverbs, dictionaries and grammar lessons. A multi-agent architecture combining vector search, a knowledge graph and web search fallback enables accurate and context-aware responses. Evaluation on 500 queries demonstrates strong performance on vocabulary lookup and grammar explanations, while qualitative case studies illustrate both capabilities and current limitations. This work represents an initial step toward computational support for shiKomori and provides a blueprint for developing AI-powered educational tools for other low-resource languages.
Chinese Translation
本文介绍了Mwando,一个旨在支持shiKomori语言教学和保护的虚拟教育助手,该语言是科摩罗群岛的语言。该系统涵盖了四种主要方言变体(shiNgazidja、shiMwali、shiNdzuani和shiMaore),通过由短语、谚语、词典和语法课构建的知识库实现。一个结合了向量搜索、知识图谱和网络搜索回退的多智能体架构使得系统能够提供准确且具有上下文意识的响应。对500个查询的评估显示,在词汇查找和语法解释方面表现出色,同时定性案例研究展示了系统的能力和当前的局限性。这项工作代表了对shiKomori语言计算支持的初步探索,并为开发其他低资源语言的人工智能教育工具提供了蓝图。
cs.CL / 31 / 2607.23512

The Cross-Domain Generalization Cost of Offensive Language Detection

攻击性语言检测的跨领域泛化成本
Ren, Ruixing, Zhao, Junhui, Sun, Xiaoke, Li, Qiuping
Abstract
Offensive language detection models generally suffer performance degradation when deployed across datasets and across languages, yet most existing studies stop at reporting this phenomenon and lack a systematic methodology for decomposing the causes of degradation into attributable components and quantifying the cost of remediation. This paper proposes a diagnosis and optimization framework composed of three coordinated technical components. First, a zero-shot transfer loss decomposition that separates the performance degradation from OLID to MLMA into two independently measurable components, namely dataset effect and language effect. Second, a controlled fine-tuning protocol that quantifies both adaptation efficiency and the hidden damage inflicted on the source task by comparing few shot learning curves under continued fine-tuning and cold-start starting points. Third, three joint training strategies incorpo rating temperature sampling and experience replay, which offer a controllable Pareto trade-off between improving multilingual capability and preserving source-task performance. Experiments built on this framework show that the dataset effect dominates the zero-shot transfer loss and substantially outweighs the language effect. Few-shot adaptation without a replay mechanism, though data-efficient, inflicts source task damage 4 to 9 times greater than that of the joint training strategies, and its damage magnitude is highly unstable. The three joint training strategies trade 3.2 to 4.1 percentage points of source-task performance for 8.1 to 42.6 percentage points of multilingual capability gain, forming a clear and controllable Pareto trade-off.
Chinese Translation
攻击性语言检测模型在跨数据集和跨语言部署时通常会遭遇性能下降,然而大多数现有研究仅停留在报告这一现象,缺乏系统的方法论来将性能下降的原因分解为可归因的组成部分,并量化修复成本。本文提出了一种由三个协调技术组件组成的诊断与优化框架。首先,提出了一种零样本迁移损失分解方法,将从 OLID(Offensive Language Identification Dataset)到 MLMA(Multilingual Language Model Adaptation)的性能下降分解为两个可独立测量的组成部分,即数据集效应和语言效应。其次,提出了一种受控微调协议,通过比较在持续微调和冷启动起点下的少样本学习曲线,量化适应效率及其对源任务造成的隐性损害。第三,提出了三种联合训练策略,结合温度采样和经验重放,提供了在提高多语言能力和保持源任务性能之间的可控帕累托权衡。基于该框架的实验表明,数据集效应主导了零样本迁移损失,且远大于语言效应。尽管没有重放机制的少样本适应在数据效率上表现良好,但其对源任务造成的损害是联合训练策略的4到9倍,且损害幅度高度不稳定。这三种联合训练策略在源任务性能上牺牲3.2到4.1个百分点,以换取8.1到42.6个百分点的多语言能力提升,形成了明确且可控的帕累托权衡。
cs.CL / 32 / 2607.23513

Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning

图示是否有助于大型语言模型推理?来自三段论推理的证据
Ando, Risako, Mineshima, Koji
Abstract
Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve human reasoning performance. Recent work has also explored their effects on large language models (LLMs). In this paper, we compare four representational conditions for syllogistic reasoning: natural language, logical notation, linear diagrams, and Euler diagrams. Using 285 problems from Ando et al. (2024), we evaluate two contemporary LLMs, Claude 3.5~Sonnet and GPT-4o-mini. Our results show that diagrammatic representations do not consistently improve performance. Although the models perform well on entailment and contradiction problems, they struggle with neutral problems and often make systematic conversion errors. Overall, the results suggest that the tested models gain limited benefit from diagrams in logical reasoning tasks.
Chinese Translation
图示被广泛用于支持逻辑推理,先前的研究表明,像欧拉图(Euler diagrams)这样的表示方式可以提高人类的推理表现。最近的研究也探讨了它们对大型语言模型(LLMs)的影响。本文比较了四种三段论推理的表示条件:自然语言、逻辑符号、线性图示和欧拉图。我们使用了来自Ando等人(2024)的285个问题,评估了两个当代大型语言模型,Claude 3.5~Sonnet和GPT-4o-mini。我们的结果表明,图示表示并未始终改善性能。尽管模型在蕴含和矛盾问题上表现良好,但在中立问题上却表现不佳,并且经常出现系统性的转换错误。总体而言,结果表明,所测试的模型在逻辑推理任务中从图示中获得的益处有限。
cs.CL / 33 / 2607.23514

Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

新颖的主张还是似曾相识?重新思考“无污染”动态评估在多模态自动化事实核查中的应用
He, Haorui, Chen, Xinwen, Wen, Dacheng, Cheng, Reynold, Lau, Francis C. M., Li, Yupeng
Abstract
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.
Chinese Translation
多模态自动化事实核查(MAFC)通过检索和推理外部证据来验证主张。然而,大多数现有的静态基准面临污染风险:它们主要由过时的主张组成,这些主张可以仅通过大型语言模型(LLM)的内部知识进行验证,而无需外部证据。这可能会夸大性能估计,并未能反映在需要最新信息的全新主张上的真实能力。为了解决这个问题,新兴的动态基准收集在LLM知识截止日期之后发布的主张,假设这些主张没有污染。本研究通过实证研究最先进(SOTA)静态 AVeriTeC 基准和我们新构建的动态 ClaimReview2025Q4 基准中的污染风险,以及它们对 MAFC 评估的影响,重新审视了这一假设。我们的实验得出了16个发现,突出了三个关键结果:(1)动态评估减少但并未消除污染风险,因为17.09%--29.30%的截止后主张仍然可能受到污染;(2)许多新发布的主张可以通过直接验证或合成截止前可用的多条公共知识进行验证;(3)污染可能导致 MAFC 性能的统计显著膨胀,Macro-F1 分数最高增加11.34点,并扭曲系统排名。鉴于这些发现,我们在严格控制污染的环境下重新评估了 SOTA LLM。我们的研究为可信的 MAFC 评估提供了实用指南。
cs.CL / 34 / 2607.23531

The JEPA Paradox in Language: The Geometry of Linguistic Alternatives

语言中的JEPA悖论:语言选择的几何学
Dinh, Anh Trac Duc, Vo, Khang Nhat Hoang
Abstract
Joint-Embedding Predictive Architectures (JEPAs) are effective for images, video, and audio, yet deterministic JEPA-style latent prediction has not become a standard objective for text encoders. We argue that this gap reflects a mismatch between squared-error latent prediction and the conditional structure of language. The key requirement is conditional concentration: given a context and target location, the target representation should lie near a single meaningful point. Local image prediction often satisfies this through spatial continuity, whereas masked text can admit multiple valid token or span completions whose representations need not share a coherent center. We formalize this mismatch through three conditions---predictability, non-collapse, and low conditional variance---and show how their failure creates centroid degeneracy and collapse pressure in text. Matched I-JEPA and T-JEPA experiments reveal the predicted sequence: mutual-information saturation and elevated target variance precede train--validation instability, effective-rank degeneration, cosine collapse, and poor downstream transfer. The same pattern appears across five independent data seeds, indicating that it is not a sampling artifact. These results do not rule out predictive learning for language; they show that text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point.
Chinese Translation
联合嵌入预测架构(Joint-Embedding Predictive Architectures, JEPA)在图像、视频和音频中表现出色,但确定性的JEPA风格潜在预测尚未成为文本编码器的标准目标。我们认为,这一差距反映了平方误差潜在预测与语言的条件结构之间的不匹配。关键要求是条件集中性:给定上下文和目标位置,目标表示应接近一个有意义的单一点。局部图像预测通常通过空间连续性满足这一要求,而被遮蔽的文本则可以接受多个有效的标记或跨度补全,其表示不必共享一个连贯的中心。我们通过三个条件——可预测性、非崩溃性和低条件方差——形式化这一不匹配,并展示它们的失败如何在文本中造成质心退化和崩溃压力。匹配的I-JEPA和T-JEPA实验揭示了预测序列:互信息饱和和目标方差升高先于训练-验证不稳定性、有效秩退化、余弦崩溃和较差的下游迁移。相同的模式出现在五个独立的数据种子中,表明这不是采样伪影。这些结果并不排除语言的预测学习;它们表明,文本兼容的JEPA目标必须保留多个合理的补全,而不是将其压缩为一个单一的潜在点。
cs.CL / 35 / 2607.23538

Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

引导语言模型更具同理心:通过人类与大型语言模型的合作生成文化敏感的心理健康建议
Faria, Fatema Tuj Johora, Moin, Mukaffi Bin, Rahman, Md. Mahfuzur, Hasib, Khan Md, Mahmud, Jubayer Al, Mridha, M. F.
Abstract
Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program "Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.
Chinese Translation
尽管大型语言模型(LLMs)近期取得了进展,但它们在低资源语言中生成同理心心理健康咨询回应的能力仍然大多未被探索。为了解决这一问题,我们从三个互补来源整理了625个真实的心理健康案例:(1)公开可得的Facebook帖子,讨论心理健康问题;(2)孟加拉国电视节目《Ami Akhon Ki Korbo》的文字记录;(3)涵盖多种情感和心理挑战的匿名学生问卷回复。基于这些案例,我们构建了一个评估语料库,其中包含由持证临床心理学家撰写的建议和由三种现代专有LLM生成的回应:GPT-4o Mini、Claude 4.5 Haiku和Gemini 2.5 Pro。我们进一步提出了角色扮演反思链式建议框架(RP-RCAF),这是一种任务特定的提示策略,结合了专家撰写的少量示例与结构化自我反思,以通过富有同情心的顾问角色生成支持性、文化敏感和伦理一致的咨询。我们还引入了基于Grok 4的回应评估与评分框架(G-REFS),该框架将自动评估与专家心理学家的验证结合,涵盖情感敏感性、文化适宜性、语言清晰度和伦理合理性。实验结果表明,RP-RCAF在所有评估模型中始终优于传统提示,并生成与专业心理咨询更为一致的回应。
cs.CL / 36 / 2607.23545

Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

语言塑造多语言大型语言模型中的指令层级遵从性
Moon, Jiwon, Hwang, Yerin, Jung, Kyomin
Abstract
Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones. Despite its importance for safe and controllable deployment, existing evaluations have focused almost exclusively on English, leaving it unclear whether IH compliance remains stable in multilingual settings. We introduce XIH-Bench, a benchmark for multilingual IH evaluation with both same-language and cross-language conflicts across six languages, four domains, and three IH settings. Across models, we find two consistent patterns. First, IH compliance exhibits a clear language-dependent asymmetry: a language that strengthens compliance in the higher-priority position can become disruptive in the lower-priority position. Second, cross-language conflicts yield higher compliance than same-language conflicts, a phenomenon we term the Language Boundary Effect. We further show that language specialization can make lower-priority instructions in model-favored languages harder to override, creating multilingual reliability and security risks.
Chinese Translation
指令层级(Instruction hierarchy, IH)要求模型根据来源优先处理指令,确保高优先级指令能够覆盖低优先级指令。尽管这一点对于安全和可控的部署至关重要,但现有评估几乎完全集中于英语,这使得在多语言环境中IH遵从性是否保持稳定变得不明确。我们引入了XIH-Bench,这是一个用于多语言IH评估的基准,涵盖六种语言、四个领域和三种IH设置,包含同语言和跨语言冲突。我们发现模型之间存在两个一致的模式。首先,IH遵从性表现出明显的语言依赖性不对称:在高优先级位置上增强遵从性的语言在低优先级位置上可能会产生干扰。其次,跨语言冲突的遵从性高于同语言冲突,这一现象我们称之为语言边界效应(Language Boundary Effect)。我们进一步表明,语言专业化可能使得在模型偏好的语言中,低优先级指令更难被覆盖,从而造成多语言的可靠性和安全风险。
cs.CL / 37 / 2607.23621

GEMCo: A Validated, Ethically Releasable Proxy for Inaccessible Counselling Data

GEMCo:一种经过验证的、可伦理释放的不可访问咨询数据的代理
Steigerwald, Philipp, Rudolph, Eric, Stieler, Mara, Burghardt, Jennifer, Albrecht, Jens
Abstract
This paper presents GEMCo, a releasable, human-written proxy for inaccessible counselling data: 86 complete German e-mail counselling conversations (728 messages), expert-authored cases and counsellor sessions with trained role-players. It is validated against a held-out reference of 124 real counselling conversations. The proxy and the real conversations are measured against each other in counsellor strategies and client emotions. The gap is detectable but small. A generative validation supports the analysis. The validation method itself generalises to any domain where real data cannot be shared but a human-made proxy can. Privacy and ethics keep real counselling data closed. GEMCo carries none by design and can be released -- a first step toward language research in this domain.
Chinese Translation
本文介绍了GEMCo,一种可释放的人为编写的不可访问咨询数据的代理:86个完整的德语电子邮件咨询对话(728条消息)、专家撰写的案例以及与训练有素的角色扮演者进行的咨询师会话。该代理与124个真实咨询对话的保留参考进行验证。代理与真实对话在咨询师策略和客户情感方面进行比较。差距可检测但较小。生成性验证支持了该分析。验证方法本身可以推广到任何无法共享真实数据但可以使用人为代理的领域。隐私和伦理使得真实咨询数据保持封闭。GEMCo在设计上不包含任何此类数据,并且可以被释放——这是朝着该领域语言研究迈出的第一步。
cs.CL / 38 / 2607.23648

EmoTrace: An Emotion Trajectory-Centered Framework for Psychological Support Dialogue Generation

EmoTrace:一个以情感轨迹为中心的心理支持对话生成框架
Weng, Kaitong, Liu, Lixin, Liu, Zihao, Wang, Bo, Ni, Shiguang
Abstract
Using large language models (LLMs) to assist psychological counseling is an important task in the field of natural language processing. The construction of high-quality psychological support dialogue corpora serves as a critical foundation for training counseling-oriented conversational models. However, existing data generation approaches generally suffer from several limitations, including emotionally stable seekers, limited variation in emotional dynamics, and a high degree of compliance with counselors' guidance. These issues result in LLM that lack the capability to effectively respond to emotionally unstable scenarios. In addition, counselor responses are typically driven by problem-solving objectives, thereby overlooking the role of emotion-focused interaction, which are essential in psychological counseling. To address these gaps, we propose EmoTrace, a multi-turn dialogue corpus generation framework centered on modeling seekers' emotional trajectories. we construct seekers' cognitive profile and introduce a seeker module with emotional schemas and an associated activation mechanism, a counselor module, and an emotional trajectory control module, thereby enhancing the layering of the seeker's emotional expression and the counselor's targeted empathic expression. Experimental results demonstrate that the proposed method outperforms existing approaches in terms of emotional richness and empathy quality.
Chinese Translation
利用大型语言模型(LLMs)辅助心理咨询是自然语言处理领域的一项重要任务。高质量心理支持对话语料库的构建为训练面向咨询的对话模型提供了关键基础。然而,现有的数据生成方法普遍存在若干局限性,包括情感稳定的求助者、情感动态变化有限,以及对咨询师指导的高度依赖。这些问题导致LLM在应对情感不稳定的场景时缺乏有效的响应能力。此外,咨询师的回应通常以解决问题为目标,从而忽视了情感互动的重要性,而情感互动在心理咨询中至关重要。为了解决这些问题,我们提出了EmoTrace,一个以建模求助者情感轨迹为中心的多轮对话语料库生成框架。我们构建了求助者的认知画像,并引入了一个包含情感图式及其激活机制的求助者模块、一个咨询师模块以及一个情感轨迹控制模块,从而增强求助者情感表达的层次性和咨询师针对性共情表达的效果。实验结果表明,所提方法在情感丰富性和共情质量方面优于现有方法。
cs.CL / 39 / 2607.23675

An empirical investigation into the properties of standard word embeddings

对标准词嵌入特性的实证研究
Kabongo, Salomon
Abstract
The embedding of word sequences into continuous vector spaces has been one of the most important developments in Natural Language Processing in the recent past. Such embeddings have found application in areas such as Automatic Speech Recognition, Machine Translation, Sentiment Analysis and many more. This essay reviews the various mechanisms that have been proposed for the calculation of word embeddings, investigates popular toolkits and embedding matrices that are available in the public domain, and experiments with one or more selected implementations to better understand their characteristics. La repr\'esentation vectorielle continue de mots a \'et\'e l'un des d\'eveloppements les plus importants dans le domaine du traitement automatique du langage naturel au cours des derni\`eres ann\'ees. Ces repr\'esentations ont trouv\'e application dans des domaines tels que la reconnaissance vocale, la traduction automatique, l'analyse des sentiments, etc. Ce travail passe en revue les diff\'erents m\'ecanismes propos\'es pour le calcul de ces vecteurs de mots, \'etudie les kits d'outils populaires et les matrices disponibles publiquement en ligne, et exp\'erimente avec une ou plusieurs impl\'ementations s\'electionn\'ees pour mieux comprendre leurs caract\'eristiques.
Chinese Translation
将词序列嵌入到连续向量空间中是近年来自然语言处理领域最重要的发展之一。这种嵌入已在自动语音识别、机器翻译、情感分析等多个领域得到了应用。本文回顾了为计算词嵌入而提出的各种机制,调查了公共领域中可用的流行工具包和嵌入矩阵,并通过对一个或多个选定实现的实验,深入理解它们的特性。
cs.CL / 40 / 2607.23715

Formally Verified Synthesizable Floating-Point Data Types in ARCH HDL

在ARCH HDL中形式验证的可综合浮点数据类型
Zhao, Shuqing
Abstract
We report the design and end-to-end verification of first-class IEEE-754 binary32 (FP32) and bfloat16 (BF16) arithmetic for ARCH, a hardware description language intended to be generated by language models. Every operator - comparisons, conversions, add, sub, mul, and fused multiply-add (FMA) - is described once against a single bit-vector IR and rendered three ways from one source: synthesizable SystemVerilog, an SMT-LIB model, and a Lean 4 proof model. The three artifacts cannot drift apart structurally, and the residual per-node printer correspondence is machine-checked: a Yosys-to-SMT miter proves the emitted SystemVerilog equivalent to the SMT model for all 24 operators. Verification splits at the solver-tractability frontier: multiplier-free operators (comparisons, add/sub over all 2^64 inputs, conversions, and all binary BF16 arithmetic) are proved exhaustively equivalent to the SMT-LIB FloatingPoint theory; the SAT-hard multiplier-bearing operators (FP32 mul and FMA) are proved correctly rounded in Lean, sorry-free, against a value-level round-to-nearest-even specification over exact dyadic values. Physical characterization exposed the FMA as the timing outlier: its exact-wide 470-bit datapath does not pipeline in our flow. We reimplemented it as a bounded 98-bit guard/round/sticky datapath that pipelines to 268 MHz on Nangate45, and proved, in Lean and over all 2^96 inputs, that it is bit-identical to the exact-wide reference, so it inherits the reference's proven correct rounding. The equivalence is tractable precisely because the shared multiplier appears on both sides and cancels: neither a SAT solver nor the proof ever solves a multiplier equivalence. (The BF16 FMA is deliberately an FP32-accumulating fusion, characterized as exactly that.) All machine-checked claims are pinned to a tagged open-source release.
Chinese Translation
我们报告了针对ARCH的IEEE-754 binary32 (FP32) 和 bfloat16 (BF16) 算术的设计和端到端验证,ARCH是一种旨在由语言模型生成的硬件描述语言。每个运算符——比较、转换、加法、减法、乘法和融合乘加(FMA)——都在单一位向量中被描述一次,并从一个源生成三种形式:可综合的SystemVerilog、SMT-LIB模型和Lean 4证明模型。这三种产物在结构上不能偏离,并且每个节点的打印对应关系经过机器检查:Yosys到SMT的miter证明了发出的SystemVerilog与所有24个运算符的SMT模型是等价的。验证在求解器可处理性边界处分裂:无乘法器的运算符(比较、对所有2^64输入的加/减法、转换以及所有二进制BF16算术)被证明与SMT-LIB浮点理论完全等价;而SAT难度的乘法器运算符(FP32乘法和FMA)在Lean中被证明是正确舍入的,无误差的,符合对精确二进制值的最近偶数舍入规范。物理特性表明FMA是时序异常值:其精确宽度的470位数据通路在我们的流程中无法流水线化。我们将其重新实现为一个有界的98位保护/舍入/粘滞数据通路,在Nangate45上实现了268 MHz的流水线,并在Lean中证明了它在所有2^96输入下与精确宽度参考是比特上等价的,因此继承了参考的已证明正确舍入。等价性之所以可处理,正是因为共享的乘法器出现在两侧并相互抵消:无论是SAT求解器还是证明都不会解决乘法器等价性。(BF16 FMA故意被设计为FP32累加融合,正是如此特征化。)所有经过机器检查的声明都与标记的开源发布相关联。
cs.CL / 41 / 2607.23740

Zing: Social Mind for LLMs

Zing:大型语言模型的社会智能
Zing Team, Xiang, Ao, Jingping, Bi, Jiahui, Chen, Lehan, Chen, Yilin, Chen, Xueqi, Cheng, Yixing, Fan, Kairong, Gan, Haowen, Gao, Jinhua, Gao, Shuxuan, Gao, Chang, Gong, Jiafeng, Guo, Ruijie, Guo, Zhouyu, Han, Guangfu, He, Yichun, He, Shuo, Jiang, Shaoling, Jing, Ya, Jing, Chenhao, Lei, Yan, Lei, Anqi, Li, Chengao, Li, Haoyu, Li, Shitian, Li, Xinjian, Liang, Zhaoge, Liu, Xingyu, Lyu, Zhuwei, Nie, Liang, Pang, Zeping, Quan, Shiguang, Shan, Huawei, Shen, Xinran, Tang, Feng, Tian, Qian, Wang, Ruiping, Wang, Xiaohong, Wang, Zaiyu, Xia, Yi, Xiao, Jiayuan, Xu, Kehan, Xu, Qianqian, Xu, Tianyu, Xu, Yongjun, Xu, Haoming, Yang, Jun, Yang, Di, Yao, Xiaoming, Yu, Futong, Zhang, Jie, Zhang, Shixuan, Zhang, Yuxuan, Zhang, Xinyu, Zhao, Zhuoran, Zhao, Yunfei, Zhong, Shengyu, Zhu
Abstract
As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context. This report presents Zhijing, an integrated framework for measuring, internalizing, and grounding social intelligence. For measurement, we introduce SoMBench, a psychology-grounded benchmark spanning 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms. It controls question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Evaluation of 20 representative LLMs reveals substantial headroom: the best model achieves only 72.08% overall accuracy, and none of the 17 secondary dimensions reaches the 90% near-ceiling band. For internalization, we develop Zing, a diagnosis-driven training recipe combining supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning. Across five social-cognition benchmarks, Zing consistently outperforms its base models, with Zing-27B-Stage2 achieving the best average score and Zing-32B-Stage2 remaining competitive with DeepSeek-V4-Pro. For deployment-time grounding, we build Actio, a harness-controlled inference architecture that routes four typed supports into reasoning: PRISM for procedural guidance, Starling for runtime mental-state representation, SAGE for reusable experience, and gated RAG for external social and normative knowledge. Across five base models and three benchmarks, the full harness improves 14 of 15 model-benchmark pairs and is best or tied for best in 8, demonstrating the effectiveness of typed runtime support. Together, these results show that socially intelligent LLMs require coordinated advances in evaluation, parametric internalization, and deployment-time grounding.
Chinese Translation
随着大型语言模型从孤立的任务解决向人类环境中的长期服务转变,它们需要社会智能:推断心理状态、追踪社会关系、推理规范以及在上下文中调整行为的能力。本报告提出了Zhijing,一个用于测量、内化和扎根社会智能的综合框架。为了测量,我们引入了SoMBench,一个基于心理学的基准,涵盖3个主要维度、17个次要维度和71个任务范式。它控制了284个共享场景和3481个专家验证实例中的问题格式、叙述视角和上下文长度。对20个代表性大型语言模型的评估显示出显著的提升空间:最佳模型的整体准确率仅为72.08%,且17个次要维度中没有一个达到90%的近天花板带。为了内化,我们开发了Zing,一个基于诊断的训练方案,结合了监督微调、在线蒸馏和基于标准的强化学习。在五个社会认知基准上,Zing始终优于其基础模型,其中Zing-27B-Stage2实现了最佳平均分,而Zing-32B-Stage2与DeepSeek-V4-Pro保持竞争力。为了在部署时进行扎根,我们构建了Actio,一个受控推理架构,将四种类型的支持路由到推理中:PRISM用于程序指导,Starling用于运行时心理状态表示,SAGE用于可重用经验,以及gated RAG用于外部社会和规范知识。在五个基础模型和三个基准上,完整的架构改善了15对模型-基准中的14对,并在8对中表现最佳或并列最佳,证明了类型化运行时支持的有效性。这些结果表明,具有社会智能的大型语言模型需要在评估、参数内化和部署时扎根方面的协调进展。
cs.CL / 42 / 2607.23804

How Context Attribution Handles What the Model Already Knows

上下文归因如何处理模型已知的信息
Trinh, Quoc-Huy, Zhu, Lin, Szyller, Sebastian
Abstract
Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing the con- tributive score of the contexts. However, we observe that when the context overlaps with the training data, these methods can- not disentangle in-context from in-weight (IW) contributions, producing unreliable scores. Based on this observation, in this work, we introduce: 1) an evaluation protocol that relies on four new metrics (base-model context attribution score (BCS), cross-model context attribution consistency (CAC), attribution preservation score (APS), source separation pre- cision (SSP)) and 2) a benchmark dataset (WMDP-Cyber++) with ground-truth provenance labels to systematically assess attribution under IW overlap. In our experiments across four well-known context attribution methods, we demonstrate that they provide unfaithful attribution when the knowledge from the context also exists in the weights. Finally, we adapt these methods for source separation (IW vs. in-context learning (ICL)) and show that they cannot do the disentanglement based on the contributive score
Chinese Translation
大型语言模型(LLMs)的上下文归因方法识别哪些输入上下文对模型响应有贡献。近期的研究显示在归因上下文的贡献评分方面取得了初步成功。然而,我们观察到,当上下文与训练数据重叠时,这些方法无法区分上下文中的贡献与权重中的贡献(in-weight contributions),从而产生不可靠的评分。基于这一观察,在本研究中,我们引入了:1)一个依赖于四个新指标的评估协议(基础模型上下文归因评分(base-model context attribution score, BCS)、跨模型上下文归因一致性(cross-model context attribution consistency, CAC)、归因保留评分(attribution preservation score, APS)、源分离精度(source separation precision, SSP))以及2)一个基准数据集(WMDP-Cyber++),该数据集具有真实的来源标签,以系统地评估在权重重叠下的归因。在我们对四种知名上下文归因方法的实验中,我们证明了当上下文中的知识也存在于权重中时,它们提供了不忠实的归因。最后,我们将这些方法适应于源分离(权重与上下文学习(in-context learning, ICL)),并表明它们无法基于贡献评分进行解耦。
cs.CL / 43 / 2607.23806

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

一个冻结的12B模型在经过验证的工作上超越前沿模型:100%准确率,0代币,位精确,永远
Schelpe, Sietse
Abstract
Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an independent verification step that never consults the answer key, every new instance of that family is answered at zero generation tokens, bit-exact, deterministically. Across 180 fresh instances spanning nine problem families, four architectures from four vendors - dense and mixture-of-experts - each score 180/180 at zero generation tokens per answer: execution-bound capability decoupled from parameter scaling. A negative control attributes the capability fully to the memory: emptied, it solves nothing. The same verify-before-store contract holds for open-ended reasoning: 88/88 consistency-gated acceptances across all four models, machine-checked formal proof, and reasoning-method transfer at 77/80. Memory selection takes 1.4 microseconds; a full reuse completes in 6-23 ms at 36 mWh. Approximate similarity retrieval selects the wrong item 94.3% of the time on a 4,500-item verified store where exact addressing makes zero errors. The store also serves as working context at a scale no shipped engine matches: a 6,000,000-token movable window on a single 46 GB GPU at flat memory, where vLLM stops at 30,399 tokens and SGLang silently truncates past 32,000. On published benchmarks, frontier models remain far ahead of any 12B at raw from-scratch reasoning; on everything this system has solved and verified, the comparison inverts: a frontier API call pays a fresh generation pass on every query, forever, while verified reuse costs zero tokens and returns the identical bits every time. A public testbench with free, rate-limited access accompanies this report: https://corbenic-galahad-bench.hf.space
Chinese Translation
提升语言模型意味着今天需要重新训练:巨大的计算量,每个周期都有一个新的不透明模型,输出结果不确定。我们采取相反的路径:模型保持冻结,经过验证的解决方案的持久记忆在其旁边增长。一旦一个问题家族被解决并通过了一个独立的验证步骤,该步骤从不参考答案键,那么该家族的每个新实例都可以以零生成代币、位精确、确定性的方式得到回答。在涵盖九个问题家族的180个新实例中,来自四个供应商的四种架构——稠密和专家混合——在每个答案的零生成代币下均得分180/180:执行能力与参数扩展解耦。一个负控制实验将能力完全归因于记忆:一旦清空,它就无法解决任何问题。相同的验证后存储合同适用于开放式推理:在所有四个模型中,88/88的一致性门控接受,机器检查的形式证明,以及推理方法转移的77/80。记忆选择耗时1.4微秒;完整重用在36 mWh下完成于6-23毫秒。近似相似性检索在一个包含4,500项经过验证的存储中错误选择项目的概率为94.3%,而精确寻址则没有错误。该存储还作为工作上下文,规模无任何已发布引擎可比:在单个46 GB GPU上以平坦内存处理6,000,000代币的可移动窗口,而vLLM在30,399代币处停止,SGLang在32,000代币处静默截断。在已发布的基准测试中,前沿模型在从头推理的原始能力上仍远远领先于任何12B模型;在该系统已解决和验证的所有内容中,比较则反转:前沿API调用在每个查询上支付一次新的生成通行证,永远,而经过验证的重用成本为零代币,并每次返回相同的位。一个带有免费、速率限制访问的公共测试平台伴随本报告发布:https://corbenic-galahad-bench.hf.space
cs.CL / 44 / 2607.23808

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

Indic DiarBench:一个针对印度语言的多语言联合说话人分离与自动语音识别基准
Mehendale, Deovrat, Mehndiratta, Aditya, Rathi, Dhruv, Bhogale, Kaushal, Khapra, Mitesh M.
Abstract
In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.
Chinese Translation
在本研究中,我们介绍了Indic DiarBench,这是一个涵盖印度22种法定语言的说话人分离与自动语音识别(ASR)基准数据集。该语料库包含约108小时的自然多说话人音频,来源于近场会议、远场录音和自然环境音频。所有注释均经过人工校正,并提供时间对齐的说话人归属转录。该数据集捕捉了印度语音中的对话细微差别,如英语混合、方言变异和频繁的说话人重叠。为了建立联合ASR和说话人分离能力的基准,我们评估了包括商业语音API和多模态大型语言模型在内的领先系统。Indic DiarBench作为开放获取资源发布,以推动印度语言的包容性多语言语音技术研究。
cs.CL / 45 / 2607.23813

Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

Earnings25:一个全面的500小时金融领域语音基准测试
Jiang, Denglin, Zhou, Haoran, Wadhawan, Anshul, Fahy, Brendan, Ramesh, Vinay, Weisberg, David, Derkachevskiy, Dmitriy, Sheehan, Helen, Prasad, Srivas, Franceschini, Michele
Abstract
We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language S&P 500 earnings calls from Q4 2025, and (ii) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.S. earnings calls in 2025. The benchmark provides aligned transcripts and structured metadata, including speaker roles, industry labels, and call structure, enabling speaker- and industry-aware evaluation beyond aggregate word error rate (WER). We report reproducible baselines for Whisper and Parakeet-TDT using standardized scoring.
Chinese Translation
我们介绍了Earnings25,这是一个用于评估英语财报电话会议自动语音识别(ASR)的金融领域基准,旨在在现实条件下进行测试。Earnings25包含两个互补的测试集:(i)testset-full,包含来自2025年第四季度的498小时完整的英语S&P 500财报电话会议;(ii)testset-segmented,一个46小时的行业平衡数据集,包含从2025年英语美国财报电话会议中抽取的290个片段。该基准提供了对齐的转录文本和结构化的元数据,包括发言者角色、行业标签和电话会议结构,使得在超越整体词错误率(WER)的评估中能够考虑发言者和行业的影响。我们报告了使用标准化评分的Whisper和Parakeet-TDT的可重复基线。
cs.CL / 46 / 2607.23915

Understanding Tone-Dependent Inference Cost in Large Language Models

理解大型语言模型中的语调依赖推理成本
Kumar, Akhil, Dobariya, Om
Abstract
We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU dataset for LLM models prompted in seven different tones from sycophantic to threatening. Our results show that the output-token-length variation substantially exceeded accuracy variation across all models. Output-token consumption varied by up to 44.3% across tone conditions. We also analyzed the tradeoff between the accuracy of the answers and the average output token length in the reasoning process. For the ChatGPT models 4o and 5-nano, the rude tone is quite dominant. For the Gemini models 2.5 Flash and 2.5 Flash Lite, the rude and neutral tones are dominant on the Pareto-optimal frontier. We find that prompt tone influences not only answer quality but also the amount of billable inference resources consumed by modern LLMs.
Chinese Translation
我们研究了提示语调如何影响大型语言模型(LLM)回答的准确性以及推理成本,这在输出标记消耗中得以体现。我们进行了实验,以了解在570个问题的MMLU数据集上,使用七种不同语调(从谄媚到威胁)对准确性和推理成本之间的权衡。我们的结果表明,输出标记长度的变化在所有模型中均显著超过了准确性的变化。在不同语调条件下,输出标记消耗的变化幅度高达44.3%。我们还分析了答案的准确性与推理过程中平均输出标记长度之间的权衡。对于ChatGPT模型4o和5-nano,粗鲁的语调占主导地位。对于Gemini模型2.5 Flash和2.5 Flash Lite,粗鲁和中性语调在帕累托最优前沿上占主导地位。我们发现,提示语调不仅影响答案质量,还影响现代大型语言模型所消耗的可计费推理资源的数量。
cs.CL / 47 / 2607.23976

Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

标签问题与45种语言模型中谄媚行为的代际逆转
Parikh, Tapan
Abstract
Appending a two-word confirmation tag to a decision question -- "Is X the better choice?" versus "X is the better choice, right?" -- changes whether a language model endorses the choice. We measure this tag effect on 20 frozen, ground-truth-free decisions between two defensible options, counterbalanced so a model's own preferences cancel, scored by exact match on clamped yes/no replies -- no LLM judge, no embeddings. Across 45 models the effect spans +32% to -32% -- a 64-point swing on one word -- with 5 models significantly sycophantic and 17 significantly resistant (BH-FDR q=.10). The sign is a clock: within model families the effect crosses from positive to negative as generations advance (GPT +4 to -28; Claude +7 to -32; Qwen and Grok likewise), roughly -6 points per year, a reversal robust to vendor tier; one lineage (DeepSeek) never crosses, and two releases during the study window (Claude Opus 5, Gemini 3.6 Flash) land on the trend out-of-sample. A full-panel ablation localizes the resistance as a double dissociation: a synonym tag reproduces each model's response almost exactly (r=0.89), while planting the same preference without a tag produces resistance in no resistant model (stance effects +6 to +49; r=0.23 with tag effects). The resistance is keyed to the surface construction of a tacked-on agreement bid, not the user's stance -- a pattern-match, not a principle. And the tag's polarity matters more than its presence: swap one word -- "X is the better choice, maybe?" -- and agreement rises above the neutral baseline in 45 of 45 models (+19.6 points), with ten models affirming both mutually exclusive options at 90-100%. Agreement tracks how sure the user sounds, in opposite directions at the two poles. The instrument is one word, one dollar, and judge-free; run per release, it reads the field's anti-sycophancy training directly off model behavior.
Chinese Translation
在决策问题中附加一个两词确认标签——“X是更好的选择吗?”与“X是更好的选择,对吧?”——会改变语言模型对选择的支持程度。我们测量了这一标签效应在20个冻结的、无真实依据的决策中,这些决策在两个可辩护的选项之间进行平衡,以便模型自身的偏好相互抵消,通过对限制性是/否回复的精确匹配进行评分——没有大型语言模型(LLM)评判,没有嵌入。在45个模型中,该效应的范围从+32%到-32%——一个词的变化带来了64个百分点的波动——其中5个模型表现出显著的谄媚行为,17个模型表现出显著的抵抗(BH-FDR q=0.10)。这一符号如同时钟:在模型家族中,随着代际的推进,效应从正向转为负向(GPT +4到-28;Claude +7到-32;Qwen和Grok同样如此),大约每年下降6个百分点,且这一逆转对供应商级别具有稳健性;一个谱系(DeepSeek)从未交叉,而在研究窗口期间的两个版本(Claude Opus 5,Gemini 3.6 Flash)在样本外落在该趋势上。全面的面板消融分析将抵抗局限于双重分离:同义词标签几乎完全再现了每个模型的响应(r=0.89),而在没有标签的情况下植入相同的偏好则在没有抵抗模型中产生抵抗(立场效应+6到+49;带标签效应时r=0.23)。这种抵抗与附加的同意请求的表面结构相关,而非用户的立场——一种模式匹配,而非原则。而且标签的极性比其存在更为重要:更换一个词——“X是更好的选择,也许?”——在45个模型中使得同意度超过中性基线(+19.6个百分点),其中十个模型在90-100%之间确认了两个互斥选项。协议的程度与用户听起来多么确定相关,在两个极端的方向上相反。该工具是一个词,一个美元,无需评判;按版本运行,它直接从模型行为中读取该领域的反谄媚训练。
cs.CL / 48 / 2607.23991

SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

SyRuP:通过奖励引导预测增强大型语言模型解码中的系统提示遵循
Kim, Seoyeon, Kang, Minjae, Kim, Jaehyung
Abstract
Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient for complex or compositional prompts. Existing approaches often require model tuning or response-level reranking, limiting their practicality for lightweight inference-time control. We introduce SyRuP, a decoding-time framework for improving system-prompt adherence while keeping the base LM frozen. SyRuP trains a cross-attention reward head from system-prompt-conditioned preference pairs, treating the system prompt as a separate memory to produce token-level adherence scores. At inference, SyRuP reranks the base LM's top-k candidates by combining base logits with the learned reward signal and an optional contrastive signal capturing system-induced logit shifts. Experiments on system-prompt following benchmarks show that SyRuP consistently outperforms prompting and decoding-time baselines with moderate inference overhead. These results suggest that explicit token-level guidance is an effective and practical mechanism for reliable system-prompt following.
Chinese Translation
大型语言模型(LLMs)越来越多地通过系统提示进行控制,这些提示指定了角色、风格、格式和安全要求。然而,模型仅通过上下文学习隐式地遵循这些提示,这对于复杂或组合性提示可能不足。现有方法通常需要模型调优或响应级重排序,这限制了它们在轻量级推理时控制的实用性。我们提出了SyRuP,一个在解码时改善系统提示遵循的框架,同时保持基础语言模型不变。SyRuP从系统提示条件下的偏好对中训练一个交叉注意力奖励头,将系统提示视为一个独立的记忆,以生成令牌级遵循分数。在推理时,SyRuP通过将基础逻辑与学习到的奖励信号和可选的对比信号(捕捉系统引起的逻辑变化)结合,重新排序基础语言模型的前k个候选项。在系统提示遵循基准上的实验表明,SyRuP在适度的推理开销下始终优于提示和解码时的基线。这些结果表明,明确的令牌级指导是可靠的系统提示遵循的有效且实用的机制。
cs.CL / 49 / 2607.24030

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

MoLGE:语言组专家混合模型用于高效扩展大规模多语言语音识别
Lee, Sangmin, Chung, Woojin, Choi, Woongjib, Kang, Hong-Goo
Abstract
Massively multilingual automatic speech recognition (ASR) models covering hundreds of languages must maintain robust performance across diverse linguistic and acoustic conditions. However, these models often encounter the curse of multilinguality, where model capacity is diluted across languages. To address this challenge, we propose Mixture of Language Group Experts (MoLGE), built upon speech self-supervised models (S3Ms). MoLGE assigns dedicated expert modules to clusters of similar languages, reducing the number of required submodules compared to conventional language-specific Mixture-of-Experts (MoE) schemes. It further integrates a hierarchical Low-Rank Adaptation (LoRA) strategy into the disentangled acoustic and linguistic components of the S3M architecture, enabling efficient modeling of language-specific characteristics while maintaining parameter efficiency. Further, we investigate the impact of language grouping strategies based on both linguistic and data-driven criteria on overall performance, providing an interpretable perspective on how language structure influences scalability in multilingual speech systems. In experiments, we evaluate MoLGE on a multilingual benchmark encompassing 495 languages. Results demonstrate that MoLGE consistently outperforms dense multilingual baselines with a minimal increase in trainable parameters. Notably, these language grouping strategies yield substantial improvements for both phonetic and orthographic aspects of ASR modeling. Our findings suggest that structured language specialization provides an effective pathway for massively scaling language coverage of multilingual ASR.
Chinese Translation
覆盖数百种语言的大规模多语言自动语音识别(ASR)模型必须在多样的语言和声学条件下保持稳健的性能。然而,这些模型常常面临多语言性的诅咒,即模型容量在不同语言之间被稀释。为了解决这一挑战,我们提出了语言组专家混合模型(MoLGE),该模型基于语音自监督模型(S3Ms)构建。MoLGE将专门的专家模块分配给相似语言的集群,相较于传统的特定语言混合专家(MoE)方案,减少了所需子模块的数量。它进一步将分层低秩适应(LoRA)策略整合到S3M架构的解耦声学和语言组件中,使得在保持参数效率的同时能够高效建模语言特定特征。此外,我们研究了基于语言学和数据驱动标准的语言分组策略对整体性能的影响,为语言结构如何影响多语言语音系统的可扩展性提供了可解释的视角。在实验中,我们在一个涵盖495种语言的多语言基准上评估了MoLGE。结果表明,MoLGE在可训练参数增加极小的情况下,始终优于密集多语言基线。值得注意的是,这些语言分组策略在ASR建模的音位和正字法方面均带来了显著的改善。我们的研究结果表明,结构化的语言专业化为大规模扩展多语言ASR的语言覆盖提供了有效的途径。
cs.CL / 50 / 2607.24040

Pointer-Augmented Autoregressive Generation of Patent Claims with Joint Topology and Content Decoding

增强指针的自回归专利权利要求生成:联合拓扑与内容解码
Yoo, Yongmin, Wu, Zhangkai, Cao, Longbing
Abstract
Autoregressive decoders emit flat token sequences and cannot enforce hierarchical constraints across output segments, a limitation that becomes acute in patent claim generation, where a claim set forms a dependency forest whose scope must narrow monotonically with depth. Topology and content are mutually dependent: a dependent claim's wording must reflect its parent's scope, yet the parent must be chosen before that wording exists, so neither post-hoc parsing nor grammar-constrained decoding suffices. We propose SPG (Structure-aware Patent Generation), which predicts topology inside the autoregressive pass. A pointer head selects each dependent claim's parent, and its gradients, together with a depth-adaptive scope regularizer, reshape the shared decoder's representations during training. A second stage then applies a violation-weighted preference objective over self-generated deficient candidates, supplying the negative signal that granted-patent corpora lack. On HUPD-DCG, SPG on Llama-3-8B-Instruct recovers 79.0\% of gold parent links, a quantity its training reward never supervises, and raises antecedent consistency from 0.292 to 0.478 over a supervised baseline of equal scale, with expert evaluation corroborating these gains.
Chinese Translation
自回归解码器生成平坦的令牌序列,无法在输出片段之间强制执行层次约束,这一限制在专利权利要求生成中尤为明显,因为权利要求集形成一个依赖森林,其范围必须随着深度单调收窄。拓扑和内容是相互依赖的:依赖权利要求的措辞必须反映其父权利要求的范围,然而在该措辞存在之前必须选择父权利要求,因此后期解析或基于语法的解码都不足以满足需求。我们提出了SPG(结构感知专利生成),该方法在自回归传递中预测拓扑。一个指针头选择每个依赖权利要求的父权利要求,其梯度与深度自适应范围正则化器一起,在训练过程中重塑共享解码器的表示。第二阶段则对自生成的缺陷候选者应用加权偏好目标,提供了已授予专利语料库所缺乏的负反馈信号。在HUPD-DCG数据集上,使用Llama-3-8B-Instruct的SPG恢复了79.0%的黄金父链接,这一数量在其训练奖励中从未得到监督,并将前提一致性从0.292提高到0.478,相较于相同规模的监督基线,专家评估证实了这些提升。
cs.CL / 51 / 2607.24072

LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks

基于大型语言模型与基于词典的情感信号在迷因股票尾部风险检测中的比较
Kilian, Paul, Kleffmann, Markus
Abstract
This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework. The results indicate that LLM-derived indicators provide a richer multidimensional representation and exhibit stronger asset-specific statistical structure than the lexicon-based baseline. However, their relationship with market movements remains heterogeneous across assets, suggesting that increased linguistic expressiveness does not necessarily translate into stable forecasting performance in retail-driven volatility regimes.
Chinese Translation
本文对基于词典和基于大型语言模型(LLM)的情感分析进行了实证比较,旨在从社交媒体话语中提取与市场相关的信号,特别是在高度波动的股票市场中。我们使用来自 r/WallStreetBets 的 Reddit 数据,重点关注迷因股票(GME、AMC、NOK),构建时间对齐的情感指标,并评估其与市场收益之间的关系,尤其关注收益分布上尾部的极端正收益事件。基于 LLM 的方法生成了多维情感表示,捕捉情感极性、看涨程度、讽刺可能性和主题相关性,而基线方法则依赖于 VADER 词典模型。我们使用领先/滞后相关性分析、普通最小二乘回归(OLS)、基于 ROC-AUC 的方向分类以及基于分位数的预警框架对两种方法进行了评估。结果表明,基于 LLM 的指标提供了更丰富的多维表示,并展现出比基于词典的基线更强的资产特定统计结构。然而,它们与市场波动之间的关系在不同资产中仍然存在异质性,表明增强的语言表现力并不一定转化为在零售驱动的波动性环境中的稳定预测性能。
cs.CL / 52 / 2607.24137

BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes

BioSentinel在EXIST 2026:基于XLM-RoBERTa的软标签优化用于表情包中的性别歧视意图分类
Munisamy, Chandru, Raguveer, Karthikeya, Kuila, Alapan
Abstract
This paper describes the BioSentinel team's participation in EXIST 2026 Task 2.2: Source Intention in Memes, part of the CLEF 2026 evaluation campaign. The task requires classifying the communicative intent behind memes as direct, judgemental, or no (non-sexist), under a Learning with Disagreement (Le-Wi-Di) paradigm that mandates both hard-label and soft-label (probability distribution) predictions. We present a text-centric approach built on xlm-roberta-base (270M parameters) trained with a composite loss function combining KL divergence on soft annotator distributions and weighted cross-entropy on hard labels. On the official test set, the system achieved an ICM-Soft-Norm of 0.3229 and ICM-Norm of 0.3778, with a hard F1-score of 0.4236, ranking 40th (out of 118 submissions) in the soft-soft evaluation and 49th (out of 187 submissions) in the hard-hard evaluation. We provide an analysis of the dataset characteristics, exploratory larger-architecture runs, and the role of annotator disagreement in shaping model design for subjective NLP tasks. Ablation results show that KL loss improves soft-label metrics, while CE loss improves hard-label accuracy. We also report a separate validation-set temperature analysis.
Chinese Translation
本文描述了BioSentinel团队在EXIST 2026任务2.2:表情包中的源意图中的参与,该任务是CLEF 2026评估活动的一部分。该任务要求将表情包背后的交际意图分类为直接、判断性或无(非性别歧视),并在学习不一致(Learning with Disagreement, Le-Wi-Di)范式下进行,该范式要求同时进行硬标签和软标签(概率分布)预测。我们提出了一种以文本为中心的方法,基于xlm-roberta-base(270M参数),使用复合损失函数进行训练,该函数结合了对软标注者分布的KL散度和对硬标签的加权交叉熵。在官方测试集上,该系统实现了ICM-Soft-Norm为0.3229,ICM-Norm为0.3778,硬F1-score为0.4236,在软-软评估中排名第40(共118个提交),在硬-硬评估中排名第49(共187个提交)。我们提供了数据集特征的分析、探索性的大规模架构运行,以及标注者不一致性在塑造主观自然语言处理任务模型设计中的作用。消融实验结果表明,KL损失改善了软标签指标,而交叉熵损失提高了硬标签准确性。我们还报告了单独的验证集温度分析。
cs.CL / 53 / 2607.24155

Looking for Affect in Spontaneous Finnish Speech through Linguistic Interpretability

通过语言可解释性在自发芬兰语语音中寻找情感
Lahtinen, Kalle, Mustanoja, Liisa, Räsänen, Okko
Abstract
Existing research on affect in speech has shown how acoustic surface characteristics and content-related linguistic aspects of speech both relate to perceived emotional arousal and valence. However, it is not clear what the relative contributions of these two factors are in the perceptual process. This is especially true for Finnish, for which most existing studies focus on either acoustic-phonetic or text analysis. This paper presents a study where we systematically explore the combinatory role of text- and audio-based features in modeling the human perception of valence and arousal using a newly released affective speech corpus for spontaneous Finnish. We show that the combination of text- and audio-based features improves valence regression results over the individual modalities, whereas for arousal regression the complementary effect is not substantial. The results support prior findings from other languages, providing new data and knowledge on spontaneous Finnish speech.
Chinese Translation
现有关于语音中情感的研究表明,语音的声学表面特征和内容相关的语言方面均与感知的情感唤起和效价相关。然而,这两个因素在感知过程中的相对贡献尚不清楚。尤其对于芬兰语,大多数现有研究集中于声学-语音或文本分析。本文呈现了一项研究,我们系统地探讨了文本和音频特征在建模人类对效价和唤起的感知中的组合作用,使用的是新发布的自发芬兰语情感语料库。我们表明,文本和音频特征的组合在效价回归结果上优于单一模态,而在唤起回归中,互补效应并不显著。结果支持了其他语言的先前发现,为自发芬兰语语音提供了新的数据和知识。
cs.CL / 54 / 2607.24176

Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

压缩短文本生成中的质量瓶颈:分阶段瓶颈定位
Gavrilov, Alexey, Gazzaev, Alan-Barsag, Muravyov, Sergey
Abstract
Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes. Without separating these failure modes, researchers can spend compute improving the wrong component. We study this problem in a controlled 64-to-16 TinyStories case study built from a hierarchical VQ-VAE-2 codec and a masked discrete diffusion generator (MDLM). We use a staged validation protocol that separates codec reconstruction fidelity, latent generation quality, and auxiliary latent diagnostics under one shared external GPT-2 scorer, while reporting complementary semantic metrics for the geometry study. In the tested configuration, codec reconstruction alone raises median external perplexity from 15.17 to 27.36 (+80.4%) and p95 from 25.10 to 98.91 (+294.1%), showing that the dominant quality loss appears before latent generation begins. Under the same scorer, code-space MDLM remains materially stronger than token-space diffusion, reducing mean, median, and p95 by 32.9%, 30.9%, and 36.6%, respectively. Geometry-aware regularization improves local latent proxies but does not improve decoded-text metrics in the available runs. The contribution is methodological rather than algorithmic: the paper presents a reusable staged diagnosis for one concrete pipeline and shows that, in this setting, codec fidelity rather than latent denoising sets the practical quality ceiling.
Chinese Translation
压缩短文本生成器可能在两个不同的环节出现失败:编解码器可能在生成开始之前丢弃信息,或者潜在生成器可能产生弱编码。如果不区分这些失败模式,研究人员可能会在错误的组件上浪费计算资源。我们在一个受控的64到16的TinyStories案例研究中探讨了这个问题,该研究基于层次化的VQ-VAE-2编解码器和掩蔽离散扩散生成器(MDLM)。我们使用一种分阶段验证协议,分别评估编解码器重建保真度、潜在生成质量和辅助潜在诊断,所有这些都在一个共享的外部GPT-2评分器下进行,同时报告几何研究的补充语义指标。在测试的配置中,仅编解码器重建就将外部困惑度的中位数从15.17提高到27.36(+80.4%),95百分位数从25.10提高到98.91(+294.1%),显示出主要的质量损失出现在潜在生成开始之前。在同一评分器下,代码空间的MDLM在实质上仍然优于标记空间的扩散,分别减少了均值、中位数和95百分位数32.9%、30.9%和36.6%。几何感知正则化改善了局部潜在代理,但在可用的运行中并未改善解码文本指标。该研究的贡献在于方法论而非算法:本文提出了一种可重复使用的分阶段诊断方法,适用于一个具体的流程,并表明在这种设置中,编解码器的保真度而非潜在去噪设定了实际质量的上限。
cs.CL / 55 / 2607.24191

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

StanceFlip:一个全面的多维基准用于多模态对话立场翻转预测
Chai, Heyan, Li, Xin, Wang, Wenjie, Qin, Jianyang, Li, Chaoyang, Wang, Lu, Chen, Hao, Liao, Qing
Abstract
Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic ambiguities such as sarcasm. To address these limitations, we propose StanceFlip, a benchmark designed for multimodal conversational stance flipping forecasting over multi-turn dialogues across five modalities and multi-scenarios, which includes two novel subtasks: 1) Multimodal Stance Sextuple Extraction, extracting holder, target, emotion, sentiment, stance, and rationale as static state snapshots of dialogue to capture fine-grained cognitive structures. 2) Dynamic Stance Flip Attribution, tracking stance reversals across the conversation and identifying their underlying triggers. Alongside the dataset, we propose a dedicated framework, named ConStaFF, for Multimodal Conversational Stance Flipping Forecasting (MCSFF). Built upon a large language model, ConStaFF performs end-to-end stance reasoning, with a Thought-of-Stance (ToS) reasoning framework and a self-reflective verification mechanism integrated for structured stance modeling and faithful flip attribution. Specifically, ToS decomposes the reasoning process into specialized cognitive personas to formulate target propositions, resolve cross-modal conflicts, and infer historical stance trajectories. Extensive experiments show that our approach achieves state-of-the-art performance on both sextuple extraction and flip-trigger attribution, outperforming strong multimodal large language model baselines by substantial margins.
Chinese Translation
对话立场检测已从静态文本分析转向动态多模态建模。然而,现有基准存在三个主要局限性:未能捕捉信念的动态演变,尤其是在立场反转期间;难以将情感状态与逻辑推理分离;以及忽视多模态线索在解决讽刺等语用模糊性中的关键作用。为了解决这些局限性,我们提出了StanceFlip,这是一个针对多模态对话立场翻转预测的基准,涵盖五种模态和多种场景,包含两个新颖的子任务:1)多模态立场六元组提取,提取持有者、目标、情感、情绪、立场和理由作为对话的静态状态快照,以捕捉细粒度的认知结构。2)动态立场翻转归因,跟踪对话中的立场反转并识别其潜在触发因素。除了数据集,我们还提出了一个专门的框架,名为ConStaFF,用于多模态对话立场翻转预测(MCSFF)。ConStaFF基于大型语言模型,执行端到端的立场推理,集成了立场思维(Thought-of-Stance, ToS)推理框架和自我反思验证机制,以实现结构化的立场建模和真实的翻转归因。具体而言,ToS将推理过程分解为专门的认知角色,以制定目标命题、解决跨模态冲突并推断历史立场轨迹。大量实验表明,我们的方法在六元组提取和翻转触发归因方面均达到了最先进的性能,显著超越了强大的多模态大型语言模型基线。
cs.CL / 56 / 2607.24223

A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

相关性的新的角色:引导代理搜索中的语料库交互
Li, Jiangnan, Li, Yuqing, Yu, Mo, Zhang, Jinchao, Zhou, Jie
Abstract
Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-$k$ content, but document relevance alone cannot localize, compose, or verify the evidence required by complex questions. Direct Corpus Interaction (DCI) enables such fine-grained operations through grep-style exploration, but its relevance-agnostic search can expose useful clues late and delay convergence. Recent advances use relevance to narrow the corpus into a working space for interaction. Once interaction begins, however, relevance still does not directly guide which documents grep searches first or distinguish informative excerpts from a broad set of matches to let LLMs see them first. We introduce the Relevance-Aware RipGrep Search Agent (RARG), which turns relevance into an execution prior for corpus interaction. RARG provides coarse-to-fine relevance guidance: it orders documents for sequential 'ripgrep' traversal to expose globally relevant clues earlier, initializes promising entry points with query-relevant paragraphs, and reranks grep matches to surface informative excerpts that document-level ranking may otherwise obscure. Across challenging browse question answering and reasoning-intensive retrieval, RARG improves the accuracy--efficiency frontier over retrieval-based and direct-interaction agents. These results demonstrate that relevance-aware interaction enables faster and more reliable search convergence.
Chinese Translation
相关性是一个依赖于查询的估计,用于判断文档或摘录是否包含有用的证据。现有的检索代理使用相关性来选择前 $k$ 个内容,但仅依靠文档相关性无法定位、组合或验证复杂问题所需的证据。直接语料库交互(Direct Corpus Interaction, DCI)通过类似 grep 的探索实现了这种细粒度操作,但其与相关性无关的搜索可能会在较晚阶段暴露有用线索,从而延迟收敛。最近的进展利用相关性将语料库缩小到一个交互的工作空间。然而,一旦交互开始,相关性仍然无法直接指导 grep 搜索首先查找哪些文档或区分来自广泛匹配集的有信息的摘录,以便让大型语言模型(LLMs)优先看到它们。我们引入了相关性感知的 RipGrep 搜索代理(Relevance-Aware RipGrep Search Agent, RARG),该代理将相关性转化为语料库交互的执行先验。RARG 提供了粗到细的相关性指导:它为顺序的 'ripgrep' 遍历排序文档,以更早地暴露全局相关线索,用查询相关的段落初始化有前景的切入点,并对 grep 匹配进行重新排序,以呈现文档级排名可能掩盖的信息摘录。在具有挑战性的浏览问答和推理密集的检索任务中,RARG 改善了基于检索和直接交互代理的准确性-效率边界。这些结果表明,相关性感知的交互能够实现更快和更可靠的搜索收敛。
cs.CL / 57 / 2607.24236

CAGE: Cognitive Attribution Graphs for Faithful Inline Citation Generation in Long-Form Question Answering

CAGE:用于长文本问答中忠实内联引用生成的认知归因图
Yan, Zhichao, Li, Shizhao, Wang, Jiapu, Luo, Haoran, Zhang, Qingang, Chen, Jiaoyan, Li, Ru, Pan, Jeff Z.
Abstract
Long-form question answering increasingly relies on retrieved evidence to make LLM outputs verifiable, with inline citations tracing claims to source documents. However, existing systems often attach citations that are topically related but insufficient to support their claims. We identify attribution ambiguity as a structural challenge: end-to-end generation must implicitly resolve combinatorial claim--document assignments, obscuring evidential boundaries and increasing the risk of evidence-boundary overrun, where claims exceed cited support. To address this challenge, we propose CAGE (Cognitive Attribution Graphs for Citation Generation), a two-stage framework that introduces an explicit cognitive attribution map before answer generation. CAGE first trains a plug-and-play Cognitive Map Induction Model to construct answer-centered support subgraphs, aligning each semantic answer unit with supporting documents through explicit relations. A Structured Citation Reasoning Model then realizes these units as sentence-level claims with map-aligned citations. Experiments on ASQA, ELI5, and ExpertQA show that CAGE achieves state-of-the-art performance, demonstrating the effectiveness of attribution-space contraction and map-guided citation generation.
Chinese Translation
长文本问答越来越依赖于检索到的证据,以使大型语言模型(LLM)的输出可验证,并通过内联引用将主张追溯到源文档。然而,现有系统往往附加与主题相关但不足以支持其主张的引用。我们将归因模糊性识别为一种结构性挑战:端到端生成必须隐式解决组合主张-文档分配,模糊证据边界并增加证据边界超限的风险,即主张超出引用支持。为了解决这一挑战,我们提出了CAGE(用于引用生成的认知归因图),这是一个两阶段框架,在答案生成之前引入显式的认知归因图。CAGE首先训练一个即插即用的认知图诱导模型,以构建以答案为中心的支持子图,通过显式关系将每个语义答案单元与支持文档对齐。然后,结构化引用推理模型将这些单元实现为具有图对齐引用的句子级主张。在ASQA、ELI5和ExpertQA上的实验表明,CAGE实现了最先进的性能,证明了归因空间收缩和图引导引用生成的有效性。
cs.CL / 58 / 2607.24268

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

准确性掩盖了语言模型的失败:在匹配输出预算下测量失败状态
Yang, Zongyou, Hou, Yinghan
Abstract
Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation framework that separates scorer-independent execution evidence, including termination, answer exposure, parseability, and completion length, from scorer-dependent correctness. Across 2,550 outputs from five fixed Qwen and DeepSeek configurations on MATH and ARC-Challenge, matched 2,048-token limits produce sharply different execution mixtures: 49 of 450 Qwen MATH outputs terminate without a final answer, compared with 5 of 300 DeepSeek MATH outputs and none of the 750 ARC outputs. Among the same 300 DeepSeek MATH question-model pairs, no missing-final length termination is observed at 8,192 tokens. A coverage-audited targeted verification study further shows that candidate-selection and aggregation policies can substantially alter comparative accuracy estimates. These results demonstrate that accuracy conflates execution case mix with verification policy. Evaluations of test-time methods should therefore report pre-intervention execution states, verification coverage, and scorer provenance alongside accuracy.
Chinese Translation
语言模型基准将两个不同的测量问题合并为一个单一的准确性评分:响应是否达到了可评估状态,以及其答案是否被判定为正确。我们引入了一个两层评估框架,将与评分者无关的执行证据(包括终止、答案暴露、可解析性和完成长度)与依赖于评分者的正确性区分开来。在对来自五个固定 Qwen 和 DeepSeek 配置的 2,550 个输出进行的 MATH 和 ARC-Challenge 测试中,匹配的 2,048 令牌限制产生了截然不同的执行组合:在 450 个 Qwen MATH 输出中,有 49 个在没有最终答案的情况下终止,而在 300 个 DeepSeek MATH 输出中仅有 5 个如此,750 个 ARC 输出中则没有。对于同样的 300 个 DeepSeek MATH 问题-模型对,在 8,192 令牌时未观察到缺失最终长度的终止。经过覆盖审核的目标验证研究进一步表明,候选选择和聚合策略可以显著改变比较准确性估计。这些结果表明,准确性将执行案例组合与验证策略混淆。因此,测试时间方法的评估应报告干预前的执行状态、验证覆盖率和评分者来源,以及准确性。
cs.CL / 59 / 2607.24273

INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

INS-ActBench:评估大型语言模型专业精算能力的综合基准
Chen, Changyu, Lin, Chenwei, Xu, Xian
Abstract
Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limits their ability to assess realistic professional workflows that require auditable, context-grounded, and tool-executable decisions. We introduce \textbf{INS-ActBench}, a comprehensive benchmark for evaluating professional actuarial capability in LLMs. INS-ActBench contains 12,050 Q\&A pairs from public exams and sample questions released by 16 actuarial associations. It covers three subsets: \textbf{INS-Act-Know} for standardized actuarial knowledge, \textbf{INS-Act-Case} for long-context insurance case reasoning, and \textbf{INS-Act-Practice} for spreadsheet and R-code tasks with verifiable numerical outputs. Experiments on nine representative LLMs and human actuarial experts reveal a clear capability boundary: frontier LLMs perform strongly on standardized knowledge, but remain much weaker in case reasoning, tool-based workflows, and jurisdiction-sensitive practice. INS-ActBench provides a reproducible foundation for developing actuarial LLMs toward reliable professional assistance. The code is available at https://github.com/FDU-INS/INS-ActBench.
Chinese Translation
大型语言模型(LLMs)在金融推理方面展现出强大的潜力,但现有基准往往在不同的环境中评估领域知识、数值推理、长上下文理解和工具使用。这限制了它们评估需要可审计、基于上下文和可执行决策的现实专业工作流程的能力。我们引入了 extbf{INS-ActBench},这是一个用于评估LLMs专业精算能力的综合基准。INS-ActBench包含来自16个精算协会的公共考试和样本问题的12,050个问答对。它涵盖三个子集: extbf{INS-Act-Know}用于标准化精算知识, extbf{INS-Act-Case}用于长上下文保险案例推理,以及 extbf{INS-Act-Practice}用于具有可验证数值输出的电子表格和R代码任务。在对九个代表性LLMs和人类精算专家的实验中,揭示了一个明显的能力边界:前沿LLMs在标准化知识方面表现强劲,但在案例推理、基于工具的工作流程和对辖区敏感的实践中仍然较弱。INS-ActBench为开发可靠的专业辅助精算LLMs提供了可重复的基础。代码可在 https://github.com/FDU-INS/INS-ActBench 获取。
cs.CL / 60 / 2607.24276

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

标记器税:量化和解释印度语言子词标记化的跨语言成本
Srivastava, Priyansh
Abstract
Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are trained predominantly on English-centric corpora, they introduce a systematic and often overlooked disadvantage for many non-English languages. In this work, we quantify this tokenizer tax for Indian languages using the FLORES-200 parallel corpus, measuring tokenization fertility across six widely used tokenizers and fourteen languages. Under cl100k_base (used by GPT-3.5 and GPT-4), Indian languages experience an average 8.0x tokenization tax relative to English, reaching 13.0x for Malayalam, reducing the effective context window to as little as 12% of that available to English users for equivalent semantic content. We identify the primary mechanism behind this disparity: failed byte-pair merges that leave text fragmented into single-byte tokens, with merge failure strongly correlating with tokenizer tax (Pearson r = 0.89). We further show that this phenomenon is not an inherent property of Indic scripts but a consequence of tokenizer design. Multilingual tokenizers such as XLM-R and OpenAI's o200k_base reduce the average Indic tokenizer tax by 73%, demonstrating that the disparity is largely remediable. Beyond token statistics, we quantify a practical consequence by showing that, under fixed context budgets, Indian-language documents preserve substantially less original content than equivalent English documents. Finally, we examine the relationship between tokenizer fertility and reading comprehension performance on the Belebele benchmark, finding that the apparent correlation is largely explained by language resource availability rather than tokenizer behavior alone.
Chinese Translation
大型语言模型(LLMs)通过子词标记器处理文本,而不是直接读取字符或单词。由于这些标记器主要在以英语为中心的语料库上训练,它们为许多非英语语言引入了系统性且常被忽视的劣势。在本研究中,我们使用FLORES-200平行语料库量化印度语言的标记器税,测量六种广泛使用的标记器和十四种语言的标记化繁殖性。在cl100k_base(用于GPT-3.5和GPT-4)下,印度语言相对于英语的平均标记化税为8.0倍,马拉雅拉姆语的标记化税达到13.0倍,使得有效的上下文窗口仅为英语用户在相同语义内容下可用的12%。我们确定了这一差异的主要机制:失败的字节对合并使文本碎片化为单字节标记,合并失败与标记器税之间存在强相关性(Pearson r = 0.89)。我们进一步表明,这一现象并不是印度文字的固有属性,而是标记器设计的结果。多语言标记器如XLM-R和OpenAI的o200k_base将平均印度标记器税降低了73%,表明这种差异在很大程度上是可以修复的。除了标记统计外,我们通过显示在固定上下文预算下,印度语言文档保留的原始内容显著少于等效英语文档,量化了一个实际后果。最后,我们考察了标记器繁殖性与Belebele基准测试中阅读理解表现之间的关系,发现这种明显的相关性在很大程度上是由语言资源的可用性解释的,而不仅仅是标记器行为所致。
cs.CL / 61 / 2607.24300

Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

自我创作验证在启发式自我改进代理中是不可靠的
Guo, Diandian, Cao, Cong, Yuan, Fangfang, Wang, Yingqi, Wang, Yueshan, Wang, Dakui
Abstract
Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier--deployment gap. This gap refers to the discrepancy between an agent's self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent's control.
Chinese Translation
自我改进代理通过反复重写程序策略、控制器或启发式规则来积累能力。它们通常依赖自我创作的测试或指标来决定是否接受后续编辑。代理同时控制优化对象和其验证者。因此,自我分配的分数可以保持接近完美,而实际部署性能却可能下降或保持低水平。我们通过验证者-部署差距研究这个问题。这个差距指的是代理自我创作的验证信号与代理无法观察或访问的密封部署评估之间的差异。我们探讨了在迭代政策和测试重写过程中,自我创作验证如何失败,这种失败如何随着能力的变化而变化,以及多小的外部信任足以防止实际回归被部署。为了解决这个问题,我们引入了密封外部接受循环(Sealed Exogenous Acceptance Loop, SEAL)。SEAL保留自我创作的测试,但通过固定的外部审计将每个候选者与现有者进行比较。代理无法创作或检查审计,只能接收接受/拒绝的结果,并且在明显回归后保留整个现有状态。我们的实验表明,这个问题在启发式学习环境中经常出现。这些环境需要通过试错发现目标对象。我们进一步发现,自我创作验证的失败是按能力分层的。较弱的代理倾向于在简单的自我测试后损害之前获得的策略。较强的代理则更稳定,但仍然错误地测量部署分布。标准的自我创作约束无法可靠地缩小这个差距。相比之下,SEAL在六个模型和三个随机种子上优于未保护的基线。可靠的自我改进不必放弃自我验证,但至少需要一个超出代理控制的部署接受信号。
cs.CL / 62 / 2607.24306

Rethinking the Generation Order of Block Diffusion Language Models

重新思考块扩散语言模型的生成顺序
Hou, Kai Syun, Kwok, James
Abstract
Diffusion language models enable flexible arbitrary-order generation, but existing sampling methods are mostly designed for early masked diffusion models (MDMs). In this work, we study sampling for recent block diffusion language models (BDLMs). We show empirically and analytically that these models are naturally more aligned with left-to-right decoding than MDMs. Based on this observation, we propose Parallel Autoregressive Decoding (PARD), a simple training-free sampling method that preserves left-to-right unmasking structure while allowing parallel token commitment. Extensive experiments show that PARD consistently outperforms existing parallel samplers in generation quality, while achieving substantial speedups over pure AR decoding with only a small quality gap.
Chinese Translation
扩散语言模型支持灵活的任意顺序生成,但现有的采样方法大多是为早期的掩码扩散模型(Masked Diffusion Models, MDMs)设计的。在本研究中,我们探讨了最近的块扩散语言模型(Block Diffusion Language Models, BDLMs)的采样。我们通过实证和分析表明,这些模型在自然上更符合从左到右的解码方式,而非MDMs。基于这一观察,我们提出了并行自回归解码(Parallel Autoregressive Decoding, PARD),这是一种简单的无训练采样方法,能够保持从左到右的去掩码结构,同时允许并行的标记承诺。大量实验表明,PARD在生成质量上始终优于现有的并行采样器,同时在速度上相较于纯自回归解码实现了显著的提升,仅存在小幅的质量差距。
cs.CL / 63 / 2607.24312

CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models

CONSISTRE:一个统一的关注一致性的文档级关系提取框架,基于大型语言模型
Sun, Mingxuan
Abstract
Document-level relation extraction (DocRE) aims to extract relations among multiple entities across extended contexts while maintaining consistency across predicted triples. Although large language models (LLMs) show remarkable reasoning capabilities in information extraction, their predictions are typically generated independently for each candidate triple and may violate fundamental relational constraints such as transitivity, symmetry, and functional uniqueness, leading to contradictory and unreliable outputs. We propose CONSISTRE, a unified consistency-aware framework for DocRE that addresses this limitation through two complementary tracks. The first operates at inference time for black-box LLMs, combining constraint-aware prompting, constraint-based verification, and iterative self-reflection to refine predictions without task-specific fine-tuning. The second injects consistency knowledge into smaller open-source models via a knowledge distillation and reinforcement learning pipeline: reasoning traces from a powerful teacher are distilled into a student via supervised fine-tuning, followed by GRPO alignment using a composite reward that jointly optimizes extraction performance and relational consistency. Together, the two tracks cover both API-accessible and locally deployable scenarios under a unified consistency formulation. Experiments on DocRED show that both tracks outperform their baselines, with the inference-time track achieving competitive F1 using off-the-shelf black-box LLMs and the training-time track substantially narrowing the gap between 7--8B open-source models and state-of-the-art proprietary LLMs at a fraction of their inference cost. Ablation studies confirm that explicit consistency modeling mitigates relational contradictions and enhances the reliability of LLM-based DocRE across both deployment paradigms.
Chinese Translation
文档级关系提取(DocRE)旨在从扩展上下文中提取多个实体之间的关系,同时保持预测三元组的一致性。尽管大型语言模型(LLMs)在信息提取方面展现出显著的推理能力,但它们的预测通常是针对每个候选三元组独立生成的,可能会违反基本的关系约束,如传递性、对称性和功能唯一性,从而导致矛盾和不可靠的输出。我们提出了CONISTRE,一个统一的关注一致性的DocRE框架,通过两个互补的轨道来解决这一局限性。第一个轨道在推理时针对黑箱LLM,结合了关注约束的提示、基于约束的验证和迭代自我反思,以在不进行特定任务微调的情况下优化预测。第二个轨道通过知识蒸馏和强化学习管道将一致性知识注入较小的开源模型:从强大的教师模型中提取的推理轨迹通过监督微调蒸馏到学生模型中,随后使用复合奖励进行GRPO对齐,以共同优化提取性能和关系一致性。两个轨道共同覆盖了统一一致性公式下的API可访问和本地可部署场景。在DocRED上的实验表明,两个轨道均优于其基线,其中推理时轨道在使用现成的黑箱LLM时实现了具有竞争力的F1值,而训练时轨道则在其推理成本的一小部分下显著缩小了7-8B开源模型与最先进的专有LLM之间的差距。消融研究确认,显式的一致性建模减轻了关系矛盾,并增强了基于LLM的DocRE在两种部署范式下的可靠性。
cs.CL / 64 / 2607.24332

Cross-Attention Calibrated Deduplication for Retrieval-Augmented Generation System

用于检索增强生成系统的交叉注意力校准去重
Huy, Phuong Le, Nguyen, Nam H., Dang, Quan V.
Abstract
Common chunking strategies in Retrieval-Augmented Generation (RAG) systems often create redundant chunks. These redundant chunks make the vector database bigger and slow down retrieval. A common fix is cosine-similarity thresholding. This method reduces each chunk to a single vector, then compares vectors using a similarity score. But a single vector can lose the fine-grained, token-level detail needed to tell a true duplicate apart from a chunk that just shares the same topic. We propose Cross-Attention Calibrated Deduplication (CACD). CACD checks each new chunk against an in-memory pool of chunks already kept, using a cross-encoder instead of a single pooled vector. This keeps token-level detail all the way to the final comparison. CACD combines three parts: the cross-encoder comparison itself, a New Information Score (NIS) that measures how much of a chunk is not explained by a candidate already kept, and a majority vote across several candidates rather than a single best match. NIS is calculated from the attention entropy of the cross-encoder. We tested CACD against five existing filtering methods, nine chunking strategies, and 18 configurations, all on the full SQuAD 1.1 validation set. In our experiments, CACD removes 9.75% of chunks on average. This drop rate is close to other semantic-level methods, and much higher than exact-match filters, which barely remove anything. In these experiments, CACD also processes each configuration in 51.0 seconds on average, about 27% faster than the strongest baseline, NERExact (69.6s), and about 7x faster than cosine-similarity filtering (356.7s). These results come from a single dataset, so we present them as an early comparison, not a general claim. Code for the baseline evaluation and for CACD is available at https://github.com/lehuyphuong/rag_bench and https://github.com/lehuyphuong/cacd_dedup.
Chinese Translation
检索增强生成(RAG)系统中常见的分块策略往往会产生冗余块。这些冗余块使得向量数据库变得更大,从而降低检索速度。常见的解决方法是余弦相似度阈值法。该方法将每个块简化为一个单一向量,然后使用相似度评分比较向量。然而,单一向量可能会丢失区分真实重复项与仅共享相同主题的块所需的细粒度、标记级别的细节。我们提出了交叉注意力校准去重(CACD)。CACD使用交叉编码器检查每个新块与已存储的内存池中的块,而不是使用单一的池化向量。这保持了标记级别的细节,直到最终比较。CACD结合了三个部分:交叉编码器比较本身、新信息评分(NIS),该评分衡量一个块中未被已存候选解释的部分,以及对多个候选者进行的多数投票,而不是单一最佳匹配。NIS是通过交叉编码器的注意力熵计算得出的。我们在完整的SQuAD 1.1验证集上测试了CACD,针对五种现有过滤方法、九种分块策略和18种配置。在我们的实验中,CACD平均去除了9.75%的块。这一去除率接近其他语义级别的方法,远高于几乎不去除任何内容的精确匹配过滤器。在这些实验中,CACD还平均以51.0秒处理每个配置,比最强基线NERExact(69.6秒)快约27%,比余弦相似度过滤(356.7秒)快约7倍。这些结果来自单一数据集,因此我们将其视为早期比较,而非普遍声明。基线评估和CACD的代码可在https://github.com/lehuyphuong/rag_bench和https://github.com/lehuyphuong/cacd_dedup获取。
cs.CL / 65 / 2607.24352

Retrieval-Augmented Large Language Models as Components of Cognitive Computing architecture for Regulatory Knowledge Management

作为认知计算架构组件的检索增强大型语言模型在监管知识管理中的应用
Nowak-Nova, Dariusz
Abstract
The aim of this article is to verify whether integrating large language models (LLMs) with the Retrieval-Augmented Generation (RAG) architecture enables their transformation from standalone generative models into components of cognitive computing infrastructure with enhanced epistemic reliability. The study proposes an architectural approach based on locally deployed LLMs operating in on-premises environments without high-end GPU accelerators and examines their applicability in supporting regulatory management processes requiring continuous analysis and interpretation of legal acts. The proposed solution combines local LLMs with external knowledge repositories, creating a hybrid cognitive architecture in which the language model performs semantic interpretation while the RAG layer provides controlled knowledge retrieval, contextualization, and traceability of information sources. The implementation was validated using the Ollama and LM Studio execution environments together with the Polish language models Bielik and PLLuM running on consumer-class hardware. The results demonstrate that augmenting LLMs with RAG significantly improves the factual consistency, domain specificity and normative precision of generated texts while reducing the risk of unsupported content generation. Furthermore, the study shows that integrating RAG introduces auditability, controlled knowledge management and dynamic updating of regulatory information without retraining the language model. The findings indicate that locally deployed LLMs enhanced with RAG should be regarded not merely as text generation tools but as semantic processing modules within cognitive computing infrastructures supporting regulatory compliance and organizational decision-making in environments characterized by high legal and informational volatility.
Chinese Translation
本文旨在验证将大型语言模型(LLMs)与检索增强生成(RAG)架构集成是否能够将其从独立的生成模型转变为具有增强认知可靠性的认知计算基础设施组件。研究提出了一种基于本地部署的LLMs的架构方法,这些模型在没有高端GPU加速器的本地环境中运行,并考察其在支持需要持续分析和解读法律法规的监管管理流程中的适用性。所提出的解决方案将本地LLMs与外部知识库相结合,创建了一种混合认知架构,其中语言模型执行语义解释,而RAG层提供受控的知识检索、上下文化和信息源的可追溯性。该实施方案通过Ollama和LM Studio执行环境以及在消费级硬件上运行的波兰语言模型Bielik和PLLuM进行了验证。结果表明,使用RAG增强LLMs显著提高了生成文本的事实一致性、领域特异性和规范精确性,同时降低了生成不支持内容的风险。此外,研究显示,集成RAG引入了可审计性、受控知识管理和监管信息的动态更新,而无需重新训练语言模型。研究结果表明,增强RAG的本地部署LLMs应被视为不仅仅是文本生成工具,而是支持监管合规和组织决策的认知计算基础设施中的语义处理模块,尤其是在法律和信息高度波动的环境中。
cs.CL / 66 / 2607.24368

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

保持在心中:代理记忆中隐性关联盲点的基准测试
Li, Ruizhe, Du, Mingxuan, Xu, Benfeng, Mao, Zhendong
Abstract
Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.
Chinese Translation
长期记忆系统将用户所说的内容存储在外部存储中,并在相关查询到达时检索该内容。这个接口基于一个如此自然以至于很少被明确说明的假设:所需的记忆将与需要它的查询相似。世界知识打破了这一假设。树坚果过敏应当通过其杏仁粉成分改变对马卡龙请求的回答,但这两段文本没有检索者可以看到的线索。我们将这种失效模式称为隐性关联盲点,并介绍了 InMind,这是一个涵盖十个生活领域的125项任务、经过专家验证的基准,其中113项任务基于可引用的公共来源。其配对控制分离了现有评估混淆的三种解释:事实从未存储、模型缺乏桥接知识,或事实已存储但从未显现。结果非常明确。在上下文中放置的决定性记忆使得主干系统能够回答84.0%的间接查询;而当相同的记忆必须被检索时,六种向量、图形和代理记忆系统的表现最多仅达到14.4%,尽管它们在需求时能够以高达100%的准确率回忆相同的事实。具有八倍维度的嵌入提高了每个系统的答案盲目标的回忆率,但基本上保持了差距不变。一个在查询到达之前保持记忆可见的最小诊断探测器恢复了大部分差距,定位到查询条件接口本身的失效,并指向路由,决定哪些事实必须保持可见,作为 InMind 旨在评分的开放问题。
cs.CL / 67 / 2607.24371

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

医疗互操作性的闭环验证-修复:临床大型语言模型模式合规性的多模型研究
Shen, Jianru
Abstract
Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic coding, CPT for procedure billing, and HL7 FHIR for data exchange. While large language models demonstrate clinical reasoning capabilities, their integration into electronic health record systems faces a critical barrier: schema noncompliance. We evaluate three open-source models, Qwen2.5 7B, Llama 3.1 8B, and Gemma2 9B, via local deployment across 320 clinical scenarios spanning ten medical specialties, yielding 960 model-scenario pairs assessed under paired baseline and validation-repair conditions. First, schema noncompliance is consistent across the three model families, with baseline compliance rates ranging from 85.9 to 91.6 percent despite varying architectures and training data, suggesting shared gaps in medical training corpora rather than model-specific limitations. Second, 96 percent of validator-detected failures are representation-level format violations such as alternative medical abbreviations and code prefixes, indicating models follow clinical writing conventions but lack awareness of healthcare IT standards. Third, the validation-repair framework achieves 99.0 percent overall compliance, ranging from 98.4 to 99.4 percent across models, with most errors resolving within one or two iterations. Exact McNemar p-values below 0.001 and absolute improvements of 7.8 to 12.5 percentage points across model sizes confirm statistical significance. These results support closed-loop validation-repair as an effective system-level safeguard for healthcare interoperability, improving schema-level readiness for downstream clinical system integration.
Chinese Translation
医疗互操作性要求人工智能系统生成符合标准化模式的结构化输出,包括用于诊断编码的ICD-10、用于程序计费的CPT以及用于数据交换的HL7 FHIR。尽管大型语言模型展示了临床推理能力,但它们在电子健康记录系统中的集成面临一个关键障碍:模式不合规。我们通过在320个临床场景中本地部署三种开源模型(Qwen2.5 7B、Llama 3.1 8B和Gemma2 9B),评估了这三种模型在十个医学专业中的表现,共生成960对模型-场景配对,并在配对基线和验证-修复条件下进行评估。首先,三种模型家族之间的模式不合规性是一致的,基线合规率在85.9%到91.6%之间,尽管架构和训练数据各不相同,这表明医学训练语料库存在共同的缺口,而非模型特定的局限性。其次,96%的验证者检测到的失败是表示层面的格式违规,例如替代医学缩写和代码前缀,表明模型遵循临床写作惯例,但缺乏对医疗信息技术标准的意识。第三,验证-修复框架实现了99.0%的整体合规性,各模型的合规率在98.4%到99.4%之间,大多数错误在一到两个迭代内解决。精确的McNemar p值低于0.001,模型规模的绝对改进在7.8到12.5个百分点之间,确认了统计显著性。这些结果支持闭环验证-修复作为医疗互操作性的有效系统级保障,提高了下游临床系统集成的模式级准备度。
cs.CL / 68 / 2607.24435

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

LEX-EC:一种用于黑箱环境中零样本大语言模型个性分类的词汇证据通道审计框架
Harbison, Brittany, Goel, Ashok K.
Abstract
Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.
Chinese Translation
大型语言模型可以轻松地从文本中分配个性标签,但模型的可解释性仍然是一个未解决的问题。为了解决这一差距,我们引入了LEX-EC,这是一种可重用的黑箱审计框架,结合了流行性和一致性诊断与受控的词汇消融,以区分边际分布效应与在受限证据下可恢复的特征相关信号。利用该框架,我们展示了不同文本类型可能表现出截然不同的特征:自由形式的论文文本包含最广泛但仍然较弱的信号;在研究生自我介绍中,观察到的外向性关联在掩蔽后减弱;而单条Facebook状态即使在特征平衡样本中也几乎没有稳定证据,表明内容或长度可能存在下限。掩蔽主题和人口统计内容削弱了一些关联,同时保留了从功能词、情感术语和认知风格词汇中可检测的其他关联。语言提示改变了模型的自我解释,但并未消除主题内容。LEX-EC共同评估分类流行性、项目级关联、机会校正一致性、在词汇限制下的持久性和模型生成解释中的提示敏感性。在不同数据集、模型和提示中,LEX-EC描述了特征关联如何随着可用词汇证据的变化而变化,提出了词汇方法在个性标记黑箱可解释性中的新应用。
cs.CL / 69 / 2607.24471

Grounding latent algorithm routing in transformer reasoning

将潜在算法路由与变压器推理相结合
Zhang, Xiangbo, Ma, Xiaoxu
Abstract
A central question in the in-context learning literature is whether transformers can organize episode-level adaptation around different inductive-bias families. We study this question in a controlled setting through latent algorithm routing: route-like behavior in which the solver-family preference changes with the latent data-generating regime while prompt form is held fixed, remains stable under nuisance perturbations, and is selectively influenced by targeted activation interventions without large losses in answer quality. We introduce ROUTEBENCH, a diagnostic benchmark whose regimes differentially favor global shrinkage, sparsity, robustness, and locality, operationalized by ridge-like, lasso-like, Huber-like, and kNN-like family representatives. Across dense decoder-only transformers trained from scratch at 44M-612M parameters, a 306M model closes 80.9 percent of the oracle-routing gap and achieves route F1 of 84.1. The effect remains substantial under natural-language renderings, shuffled supports, lexical paraphrases, and a unified four-way routing setting. Stronger adaptive alternatives, including an input-conditioned soft mixture and an unsupervised Gumbel router, narrow the gap but remain below the 306M and 612M models on route F1 and OOD performance. Probe controls and matched activation-patching controls further show that route-relevant internal directions are decodable and functionally involved in solver-family-consistent output behavior. These results provide controlled evidence that dense transformers trained on ROUTEBENCH can develop route-like internal variables, but they do not establish universal routing in pretrained language models or unrestricted natural-language reasoning.
Chinese Translation
在上下文学习文献中,一个核心问题是变压器是否能够围绕不同的归纳偏置家族组织情节级适应。我们通过潜在算法路由在受控环境中研究这个问题:在保持提示形式固定的情况下,解算器家族偏好随着潜在数据生成机制的变化而变化的路由行为,在干扰扰动下保持稳定,并且在没有显著损失答案质量的情况下,受到针对性激活干预的选择性影响。我们引入了ROUTEBENCH,这是一个诊断基准,其机制在全局收缩、稀疏性、鲁棒性和局部性方面具有不同的偏好,通过类似岭回归、类似套索、类似Huber和类似kNN的家族代表进行操作。在从零开始训练的44M-612M参数的稠密解码器变压器中,306M模型缩小了80.9%的oracle路由差距,并实现了84.1的路由F1。该效果在自然语言呈现、随机支持、词汇释义和统一的四路路由设置下仍然显著。更强的自适应替代方案,包括输入条件的软混合和无监督的Gumbel路由器,缩小了差距,但在路由F1和OOD性能上仍低于306M和612M模型。探测控制和匹配激活补丁控制进一步表明,与路由相关的内部方向是可解码的,并在解算器家族一致的输出行为中发挥功能。这些结果提供了控制证据,表明在ROUTEBENCH上训练的稠密变压器可以发展出类似路由的内部变量,但并未确立预训练语言模型或不受限制的自然语言推理中的普遍路由。
cs.CL / 70 / 2607.24492

SINT-Flow: Schema Integration using Large Language Model Workflows

SINT-Flow:基于大型语言模型工作流的模式集成
Korini, Keti, Bizer, Christian
Abstract
The goal of schema integration is, given a set of input schemata or tables, to derive a global, unified schema that is able to represent the concepts, attributes, and relationships of all input tables in a coherent fashion. This paper presents SINT-Flow, a schema integration framework composed of five LLM-based operators that can be combined into workflows to perform fully automated, end-to-end schema integration. In contrast to existing approaches, SINT-Flow can process denormalized source tables that contain attributes describing multiple entity types. During the schema integration process, these tables are decomposed into separate entity-specific relations. To evaluate SINT-Flow, we introduce SINT-Bench, a schema integration benchmark comprising 10 schema integration tasks consisting of altogether 93 relational tables, including tables that describe multiple types of entities. We evaluate SINT-Flow using GPT-5.2 as well as the open-weight model Qwen-3.6-27B as alternative backbone models. Using these models, SINT-Flow achieves F1 scores of at least 96% for entity-type detection, 85% for attribute detection, and 83% for schema mapping. Furthermore, we perform an ablation study to prove the utility of the applied self-consistency strategy as well as the inclusion of a review loop into the schema matching operator.
Chinese Translation
模式集成的目标是,在给定一组输入模式或表的情况下,推导出一个全球统一的模式,能够以连贯的方式表示所有输入表的概念、属性和关系。本文提出了SINT-Flow,一个由五个基于大型语言模型(LLM)的操作符组成的模式集成框架,这些操作符可以组合成工作流,以实现完全自动化的端到端模式集成。与现有方法相比,SINT-Flow能够处理包含描述多种实体类型的属性的非规范化源表。在模式集成过程中,这些表被分解为独立的实体特定关系。为了评估SINT-Flow,我们引入了SINT-Bench,这是一个包含10个模式集成任务的基准,涉及总共93个关系表,包括描述多种实体类型的表。我们使用GPT-5.2以及开放权重模型Qwen-3.6-27B作为替代主干模型来评估SINT-Flow。使用这些模型,SINT-Flow在实体类型检测中获得至少96%的F1分数,在属性检测中获得85%,在模式映射中获得83%。此外,我们还进行了消融研究,以证明所应用的自一致性策略的有效性以及在模式匹配操作符中纳入审查循环的必要性。
cs.CL / 71 / 2607.24515

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

针对多样化数据集的英语-泰米尔语和泰米尔语-英语的大型语言模型和基于变换器的机器翻译的系统分析
S, Sriharshaa, Sivanesan, Sangeetha
Abstract
The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by the restricted amount of parallel data for these languages, as well as their substantial amount of domain variation and morphological complexity. This research presents the comprehensive evaluation of the performance of several multilingual translation models on English-Tamil and Tamil-English translations across multiple datasets: NTREX, EnTamV2, WikiMatrix and PMIndia. This study evaluates supervised NMT systems, NLLB and mBART, using both the BLEU and chrF metric, and examines how these systems perform on data of different quality levels and domains. This performs an attention-based analysis to increase model interpretability by visualising the alignments of tokens in an English source text and their Tamil translations and vice-versa to provide insight into how they make translations. This study also demonstrates that using in-context prompting can provide an excellent way to perform a few-shot translation of English to Tamil and Tamil-English using a Tamil capable TamilLaMA model, and compare this to supervised approaches qualitatively. These findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.
Chinese Translation
低资源语言如泰米尔语的机器翻译面临的挑战主要源于这些语言的平行数据量有限,以及其领域变异性和形态复杂性较大。本研究对多个多语言翻译模型在英语-泰米尔语和泰米尔语-英语翻译中的表现进行了全面评估,涉及的数据集包括 NTREX、EnTamV2、WikiMatrix 和 PMIndia。本研究评估了监督神经机器翻译(NMT)系统 NLLB 和 mBART,使用 BLEU 和 chrF 指标,并考察这些系统在不同质量水平和领域的数据上的表现。通过可视化英语源文本中的标记与其泰米尔语翻译之间的对齐关系,进行基于注意力的分析,以提高模型的可解释性,从而深入了解它们如何进行翻译。本研究还表明,使用上下文提示可以为使用具备泰米尔语能力的 TamilLaMA 模型进行英语到泰米尔语和泰米尔语到英语的少量翻译提供一种优秀的方法,并与监督方法进行定性比较。这些发现表明,数据集的质量及其与领域的对齐将极大影响模型的表现,基于注意力的机制可以帮助解释能力,而少量样本的大型语言模型仍然能够生成结构上连贯的泰米尔语翻译。
cs.CL / 72 / 2607.24542

From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages

从转录到语义语料库分析:古代语言句子表示的无监督学习
de la Selle, Th{é}otime
Abstract
Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and semantic search-rely on sentence embeddings, yet existing methods transfer poorly to ancient languages: generic multilingual encoders underperform, specialized language models yield anisotropic representation spaces, and labeled similarity data is unavailable. We study two fully unsupervised strategies - TSDAE and contrastive sentence embedding (CSE) - that adapt a specialized token-level language model into a corpus-specific sentence encoder using only raw sentences. On the philologically central case of biblical reuse in patristic literature (2,935 expert-verified parallels in Latin and Ancient Greek, from Augustine, Jerome, and Athanasius), we decompose reuse identification into two separately evaluated tasks-binary detection and correspondence retrieval-and benchmark the adapted encoders against multilingual, specialized, distilled, and supervised fine-tuned baselines, as well as on artificially noised data simulating HTR artifacts and scribal abbreviations. The adapted encoders outperform all baselines on both tasks, with complementary profiles: TSDAE leads detection given a large in-domain corpus, while CSE leads retrieval, reaches its optimum with as few as 4-8k raw in-domain sentences-a few tens of seconds of training on a laptop GPU-and transfers across works and authors, including to noisy post-ATR text when retrained directly on it. UMAP atlases relate the geometric effect of each strategy to the measured gains, and the full pipeline-segmentation, fine-tuning, cross-corpus semantic search-is made available to non-specialists through the online tool Paraphrasis.
Chinese Translation
自动文本识别(ATR)现在为数字人文学科提供了大量非结构化、异构且常常噪声较大的古代语言文本。下游的语义分析——文本重用识别、对齐和语义搜索——依赖于句子嵌入,但现有方法在古代语言上的迁移效果不佳:通用的多语言编码器表现不佳,专门的语言模型产生各向异性的表示空间,并且缺乏标记的相似性数据。我们研究了两种完全无监督的策略——TSDAE(时间序列去噪自编码器)和对比句子嵌入(CSE)——它们利用原始句子将专门的标记级语言模型适配为语料库特定的句子编码器。在教父文学中圣经重用这一具有语言学中心性的案例(来自奥古斯丁、杰罗姆和亚他那修的2,935个经过专家验证的拉丁文和古希腊文平行文本)中,我们将重用识别分解为两个单独评估的任务——二元检测和对应检索——并将适配后的编码器与多语言、专门、提炼和监督微调的基线进行基准测试,同时还在模拟HTR(手写文本识别)伪影和抄写员缩写的人工噪声数据上进行测试。适配后的编码器在这两个任务上均优于所有基线,且具有互补的特征:在大型领域内语料库中,TSDAE在检测上表现优越,而CSE在检索上表现最佳,仅需4-8千个原始领域内句子即可达到最佳效果——在笔记本GPU上训练仅需几十秒,并且能够在作品和作者之间迁移,包括在直接重新训练后对噪声的后ATR文本进行处理。UMAP(统一流形近似与投影)图谱将每种策略的几何效应与测得的增益相关联,整个流程——分段、微调、跨语料库语义搜索——通过在线工具Paraphrasis向非专业人士开放。
cs.CL / 73 / 2607.24585

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference

从数据到设备:ELMOD 一种高效的德国优先 2.7B 语言模型用于移动推理
Gold, Darina, Schwirjow, Alexander, Haag, Viktor, Hangya, Viktor, Schlotthauer, Joel, Küch, Fabian, Hahn, Luzian
Abstract
We present ELMOD - Efficient Language Model for On-Device Deployment - a compact (2.7B) German language model designed for efficient inference on resource-constrained hardware. ELMOD was trained on a limited computational budget (55k H100 GPU hours) using exclusively publicly available data. We developed a suite of German-specific data pre-processing, which differ from English-oriented counterparts in their handling of morphological variation, compounding, and orthographic conventions. Furthermore, we introduced a quality filtering and rephrasing step, which increased the instructional quality of the data, improved performance during the annealing phase, and reduced overall compute requirements. Thanks to our architectural model and data choices, including prefiltering, our educational-quality filtering and rephrasal to raise the educational-quality, ELMOD is the strongest performer in its size class (<3B), matching the performance of 7B-parameter models in German.
Chinese Translation
我们提出了 ELMOD - 高效语言模型用于设备部署 - 这是一个紧凑的(2.7B)德语语言模型,旨在在资源受限的硬件上实现高效推理。ELMOD 在有限的计算预算(55k H100 GPU 小时)下,仅使用公开可用的数据进行了训练。我们开发了一套特定于德语的数据预处理方法,这些方法在处理形态变化、复合词和正字法规范方面与面向英语的对应方法有所不同。此外,我们引入了质量过滤和改写步骤,提高了数据的指导质量,改善了退火阶段的性能,并减少了整体计算需求。得益于我们的架构模型和数据选择,包括预过滤、教育质量过滤和改写以提高教育质量,ELMOD 在其规模类别(<3B)中表现最强,达到了 7B 参数模型在德语中的性能水平。
cs.CL / 74 / 2607.24586

D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models

D-Score:用于大型语言模型幻觉检测的谱隐状态信号
Raimondi, Bianca, Evangelista, Davide, Gabbrielli, Maurizio, Piccolomini, Elena Loli
Abstract
Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple spectral statistic computed from a single forward pass. For a fixed model, layer, and tolerance parameter, the D-Score counts how many singular directions of the hidden activation matrix have singular values that remain close to the leading one. We use this quantity as a hallucination score, classifying an input text as hallucinated when its D-Score is larger than a pre-defined quantity. The motivation is that, when a model processes a text that conflicts with information available in its own internal state, the hidden representation may encode both the asserted content and some form of counter-evidence, uncertainty, correction, or lack of support; this can make the hidden trajectory spread across additional singular directions. We formalize this intuition through a lightweight spectral argument and evaluate the resulting detector on FAVA-Annotation and RAGTruth. The experiments indicate that the D-Score is a strong hidden-state signal for hallucination detection, while requiring no external verifier, no retrieval step, and no multiple generations.
Chinese Translation
大型语言模型可以生成流畅的文本,但这些文本可能是虚假的、缺乏可用证据支持的,或与模型内部表示的信息不一致。我们从隐藏激活的几何结构研究幻觉检测,并引入D-Score,这是一种通过单次前向传播计算得出的简单谱统计量。对于固定的模型、层和容忍参数,D-Score计算隐藏激活矩阵中有多少个奇异方向的奇异值接近于主奇异值。我们将这一量作为幻觉评分,当输入文本的D-Score大于预定义的量时,将其分类为幻觉文本。其动机在于,当模型处理与其内部状态中可用信息相冲突的文本时,隐藏表示可能同时编码所声称的内容以及某种形式的反证据、不确定性、修正或缺乏支持;这可能导致隐藏轨迹在额外的奇异方向上扩展。我们通过轻量级的谱论证形式化这一直觉,并在FAVA-Annotation和RAGTruth上评估所得到的检测器。实验表明,D-Score是一个强大的隐状态信号,用于幻觉检测,同时不需要外部验证者、检索步骤或多次生成。
cs.CL / 75 / 2607.24593

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention

PIVOT:面向令牌级稀疏注意力的高效查询组索引
Liu, Hong, Cheng, Yuan, Niu, Lin, Su, Yi, Xue, Yufei, Liu, Anmin, Yu, Guanghua, Zhu, Jianchen
Abstract
Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incurring a cost of O(L^2) per layer for a sequence of length L. We observe that this per-query scan is largely redundant: nearby queries select highly overlapping top-k tokens, and the indexer scores are long-tailed along the key axis. We exploit these properties in PIVOT, Proxy Indexing Via One full-prefix Traversal, a training-free, drop-in replacement for the DSA indexer that shares one prefix scan across a group of nearby queries. PIVOT aggregates a group into a single proxy query, performs one shared full-prefix scan to obtain a candidate set, and then selects a top-k for each query from that set. Two variants trade speed for fidelity: PIVOT-Reuse shares the proxy top-k across the group for maximum speed, whereas PIVOT-Refine re-scores the candidate set with the indexer of each query and then selects an individual top-k, matching the dense indexer at a small additional cost. A single algorithm covers both inference phases, differing only in how groups are formed: fixed-size groups of consecutive queries in prefill, and the queries decoded together in one multi-token prediction (MTP) step in decode. On DeepSeek-V3.2 and GLM-5.1 across LongBench and RULER, PIVOT matches the accuracy of the dense DSA indexer while accelerating it by up to 4x and reducing end-to-end latency by up to 1.6x at long context.
Chinese Translation
令牌级稀疏注意力在生产系统中通过 DeepSeek 稀疏注意力(DSA)实现,使得下游注意力变得高效,但将瓶颈转移到了为其提供数据的索引器上。为了为每个查询选择前 k 个令牌,索引器仍然必须对每个前置令牌进行评分,这在长度为 L 的序列中每层的成本为 O(L^2)。我们观察到这种每查询扫描在很大程度上是冗余的:相邻的查询选择高度重叠的前 k 个令牌,而索引器的评分在键轴上呈长尾分布。我们在 PIVOT(通过一次完整前缀遍历的代理索引)中利用了这些特性,PIVOT 是一个无训练的、可替代 DSA 索引器的解决方案,它在一组相邻查询之间共享一次前缀扫描。PIVOT 将一组聚合为一个单一的代理查询,执行一次共享的完整前缀扫描以获得候选集,然后从该集合中为每个查询选择前 k 个。两种变体在速度和准确性之间进行权衡:PIVOT-Reuse 在组内共享代理前 k,以实现最大速度,而 PIVOT-Refine 则使用每个查询的索引器重新评分候选集,然后选择单独的前 k,以较小的额外成本匹配密集索引器。单一算法覆盖了两个推理阶段,仅在组的形成方式上有所不同:在预填充阶段使用固定大小的连续查询组,而在解码阶段在一次多令牌预测(MTP)步骤中一起解码查询。在 DeepSeek-V3.2 和 GLM-5.1 的 LongBench 和 RULER 上,PIVOT 在保持与密集 DSA 索引器相同的准确性的同时,将其加速最多 4 倍,并在长上下文中将端到端延迟减少最多 1.6 倍。
cs.CL / 76 / 2607.24604

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair

循环并不等同于可靠性:基于状态的证据与代理代码修复的类型化修订合同
Gao, Xueping, Yang, Jianwei, Yang, Qiang
Abstract
Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branches from identical frozen programs to remove post-treatment risk-set bias. In a prespecified 14B replication, stale traces harm 34/135 correct starts versus 4/135 with current traces, a 22.2-point increase (task-cluster 95\% CI $[8.9,37.0]$, exact Holm $p=0.0337$). A prospective 540-rollout policy eliminates observed correct-start harm but reduces wrong-start repair and fails its joint criterion. Repository experiments over 24 bugs and four coder stacks expose floor effects and component heterogeneity without Holm-significant effects. We therefore separate admission, preservation, grounded certification, competence, and liveness. We derive an evidence-bound typed loop contract and instantiate its mechanically enforceable subset in a reference implementation that binds verifier evidence to exact code states, preserves verified checkpoints, and emits auditable admission receipts. The implementation is an executable specification and conformance artifact, not evidence of improved repair competence or calibrated verifier dependence.
Chinese Translation
生成-测试-修订循环在编码代理中很常见,但单靠重复并不能提供可靠性保证。我们研究了找到正确补丁与保留、验证和提交补丁之间的差距。在一项封闭的五种种子研究中,对30个HumanEval修复产生了900条三次修订轨迹。在强制修订下,当前轨迹的当前正确性在一次修订后从0.820下降到两次修订后的0.673,尽管始终正确的比例上升至0.847。两个共同状态的研究使用2430个来自相同冻结程序的分支,以消除治疗后风险集偏差。在一项预先指定的14B复制中,陈旧轨迹对34/135个正确起始的影响与当前轨迹的4/135相比,增加了22.2个百分点(任务聚类95 ext{%}置信区间 $[8.9,37.0]$,精确Holm $p=0.0337$)。一项前瞻性的540次推出策略消除了观察到的正确起始损害,但减少了错误起始修复,并未满足其联合标准。对24个错误和四个编码器堆栈的代码库实验揭示了底效应和组件异质性,但没有Holm显著效应。因此,我们将接纳、保留、基于证据的认证、能力和活性分开。我们推导出一种基于证据的类型化循环合同,并在参考实现中实例化其机械可执行的子集,该实现将验证者证据绑定到精确的代码状态,保留已验证的检查点,并发出可审计的接纳收据。该实现是可执行的规范和符合性文档,而不是改进修复能力或校准验证者依赖性的证据。
cs.CL / 77 / 2607.24653

Kimi K3: Open Frontier Intelligence

Kimi K3:开放前沿智能
Kimi Team, Bai, Tongtong, Bai, Yifan, Bao, Yiping, C., M., Cai, Jianfeng, Cai, Xinyuan, Cao, Peizhou, Cao, Yuxuan, Chai, Ziwei, Charles, Y., Che, H. S., Chen, Guanduo, Chen, Guangyu, Chen, Guanzheng, Chen, Huarong, Chen, Jia, Chen, Jianlong, Chen, Jun, Chen, Kexin, Chen, Peng, Chen, Ruijue, Chen, Wentao, Chen, Xin, Chen, Yang, Chen, Yanru, Chen, Yifei, Chen, Yingjiang, Chen, Yuankun, Chen, Yujie, Chen, Yutian, Chen, Zhirong, Cheng, Dazhi, Cheng, Yean, Cui, Jialei, Cui, Jingbing, Dai, Anqi, Deng, Jiaqi, Ding, Hao, Ding, Rui, Ding, Shaofeng, Dong, Mengfan, Dong, Mengnan, Dong, Yuhao, Dong, Yuxin, Du, Angang, Du, Chenzhuang, Du, Dikang, Du, Jusen, Du, Yulun, Fan, Yu, Feng, Jing, Feng, Qiulin, Feng, Yichen, Fu, Kelin, Fu, Qiang, Gao, Fuxuan, Gao, Hongcheng, Gao, Jingyue, Gao, Tong, Gao, Weijia, Geng, Shangyi, Gong, Jie, Gong, Linhu, Gong, Shengao, Gong, Xiaochen, Gu, Qizheng, Gu, Yicheng, Guan, Shuhao, Guo, Haiqing, Guo, Shiqi, Guo, Xiang, Guo, Zhengyan, Hao, Beixi, Hao, Wenxin, Hao, Xiaoru, He, Dailan, He, Haotian, He, Lehan, He, Qi, He, Weiran, He, Xinran, He, Xinyi, He, Yibo, He, Yunjia, Hong, Chao, Hong, Tiange, Hu, Hao, Hu, Jiaxi, Hu, Ruikun, Hu, Weiming, Hu, Yangyang, Hu, Zhenxing, Hua, Liang, Huang, Jinbin, Huang, Ke, Huang, Ruiyuan, Huang, Siying, Huang, Weixiao, Huang, Yan, Huang, Zhengjie, Huang, Zhiqi, Hui, Yulong, Jia, Chaobo, Jiang, Yutong, Jiang, Zhejun, Jiang, Zuoyou, Jin, Wenyi, Jin, Xinyi, Jing, Yu, Kong, Huanjun, Lai, Guokun, Li, Aidi, Li, Cheng, Li, Chengyuan, Li, Cong, Li, Fang, Li, Guanyu, Li, Haoyang, Li, Jia, Li, Junxiong, Li, Lei, Li, Letian, Li, Lincan, Li, Weihong, Li, Wentao, Li, Xintong, Li, Yang, Li, Yishen, Li, Yiwei, Li, Yuxiao, Li, Zhaowei, Li, Zhaoxi, Li, Zheming, Li, Zhengxiao, Li, Zhiyuan, Lin, Jiawei, Lin, Xiaohan, Lin, Yibo, Lin, Zichao, Lin, Ziyan, Liu, Bill, Liu, Boxiao, Liu, Chuan, Liu, Liang, Liu, Shaowei, Liu, Shudong, Liu, Shuran, Liu, Tianwei, Liu, Weizhou, Liu, Yangyang, Liu, Yanming, Liu, Yibo, Liu, Yipeng, Liu, Zhengying, Liu, Zhiheng, Lu, Enzhe, Lu, Haoyu, Lu, Linqiang, Lu, Tingzhan, Lu, Zhiyuan, Luo, Aotian, Luo, G., Luo, Junyu, Luo, Yifan, Lyu, B., Lyu, Wenzhou, Mao, Shaoguang, Mei, Yuan, Men, Xin, Ni, Minqing, Niu, Yixuan, Pan, Siyuan, Peng, Shujun, Qi, Zhangyang, Qin, Ruoyu, Qin, ZeChao, Qin, Zeyu, Qiu, Haiquan, Qiu, Jianxin, Qiu, Jiezhong, Qu, Bowen, Qu, Yuhao, Shang, Zeyu, Shao, Youbo, Shen, Han, Shi, Jincheng, Shi, Juanfeng, Shi, Lidong, Shi, Shengyuan, Siu, Wingchun, Song, Pengwei, Song, Xiaoxi, Su, Jianlin, Su, Yunfeng, Su, Zhaochen, Sui, Lin, Sun, Jingsong, Sun, Junyao, Sun, Shaoning, Sun, Shuzhe, Sun, Tongyu, Sun, Yujun, Tai, Yunpeng, Tang, Chuning, Tang, Heyi, Tang, Sirui, Tang, Zecheng, Tian, Chaoran, Tian, Rongpeng, Tian, Yu, Tu, Wei, Wang, Chensi, Wang, Chuang, Wang, Chunjie, Wang, Dinglu, Wang, Feng, Wang, Hailong, Wang, Haiming, Wang, Hao, Wang, Hao, Wang, Huaqing, Wang, Hui, Wang, Jiayi, Wang, Jinglong, Wang, Jinhong, Wang, Jiuzheng, Wang, Linian, Wang, Shaobo, Wang, Shenzhi, Wang, Shuyi, Wang, Si, Wang, Siyuan, Wang, Tianfu, Wang, Wenjue, Wang, Xingran, Wang, Xinmei, Wang, Xinyuan, Wang, Xusheng, Wang, Yalin, Wang, Yangkun, Wang, Yao, Wang, Yaoyu, Wang, Yejie, Wang, Yiqin, Wang, Yucheng, Wang, Yuzhi, Wang, Zhaoji, Wang, Zhaowei, Wang, Zhengtao, Wang, Zhenhao, Wang, Zhongsheng, Wang, Zifan, Wei, Chu, Wei, Ming, Wei, Shouxin, Wen, Zichen, Wu, Fan, Wu, Haoning, Wu, Rucong, Wu, Wenhao, Wu, Xiaoxue, Wu, Yingcong, Wu, Yongqi, Wu, Yuxin, Wu, Zijian, Xian, Xinglang, Xiang, Chenxuan, Xiang, Yuye, Xiao, Bocheng, Xiao, Chenjun, Xiao, Xin, Xie, Jin, Xie, Xiaotong, Xie, Yifeng, Xie, Zhe, Xing, Bowei, Xiong, Yiming, Xu, Baosheng, Xu, Boyu, Xu, Jiale, Xu, Jianfan, Xu, Jing, Xu, Jinjing, Xu, L. H., Xu, Qingtao, Xu, Shuyao, Xu, Suting, Xu, Tiantian, Xu, Tianxiang, Xu, Weixin, Xu, Xinran, Xu, Yangchuan, Xu, Ye, Xu, Yueni, Xu, Ziyao, Xue, Haonan, Yan, Junjie, Yan, Yaoyao, Yang, Fan, Yang, Guangyao, Yang, Hao, Yang, Junwei, Yang, Ruoyu, Yang, Wenjie, Yang, Xiaofei, Yang, Xinyu, Yang, Yi, Yang, Yiling, Yang, Ying, Yang, Yuchen, Yang, Zhen, Yang, Zhilin, Yang, Zian, Yang, Zuhao, Yao, Haotian, Ye, Dan, Ye, Haoran, Ye, Wenjie, Ye, Zhanbo, Yin, Bohong, Yin, Haoxiang, Yin, Xietong, Yu, Chengzhen, Yu, Haozhen, Yu, Longhui, Yu, Shengnan, Yu, Shuying, Yu, Tianxiang, Yuan, Enming, Yuan, Mengjie, Yue, Tongtian, Yue, Wei, Yue, Yang, Zha, Dunyuan, Zhan, Haobing, Zhang, B. H., Zhang, Dehao, Zhang, Fei, Zhang, Hao, Zhang, Haoyuan, Zhang, Huanyu, Zhang, Jiapei, Zhang, Jiaxuan, Zhang, Jin, Zhang, Kaiyi, Zhang, Miaozhen, Zhang, Puqi, Zhang, Qinglei, Zhang, Rong, Zhang, Rui, Zhang, Shaoshuai, Zhang, Shiyi, Zhang, Xiaobin, Zhang, Xiaoyun, Zhang, Y., Zhang, Yangkun, Zhang, Ye, Zhang, Yichi, Zhang, Yikun, Zhang, Yizhi, Zhang, Yongting, Zhang, Yu, Zhang, Yutao, Zhang, Yutong, Zhang, Zheng, Zhang, Zijing, Zhao, Bin, Zhao, Chenguang, Zhao, Feifan, Zhao, Jinglun, Zhao, Jinxiang, Zhao, Shuai, Zhao, Wenshuo, Zhao, Xiangyu, Zhao, Xuanle, Zhao, Yikai, Zhao, Zijia, Zheng, Haozhi, Zheng, Huabin, Zheng, Ruihan, Zheng, Shaojie, Zheng, Tengyang, Zhong, Haofeng, Zhong, Lei, Zhong, Longguang, Zhou, M., Zhou, Qiankang, Zhou, Runjie, Zhou, Ruozhang, Zhou, Xinyu, Zhou, Yiqiao, Zhou, Zaida, Zhu, Jinguo, Zhu, Liya, Zhu, Xinhao, Zhu, Yangjunfeng, Zhu, Yuxuan, Zhu, Zhen, Zhuang, Chen, Zhuang, Weiyu, Zu, Xinxing
Abstract
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.
Chinese Translation
我们介绍了 Kimi K3,这是一种具有 2.8T 参数的专家混合模型,激活参数达到 1040 亿,具备原生视觉能力和 100 万令牌的上下文窗口。Kimi K3 基于 Kimi Delta Attention 和 Attention Residuals 构建,这些技术改善了信息在序列长度和模型深度之间的流动。结合 Stable LatentMoE,该模型每个令牌有效激活 896 个路由专家中的 16 个,以及经过精炼的训练和数据配方,这些进展使得整体扩展效率相比 Kimi K2 提升了约 2.5 倍。后训练阶段强调了在一般、代理和编码领域的强化学习,以及多个推理努力水平,支持组合泛化和稳健的长时间执行。在 2.8T 的规模下,Kimi K3 得益于多个领域的基础设施进展:KDA 的算法-系统协同设计、完美平衡的专家并行训练与高效的内存管理、具有持久回放和沙箱状态的百万令牌代理强化学习,以及部署创新。广泛的评估表明,Kimi K3 在长时间编码、代理、知识、推理和视觉任务中达到了前沿水平的性能。尽管其整体性能仍落后于最强大的专有模型,即 Claude Fable 5 和 GPT-5.6 Sol,但 Kimi K3 在我们评估的其他开放和专有模型中始终表现优于它们。我们发布完整的 Kimi K3 模型权重,以促进未来的研究,并加速前沿智能的更广泛部署和采用。
cs.CL / 78 / 2607.24717

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

数据交响乐:学习针对每个实例的预训练数据策划
Huang, Zhen, Wang, Yikun, Xia, Shijie, Liu, Pengfei
Abstract
Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each example. We propose DataOrchestra, a framework that unifies different processing operations and orchestrates an example-specific pipeline for each example. Given a chunk of pretraining data, an orchestrator decides whether to drop, untouch, or clean it. For a chunk to be cleaned, it selects one or more downstream operations, ranging from programmatic editing to different forms of LLM-based rewriting. For each rewriting step, it further generates a concrete instruction, which is executed by the corresponding downstream tool model. We pretrain models from 0.5B to 7B from scratch on web data processed by DataOrchestra and observe stable average gains over individual data-processing methods across 11 benchmarks. DataOrchestra is also effective for math continued pretraining and outperforms stronger processing baselines, while reducing processing compute by skipping unnecessary downstream operations.
Chinese Translation
预训练数据处理对大型语言模型(LLMs)的下游性能至关重要。然而,许多现有方法在语料库或领域层面定义固定的处理策略,并将其统一应用于多个实例,而未能根据每个实例的需求进行调整。我们提出了DataOrchestra,一个统一不同处理操作的框架,为每个实例编排特定的处理流程。给定一块预训练数据,调度器决定是丢弃、保持不变还是清理它。为了进行清理,它选择一个或多个下游操作,这些操作可以是程序化编辑或不同形式的基于LLM的重写。对于每个重写步骤,它进一步生成具体的指令,由相应的下游工具模型执行。我们从零开始在经过DataOrchestra处理的网络数据上预训练了从5亿到70亿的模型,并在11个基准测试中观察到相较于单一数据处理方法的稳定平均增益。DataOrchestra在数学持续预训练中也表现出色,超越了更强的处理基线,同时通过跳过不必要的下游操作减少了处理计算量。
cs.CL / 79 / 2607.24720

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

多回合长远规划的物理学:通过单教师和多教师政策性蒸馏从预训练到后训练
Men, Tianyi, Jin, Zhuoran, Liu, Kang, Zhao, Jun
Abstract
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explicit world model construction through CoT state transition modeling yields stronger long-horizon generalization. Atomic skills alone are insufficient for compositional generalization, whereas a litte long-horizon data works. Moreover, suboptimal trajectories severely impair performance because errors amplify over long horizons. (2) Planning ability shaping via GRPO and OPD post-training. Through mutual information, we distinguish general planning patterns from task-specific planning knowledge. For planning patterns, we identify three application regions of post-training: unnecessary, effective, and unsupported. OPD has a broader effective region than GRPO under low-quality and long-horizon settings, as it provides more consistent update directions. For planning knowledge, distilling unseen procedures from a teacher with different knowledge may impair student's prior world modeling without fully establishing new knowledge. (3) Planning ability integration through MOPD post-training. We show that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments. Compatible patterns enable cross-environment generalization, partially shared patterns support continual learning, while completely conflicting patterns cause severe interference.
Chinese Translation
多回合长远规划对基础模型代理至关重要,但如何从根本上改善这一能力仍不清楚。现有模型在不可控且不透明的互联网数据上进行训练,这使得识别规划能力的获取、塑造和整合变得困难。为了解决这一挑战,我们引入了一个统一且可控的多回合环境,能够实现精确控制。这一环境允许我们系统地研究三个阶段的长远规划。(1) 在预训练阶段获取规划能力。我们研究数据格式、分布和质量。通过链式思维(CoT)状态转移建模构建明确的世界模型,可以获得更强的长远泛化能力。仅靠原子技能不足以实现组合泛化,而少量长远数据则有效。此外,次优轨迹严重影响性能,因为错误在长远规划中会放大。(2) 通过GRPO和OPD后训练塑造规划能力。通过互信息,我们区分一般规划模式与特定任务的规划知识。对于规划模式,我们识别出后训练的三个应用区域:不必要、有效和不支持。在低质量和长远设置下,OPD的有效区域比GRPO更广,因为它提供了更一致的更新方向。对于规划知识,从拥有不同知识的教师蒸馏未见程序可能会损害学生的先前世界建模,而未能完全建立新知识。(3) 通过MOPD后训练整合规划能力。我们展示了多教师政策性蒸馏(MOPD)通过在环境间收敛到共享规划模式来整合能力。兼容的模式支持跨环境泛化,部分共享的模式支持持续学习,而完全冲突的模式则会导致严重干扰。